May 11, 2024 · Past Event
The roundtable explores the interconnection between vision and attention, examining how these faculties shape our perception of truth and falsehood. Core topics include animal consciousness, sensory perception, and cultural evolution. The discussion addresses how vision synthesizes other senses and provides evidence for claims of truth or falsehood -- connecting eyewitness testimony to certainty. The event also considers vision's role in creating images, written language development, and historical cultural evolution.
This roundtable assembles vision scientists, a depth perception researcher, and an AI engineer to explore how vision works and its relationship to consciousness and artificial intelligence. The panelists examine why primates evolved such dominant visual systems, with each expert initially trying to champion other senses before acknowledging vision's exceptional information bandwidth and its unique capacity for active exploration through eye movements and scanning.
A major theme is the gap between naive assumptions about vision as camera-like recording and the reality that the brain performs massive inference and interpretation. The famous dress illusion, inattentional blindness experiments, and studies with congenitally blind patients who gained sight in India all demonstrate that vision is fundamentally constructive. Paul Linton challenges the panel with evidence that distance perception relies more on top-down cognitive processes than stereoscopic depth cues, while Kevin Chan discusses how even early visual processing in the thalamus involves substantial top-down feedback rather than pure bottom-up data flow.
The conversation bridges naturally to artificial intelligence, with Andrew Shum explaining how convolutional neural networks initially mimicked biological visual hierarchies and how attention mechanisms became key innovations for AI. The panel grapples with deeper questions about consciousness, mental imagery, and dreaming, noting the recently popularized phenomenon of aphantasia and debating Kenneth Miller's characterization of vision as controlled hallucination. The discussion concludes with honest acknowledgment that neuroscience remains far from explaining how visual experience comes together or whether AI systems could achieve genuine consciousness.
00:00:00 will be the last until the fall and today we had a meeting in fact discussing the programs we will have in the fall and uh Dr haritz is the associate director and he will introduce the participants and then we'll get going good afternoon everyone thank you for being here on this glorious um spring day uh yes this is our last uh Round Table on the topic of the senses
00:00:32 and this is on vision of course um before we get started I want to introduce our panelists um so you can if you could raise your hand when I call your name Alysa amof is an associate professor of psychology at forom University prior to joining The Faculty at forom she was a research scientist at Carnegie melon University in the department of psychology and was an adjunct faculty at the robotics Institute at carnegi melon University Dr amof received her PhD from the
00:01:03 Department of psychology at Harvard University she then went on to a postdoctoral fellowship fellowship at the department of psychological and brain Sciences at University of California Santa Barbara Dr amanov uses an interdisciplinary approach by employing complimentary methods to explore the cognitive Neuroscience of visual scen understanding her research is at the intersection of vision and memory exploring how the the visual world is interpreted based on experience she uses fmri to explore how the human
00:01:35 brain process processes and represents objects scenes and the relationships between them Dr Kevin Chan is an assistant professor of Opthalmology Radiology neuroscience and biomedical engineering at the NYU Grossman's school of medicine and the New York University tanton School of Engineering he completed his doctoral studies in biomedic engineering at the University of Hong Kong and was awarded the Lee K Shing prize for the
00:02:08 best PhD thesis at the University studying Imaging of the visual system at NYU his laboratory focuses on new non-invasive methods for Imaging neurod degeneration neurodevelopment and neuroplasticity in both humans and experimental animal models of vision related diseases and injuries to guide Vision preservation and restoration his team combines the use of optical Optical coherence tomography magnetic resonance
00:02:39 imaging neuromodulation and psychophysical assessments to determine the processes underlying the interplay among eye brain and behavior in health and disease Dr Paul Linton is a presidential scholar in neuroscience and society and a fellow of the Italian academ Academy at Columbia University specializing in human 3D Vision he is the author of the perception and cognition of visual space
00:03:10 and the lead editor of the Royal Society volume new approaches to 3D Vision he has made significant contributions to our understanding of stereo vision of course how we see in 3d with both eyes and is developing new approaches to visual scale the perceiv size and distance of objects in VIs ual shape the perceived 3D shape of objects he has worked on virtual reality as part of the deep focused team at meta reality labs and taught philosophy at both Oxford and
00:03:46 UCLA Kenneth Miller is the Peter Taylor professor of Neuroscience co-director of the center for theoretical neuroscience and co-director of the neurobiology and behavior graduate program at Columbia University he received his PBA from Reed College his Ms and PhD with distinction from Stanford University and completed his post-doctoral work at UCSF and calac he is a founding member of the editorial board of Journal of computational
00:04:18 Neuroscience and has served as faculty for many years at various summer schools in theoretical and computational Neuroscience he received the Schwartz prize in theoretical and computational Neuroscience s given by the society for neuroscience and has also been recipient of the Alfred P slan research Fellowship sural Scholars award and National Science Foundation graduate Fellowship Andrew shum is an entrepreneur with a background in software engineering and artificial intelligence he was co-founder of thread
00:04:50 genius which built computer vision algorithms for retail and e-commerce applications including visual search and discovery thread genius was acquired by Auction House Souther bees where Andrew focused on expanding the auction business through machine learning technology his team pioneered efforts to incorporate Vision models for understanding consumer taste and price estimation of artwork before founding thread genius he was a software engineer
00:05:21 at Spotify working on music recommendation problems Andrew completed his undergraduate studies in computer science and physics at MIT okay sorry okay well again welcome everybody um we um really interested in hearing what our panelists have to say today about Vision I thought one way to start off was a question that was raised while
00:05:51 we were uh getting together here which is related to why was the sequence for our course for our course of roundtables on senses sequenced the way they were we went from touch to taste to smell to hearing and then vision and I think in some sense people feel Vision may have a sort of a a top spot in that hierarchy of Senses so I wondered if anyone wanted to take a shot at that and take it from
00:06:22 there I don't know if he really Top Shot but if we really want to quantify like the bandwidth of uh the five senses there more than five senses um well for smell and taste probably is hard to quantify but we may have roughly five tastes and then smell we may be able to uh identify a couple of different chemicals like I mentioned last month if I remember corly uh in a round table uh touch um is Al there are some ways you can quantify the um like this um able
00:06:53 the ability to differentiate texture and uh you can differentiate from roughly uh 50 Herz all the way to uh 300 Herz sound um I think uh two weeks ago um one of our speakers mentioned uh we can actually detach from 200 200 Herz all the way to 20K Hertz now Vision um we actually have a pretty wide band with from uh in terms of nanometers and we can detect from Red all the way to VI Violet some people may have a little bit
00:07:24 of sense of beyond that um so the the bandwidth and you can also process everything in parallel of different features right the color the texture um different uh scenes that Alisa is very expert of it um so maybe that's the one of the reasons that we try to phrase this way and human actually occupies almost onethird of our brain dedicated Visions actually I wanted to inquire more about what you said about bandwidth one and um and the fact that there are more than five um I guess yeah what how
00:07:55 I interpreted was just sort of like the information density of like the content you're dealing with whether like I know to me it's just like touch occupies like just less it can be represented with like less bits than say like like visual content but what you had said about more than five sents is like are you sort of I was curious how you would like Define a sense using like using those I guess Dimensions or
00:08:26 features I mean another another aspect is just a question of science right so how do you test these various different senses um and with vision you can use displays for for smell um for touch for Taste it's always been very difficult to to sort of isolate particular signals in that way so I I think historically as well there's been a a good reason why the last 200 years of of sort of Behavioral work has tended to focus on primarily on Vision I think also the reason why we feel vision is the most important is because we're primates and
00:08:57 primates evolve to heavily specialize in Vision our our cortex developed multiple areas to process Vision far more than for any of the other senses um so um if you're a monkey and we're basically monkeys then um Vision vision is at the top and and dominating and captivating in that sense right the vision is one of our most Salient resources of our attention um and vision will get us more so than um what's happening in our
00:09:27 visual world would dominate how we experience the rest of our senses but for the rats on the street it vision is pretty low down you know touch and smell are what much higher so it really depends on how your sensory systems have evolved yes even though R and and mice are relatively poor in Vision it's actually amazing to know that for rodents they are evolutionally also have some preferential sensitivity life um for rest they usually um in the past they were in the desert and they have to
00:09:58 avoid praise from the top of the surroundings so they you actually have a relatively High sensitivity towards the top of our visual field whereas for humans actually we are more a little more sensitive towards the lower B of the field probably because we want to avoid the obstacles on the floors so uh even the resolution is indeed different but the fishing is still playing some roles uh that may be important for different species and if we extend further egos actually they have even sharper visions and uh for humans normal vision is 2020 that mean means we can
00:10:29 see Vision um uh for number subject you can see subject an object that is about 20 ft uh but for ego uh we they can actually see something that is 25 that means uh you have to go all the way to 5 ft in front of the subject you know to see something that a ego can see 20 ft from that so it is really amazing to uh understand how the world of vision is really evolving um and to explain further jellyfish at 24 eyes for some of
00:11:02 the species so a related question is what do we all mean by Vision right so we're s contrasting Vision against these other senses but I I guess that we're probably all doing slightly different things with a slightly different Focus so if if each of us had like a sentence to sort of highlight what what we take to be the central question that that we're sort of excited why does vision get us out of bed in the morning right I would say that it starts with light
00:11:32 and it's about how our how we've learned to interpret that light and it can be very different for you and I um uh but and the how that experience helps us to filter all the information that comes to our eyes which is enormous so talking about just the bandwidth um and how much information we're getting through the light that hits our eyes is enormous and our experience helps us to filter it and again concentrate on what is sing are most important to us so my my thought
00:12:03 about vision is is the perception of how we interpret the light that falls on our eye so so in a brief um few sentences I think fishion can be uh separated two components for at is for humans one is to form images so that you can really interpret it uh surroundings and environment like Alisa mentioned um there's also another important uh feature that is to process non image um forming um features for example I just
00:12:34 came back from um Seattle yesterday actually and on Tuesday I was actually back from Singapore so all these light changes actually would affect our Cadian rhythms and all these unconscious changes and yeah I'm half doing but at the same time all these are actually uh important but we may not be aware until there's some changes uh uh from the surroundings there's also something important about the formation of objects how we segregate our world into different objects and the role that
00:13:06 Vision plays in that which I think may be pretty distinct compared to the other senses I don't think any sense you do that with I mean if you think about when you have your eyes closed and you're feeling something you're you're inferring what kind of object it is it's a it's a glass it's a bottle it's a piece of paper um and I think well it it depends on uh you know exactly what their nervous system wants to figure out about the world but I think at least for many creatures certainly for mammals um whatever
00:13:38 sensors they're using they're forming objects uh they're forming a picture of the objects in the world at least the ones that they relate to um we just happen to really uh lean on Vision to do that but but I would argue it's definitely not vision and isolation we understand objects by being in an environment in which we interact with things that is how we understand objects how we can um work with them the affordances that they give us how we can separate them from the background and if
00:14:08 that is important to us um so vision is largely I think impacted by all of our other senses in order to segment objects from the background based on importance of our other senses I think Helen Keller has always been a good example of just understanding that we don't need Vision that to segregate objects I mean our brains make sense of the world with whatever senses they' got and the fact that she could do it you
00:14:39 know without Vision or hearing is just it's extraordinary but it makes you realize you know what it is our brains are in the process of doing with whatever sensory information it gets she was part of a cultural she was she was part of a general culture already right so even though you're right that's miraculous that she was able to create the sense of the world without vision and sound but still she was learning from people who had those experiences right that um well surely
00:15:11 it's interesting I think we have this panel of people in Vision I'm trying to get you all to Crow about vision and now everyone's saying it's no big deal all the other senses to to provide it but so what's distinctive about what would you say so what would you all say makes it distinctive visual experience I mean it's very different to see something visually than to than to touch it or to taste it or to smell it right and so I one thing I I I think we have to be cautious about is there's a certain Trend in cognitive science of just thinking of vision in terms of information processing right and so what
00:15:41 you'd say is well Vision gives us certain information touch give us certain information smell Etc and then the brain integrates all that information into a coherent sort of percept multisensory percept but and that that's an account you know and and the brain's cly doing something on that level but you wouldn't want to let the term information loose get rid of the fact that we have very different experiences right so if you just think in terms of information information processing and information integration you're losing that sense in which the
00:16:13 color red looks red right and it's very different from the sound of a bell to to use an example um so yeah one of the challenges I think I think cognitive science and Neuroscience is facing is how do we bring that experiential aspect back into the subject um it had always been there it been a sort of a key feature of the work in the 19th century um and then I think in the 20th century it sort of disappeared in the early 20th century with a focus on behaviorism uh it didn't in other subjects like Philosophy for instance it
00:16:44 was always focused very heavily on on Raw experience um and so yeah I I think one of the challenges for us is is to bring that experiential aspect back into cognitive and Neuroscience and in terms of information processing actually there are some basic r rules of how to define Vision as from the psychological Consciousness point of view um a few of them may include like for fishion you have a limited field of f so your objects may be uh not perceived if you're outside of this field of f if you put like a a bottle versus uh table
00:17:17 together you'll be able to perceive uh the front and the back the depth perception and then if you move around or you rotate the object you may see perceive is the same object or you can also see different sizes when you move around uh to the close t or more near side space um so um if we can use uh sound or touch to understand these features could that also be at least indirectly Vision that is something perhaps we can discuss further uh and
00:17:47 our lab actually has been trying to help some of our blind subjects to um to use remaining senses uh like the sound and also the touch you know to interact with the visual environment and we are trying hard to uh for example use some video cameras to convert visual image into these sound and touch features and see how we can try to um um further uh interact with these um low level features of the fishion even though they don't have input from the iOS yeah I think if you were ask people
00:18:20 like if you were only choose one sense I think most people would choose Vision to to sort of keep but here asking the question like which is the most pleasurable too I think you would get uniformly distributed answers at least that's what I've I've found in my experience I think one of the challenges of thinking about information processing is that it can change very much with a goal um and what you're trying to get out of the visual information um and our brain segments visual information based
00:18:51 on those goals you know you can go as broad as thinking about just understanding where things are in space versus identifying what it is and uh so there's not one type of information processing um and I think we don't know actually all the different types of information processing that we're doing with visual information and I think that brings us to a complexity of trying to understand it computationally um because there's so much about a human using this um visual information rather than something like a
00:19:25 computer so so how are we going to answer the questions of vision then I think I think everyone around the table use very different techniques and approaches um why why do people favor C certain approaches over others what just not sure it's exactly addressing that but I think part of I like I was trying to think about well why did Prim AG focus on vision and I don't think we know exactly the answer but if you think about the different senses Vision lets you scan the world you can look out at
00:19:55 distances you can look out and and see what is out out there um hearing you can't scan the world something has to make a noise in order for you to hear it something doesn't have to make a light in order for you to see it um smell again you can only sample the smells that come to your nose um you can't just scan the world with smell or with taste or with touch um and so I imagine that that ability to scan the world for things for forces for things that you that the animal wants to make use of
00:20:26 must have been somehow driving primates to specialize on Vision but it's just a speculation but that assumes that we're in light that yes in the dark obviously it's not so good yep and also echolocation gives you the same ability which you know so a number of animals have so they can also scan the world and form images of you know all the objects around them but which sense is most essential for life for life I don't think it life just evolves in so many different forms and and so many different niches and each
00:20:56 one has different sensory requirements how about for human human life say it again for survival well again that's life but I mean it depends on which ecological mits you're trying to survive in and how you evolved for us clearly Vision I think but there is actually a few surveys uh trying to um understand uh if someone got cancer if someone lose their arms L their sight's hearing which one they fear the most um sight is actually the most um important one that they will fear the most and um and we
00:21:27 when we talk to the blind patient or people especially when they um were um acquiring dead blindness late actually um they uh always say they the first thing they want to know is really like the um dness similar to the echol location because they do not really develop those echol location intrinsically but once they lose their site the the thing they can really interact is actually something that's within their hand reaching distance so they want to know further what in when you get into this room for example uh if
00:21:57 you have no light uh is total darkness so what is happening how many people around will we bump into a chair or a table so all these actually get scanning of The View with a few shots of fishion actually is really important uh compared to other senses I would say I'm curious I think this is probably um an answer you would know but like in terms of information processing time with with the with the fastest speed vision or another sense
00:22:28 not sure about this actually yeah uh but in fishion there are also uh slow conduction velocities for like color processing there's also fast conduction velocities for like changing in contrast so how are they directly compared to sound I guess it really is more complicated if we want to directly compare them and what's the end go there like what at what you know you can count how long it takes for a visual cortex primary visual cortex to respond to light but that doesn't tell us very much that's not you know if we stopped vision then we wouldn't understand very much
00:23:00 about what we're seeing right there's like the the process of internalizing that information and making sense of it yeah yeah with that that there is also um conscious um Vision that is really through the entire facial pathway including the CEST but uh there are also patients who have their CEST actually got removed of um damag uh because of stroke for example they caus cortical blindness um these people even though they cannot really interpret or perceive Vision uh consciously uh if we ask to guess uh whether they see something um
00:23:31 they could often um be able to get the right um Choice by um above chance so somehow we call it blind sight or yeah cons unconscious uh Vision some people argue about this term but there's actually uh some pathway that may be related to this kind of uh unconscious or subconscious uh Vision this may be a little faster than a conscious Vision um or maybe their pathway that are connecting to different parts of the brain when their quartes is damaged it
00:24:02 so it really depends on the situation I guess yeah there was a a recent uh this I think this past week or so is an article in nature about um it's about fish and uh the the fish they found in this whatever ecosystem it is the the ones that were exp explored more if they the species that were engaged more in Exploration seem to play a greater role in the further development evolutionary
00:24:33 development so there was more diversity in those types of fishes that were exploratory as opposed to the ones that were less exploratory and almost as if as if it accelerated natural selection process so was interested in this idea about scanning which is of course one of the ways we explore the environment I'm wondering about scanning and what that might have to do with our attention and our memory and things like like that I wonder if so um I had a very simple task in my lab where we just had to look at
00:25:04 scenes and we just had to label all the objects in there there was no time limit we were doing this to get a database of images and the objects in them and the amount of individual variability there was was astounding to me actually um that you know I saw a tree somebody else totally didn't see a tree somebody else saw a chair I didn't see the chair and it was It was a astounding to me how much what I see in that scene is different um than compared to another person and then once I see okay I should
00:25:35 be looking for a chair where's that chair I see the chair but attention is everything yeah that's also an interesting um experiment called Invisible Gorilla right so uh that yeah you may have heard of it like people are passing the basketball uh for people who are wearing white shirts versus black shirts and then the the task was to count how many passes were for the white shirts in the end most people can count it correctly but uh actually the pun is there's actually a gorilla which is black walking around the scene and they were not able to I'm sorry I'm sorry
00:26:07 yes right yeah so this may that was the second yeah that's right right right yeah that was Dan Simons and um Chris shab yep exactly and it's amazing I mean you do it every time like every class I I show that it you know everybody falls for it every single time yep yeah if we move even closer to recent a social media uh sharing there there's also a um a post that is comparing asking you about the dress right there's a blue uh dress that has
00:26:39 some blue and uh dark bands and people ask you what do you see what color you see actually um a lot of people see the blue and the uh black band but quite a lot of other people also see this is a gold dress so even though it's the same picture you can actually perceive differently um maybe it's Rel to the brain experience maybe is R to cognition um I'm not experiment there there's a lot of date on that to say say that it's about how you interpret um the scene and if you're looking at the dress within a
00:27:10 shadow or within direct light and that is how you interpret the color so there's been a lot of research like whole issues of journals on interpreting Wonder um love to read that how we how we uh perceive the dress but I think what that's getting at is is that you know we our naive experience is that we're sort of you know our eyes are like a camera we open our eyes the light comes in and there's a picture uh and it's not like that at all the the brain has to infer
00:27:40 the picture and paint the picture um and one of the inference that it does is is it infers the color of an object so if if the light coming from an object is yellow light that doesn't mean that the brain perceives it is yellow because the brain has to compare it to the illuminant if the illuminant is almost all yellow then all the objects are going to be primarily yellow but the brain tries to infer well that one is reflecting more green well that one is reflecting more blue and that's how your brain assigns colors and that's why the the blue dress looked different colors
00:28:12 to different people because it just happened to be an unusual picture that brains could interpret in two different ways in terms of what was the illuminant and what was the object and your brain is conf is Computing what's the reflectance of the object and if your brain makes a different assumption about what the aluminum was then it makes a different conclusion about what the reflectance was and what the color is that dress that picture just happened to be right at the cusp between two different interpretations there's there's a great
00:28:42 picture of um a bowl of strawberries there's a lot of people that study color and they they um one of them has this picture of a bowl of strawberries and it's actually black and white um but you can swear that those strawberries are red um and you really interpret those strawberries as red and then when you isolate just a patch of that picture you can see that it's completely gray um and all my students fall for this and and are astounded that their Vision tricked them so much but it's it's it's how we
00:29:14 interpret the world and if we didn't do this to help facilitate Vision it we would it would take forever to put one's foot in front of the other because we would take so long to process everything um we have to use our experience to help interpret in order to move things along well I don't think it's I mean we do use our experience but I think a lot of this is built in by Evolution I mean Evolution knew how to do the computation of the reflectant um and similarly the the size of objects you know if you're on a beach and you see people far away
00:29:44 if you perceived them at the size they are on your retina you would think they were ants but you know they're people because your and your brain actually blows them up and makes them bigger in your in the image you perceive than the than they are on your retina because your brain computes it they're far away and therefore makes it bigger and that's believed to be why you know the Moon looks much bigger when it's on the horizon because on the horizon there's a lot of things between you and it and your brain says oh that's far away and it makes it bigger this this is actually a like a very topical and very
00:30:15 hot problem in like in self-driving cars right now um just how to deal with these edge cases by providing experience or data like there's a famous uh incident which like uh has a autopilot like crashed into a truck cuz it was stopped at a you know intersection and a Truck Rolls by and it's like the same color as the sky and so you know that's that's a particular Edge case cuz it hasn't SE in trucks that look like that apparently so yeah no it's interesting to see the parallel between like you know The Human
00:30:45 Experience and using that part of the brain to kind of you know merge what what you perceive and I think that kind of relates with the um the case with the the blue black dress it's like I think the people who are kind of um you know people who know anything about a bit about photography or how kind of exposure Works they might be able to kind of adjust the blue color by you know knowing that you know that camera exposure was probably way way way high and if they were taken at a normal exposure it would be like look at black dress um but it's it's interesting that we're
00:31:17 we're discussing images right you know in terms rather than sort of real world per or objects um so I I I think you know 20th century in terms of vision science has been very dominated by images on displays um be very interesting to see how that changes as as we go further towards virtual and augmented reality um yeah we're getting what's known as more ecologically valid so you can sort of move through a scene you can you can uh you know understand the scene
00:31:48 slightly better so in the case of the dress for instance the illuminant would be there in the scene with you in addition to the dress in addition to all the other things um so it's it's an interesting question right you know have we been studying Vision or have we been studying picture perception um and and yeah so just for two things to that is um so we are dominated by studying images for example in the scanner right looking at using fmri to understand how the brain processes Vision we've been stuck with images right um uh Jodie um C
00:32:19 has been able to bring 3d Real World objects into the scanner um and has people looking at real objects versus picture and how much more that elicits activity in the brain and how much more we the brain is active and how the big differences is important um and astounding in a lot of ways um so I think it is really important and there's a lot of um work being done about understanding vision and making judgments about Vision when you're
00:32:50 actually standing in the environment which is different than looking at a picture me how much in this in these in these instances is that you know you can you might have the potential to interact with the object versus where you know right from the beginning I cannot interact with it what role that plays in our visual understanding I think it's huge I don't think that we trick the brain by saying that an image is the same thing as you know holding a chocolate bar versus
00:33:21 seeing a chocolate bar is very different right and this goes again to the fact that vision is not done in isolation and we don't see things um in isolation of the rest of human cognition um and so I I I don't think we trick the brain and uh but it's a kind of a necessary design in order to have a lot of control to be able to study different aspect of vision there's a very um interesting project similar to what you said but the project is called project prash and it's
00:33:51 actually a very meaningful Outreach to the um Villages that are relatively developing on dev in Northern India and they uh the Physicians try to scream people who are congenitally blind but potentially treatable um during towards the adult toe or teenagers time so uh when they were able to for example uh replace their uh character when they born blind uh during adulthood um and then we gave the um for example an a card of the same
00:34:23 object versus a real objects uh these patients were only able to use their touch in the past right before the treatment uh when we're immediately showing the two together and put two uh different objects together for example they were not able to recognize the from the image without interaction but with some more um experience uh they will be able to gradually really uh identify using those um images so there's something really relate to interactions experience at the same time going back there the brain may also be hotwired
00:34:54 evolutionarily even if someone is conally blind um and we try to use uh the technique we mentioned using sound to interpret some words or maybe use Touch to read Braille uh there are some visal word form areas that means uh brain regions that originally for thought to be uh for interpreting word would also be recuited so hope there also there both nature and natur putting together I mean on the flip side uh one of the questions I work on is is
00:35:25 distance perception so how do I know the distance of an object object um and the more and more I work on distance perception the actually funnny enough the less and less I'm convinced that it's that important that you're in the scene um so people would talk about sort of the motion of of your body but actually when you isolate that in experiments that doesn't seem to make a massive contribution um people talk about the the angle of inclination um of of the floor or of a table again uh it's it's sort of quite debatable in experiments and then you and you think well we spent sort of two years in the
00:35:56 pandemic um you know looking at each other on video screens um very rarely making gross areas in terms of this the scale of scenes um so that that's one context where I I I I the more I work on it the more I think actually quite similar uh judging scale in see in in the real world and also in in in pictures thanks good question it doesn't matter it doesn't play a role like knowing your size in relation to everything else so like I know the distance of how far my arm goes out
00:36:28 doesn't that help me interpret everything else uh sure sure but but in terms of I'm just trying to think so so there's some there some definitely some benefits of stereo Vision so having having the two eyes so you can increase or separate the separation between the eyes and that will affect uh scale judgments but actually when you even when you look at that that seems to be a very very high level process um that doesn't rely on sort of basic mechanisms so for instance was thought the visual
00:36:58 system just had a a rangefinder from the two eyes rotating and fixating on it that doesn't seem to work um so it's it's interesting because you look at the textbooks they say well look you know you've got all of these sort of what they call triangulation cues right it's sort of different different viewpoints uh so from which you can sort of infer like you would in computer vision the position doesn't seem that human Visions doing that it seems like it's treating even the real world scene in my research at least right and I I can be wrong um you know bit bit bit more like the
00:37:30 pictorial case yeah how similar are you sort of comparing these tasks CU I think it's like to me like you you task a machine to like estimate distance like it's going to give you a pretty precise answer versus asking a human I mean they could probably tell you I mean this is not my my fa research but they could tell you like relative distance but you know kind of giving you the answer of oh this is like one meter versus like 1.1 I think it's probably a bit more hairy I mean it
00:38:00 falls off distance but but I think it's fair and where I tend to test things is really en reaching and grasping um uh and they you really you know people they have a bit of online control so you can sort of see your hand but actually that I I don't think that plays a massive role somehow if we move from the zoom um Tod images back to a little bit uh in the early 90 uh centuries um I guess U one of the famous painters Picasso is
00:38:31 actually also interested in those death perceptions but um if you look into his own portrait uh his right eye is actually kind of misaligned U with the left eye so potentially he may not be able to really focus on the death perception but that actually gives him advantage of seeing things in 2D like in a zoom View and he he has a very strong visual memory he can stare into an object very long time time and after after going um this subject or object is gone um he can still repaint it uh and
00:39:03 add his own interpretation so uh if we want to go back to the question is it just fishion from the eye or there's something interpreted by brain maybe visal memory also plays a very important where um along with this experience it sort of suggest doesn't it that there's some more holistic processing or and top down processing that accounts for what you're finding in your research yeah yeah so so definitely not you know and I I'm happy to be wrong
00:39:33 about this right but for my own stuff the the way I think a visual scale is is is much much more I would say quote unquote cognitive in its its relation than than the 3D shape of an object which I I think is provided by by stereo Vision pretty pretty directly um which which is the depth perception from the difference between the two eyes um and that's that's interesting because you can you can mess around with other sources of visual information like shading and perspective and very rarely does that seem to to impede that just
00:40:03 that immediate percept of 3D shape from stereo so in my own work at least I I draw a sharp distinction between 3D shape and and scale um but I can imagine other people don't well regarding top down um attention of feedback uh goal oriented stat um um we were also interested in the meaning of the fish ctis that means the brain areas at the back of the brain and are they actually playing some roles from the eye towards the brain that is
00:40:34 more top down or if there is nothing inputting from the eyeballs uh is the brain uh So-Cal visual brain is it still using it use it or Lucid principle right uh it turns out that uh if there are patients who have gone blind either congenitally or late um we were able to still recuit the So-Cal visual area of the brain uh and make uh uh at least from the functional Imaging point of view they will become active when we perform some other tasks from the other senses um one of the um potential
00:41:07 meaning we looking into it is maybe related to the top down attention perhaps even uh even for normal vision there is the So-Cal goal oriented task but even if you have no more um input from the eye it may still be running um in terms of these selective uh attentions tases or but using other senses Maybe some repurposing and what and and a lot of people were saying that um for blind patients they may be having super Supernatural senses of the sound or
00:41:38 touch does that mean this is because they have a spare visal area that can help to run these um so the brain is fantastic we're still looking into the meaning of it right now yeah and just to add to thinking about top down and thinking about how the brain is wired um there's been a lot of work to say that we have a lot more more connections going feedback and going from more higher areas to lower areas than bottom up um which says that there's a much stronger need for and influence of kind
00:42:09 of top down information rather than bottom up even at um even at the thalamus where we would think the phalus is one of the first places that um the information from the eye goes and we would think and then that goes to visual cortex and we think that should be a pretty bottomup place and it turns out actually that there's more feedback connections um going into the lgm the thalamus rather than um the other way which s again suggests that so much of what we see is interpreted through our
00:42:39 top- down um expectations and uh feedback so so to to sort of build on that there there's a there's a big debate in the field I mean but different people take different positions um and even with Invision I think there's sort of schools who would focus on one aspect rather than others so if if you you have light coming into the eye landing on the retina it it goes in into the middle which is the the lgn you just discussing then it goes to the the back of the brain the primary visual cortex and then
00:43:10 the suggestion is is called V1 and then the suggestion is it goes in in sort of two visual streams so so one's called the dorsal visual stream which is along the top which people associate primarily with action and then the other is the vental visual stream which goes along the bottom which is primarily people are associate with object identification and and sort of s similar kind of judgments about about the scene so you've got all of that that's a a general model people emphasize certain aspects to downplay others then youve got a separate question which is visual
00:43:41 experience right we just have visual experiences and so there's a real debate of of you know if we have visual experiences where where should we think about them sort of being processed primarily should we think of them as being towards the front of the brain which is where you know if it travels this way prefrontal qu text where you're making more sort of cognitive heavy cognitive inferences or should we think it being primarily back of the brain phenomenon um and that's I find that really fascinating I think that's really interesting um there's just no real consensus um as to how we should think
00:44:13 about those questions I agree and uh well there there are some sort of uh thoughts of people saying maybe the brain is only 10% used it but actually is probably not true in our brain Imaging we found that there's a task involved or even when they just um they dreaming the whole brain is actually coing with each other um it is probably involving those bottom up um pathway from the ey towards the fe1 and then those also ventral stream but it is likely also involving a lot of attention networks Thea mode Network the second
00:44:44 networks all together running depending on the situation I mean just add so quite quite useful because you you might be thinking oh we're all discussing human Vision right but but what what about AI right and and the interesting things from that what I just briefly note is that that vental part stream I was I was discussing from the primary visual cortex down the bottom of the brain that sort of way of of carving up the the processing is actually what was the inspiration for for the most recent sort of AI models of
00:45:15 of vision um yeah they uh referring to convolutional neural n yeah it's actually from what I call it's exactly mimics how um the visual areas are linked together with um kind of the base layers processing things like uh or uh detecting things like the edge of certain objects working way up it will uh then detect textures and then different features like eyeballs or the presence of ears and these are all kind of linked together to be used to you
00:45:47 know um Bally detect certain objects like animals and machines and stuff like that um yeah and then it's it's sort of interesting how that model then kind of improved over the years I think the first that really kind of um made headlines was in in 2012 it was called Alex net um and that was kind of the first real implementation that um uh kind of made a huge difference in sort of all the machine learning models that preceded it and then all the kind
00:46:17 of improvements that were made on top of that were quite interesting because they all had to do with sort of like uh how much information do you retain as you go from the base layer or you know VIs area I think it's five or four and all the way up um uh it's also interesting kind of seeing um just where the limitations of where you can draw the parallels between you know what's going on in the visual cortex and how the machine learning models work um I think open AI had this
00:46:48 example where they would take an image of like a gorilla um and make these very minute changes just on the pixel level um essentially to our eyes it would just appear as like an image of of white noise they would then add it onto this image and it could actually trick the model to believe that it was not a gorilla but indeed um another sort of primate uh but just from from a human perceptual level they looked identical um so it's interesting I think um kind
00:47:19 of from the computer model level you know it's it's I think what they had to do was that what's going on on kind of the neuronet level it's like processing very much pixel by pixel whereas humans kind of take a more holistic approach to object detection and object recognition so I think just going back to the top down discussion and um convolutional models and the evolution of convolutional models since 2012 is incorporating top- down feedback
00:47:49 actually and having those connections go from the higher levels of the um dnns to talk to the lower levels to segment what features and what visual information should be important hopefully kind of building on ra building on just a pixel by pixel model and if you right uh the title of this uh Round Table is synthetic Consciousness right um if we go this term is probably analogous to the one AI term right called artificial Consciousness I guess right now we at a
00:48:21 machine learning stage uh which is uh kind of narrow um um non-age learning but there could be a potential move towards the um machine um conscious stage what do you think about this it's it's incredible seeing how fast um this all this changed since 2012 um because basically the state of machine learning a and deep learning from like 2012 to 2017 was that is basically we we had to
00:48:53 figure out how to like mathematically model visual cortex and mimic what you know what's going on back here um you know we were able to uh you know teach machines to recognize items and kind of detect where items are and stuff like that and then all of a sudden in 2017 uh we had come up with this thing called attention um I think it's modeled after the attention the concept of attention in in the the psych world uh but essentially um uh it allowed uh the
00:49:25 machines to basically retain the information that it needed to then solve more complex problems and that was actually the uh key uh technique that led to the invention of models like chat gbt and llms stuff like that uh were then able to allow these machines to solve like much more complex tests and also uh you know generate just way more in terms of language output um and actually like that plus a bunch of other
00:49:57 key Innovations like led to this kind of um this huge uh uh huge amounts of research in in sort of like generative uh AI um which you know back you know during the 2012 to 2017 error uh it was is extremely inent like if you like I think the most advanced technology um where generative adversarial networks AKA Gans and if you compare the results from that error to like now it's it's a night and day and that I think one of
00:50:28 the the key Innovations was um incorporating uh a number of tech techniques one of which is is is the is attention yeah oh I think what's called attention and deep that said what's called attention by psychologists are yeah it's not it's the same word but it's not clear it's the same concept at all yeah the attention in the the sort of uh deep Nets world is sort of uh I guess one way to put it is just figure
00:51:00 out a more efficient way to retain like um retain State and the areas that matter um yeah previously there was uh um you know the the sort of like models that uh had been kind of cutting edge for natural language processing were recurrent neural Nets AKA rnns and they would run into this problem called like the managing uh gradient problem and is basically when you're feeding
00:51:30 these neuron Nets with like just a lot of text based information it will try to just with uh hold way too much uh State about all the words and in fact uh it would get really good at kind of just the words at the end of like say a paragraph and kind of forget a lot of the stuff uh in the beginning and and suddenly with with attention with this attention um mechanism we could actually you know figure out the right things to focus on when we feed this data into into neural Nets I was um I want just wanted to go
00:52:03 back to this these two Pathways as you mentioned Paul those are the of what and where Pathways that's how they're commonly referred to right the dorsal being the the uh where and the vental being what right just want to had that for people who were interested in looking for it um I think it's really interesting that we're talking about this top- down approach to understanding and and we started off also uh for for a while uh discussing how many in in in the ways that our vision is not like a camera it seems like we'd like to think
00:52:34 of it being built up of these elements that are camera- like elements where there's data and you build up a picture we realize well that doesn't quite work and now we see neural network neural net approaches trying to model attention let's say whether they do that well or not is I guess still being worked out do you think that the ambiguities in these lower levels is an advantage I mean these these ambiguities like for example in the color confusion
00:53:07 and um the ways in which you we can't quite figure out why do we know all the elements of distance it seems neuroanatomically and yet it doesn't add up to a distance so do you think some of the elements being looser or a little bit more ambiguous plays a role in how well we see I think it's more an issue of um we don't perceive the world in terms of the light at each pixel we perceive the
00:53:38 world in terms of objects and so the brain has to infer what are the distinct objects what are their properties what's their color what's their shape what's their size um where you how are they where are they um and I think it's you know you start with the raw information and you have to a lot it to interpret it at that level of objects is very ambiguous it requires quite a lot of inference um
00:54:09 which the brain has to do in or before you can see anything before the picture gets painted uh that you see um so I don't think that there's an advantage in being ambiguous it's just that it's inherently ambiguous the the the the the light that comes into the potentially has there's many interpretations that are consistent with that pattern of light your brain has to figure out what's the most likely one which it usually gets more or less exactly right I mean it's really good uh but there are cases you know of Illusions where it's
00:54:42 ambiguous even before getting into the brain uh if you you uh refer to what you said about the camera structure you have a lens similar to the ey lens uh in the camera versus the eye but uh in the retina is different uh in for for for camera you probably would get a high definition picture of every places but for Ro sorry for eye human eyes uh we have the volier or maula which is corresponding to the central vision you have very dense photo receptor responsible for colors called the cones
00:55:13 whereas for the surroundings of these Central Visions they are relatively low resolution we use them for Low Vision or low luminance Vision they are um highly populated by RW uh they may be of different purpose this is for example for the V high density cones they may be ready for Sharp uh images but uh in terms of ambiguity maybe if you want to at least nav Navigate in dark room or if there is something suddenly from the surrounding just slap over and you want to avoid it uh in terms of these motion
00:55:44 um uh sence um entries then maybe you don't really need first sh um fishing and uh and perhaps that would give you some rooms for you to uh compare between more efficient possibly to to have a phobia and then po resolu there's there are many various conjectures and theories about that I don't think it's there's any definitive uh answer um but I guess it is another way just one more illustration of the
00:56:16 fact that your brain has to uh do a lot of inference to paint the picture that you see which is that in reality there's only what like a thumb a thumbnail width that a you know at an arms length that you really see clearly and the rest you don't see very clearly at all but you don't know that because from your from your eyes moving around and seeing things clearly your brain Paints the picture as if you're seeing everything clearly at once and you're not and included like all the way out here we actually don't see color right we think we see color and we fill in color but we
00:56:48 don't at all um and it just shows yeah how much is being filled in that is beyond what a camera would do what camera would pick up so and in terms of the top down and bottom up I should say you know it's true that the there's generally more top down and bottom up projections but there's a long-standing discussion debate uh what that means because the bottom up projections tend to be driving there m Murray Sherman from Chicago you Chicago came up with
00:57:20 this distinction between driving and modulator drivers and modulators and they're actually different properties are the synapses that make them drivers or modulators but so the top down tend much more to be modulators there's a lot more of them and yet they sort of have less control over what how the neuron responds it's much more bottomup driven but clearly you you know the modulation is doing a lot because in the end we have to do this inference that requires putting a lot of things together but how that works is something that's really
00:57:52 very unclear yeah and something you may be aware is that everyone not only have a yellow spot as the central vision everyone may also have a blind spot actually should should have a blind spot um um but uh our brain is able to interpret and try to fill up those spin for by using your other eye or uh in some cases when you have a disease like gloma sometimes you have a localized facial deficits the brain may also be able to help adapt and fill out the uh deficits using the other eyeballs this
00:58:24 may be something uh ambigous but helpful to understand the surrounding having said that we also need to be cautious because if you are really driving and uh it happens maybe maybe there's a kid coming over and it's in your blind spot and the Brain actually fills that Gap with other things then you need to be cautious if uh to to avoid any husters yeah and just fun fact we talk about FIA um and and we briefly touch upon Egos
00:58:54 and actually egos has two f years what but what's the purpose probably um is the top down for that I I don't know these but it's just fun fact yeah the two fos converge and oh no one of them is more 45° with the and um in the deepest end of the retina the other is 15% more shallow so maybe they are really focusing on different aspects when catching the mice maybe and then the mice are trying to avoid them from the top The
00:59:24 View there's some unexpected implications for having these high resolution phobias that you've got to move your eyes to fixate on different things for stuff like depth perception so if if you if you think that stereo Vision which relies on the positions of the two eyes is is taking points on on on the retina on the back of the eye and projecting out into the world and working out where these different points intersect in the world well the problem is is that now your eyes are are moving all over the place so you you no longer have a real sense of how to go from points on the retina on the back of the
00:59:55 eye to points in the world unless you know what the rotation of your eyes are um so and similar things so uh I talked about motion parallax so the information as As you move in the scene things closer move more than things in the back but now if you want to fix a on something of Interest with your high resolution phobia actually your eye is going to stabilize it and that's also going to complicate the the motion parallax calculations so all all of these little things feed in and make everything super difficult it's a wonder that we see it
01:00:25 all so one other thing I want to mention like in comparison to camera in comparison to Ai and computer vision is that um we don't have one stream of processing visual information right in the brain we have multiple streams one of them like the the kind of the biggest way to to break it up beyond beyond the what in the how pathway is um but is uh motion and very crude processing of things are moving there's contrast I need to orient towards that versus very
01:00:56 high resolution color and detail and we are processing Vision simultaneously in these two Pathways and I think that poses something very different than as far as I know what is implemented in computer vision these days um and clearly what a camera would pick up so which says that I think human animal V vision is is has again doing multiple things with visual information and again what's this information processing there's not one answer
01:01:29 we've talked a lot about the uh occipital loes and the ways in which this information gets back there to the visual cortex and then may come forward we haven't spoken I was eager to hear a little bit more about attention and let's say the frontal eye fields which haven't come up in our conversation yet so I wonder if any want to take a whack of that how that may because there's this other interesting issue of attention generally um where we um seem to be able to focus
01:01:59 our Mind's Eye internally so I was just curious about that idea is like a more how Vision may be yeah maybe um like a general template for attention for all I will say when you look at attention in the brain the Maj one of the major areas is the frontal eye Fields because where we look is so tied to where we're attending um you can't really differentiate that very
01:02:31 well and just to add on top we have actually touch a bit um there's a region called Superior ccus which is below the CEST uh and it's actually a multi sensory and multimotor area it coordinates between efficient sound and touch and also the eye movement it may also coordinate with the frontal eye field it may be related to the attentions that's all I know well that's also part part of why people that plays a role in Blindside I think also right because people respond to things through that
01:03:03 pathway yeah it it was I mean I'm I'm no expert on this but it was surprising when people found that the superior calculus which is a subcortical structure plays a major role in in determining where your attention is and this wasn't expected it was it was always assumed it was in the cortex and it it is I mean it's shared but um I wish I could remember the experiments but Rich Cress did these experiments where he really showed that uh by some measures
01:03:35 that I can't remember that the superior calculus is kind of U the the much bigger factor in where your attention is and that's the spiculus is sort of uh probably somebody could really argue with this and tear it down but you know roughly it's it's what the frog sees with you know and then we've replaced that with the or we we we put the cortex on top of that and that's most of what we see with and uh but there's still this old pathway there that's doing a
01:04:06 lot of our orienting and attending with that said uh if we go back to some P patients with facial uh deprivation the blind patients uh we when we look into the brain when they're at rest um the visual cortical area again uh they may actually coordinate A Little Bit Stronger with those attention Network so um there's something happening in those areas that are contributing to attention um they could be uh Beyond those Cal Superior ccul
01:04:37 areas yeah what you were talking about before I mean intentional blindness and and it just basically if you're if you're not attending to something you don't see it uh roughly I mean you do you know something's moving it catches your attention so there you have to be able to see things for them to catch your attention but but if you're otherwise yeah kind of occupied you you're not going to see it this goes back to that uh that example you gave with the gorilla or man in the gorilla
01:05:09 suit right if we're like occupied it's like a with a task we like hone in on a we attend to I guess like um you know certain certain parts in the in the in the scene rather than um others that visual we have a many of us anyway can imagine visual scenes without actually seeing them and there are others who apparently aren't good at doing that I mean how do what what role do you think that plays in let's say
01:05:40 Vision generally and then also imagination I guess if that's anyone wants to take a whack at that can can I piggy back off the top of that to ask a further question more controversial I mean I really I really think these those kinds of questions go to the very nature of of visual Consciousness right so um one view of what it is to be conscious have visual experience is that there's like an internal movie going on in the head somewhere right and and and it's funny you you talk to a bunch of scientists about say 50 to 70% say that that's
01:06:12 utter rubbish but then sort of 50 to 30% say well maybe there's something in that and I think to the extent people say that maybe there's something in that that that does seem to be a fairly accurate description of what goes on when we dream right right so you know is there is there a parallel between visual visual experience quote unquote when we dream and visual experience in you know when we're here in the real world type thing right so what what that goes to is is you know title of the the thing is synthetic Consciousness seeing and
01:06:44 believing you asked the question about AI Consciousness I mean before you can even get to that you've got to get to human consciousness right and and and then the question is well how do you even begin to sort of say ize categorize think about um you know what it is to to to be a conscious being um and I I I ultimately think that's that's probably the reason well certainly the reason why I study the brain right you know I I don't study the liver for instance and that's because it has this
01:07:14 quality of visual experience which just doesn't exist in in in any other organ so yeah not not many species Can Dream first of all and secondly when we are dreaming actually that there is rapid eye movement there is also some metabolism happening one of them is called the ceric nervous system that's involve and uh one of the major brain regions involv is the Bas of for brain all these can contribute to attention dreaming all these um and uh what happen when you are really strongly focusing on
01:07:46 uh certain event uh perhaps you're really training up or really modulating these um neurom metabolized and transmitting uh substances um so how how these interact with each other when there is sufficient related disorder U very often we our lab and others also observe these neurochemicals can be altered so uh when you when one person is perceiving the world of the same scenes with another person be also seeing slightly different views because of the neurochemical changes
01:08:18 um is still a lot of work we have to do with uh and um um how can we modulate better to make it um with the the title is seeing is believing but if everyone is seeing different things um what's the right way to do to be eyewitness for example going back to the uh theme of this topic sorry Paul I just want to go back to are you trying to say if people don't have mental imagery they're not conscious no no no not not at all but
01:08:49 what I'm saying is is that that mental imagery or i' I'd like to focus more on dreaming right is is uh me do you mean in images yeah I mean what I experienced you know when I wake up having had a dream at night is is is a a vivid I would say three-dimensional visual experience have that oh for sure for sure but what I'm saying is given that that exists in a lot of people is that an interesting way of thinking about everyday Vision right is it's a sort of
01:09:19 internally generated 3D movie right um some people call that the cartisian theater and then they add it's a fallacy right it's a cartisian theater fallacy Dan danit work and stuff like that I think it was very much dismissed in the '90s but increasingly as people think in in a number of different ways about Vision so predictive processing theorist about vision would would I think be be inclined to thinking of vision as a sort of controlled hallucination or something like that um I you know I just think I just think
01:09:49 it's fascinating forget the the theory of vision the fact that we have dreams um you know if if the brain's able to generate an an internal 3D experience of objects like that in dreams right feels that that Vision must be must be fairly close would yeah I think that uh uh imagery has become um a very kind of Hot Topic recently and there's this um there's these cases of aphantasia where people
01:10:20 don't have any mental imagery and it turns out that and so once something like that comes into more popular attention then people are like oh wait that's me and it becomes much more apparent so um and this is happening with uh visual mental imagery and we we actually just did a study on on mental imagery just getting people the of the general public to do a study and the variability on mental visual imagery is astounding right I keep on saying
01:10:51 astounding because vision is astounding but um there's we're we know actually I think very little we kind of we assumed mental imagery and then studied mental imagery without thinking so much about it being on a spectrum yeah I think actually that I'm I'm either a fantas I either have a Fantasia or I'm very close to it I have very very little visual imagery um but I have absolutely normal dreams I I see everything clearly in dreams so my brain
01:11:22 is capable of of generating uh images on its own but it it won't do it you know on command maybe so so just jumping I mean that also feels like the the as you were saying the top down is is modulatory but not not necessarily excitory and so if if you've already got a visual stimulus I guess well I guess in the imagery case you can close your eyes right and so that would remove the the visual stimulus to some extent but uh yeah it's it's it's an interesting yeah and I do think I guess I'm along the lines of
01:11:53 it's a controlled hallucination and that you know what what I've been emphasizing a lot is that your brain has to compute and paint this picture you don't just open your eyes and see and um in normal vision that the picture that gets computed is very tightly anchored to the incoming information um but it still has to be inferred and computed and we still have very I mean there's lots of visual information in there that you don't consciously see and we have basically no idea what it is in
01:12:27 the neural activity that makes you know the part that we are conscious of and where all the rest goes and um we just have no idea uh but but it's very clear that that same hallucinating process can run quite free of incoming information during during dreams for example or or for people who can who can close their eyes and and mentally image I mean could I follow up yeah um so I mean a related question I think the answer is going to be we just don't know but um is there a place you would
01:13:00 hypothesize right so so rather than we know where it all comes together right because it feels like in our visual experience it all comes together in this beautiful percept but then if you read the textbooks often they'll be like okay but color process here and this is process here and ex right it all seems distributed throughout the brain so I mean I think it's it's a bigger issue of how the of how the brain works and that we have a unitary experience um you know we um the world is ambiguous but you see
01:13:32 one thing you hear one thing um you uh you take one action you know you you may plan multiple actions but you don't you're not you know unless in some very abnormal States you're not fighting with yourself across many different actions trying to be execut um you always arrive at a unified interpretation of this ambiguity and um you know and there's even cases of of like binocular rivalry where where the two eyes are getting two different bit bits of information two different
01:14:04 images and you could see either one or you could see one on top of the other but at least in some circumstances you first you see one then you see the other then you see one then you see the other your brain is always deciding on one even though there's two things clearly there that it could see um and that whole process I mean is much more than Vision it's it's it's it's our whole experience every sense every action uh there's a unity and uh there's a single thing that the brain decides on and we have no idea
01:14:35 how that happens and I I would say yeah there's some areas that we consider to be Vision but that's in that's small compared to the amount of brain that is actually processing visual information um once you get past primary visual cortex these regions can be activated by lots of different information um that doesn't have to be Vision specific so I I think most of the brain is actually integrating Vision with kind of
01:15:06 the rest of cognition memory language conceptual understanding navigation spatial understanding and there actually bring circus not only connec primary future accordance to these areas there are some recent studies suggesting that I may also have some interets or direct fibers direct connected to these memory um emotion control uh circadian rhythm etc etc so and and since we are crossing between um
01:15:37 senses actually there's also an interesting phenomena called sinesia so when you are simulating one uh particular sensus you may be able to activate concurrency the other senses how these are actually working I don't think we have a per clear understanding um right now maybe there are some sort of cross model inhibition being disinhibited um but this could be very hard topic you can try to disintegrate and understand how things are integrated but you you both avoided the question is is there going to be are we going to
01:16:08 ultimately find one place where it all comes together or I mean the old problem with homunculus I mean is there you know is there is there the little man or woman inside you know steering things looking at the screen uh you know we don't think that's there somehow the whole thing just works by itself without referring to some Center that's deciding on everything but you know but we don't know how that that's probably something towards the front of the brain that work doing these integrative stuff like basil G gear
01:16:38 austrum but these areas are still be develop so it's it's not one area it's a network of areas and it's how these areas communicate I was just to maybe finish up this conversation about imagery um I'll throw this idea out here because it just came came to me and it could be completely ridiculous and wrong but um I wonder whether and to what degree visual imagery um permits humans to create counterfactual thoughts um that may also
01:17:10 contribute to the INF the inferences we make and whether that means we can make better inferences than let's say other organisms that don't have that capacity is that a visual property or and well so what's an example of what you're thinking of for the counter well if you counteract would be like you you have in your mind's eye an event occurring in a way different or hypothetical way in which an event May take place it's not taking place yet so again I think that's a much more General thing
01:17:41 that the brain does that it you know it can use vision for but Helen Keller can do it too um it uh for us a lot of it might be visual because we a lot of what we do with our Brands is VIs visual but that's not but there's a a more General process that's using anything using Vision using hearing to to make a picture of the world or of a possible world or I say picture that's visual but you know to make a to make an imagination
01:18:16 of I mean this somewh goes back to the you know debate of the 80s with poition and cin of of is our Mind's Eye does it have to be visual in order to um answer visual questions or can it be some type of linguistic form and you know I I think the fight is is probably still out there and I think there's not one answer to it okay well I think we should open up the to the to the audience if they want to please go up to the microphone and and folks we insist on you going to the
01:18:46 microphone just because it otherwise doesn't get picked up for our broadcasting so thank you okay uh so something that occurred to me after listening to this panel was that I mean it was addressed to some extent but I would have thought that it would have been addressed more is the issue of blindness so for one thing I'm wondering people who who never saw and are blind when they dream is it primarily auditory or do they have dreams in taste and perception as well and hearing and the other thing that it makes me think of is
01:19:18 that since there was a relationship between um between vision and and concentration as well as emotions has there been any studies in which it's been studied whether blind people have a different level of intelligence because concentration is a part of intelligence and also maybe blind people having different emotional responses since they don't have the vision um input into um
01:19:49 developing their emotional connections as well yeah uh so how about intelligence there were some studies trying to differentiate or compare between congenitally blind patients of people uh versus acquired blind people it turns out that in uh well the early prelimary data suggests that uh the con people may be able to perform arithmetics a little bit better than the acquired blind Pap what does that mean they have some intelligent difference
01:20:21 with that facial areas Ur not not no longer assigned to Vision but they are you doing it um for intelligence um those are hypothesis we need to test further uh regarding the dreams um because con p uh subjects do not have uh visual experience I don't think they can really explain if they have visual dreams um but uh when we give the so-called sensory substitution devices to our acquired blind patients sometimes they would say they can see some lights some falce FS soal um would they be able
01:20:54 to see like a visual dreams uh I I guess they may be uh because they have visual experience but we have a study that there may be lature out there I'm not aware of yeah this is only applying to congenital blindness right uh those were the studies trying to separate the two um and they're just comparing the two whether yeah go ahead well that also reminds me of the observation expressed last time and
01:21:25 hearing that can generally deaf patients who get uh cckar implants or other sorts of U implants to restore hearing uh have cannot process speech if they've been exposed to speech and then lost their hearing they can recover that and they can they hear sounds that can be organized into speech patterns hi I had a question I wanted to address synthetic and Consciousness in two questions I wonder if you could compare the computation efficiency as
01:21:56 well as uh level of Consciousness between artificial and kind of human intelligence as we understand for to give an example uh artificial neural networks if you talk about uh vision and language kind of combined diffusion model uh as opposed to just a language like a chat jpt there are a 100 times or two orders larger for language and yet in the human brain we have 20 30% of our cortex a lot of it integrative but still much more dedicated to Vision as opposed
01:22:27 to language which is the opposite case in AI so that speaks to the question of computational efficiency of the human brain has Evolution kind of optimized for uh kind of a generality of vision at the cost of language efficiency and the other question is on Consciousness two-part question the other thing is if you a bladed a human mind and you couldn't you for example lost off uh Vision modality um how would that affect your Consciousness as opposed if you ablated your language centers in terms
01:22:59 of Consciousness because you know there is that sense that most of our information comes from the world through vision and so you imagine vision would be much more important and yet when I think about losing one of those two senses I much rather be blind and be able to think about PR and Mathematics and history uh than I would be like my chicken or my dog who have really acute vision and emotional experience but don't really can't really cognate on it think on that so Consciousness and efficiency could could the panel speak to those two questions I
01:23:30 just want to clarify your chicken we raise chickens and I have a golden retriever okay I just want to make sure yeah I on the topic of uh computation efficiency I I suppose it is possible that we're not developing data structures to most efficiently you know one represent the the you know the data that comes in whether it be like natural language or uh imagery or a video and
01:24:01 and on top of that like processing it and Computing what I mentioned earlier attention like like the attention that we know today and like sort of the the current research of like gender of AI and llms it's you know essentially it's like a gigantic matrix multiplication problem right so you have to represent you know all your text as sort of um you know if you if you have like a dictionary with like a thousand tokens like that's actually how how many um
01:24:31 that's how big your your vectors are and that then you know that then affects how many or sort of the computational um Power that's needed to then uh compute this gigantic Matrix um uh and so you know essentially what's what's happening in the brain is probably not that um but it it's the representation that we need to feed into like our the the the chips that we use to then um you know uh make these computations that then generate natural language so yeah it's very possible that you know what's going on
01:25:02 in here is is exactly not how we're actually modeling it mathematically but but I think you a good deeper so so this question of whether the computations are the same but even deeper question is whether the aims are the same right so I I I think with take take 3D Vision for instance if if you're thinking about computer vision what you do typically want is a what's known as a 3D depth map which is essentially the distance of each pixel in the image um and and I I don't think
01:25:33 the human visual system produces anything like that um I I think it's it's a very Rough and Ready approximation for for for whatever it needs in in that particular moment um it's weird cuz we are seeing there's like the the the sort of Next Generation that some people want to think about AI building on human vision is to say well actually there's a sort of a graphics engine in the head right so you've got these retinal images there's a graphics engine in the head that sort of simulates all these different images
01:26:03 based on the illumination based on the objects based on all these things then Compares them to the retino image and works out these things and that sounds like a potential approach for computer vision but it just for me at least it sounds a million miles away from from what I think human vision is doing but as a neuroscientist from the biology of it why would humans be the inverse of AI in terms of dedicating so much processing power to Vision when in AI it seems relatively trivial compared to language
01:26:35 I no but I think I think the difference also is that in AI we're talking about like statistical regularities to get us to our understanding whereas with humans language is much more Rich than that um and I think it's a little bit of comparing apples to orang es yeah there a commercial interest also I mean the interest in visual uh artificial Vision I think a lot of it relates to driving cars right that's
01:27:06 driving a lot that's driving a lot of the uh research and uh for language models it's a way to engage people and I think that's obviously we see this is very hot right can can we fake people into thinking these are conscious beings that you're talking to and you know there's money there I think that may also explain the sequence well the other more interesting question is on Consciousness that is um when I think about vision and I think about Consciousness I think about integrating one definition of
01:27:36 intelligence is be able to predict and so I think of a physics engine when I think of visual I think of can I predict what's going to happen if that person drops that so that's kind of the Consciousness maybe there's something beyond that that you could tell us about but when I think about language and I think about Consciousness I think about really the definition of Being Human so how would you compare those two modalities in terms of Consciousness Vision versus language and what are kind of the ceilings how do they interact I mean I I I think as as
01:28:09 scientists we have we have to be honest which is if we were looking at the human brain and human behavior from the outside I don't think we'd hypothesize Consciousness right I I think it has to be the fact that we experience it ourselves um and so I just I I I I'm going to plead ignorant at this point and and say we just don't have the conceptual tools at at the moment to to I mean look for myself I I think there's some interesting questions in 3D Vision so we have 3D visual experience so that that for myself that would rule out the retina and and the next stage the lgn
01:28:41 and knock us up to the visual coures but then after that you know it's a real um you know all bets are off um so yeah yeah I I I wouldn't be inclined to try to compare the raw computational resources used in in AI doing something versus the brain doing something because it's just not at all clear that the algorithms are the same I mean in in Vision at least there's more quantitative measures of to
01:29:12 what extent the representations are the same and you know which is sort of half full half empty kind of question but there is at least a fair amount of similarity in the representations but in language I think all bets are off I mean as as you were saying it's the large language models are purely statistical and and frankly nobody knows how they work I mean it was the people who who developed the large language models did not anticipate they were
01:29:44 going to develop large language models that you know that were that powerful I mean they were just trying to predict the next word and the damn thing could could write essay you know it just nobody expected that and and I don't think anybody knows how it works you know but our language is much more tied to whatever our brain thinks is reality um and not just not just statistics and putting words together and that's just a completely different computation is the is my conception of Consciousness in the visual realm is that is that kind of a simplistic idea
01:30:15 about you know a physics engine and predicting what happens in the next frame is there a higher level of Consciousness that I'm missing yeah I don't think we have the conceptual tools to talk about godness okay that the hard questions thank you anyone else okay well wa up please there's something about this that
01:30:48 I find very frustrating uh coming as I do from very different disciplines and uh it it seems to me that what frustrates me most is the need for some kind of unification and in other words you know I get a lot of really good interesting pieces right but I can't put them together by the way I came the this ones I came from well
01:31:20 Psychotherapy and the other one was fine arts right so in other words being so uh involved in the world being so involved in experience through something subjective right is what I find so uh frustrating uh about something that does a terrific job of taking things apart but doesn't quite put them together that and and uh the one discipline that I see
01:31:54 is trying to put them together and I'm not a friend of this discipline is uh AI uh you know so uh this this is my difficulty and you know and you know how how do you try to you know bring that Unity to to the to the to that experience
01:32:27 I I I so I I I think there's something to be said for for for divide and conquer um in so far as um so so the kind of stuff I that interests me is very lowlevel visual processing of of of stereo Vision in 3D and um and I think I think part of the reason it interests me is that actually so long as I know the separation between your eyes so so long as I know know that your eyes um you know particularly good resolution um and
01:32:59 and some people have some cortical some some brain deficits with stereovision but on the whole people generally have it given those things given those basic assumptions then actually the the kind of effects that you see in people are pretty generalizable right um they're not they're not particularly subjective um but I presume when we go into to other areas of vision that get associated with the more subjective then you you really do sort of go diverging into to to very different sort of responses depending on who you ask
01:33:31 which is what you alluded to um but you but for myself I'd also presume that those conclusions build out on the sort of lower level earlier visual processes um so yeah in terms of putting it all together and and and how we do it you you you've got an option right you can either stay at the very low level like I like to to do and and and hope to focus on those in which case you can get a I I hope at least a fairly unified account of of visual processing visual experience or you go up to to your kinds
01:34:01 of questions I think yeah I mean I I would say you find it very frustrating I find it really exciting I think this is where the magic is um and I think that the more that we can study this the more we understand how humans process the world so uh I I think that's what makes humans amazing actually and makes us so different from computer and also I I think I think that our goal in Neuroscience or in any science is to is to um develop you know a deep unified
01:34:36 understanding of of the process but the frustration we all have the frustration you know we would like to understand how the brain works and we don't we'd like to understand how uh we have a unified experience and uh from you know whatever sensory information we happen to have access to uh and whatever history and memory we have access to and and but we don't and so you know we it's it's a long arduous process I mean physics had to understand you know a lot
01:35:07 of very specific things before they finally got to you know a grand unified theory almost not not quite but they're you know they're a lot closer than they used to be um so it it's just a long process in science to to get to the unified answer that we all dream of and I'm optimistic there are different puzzles just have to live another 100 years yeah I guess I guess we are now currently touching different parts of the elevant but at one point we will be
01:35:38 able to see the whole pictures and uh together with uh even along the whole lifespan if we want to touch back to the Fine Arts thing or Picasso thing uh actually for um youngsters who can draw well they tend to have good visual memory uh and all these are flexible they can be developed but how these development thing pisting and then the pieces of uh uh I versus mental uh imagery and other um processing can be put together into a unified version uh we're really looking forward to that uh
01:36:10 maybe AI currently is uh in the stage of um artificial narrow intelligence is is of similar stages we are pushing towards artificial general intelligence or even artificial Consciousness intelligence but there's big gap uh before we can Reep that you know the stories of the you know the cathedrals in Europe that took hundreds of years to to build and you know each generation would put some bricks in and then their their children would put some more bricks in and and it would go on like that for generations and generations to generations and
01:36:41 that's what science is you know we're not we're not going to build the cathedral in our own lifetimes I think I please please I I mean I think just as you know that that's what I tell my students all the time the more that I know the more that I realize I don't know right what reminds me that's exactly what I was going to get at is that well first of all I was very impressed by how everyone in the panel was able to say I don't really know or we can't fill those gaps right now and I think that's wonderful scientific honesty um also it's so nice to hear you say that it's part of the what you said
01:37:13 magic or mystery of You Know The Human Experience and of course this is the space that not only scientists can explore in their creative and imaginary ways but also of course visual artist I mean in Vision in terms of vision so it's an opportunity for artists also to get in there and to work that space and uh improve our Enlighten Us in some other sorts of ways that we can't maybe exactly put language to so I want to thank everyone again for this excellent
01:37:45 panel and uh would look forward to our audience coming back in the fall for our uh roster of uh new round tabls thank you again [Applause]
01:38:23 e