October 16, 2022
The program GPT-3 can create language that gives the impression that it is thinking. What will our interaction with robots of greater and greater verbal agility mean in the near future? What sort of Other will these robots become, evolve to? Is awareness of a code incompatible with any form of realism, and what does this mean for epistemology and ethics?
This roundtable brought together philosophers of mind, computer scientists, literary scholars, and AI ethicists to examine whether large language models like GPT-3 genuinely think or merely produce convincing text. The discussion began with an explanation of how these models work through prediction on massive datasets, then moved to the central question of whether their outputs constitute real creativity, understanding, or intelligence.
A significant portion of the conversation centered on the philosophical implications of LLM capabilities. Panelists debated functional role semantics, whether these systems possess implicit world models despite lacking explicit ones, and what their characteristic failures reveal about their inner workings. The distinction between behavioral criteria for intelligence and questions about inner experience echoed classic debates in philosophy of mind. Some panelists argued that the models' ability to generate contextually appropriate language across domains suggests emergent understanding, while others maintained they are sophisticated pattern matchers without genuine comprehension.
The discussion also addressed practical and ethical dimensions, including the importance of AI literacy for humanists and social scientists, the history of human-computer interaction, and concerns about treating LLM outputs as authoritative. Audience questions pushed the conversation toward code generation, structured outputs beyond literary text, and whether the distinction between outward performance and inner experience matters for how we should relate to these systems.
00:00:00 theories and but you still need a theorist attacks you still need the theories to test well sometimes oh yes okay hi good morning everyone welcome back uh the Helix Center this is a round table number four in our series on coding in the new human phenotype the topic for
00:00:32 today's talk is our natural language generators for real we've uh have a wonderful panel of of uh experts that talk to us about this and talk us through this if you didn't see any of the three talks from yesterday I really do recommend you take a look at them on YouTube where if you look up Helix Center you'll find them they really were terrific so I want to without further Ado go on and give a
00:01:03 brief buy of each of our esteemed panelists first Francesca Rossi is an IBM fellow and the IBM AI ethics Global leader she is based at the TJ Watson IBM research lab in New York where she leads AI research projects she co-chairs the IBM AI ethics board and she participates in many Global multi-stakeholder initiatives on AI ethics such as the partnership on AI the world economic Forum the United Nations
00:01:36 itu AI for good Summit and the global partnership on AI she is the president of aaai the worldwide Association of AI researchers Dennis e tennan is an associate professor of English and comparative literature at Columbia University his teaching of research happened at the intersection of people text and Technologies a longtime affiliate of Colombia's data Science Institute and
00:02:06 formerly a Microsoft engineer and a burpman center for internet and social Society fellow his code runs on millions of personal computers worldwide Tennant received his Doctorate in comparative literature at Harvard University under the advisement of Professor Elaine scari and William Todd a co-founder of Columbia's group for experimental methods and humanistic research and the editor of the on method book series at Columbia University press he's the
00:02:38 author of plain text the Poetics of computation 2017. is an associate professor of computer science and data science at New York University and cifar fellow of learning in machines and brains he's also a senior director of Frontier research at depression design team within Genentech research and early development he was a research scientist at Facebook AI research from June 2017 to May 2020
00:03:11 and a postdoctoral fellow at the University of Montreal until the summer of 2015 under the supervision of Professor yahshua bengio after receiving his PhD in MSC degrees from alto University in April 2011 and April 2014 respectively under the supervision of Professor yuha carjones Catherine Elkins has written over a dozen articles on memory Consciousness and embodied aesthetic experience in a
00:03:43 wide range of writers from Plato and sappho to Wordsworth and wolf improves In Search of Lost Time philosophical perspectives she reframed proof's exploration of Consciousness in light of integrated information theory in the shape of stories she used the AI software sentiment arcs to develop the first robust methodology for exploring the emotional arcs of stories her audible.com lectures on the Giants of French literature and the modern novel
00:04:15 have won her in international audience Noah Jean syracuso PhD in math from Brown University is an assistant professor of mathematics and data science at Bentley University Noah's research interests include algebraic geometry the abstract study of systems of polynomial equations and their Solutions machine learning especially topological and geometric data analysis artificial intelligence empirical legal studies phylogenetics and misinformation
00:04:49 net block is silver professor of philosophy Psychology and Neuroscience came to NYU in 1996 from MIT where he was chair of the philosophy program he works in philosophy of mind and foundations of Neuroscience and cognitive science and is currently writing a book on attention he is a fellow of the American Academy of Arts and Sciences a fellow of the cognitive science Society has been a Guggenheim fellow a senior fellow in the center for the study of language and information a Sloan Foundation fellow a
00:05:22 faculty member at two National Endowment for the Humanities summer institutes and two summer Seminars the recipient of fellowships from the National Endowment for the Humanities the American Council on learned societies and the National Science Foundation and a recipient of the Robert a mu alumni award in the humanities and social science from MIT he is a past president of the society for philosophy and psychology a past chair of the MIT press cognitive science board and past
00:05:54 president of the association for the scientific study of consciousness okay it's quite a mouthful anyway speaking of mouthful so we're going to be talking today about a language uh generation artificial language generator and what that means for us uh in our near and more distant future so welcome everyone okay so apparently all I need to do is start to say a few words and you're all
00:06:25 going to fill in the rest uh the the way natural the way the natural language generators work as some of you may know is that uh after giving given a prompt of a few words an entire story will come out like in remembrance of uh In Search of Lost time so um I wonder if any of you want to take a shot at giving a general orientation to our General audience about what for example the gbt3 generator is and what we might expect from it in its future uh
00:06:57 iterations I can give it a start and then people can chime in so so until few years ago artificial intelligence was becoming very good at interpreting content that we were producing like images and text and other content so natural language generators are instead the most recent advances of AI where AI is becoming good also at
00:07:30 generating new content rather than just interpreting the content that we are producing and in order to do that these and and I'm saying this content and not text because it's not limited to text but it can also be videos it can also be images for example so different kinds of data that is content that is generated so so-called generative AI so AI that can generate new content besides being able to interpret content that we
00:08:01 produce and the way I mean this I leave out a lot of details but the way it it can do that is by being trained on a vast amounts of data unlabeled data is called so data that is found on the web without any um any curation from human beings any labeling is called from human beings so this uh data that is found on the web in
00:08:33 vast amounts that that is used to train this AI systems and for example text that is found on the web that is trying to use these AI systems that then can again as you said you know respond to a prompt in an appropriate way this data is the source of knowledge if you wanted knowledge of the system but there is also knowledge that is given in the prompt itself so there
00:09:05 is an area called prompt engineering because writing a prompt also the way you write the prompt of how long the prompt is what you put in the prompt can also trigger a different more informative or less informative response from the natural language generator and and if I now put the heart of a company like IBM or others they may want to use this for some applications one possibility to use them is to take
00:09:38 them as a trained by these vast amounts of data and then further tune them with supervised and label data on a specific task to but very little amount of data to solve a specific task but with the general knowledge that is given by this initial phase of training over vast amounts of data and this allows you to have a specialized solver for one
00:10:09 particular problem but with this more general knowledge which is needed usually to solve well even a specific task and especially when the kind of label data that you need for respective is not is kind of limited so you have a lit you just need a little bit of this label data for the specific task because you rely on this vast amount of General let's say data so that's some initial thing but feel free to chime in I would
00:10:41 add that I mean I think the probably most important thing for many listeners to realize is just how amazing the Productions of some of these programs have been um I think you know they've been so incredible that two years ago nobody would have predicted that they could or three years ago so gpd3 came out in 2020. um so a little over so around three years ago so um well at the time of gpt2
00:11:16 um I don't think anybody would have predicted what they could do it's just really if you haven't seen any of these things like there's been a couple articles in the New York Times Etc it's kind of amazing but it's important to realize what they're good at and what they're not good at so what they're good at is open-ended contacts where there's no very specific right answer where what counts as style and creativity that's what's kind of amazing creativity these things are good at the they're the
00:11:47 opposite of what everybody would have expected they're so they're really good so for example the New Yorker published a poem in the style of Philip Larkin uh it was actually pretty good um and um you know people could get them to you know write news articles and uh uh you know and uh give interviews in the style of a particular person with with the right to fine-tuning that Francesca mentioned
00:12:18 um so that's kind of amazing I think that's the first thing to realize is that it's just how astonishing they really are but the second thing is the severe difficulties so the main one really bad severe difficulty is that they have been pointed out by many people no world love so they just continue a style in a certain way
00:12:49 um and with ignoring the actual facts um uh so they'll spin a a web of of a text uh along a certain line but then it can just completely contradict what's what's really true even though you could get the real information from Siri or Google search to importantly uh Francesca mentioned this but they don't actually I mean you can add a Google search to one of these
00:13:21 large language models but it's then you have a problem with the interface so in just the operation of the large language models they don't have access to the internet they're trained on the internet you know with enough electricity to power a small City and then you can run the trained model on a smaller computer but they don't actually have the ability to look things up so you'll get a better answer from Siri than you will from that
00:13:52 um now with all their failings people have tried to hook up more standard systems to them and then there's the issues of how that interface is supposed to supposed to work so the big negatives are no world model um in the in the generation of the language no understanding of long-range dependencies um so um you know some of their critics have been fond of pointing out that they treat words as like a stream
00:14:23 uh without getting the hierarchical structure and there was a paper that came out I believe just yesterday by Stan dehan's laboratory looking at short-range dependencies and long-range dependencies with a relative clauses so um the case the example they used I think was the key um uh the short range would be the key is green or the key the man had is green and then you keep adding relative clauses like the key that the man in the
00:14:55 corner had was green and as soon as you get to a fairly long relative clusters forget it they don't know which things did what to whom um so but that is a general feature of their uh not understanding the structure of language language is what they're made for but they keep getting things wrong and then the real I guess the most significant issue is you know some people think okay you just need bigger ones um and then others think no there's
00:15:26 something really principled missing and that I think that's the key to bait another thing they're really bad at is um logic and arithmetic again like the opposite of what people think so I think you know people in the um who aren't familiar with the these things have to realize they're quite different from their strengths are different from what everybody would have expected and their weaknesses are especially different from what everybody
00:15:58 would have expected so it's a pretty pretty peculiar thing and you know for me I I'm more interested in the mind that I have and lots of other things that I want to know what is that tell us about the mind I met it I can't actually put something on top of that yesterday so one of the issues I I see you know the US looking at all these amazing language generators like the gpts and whatnot is the fact that they are doing something amazing which is very obvious because we can just see them doing amazing steps but we
00:16:30 actually don't know exactly how those amazing things are happening so let's just pointed out the creativity which is amazing right so these things are actually creating something new that it has not seen before during training time and then in the term of the machine learning that's called generalization so how can these models be able to do something that it or it was not trained on explicitly and then how it happens in the statistical scientist that it does all those counting of all the parents they see during training and then what it does is that because it has a limited capacity it needs to compress all the
00:17:01 things that it has seen and then while doing so it loses some information but the information lost is the information gained on the things that it has not seen before and then what are these are the new parents that these compression mechanisms more the training algorithm are actually focusing on that we find really amazing so the process of let's say generation generalization or the creativity by compression is a complete mystery at the moment a lot of people you know a lot of theoretical theoreticians actually do work on it from the perspective of well is there
00:17:33 some kind of this implicit let's say regularization happening that is you regularize how learning happens and thereby encouraging these models to do something that it has never seen before annuity how does that actually connect to the amazing nature of the generalization of the creativity that we see and then this lack of our understanding of this very simple fundamental things essentially we're saying that they will these models are counting and compressing and somehow magically thereby does generalize to a
00:18:04 completely unseen or the new things that look amazing to us so what is this right it's a very simple thing right the question is like out there but we have absolutely no idea what is the even right way to approach answering the question now let's stop for a second about the word creativity I mean there's so many different branches to you know to look into and we'll hopefully do that it's amazing uh so I'm going to stop about the word creativity because let's stop and imagine that I would prefer we use creativity as demonstrated by these
00:18:36 as like creativity with an asterisk and not assume right off the bat that it's creativity is it's the same property that we have we instantiate when we're being creative is it just a very Advanced way of being a dancing bear are these very very clever things that do wow us because I didn't think a computer could write something inventive like that and how deeply does it go and and as as you were saying Ned in some ways what does it mean about the mind and
00:19:07 creativity in humans so does anyone want to take a a little whack at that I think it's helpful to zoom in for a second on the training process itself so we talked about how it scans this text it's actually a very simple process that it's undertaking while it does this it doesn't read the text but as it processes the text the computer algorithm just masks or it sort of hides random words and it acts it asks the the neural network to try to predict the missing words and that's the whole process it just reads along so imagine I'm talking to you and I say my dog
00:19:39 likes to blank everyone in this room I'm sure it's not your head you heard sort of an autocompletion maybe my dog likes to play my dog likes to walk and that's all we're asking the computer to do we give it some text or it reads some text and then it tries to predict the next word the missing words and as it does that it just develops this process of being able to predict the next word and I think um tying into some of the earlier things as far as like how it's surprisingly bad at things like arithmetic and logic when
00:20:09 you think about that process of course it is it it'll do well of things like what is you can prompted mask of what's two plus three well because it's seen that in training text so it can predict the answer is five but if you give it more complicated numbers that it hasn't seen it's a little bit I think like some children learn to spell by memorizing the spelling of words rather than phonetics and this this algorithm is very much the non-phonetic version It's just memorizing a bunch of correlations and patterns without developing that understanding what's surprising I think
00:20:40 is that it does have some hints of understanding that it shouldn't from such a basic process and going back to your your question about creativity I think one thing that's helpful to think is you know if you have your phone and you you start typing a text message and it suggests words you can just keep clicking those buttons you're basically running something like gpt3 based on the text that's there it'll just sort of randomly predict the next word it's very improvisatory so I think it helps to think a little bit maybe like jazz where by Nature it's just kind of rambling and
00:21:12 improvising and making up as it goes which actually can make it seem more creative than a very structured rigid thing where it's trying to articulate and express an idea it doesn't have the idea but it can just kind of wing it and improvise Words which I think gives it a kind of like local aspect of creativity but not the global there's not a creative idea it's expressing but the wordings are creative for the lack of idea that it has well this may be a little bit of what you as you said earlier and Ed if you follow it long long enough it sort of it seems to go off the track a little bit right yeah
00:21:44 well a complete change Persona et cetera but you know your your questions suggested that we're amazed at how creative is is for a machine I don't think that's right I can't write a poem in this in the style of Philip Larkin I can't draw the amazing pictures it draws and the other things that it it does also or can be well beyond what many people could do I mean it explains jokes for example it looks The Economist published a series of of
00:22:16 little bits of of where they showed an economist covers they explained the covers I mean it did a really good job I don't remember which system it was but it did an amazing job at the same time that same issue of The Economist had a wonderful little article about Douglas hofstadter where it asks questions like um um you know the last time when will the
00:22:49 next time that Egypt be transported over the Golden Gate Bridge and it gave an answer assuming that Egypt could be transported over the Golden Gate Bridge so that's the lack of a world model so creativity I think it's I mean look I don't know how to define creativity it isn't what we do probably but um it's it's very very impressive combined with these utter lacks of logic reasoning in a world model so
00:23:20 can I kind of push back on this word amazing that we keep using is that so so my work specifically with language generating is historical and some of the earliest materials that I found were so Ramon loli or Ramon UE was a medieval medieval majorcan Theologian who created these paper Machines that were prototype language simple combinatorial language generators where you rotate the circles and combine all possible truths
00:23:52 about God and that work and those devices were so amazing in a sense and so kind of uh unreasonably effective that it spawned a number of Cults all across Europe where lalians are the theologians Lolly and Poets that persisted for centuries and and I think they asked some of the same questions that we are asking of these somewhat more sophisticated tools but the thing is this kind of kind of create creativity that's combinatorial that's
00:24:23 mathematical that's statistically driven has been with us for a very very long time it's just we tend to kind of forget that history and then ReDiscover those devices and be amazed again and kind of be discomforted again by their presence in our midst but that's why I would a little bit say like you know what you know is this is this a new kind of phenomenon is this something that we're continually struggling with I will say we were teaching uh earlier form of deep learning that would generate text and it
00:24:54 worked pretty poorly it was somewhat word salad it seemed creative but it didn't really make sense all the time and I still remember the day that gpt2 came out and we started working with it with students and we taught it to write like Oscar Wilde and like check off and like all these writers and do things exactly as you said that my students have trouble with I asked them okay you read Virginia Woolf write a passage like Virginia Woolf you know they can't right and so that's more of EX Paradox right that it can do things that are very difficult for us but it can't do things that are easy right for us and so people
00:25:27 get very confused about it because they say oh I can't do this right and therefore it's dumb but it's part of that Paradox and even in terms of counting people have found that it counts like little children count right as they're learning to count so if you haven't worked with it it can be kind of hard to understand because you think how can this work the other thing that I would say is for decades people wanted to teach computers how to process language and generate language based on rules right this was chomsky's
00:25:58 Universal grammar and we thought if we just gave it enough rules and enough edge cases somehow that would work and it really hasn't worked well at all and no one really expected that if we just gave it a massive amount of language it would be able to do such a phenomenal job so we're all still trying to figure out what does that mean and what does that mean about how language works and what does that mean about the nature of meaning and what does it mean about the nature of our own mind and creativity so it has a lot to teach us but it doesn't fit neatly into human categories and and
00:26:29 that's kind of what we're experimenting with right now to try to figure out how does it work and what does this mean yeah it might be also I mean connecting also to what was said earlier you know many of these behavior that we see as amazing is historically and say that this happened many times are is a bit Cherry Picked because you know then you generate a lot of different texts from my prompt and some of them are amazing
00:27:01 and some of them are amazing in the negative and wrong way you know because it's completely out of track so so we have to be aware that there is no real reliability there so there is a problem with reliability and and then there is also a problem connected you know going back to creativity is how do we want to use these systems you know to replace human creativity or to augment and support and expand human creativity
00:27:34 um like my one of my recent talks I used I did this PowerPoint slides all the pictures in my PowerPoint slides were generated with Dali which is a not a tax generator but the image generator from a textual description of a scene so all my images were they were generated using this algorithm and they were beautiful and I was you know oh my God you know a lot but then at the end I said okay wait a minute I didn't use any
00:28:06 graphic designer I didn't pay any copyright because you own the images that you generate with Ali so what does that mean if everybody would do like that then graphic designer would be out of a job that would be possible consequences and that's maybe economically in their business model is going to be a damage for them but most importantly if everybody would do like that what will happen to the creative process not the outcome but the process
00:28:37 of creation that Society needs to have and people need to have if you have a society where nobody follows the creative process anymore then what kind of society is that going to become so really the question about how we want to use these um these new techniques and and uh within our society so we have to be more conscious about besides being excited about the novelty
00:29:09 of the thing which maybe is not really a novelty as you said but also about more conscious about how we want them to be embedded in our job or in our interaction with each other in our creative process and so on so what is the role of this technological progress for example by these techniques in the progress of society so let me also give you another historical anecdote that may answer a little bit you know we can look back and see how did we deal with with this before so
00:29:39 um in my book that's that's upcoming I talk a lot about the previous sort of flourishing of text generators which may surprise you was late 19th early 20th century which were exactly the template based generators were which work like this Mad Libs you know something something blank fill in the likely blank and there were devices and charts and tables that were used on a massive scale to generate movie plots to generate theater plays today when you're looking
00:30:12 you know at the Netflix show it's the the practice of using template generators um of you know particular lineage that go back to like the plot Genie and these are like best selling best-selling writer 8 books that are completely integrated in contemporary writing creative writing programs right and what they've they have contributed to the mass production of art you know so if we think about the flourishing of Art in the 20th century
00:30:42 and the massive art popular art it's kind of completely integrated the particular form of writing you know when you see it early on people are actually saying like Okay we need I need to write more I need to sell more little you know like a pulpy pulpy plots to the magazines I need to sell more pulpy plants to the to the Hollywood studio and they and we are living in a kind of uh cultural environment where this algorithmic combinatorial template based
00:31:14 a little bit maybe statistical to a lesser extent but is very well integrated to our processes but to the extent that we don't normally see it right when you see that those credits roll at the end of a Netflix show you're not thinking like oh dude wait wait a second they didn't actually invent this they you know did they use a particular you know like writer's Aid that helped them in the summer like unfair and creative places but there are complaints that a plot seems formulaic or it seems you could people can perceive that from time to time like this this plot seems
00:31:44 have been written by a computer for example I wanted to follow again along this idea that what sorts of creativity are these different sort of properties or categories amazement and has been mentioned a lot so when you look at Art visual art and maybe it's easier in visual art to come up with this we have a history of being Amazed by certain masters of art so Da Vinci draws something and it looks amazing we
00:32:15 part of our amazement is how did he do that well what was it what was the manual in the dexterity and the what and whatever else goes into it I I'm not equipped to say all the things that go into it but I do know that the human factor is absolutely a part of what amazes uh me and most of us now when a computer generates a da Vinci like sketch it's not doing the same thing and we might say oh it's amazing because wow what an incredible facsimile but it doesn't it didn't do the steps that
00:32:46 amazed us historically what what historically accounted for our amazement and seeing art so that's the that I think expresses the difference between sort of formulaic and algorithmic generated creativity and the creativity that we're accustomed to and is that something that's going to be bridged something let me push back a bit on the setting yes some of this as a very basic principles on top of which we have built this large-scale language models have been here for centuries and
00:33:16 in particular you know if you read the book from 1950 by Paul Shannon in fact we actually do see the perplexity as well as the low probability and how we want to actually maximize the low probability of the correct following words given all the context in order to in fact model the distribution it's all written there yes that's true but that doesn't necessarily imply that we are only relying on them in order to build these large-scale language models to build them anything like that was one of those people who actually built a very large ones in the first place and it
00:33:47 takes a lot more so than simple principle of counting and compression unfortunately it's not that simple there are a lot of let's say ways in which we can parametrize so any idea why everyone is going crazy about Transformer which was only proposed about five years ago is because it took us half a century to get to that point in order to come up with all those simpler simple but algorithmic techniques and tricks that we had to develop and then all this optimization algorithms that we use yes we can go all the way back to again 52
00:34:18 or so to talk about the stochastic grad in this sentence all but to figure out the what are the right way to you do the optimization took us another let's say half a century so what that means is that the what is amazing is not the fact that you're okay these aren't tempted based as a generators that have done amazing compression of the amazing amount of data that looks to do something amazing but how we were able to actually build this so it's actually not that different from let's say you know being Amazed by let's say all those monsters it is doing something really really
00:34:49 amazing inside that we don't know understand how it's doing it probably not in the way that human Masters are doing but still it is doing something that we didn't know and that we still don't know exactly how we doing so I think we're supposed to be amazed by this now of course amazed and we have to try to figure out how not to be amazed as well yes and that's I guess our job but yes so connect you to this I mean I think that it the fact that this is kind of a black box that we don't know how it
00:35:20 generates this output from The Prompt um also leads us to so computer science coding programming computer science and AI used to be a discipline where you the researcher the the software engineer the programmer we're having some goal in mind for a machine to do that and then you were coding that right so and then there were verification validation
00:35:51 techniques and so on but you knew what the machine was going to do you know so and then you would check that it would do it they didn't put any bug and so on now instead so that was the computer science exact science kind of approach now it's becoming with data-driven approaches more like our Natural Science so you build this thing but you don't know how it works and then you start experimenting and testing you develop some hypothesis and then you
00:36:23 test whether deposit is true or not so it become becomes closer to being a natural science rather than a computer science right so so it it's now merging these two approaches in that were typical of these two kinds of Sciences and disciplines because of the nature of these machines that we build but we don't know how they work so then we treat them as we treat the laws of you know laws of physics you know then we
00:36:55 try to find some properties of what we built because we don't know how they're going to behave foreign so there are problems small subset of problems that we know how to solve by
00:37:27 specifying what we want to do and then that's you know what the traditional computer science has focused on over the pesticide of centuries right and then there is a slightly bigger subset of the problems for which we know the nature or the evolution has figured out how to solve the problem so yeah for instance us driving right so any person you take give them 20 hours of the driving lesson we know that person will be able to drive in the New York City without animation so there there's a subset of the problems where we know solution exists but we don't know how to implement the solution or the implement
00:37:58 the solution and then for and then there's a slightly bigger subset of the problems where we don't yet know whether nature Revolution or the universe has speak about an answer to but then you know we want to solve those problems because so if you solve those problems it's going to be very very practical and whatnot and then of course there's a bigger sales of problems where maybe there's not even a solution to start with but then you're at the machine learning mode it's kind of the same AI is there to solve the slightly larger subset of the problems that we were we're not even supposed to
00:38:28 know how to answer but we know that the solution exists and we are going to build an algorithm to come up with a solution to that one so yeah in a sense if you knew how these solutions that were given out by this machine learning or the area algorithms were working then we probably would not have needed AI learning so let me give also an oversimplified view of my answer to your question so when you know how to solve a problem you can define an algorithm which just means a sequence of steps so
00:39:00 you do this and then you do this like a recipe no you take the ingredients then you mix the eggs and then you know whatever then you do this and you do this and in the end you get what you wanted as a result so if you know how to solve a problem in that way you code these steps into you know with whatever programming language you have you put this into a machine and the Machine will follow the steps so then you don't have a problem or not knowing how the machine goes to the result because the machine followed exactly the steps that you told
00:39:30 them to to get to the result but when you don't have these algorithms very clearly algorithm to find a result of a solution to a problem like in you know recognizing the phase of a person or driving or whatever this you don't have these very recipe very nice recipe because there are so many variables to consider so many insert so much uncertain you know so what you do you just give a lot of examples of problems and solutions problems and solutions and
00:40:02 then they let the machine learn from these examples by looking at the the problem finding something in output and if the output is different from the solution that you gave just modifying a little bit the parameter so that the the distance between the two is smaller and then with the next problem next problem so at the end you get the machine that by testing it it behaves rather well in that problem of finding a solution but you don't know that you don't you cannot see there is no
00:40:34 sequence of steps so that tells you what the machine is doing exactly how to solve the problem because there is no sequences as he was saying if there were a sequence of steps to solve that problem you would use it and you would not use these other algorithm but to recognize the face for example there is no one sequence of steps that you can use so you have to use this other approach but then you're not sure exactly what comes out and why comes out let me get an illustration of how little we understand so all all the big famous
00:41:05 successes have been driven by a machine architecture called the Transformer architecture however so you know I think initially people thought well look it's this architecture that's really responsible for all this stuff but now it looks like people are moving towards getting similar successes using other architectures so it's not even clear that it is the Transformer architecture that is really responsible for the big successes so I think what
00:41:37 she said about natural science it is much more like a natural science now you can build the successful models but understanding what they're doing is like we're examining it like we examine the nature of atoms or something so it's a it's a real puzzle oh we don't know how we ride a bike real I mean in the in the sense in which we don't understand computers and how they figure out problems we don't know how we ride a bike I mean we know we train ourselves careful with that I suppose you know
00:42:08 there are things that sound similar where there is a real simple algorithm like catching a Fly ball right there's a stateable algorithm yeah it seemed mysterious like riding a bike right there was turns out there was a stateable but following along with Francesca's idea it required Natural Science investigations to come up with that solution because the person on the street who rides a bike or the outfielder who catches a pop fly can't tell you well this is how my brain figures it out right it's it's a scientific
00:42:39 um problem to look into right so another example like that is chicken sexing so for many years that was you know so you I think that was illegal Pizza it was a cargo School and you didn't know that they were hens or going to grow up to be hens or roosters so commercial purposes you need to separate them and for for years there were these trained people who did and
00:43:11 nobody knew how they did it and somebody did you know investigated a little more and they figured out exactly what they were doing and now they have a very simple method so you can have a lot of things that people can do through training and you don't understand them and then if you investigate sometimes you understand them and that may well happen with these machines we we you know if we get many different architectures that could do the same thing look it could turn out that the fundamental basis is just
00:43:41 prediction that you know people didn't used to think that the way the mind or was so heavily reliant on prediction and um you know for some years now there's people who study perception not cognition but perception have been focusing on prediction maybe that's the key here is that however you can do the prediction that's what makes the thing run if I can just add to that so what
00:44:12 we've been doing in the lab for a number of years is precisely this kind of experiment and it is extremely tricky because it is sampling from this distribution a probability but you can actually tune it to get more surprise or less surprise it doesn't work like a Google search so many times when our students start to work with it they think it's going to work just like using Google search in fact the interface is very different because it's going to continue with whatever you give it if you give it very general language it won't give you anything back except generalities you have to give it the
00:44:43 texture you have to think about language in terms of complexity in order to get it to give you that complexity back and so and then you have to think if you train it to write like a writer how long do I train it you can over train it you can under train it so you're in this entire search base and then there's a question of how many times do you try it do you try it a hundred times and how do you figure that out so it's actually extremely difficult even to experiment with it and just to give you an example we created this Diva bot an AI that the Improv we worked with an actress in LA
00:45:15 and she had a very difficult time doing improv in the typical way which is that you give a kind of General response to your fellow improv improviser so that they can do something cool with it well if you give a general prompt to the AI it gives a general prompt back she had a better time giving very specific things like she was comforting a crying baby and I hope I can say this on YouTube she says how do I stop my crying baby well the diva bought the dpt2 which is a fairly small model said wear a condom
00:45:50 [Laughter] I will ask the computer to explain why that was so funny we were talking about slightly different forms of not knowing so there is not knowing there's a part inside of the process and there's a part which we can't quite explain and that doesn't
00:46:22 seem to be that unusual so there's plenty of engineering Solutions where we sort of know that they work and we design them but like if so for example like slash memory uh flash memory works by Quantum tunneling if you ask somebody like show me exactly how did the electron penetrate The Blob of you can't it's it's unpredictable and but we know it's like a very predictable process like we're comfortable with it it doesn't seem to be that right but there is a part of it part of the chain of explanation which is missing right and
00:46:52 it's it's uh we're okay with that so that's one type of annoying the other type of annoying I think very important that you mentioned is we don't know that the political social cultural effects right so we don't know what that technology will will do to us uh and I think you're very right it's very important to kind of like experiment and and think and and reflect and uh about what the effects will be you know when I looked at um I was reading the papers of liquider so remember lick lighter was a
00:47:23 Palo Alto research lab they did the joystick and they did a lot of early also early some of the first word processors were developed in this lab and so they developed so for example I remember they developed the copy and paste function and they're like our mind is blown you can take a whole chunk of text and like move it to another place but what I loved what they did and which we don't do is they said let's experiment with this and have a diary of what that does to our process of creativity because we don't know we
00:47:53 don't know what this weird topic and pasting will do to our process of writing and then there's like this purposeful let's integrate this into our research let's treat it like a natural it's it's a system that has complicated and unpredictable social so common can't live effect on writing and so they're like let's have there are these cognitive Diaries of like here is how my writing changed because I'm able to very fluently uh take a piece of text and move it to a different place I can cut it apart I can I can separate it so
00:48:25 that's and I'm like I love it and we should do more of that we should kind of have a lab which just does the social cultural political experimentations with these new new systems well that raises an issue that we really haven't talked about which is are there going to be bad effects of these these systems and can we think about that and think about what to do about them I mean the thing that worries me the most is that um you know with a huge success of open AI every big tech company is making
00:48:56 humongous machines you know so GPD 3 and 175 billion parameters and the new Google one is 500 million gpt3 I'm sure gbt4 will have sure be even bigger and the thing is the bigger they get the more hidden are the failures so they you know they'll be able to you know so we they won't the failures won't be just so much they're out in the open but yeah with all the money invested in them
00:49:26 right there's going to be a lot of pressure to use them and people are experimenting with um a pixel of inputs and robotic outputs now and I think with all the so much money in this stuff now there's gonna be you know robots with eyes and and you know could do things in the world and they're going to be powered by these machines that in some fundamental way or unreliable yeah
00:49:57 they don't know they don't have a world model they don't know what the facts are and they can't count very well they can't do you know I mean I'm sure that the you know GPT 4 gpd5 will be really good in small arithmetic problems but maybe there'll be some other arithmetic problem that it'll where it'll completely fail and those things will be used commercially and yeah and and it's it's not just a matter of being correct or failing is also that these models
00:50:29 that that they don't know exactly what it means to fail and be correct but it's also that these models are trained in this way they don't have really a clue of the values that we want to have in our technology when we allow the technology to make decisions or to recommend decisions to a human being so they don't know about fairness so it has been shown that in many examples these models are biased so they make they they say things they write things
00:51:02 in a way that shows that they are biased and and that's because there was really no curation of the data that they were trained on and so and and it's not possible to have a curation of that data maybe it will be possible by using these two steps approach like first you'd build something like gpt3 and then you further tune it for a specific task and then you you try to do the best you can to do bias mitigation in that second step but in any way you have to find a way to make sure that they do things or
00:51:35 right things in a way that embeds our values otherwise they got not only they they tell you wrong things but they tell you things even when the problem is not right or wrong but they tell you things in a way that is I mean the gbd3 has been tested in in a test the suicide hotline and and somebody interacted saying that he wanted to he was thinking about committing suicide
00:52:06 and gbd3 I don't I don't know if it was gpd3 but the language model responded said you should so we have to be careful about deploying them first we have to understand how to embed our values and I have this uh you know feeling that to a bad values in these models you cannot just do it with more data more computing power and only a data driven approach but you have to combine data-driven approaches with rule-based approaches
00:52:36 because you know bottom up and the top-down approach you cannot I mean in you cannot it's a big word I'm saying I don't see that emerging for now from a data driven approach only well I can push back a little bit because especially how we should we cannot actually solve these problems from the using let's say data driven approach in the sense that the the set of approaches that we have used so far in order to build these large-scale language models right it's extremely extremely narrow in particular
00:53:07 when we view the entire field of machine learning more artificial intelligence so let me just called it let's say more of a passive let's say efficient learning where the data was somewhat collected with purpose they may not necessarily align well with the purpose of how we use data and we try our best with the statistical approach in order to let's say count better compress better so that the work is going to generalize better however that's just a one very narrow sense of machine learning in fact the another side that I would call as let's say active machine learning is where we
00:53:39 introduce the assumption that these systems are going to actually enjoy with the other systems or even with the environment and then with that of course we can either make them actually interact with environments or we can also say that the well here's a set of data set that was collected based on the assumption that there were some kind of interaction that is going to happen in the future then we can now start using all those offline reinforcement learning algorithms Active Learning algorithms and even some of the Notions from causality and things that are beyond the
00:54:09 simple let's say statistical approaches coming and then my sense is that it may not have to be rule based but something that is beyond but already there are a lot of let's say candidate algorithms as well as these self-disciplines of the machine learning that are being studied and then they have been studied for many years like okay historically saying you know the electronic already wrote a lot about the machine intelligence and one of the things that Alan did mention and emphasize back then and let's say in his writing before this is the necessity of the reinforcement so the system interacting with the users or their
00:54:40 environments and then get the signal that actually tells the system about what it observes is aligned well with what it's supposed to do so I I almost feel like it's not really about the overall data approach but more like the narrow subset of the algorithms that we have used so far so what is the limitation of that and is it possible that we already have some solutions to many of the issues that have been raised with the current generation of language models maybe gpt4 or is going to be trained with something else or five
00:55:11 before I heard that they already trained this on everyone I like what you're saying and one of the reasons why I teach humanists and social scientists and artists all about AI is because I want all of us to have a seat at the table and discuss these things but one of the most interesting things to me so far is I ask students to come up with these rules and they're very uncomfortable and most do not and they find that their ethics and their value system cannot be put into the rules and and so we're going to have to deal with this in an interesting way and then the question is whose rules and who decides
00:55:43 right and can we even decide amongst ourselves if we agree on what these rules are so I agree with you that we want to have a human-centered AI but I'm not sure it's as easy as just coming up with rules but isn't it the question to all technology right so and what what you're saying and what what you've mentioned is embedded values are embedded okay what are the embedded values in the automobile right like look at the effects of roads and cars have had you know so it's interesting how we're in
00:56:15 this moment I think because it's called artificial intelligence where like we expect more from you but other kind of uh uh fantastic Technologies like driving right they seem uh banal and mundane but yet they have you know they like reform the world right the from from the politics of energy to the way our cities are structured and and also not I you know I can't say that that Accord to my values maybe you know it's easier to commute but then at the same time there's pollution and again it's
00:56:45 the problem where the value was not necessarily embedded in engineering itself because the engineering often has a specific problem whereas the political and the social impact has either it's immersed in the complexity of like human existence which is not you know it's not rule based and it's not kind of resolved yeah I think so I'm glad you've raised the point about just that using the phrase artificial intelligence makes us feel it one way one thought experiment I found that kind of helps for me at least is everything gpt3 in
00:57:16 these languages these models know of course is things that picked up from Human data right it's not like a chess program where an AI plays an AI this is a computer learning entirely from Human data so one way to think about it if I think it knows is from humans it's really a vehicle for providing access from to a human to our sort of a massive of human knowledge but we have other technologies that do that like Google search even Wikipedia Wikipedia is a distillation of
00:57:46 collective human knowledge that's accessible to an individual so somehow you know I know that Google search actually does use neural networks but no one thinks of searching Google as talking to an artificial intelligence and certainly no one thinks of Wikipedia as an artificial intelligence but somehow thinking of all of these things as ways of taking the massive human knowledge and projecting it to an individual gpd3 is a kind of more stochastic version of that but at the end of the day somehow that helps and helps me not think of it as an AI and just think of
00:58:17 it as a distillation tool but it also explains why working with it can be so Eerie because you do feel like you have access to this kind of collective mind of humanity in a in a really interesting way but Google searched I could say listen you get all the toxicity the misinformation the mystery I mean Google search is amazing because it's this bizarre window into human creativity I do think there's also a kind of an existential question that I see with a lot of my students who who want to be artists they want to be
00:58:48 writers and they move through these stages of grief almost in working with this right and you know we thought that robotics would be further along I would love a robot to clean my toilet and make my scrambled eggs but in fact that has proven very difficult and what we found is that you know disembodied AI has done and now can do all these intellectual tasks creative tasks so I think what is most unnerving that we have to somehow come to terms with is that we thought we would have the Jetsons with the robots first and instead many of us whose
00:59:20 entire way of being is based on kind of intellectual work are having to come to terms with the fact that we've actually succeeded there first well I can think of a new rule you know all robots must wash their hands before leaving the path um I wanted to say you know we all uh use rules when we speak to each other right so language already has rules embedded in it but one of the things that hasn't come up so far is that uh in the history of human discourse people make promises and vows and they swear
00:59:53 Oaths and things that those are all basically language-based behaviors that we don't see or hear so much about and maybe maybe there maybe it's happening I just don't know about it but I think that's such an important aspect of human language generation that we we make promises well one of the points that people often make about the large language models is they have no goals um so you know you can try giving them a goal you can say you know you can you
01:00:25 know use text to tell them their purpose but then that's just more text that they will then follow with more predictions so um you know it's another way in which they're really fundamentally unlike human or whatever kind of mind they have it's fundamentally unlike a human mind okay I have a question kind of for everybody because you know when I'm looking at the problem that we were given and you know there's nothing in there that that has the words intelligence or mine and how and it's
01:00:56 interesting to me that throughout like those words immediately enter their conversation and in fact in the research in the AI research program from the very beginning language and intelligence and mind so language is kind of one of the most marked kind of uh feature of intelligence but you know we can I mean we talked a little bit about Vision we talked about I mean there's a whole bunch of other so I feel like there's a cognitive linguistic slide that we are engaging in where we're beginning to
01:01:27 speak about compelling language generators but right away we're saying okay is there a mind there or is this you know is this intelligent where is where is again we wouldn't neces you know I don't have the same question to a calculator or to like a more complicated statistical model that whatever predicts the weather I don't go like oh does it actually feel the weather or does it you know is it intelligent in their play so there's something I mean I guess it's a question to everybody is why you know what role does the how is the connection between language and mind and
01:01:59 intelligence and why are we naturally sliding immediately from language and text into an into intelligence in mind so there are two questions that have been very much prevalent in the literature of the stuff that have not come up up here and I think there's a good reason for it one of them is that is it really intelligence I I don't myself think it's a very useful question because it's such a vague notion but here's the other one is kind of interesting which is
01:02:29 a lot of the discussion has stemmed from a paper by bender and Kohler or something they call the octopus Dad how can you know can they they call these machines stochastic parrots so just trade out words and as out comes more words so you know you're never getting to the world and I think for good reason that hasn't come up here because it's really completely irrelevant um you know the um what we've uh the the issue of can something trained on more
01:03:01 words be in some reasonable sense intelligent creative solve problems do things it's a kind of non-issue really um so uh I think you know there's a good reason why that hasn't come up here but it has dominated a lot of literature well but doesn't uh having the failings that you mentioned at the very outset the sort of glaring failings where it seems computers don't have a world view isn't that another example of how intelligence would some some definition
01:03:33 of intelligence would would perhaps address that problem well I don't think you could solve it by definition there are certain um maybe we could use a neutral term intellectual capacities that they do extremely well better than us and other and others where they don't so I think we have really focused on on that issue here of and also trying to understand what it is they are doing I do seem to know things a little I mean
01:04:03 we you mentioned that as well um and that's part of when we do experiments we're trying to figure out what does it know now do I mean no in a human way absolutely not when you see it fail those are Clues so gpt2 is a much smaller model we had a student train it to write MTV Daria episodes it seemed to know the characters they seem to know the kind of plots and the ways that people interacted but it would make really bizarre Goose a person would pick up the phone and then pick up the phone later in the scene uh somebody stirred lasagna you know and you go it's ridiculous
01:04:36 the fact that you notice means most of the time it does actually though and and so that's the fascinating thing and we had a creative writer give it a prompt about an Inky black C it immediately knew that these people were probably on a boat that they might be fishing that when they pull up a net usually what it has in it it's unusual to have a body which it did it knows that the ocean is not made of ink right so it does seem to know no stuff and that's what we're experimenting with when people get upset because it's not knowing like a human
01:05:06 but we're still trying to figure out what it knows and what it doesn't know coming back to the steering the lasagna yeah but you never know maybe on the web but there was somebody's TV you never know what people can do to this recipe yeah but because people actually make an argument in fact this is the argument there's quite a few people have been making over the past few years and just two days ago David Chalmer said that well you actually made the same argument is that it is it possible if the system
01:05:37 had a work model we would expect it not to save lasagna or once he says that you know it was steering something right but then another way to think about is that if perhaps that actually tells us that if the model had the perfect predictive model then you added a probability assigned to lasagna in that context would be zero and then in that case can we actually say the other way around saying that you will look at this amazing predictive model it has that probably implies that it has the word knowledge or the word model in it already so then perhaps by simply making
01:06:08 prediction better and better or the model's predictive capability better and better maybe it's going to automatically we'll have to come up with a word model that's going to reflect how word works and how we think and so on right so yeah is it really yeah no I I agree that's why when I say it cannot I mean you cannot say this approach cannot build a world model Maybe by building a better prediction model then that would build not an explicit word model that you can
01:06:39 see but then in as a result the result would be as if it were it had an explicit word model uh but but yeah but I don't know I don't know is the big difference we can make explicit we can we can formulate explicit ideas right keep a record of explicit facts and that's not what these machines do and I I think the question uh is
01:07:11 can the kind of compression training that they get give that effect I mean I'm betting on no myself I'm betting that there that there will that there is something about being able to be ex being able to be explicit yes but of course one could say I mean being David's Advocate that if one opens my skull and looks inside it doesn't find any explicit things there any explicit rule or logic rules or
01:07:41 whatever but then I verbalize my explicit model in some way that says okay I know about these rules and this and that but you don't find the rules by looking inside so one could say this is similar to this huge machines with the huge number of parameters that what you see if you look inside this Justice billions parameters and the values of this parameters and Tom tell you don't tell you anything about this machine having a model or not but but then by
01:08:14 generating the output maybe you realize that this machine is a world model so so in some sense I see that this doesn't roll out having a role model even without an explicit characterization of the world model of the rules that are learned but um well we just had this example of lasagna stirring here and it's I think it's interesting the way in which human humans sort of curate their own information in a weird way because it may depend on
01:08:46 who the expert is that's teaching them about it so you had this wonderful reaction I hate to sound like I'm biased by Italians and what they think about things like lasagna but when you go you don't start a lasagna no that was much more weight to me I'm saying that I'm sure there must be somebody in the world I'm sure of that no exactly maybe in this country if your parent says something to you or someone who you know somehow is as a
01:09:16 human this person's advice here means a lot I'm going to pay attention to it yeah this other person's advice does you know you learn these rules in a very different way right because of the sort of emotional balance and the understanding the world view about who the expert may be yeah I mean what's interesting to me is again we are the kind of problems uh is being that are being identified in this conversation they kind of assume like a totalizing intellect and then like oh like these are the things that are missing to get to this notion of like
01:09:47 the perfect intelligence and and again there's something about language because the language pulls in all of the world and then so we expect it kind of to do better and we we see this kind of but they're Universal failings we don't expect the same kind of sort of the same kind of um what should I say not hubris but they say the same performance from other robots you know so you know they're amazing robots that build machines and you know they're and they're using complex statistical Vision techniques we never go like oh why
01:10:18 doesn't it have a world model of a lasagna you know why doesn't the robot in the you know whatever Tesla Factory uh and you know so my question will be why why is it so but as soon as we start talking about language generators which which are often built with the print also it's a machine that's built with a particular purpose we right away want to say does it have emotion does it have intelligence why doesn't it understand lasagna right in a way that we don't demand a lot of machines but by the way also with with people no like I was
01:10:49 raised in Italy so in a very biased way about what you should do with a lasagna for example so so in some sense I built during several years I built like a model of what you should do with that object and what is appropriate and not appropriate to do I was not raised or trained you know with the data coming from all over the world maybe if I were trained like that or if I grew up with data coming from all over the world for
01:11:20 me it would be equally good to steer or to not steer a lasagna maybe you know so so and in some sense we are expecting from an object that is framed with equally important data from all over the world all the different cultures all the different regions and maybe there is a lot of contradicting evidence of what you should do with an object and then of course I I as a person I could say well you could steal or not steal lasagna if
01:11:52 I was raised with experiences from all over the world which I I did not you know I was just raised with experiences one region of the world so in some sense it's not surprising that there are you know contradicting pieces of information that are collected by and and are helping training these machines and then the machine pits out pieces that are maybe consistent with the sub part of its information that it was trained on and
01:12:23 not the other one well and to add to what you're saying we have creative writers who are trying to get it to be creative so we are tuning the hyper parameters to get creativity and then we get stirred lasagna so it could be that that stirred lasagna is sampling from the surprise right as well I'm gonna try it yeah I want to push back on Dennis a little bit so you're you're sort of blaming the human saying why do we keep asking for more and expecting more from gpt3 than other things and I want to say it's its own fault it's the computer not
01:12:55 the humans it's because it bullshits so much right if you have a robot that's designed to assemble a car that's what it does but gpt3 set makes up things about everything and it acts like it knows so much and it just bs's so yeah there's some responsibility human that we look for more but it's responsible it vastly outsteps it's not trained for a purpose it tries to do everything and It embarrasses itself isn't it doesn't that make it very human in a way yeah but with the horse
01:13:26 what about these failures that are so funny you know we're getting a big laugh out of some of the ways in which the the generator comes up with these gaffes like what is that a clue to anything do you think about creativity because so there are actual genuine let's say uh degeneral cases that arise from the our current let's say practice of trading these models in fact there are quite a few people including my own lab where we actually look into those that say failures and try to come up with the let's say mathematical or statistical you know say justification why those
01:13:57 failures happen and they have to fix them but there are so many of them at the moment to the point that they were fixing one at a time but yeah there is a chance that the what we really need is a new paradigm of how we train these models rather than fixing every single digital cases at a time because we're just adding in new term to the loss function every time because I saw you right in lasagna on your schedule it's the first thing I'm doing when I get back and just to add to your idea why we are fascinated by language or the language
01:14:28 generator over the robots and one another city you know like do we perceive or detract with the world in many different ways you know do we perceive the word by looking at them hearing about it sometimes you will touch you have to remove things but then languages actually yet another medium by which we can interact with the environments right so by interacting with the other Asians or the interacting with the other computers whatnot and then I think that one uh one unique aspect of the language is that it actually expresses very uh diverse spectrum of the abstractness so let's
01:14:59 say what we see it's extremely concrete we see what is often there I mean there's a bit of a hallucination or another but generally we see what is there we hear what is actually being what is hitting our actual electron right and then we touch things that are like here again beside all those hallucinations but languages where we can actually Express all those as extremely abstracting as well as extremely concrete thing in a single let's say sentence single phrase and I think that that actually makes this medium a very very unique and fascinating compared to other things
01:15:29 they're being done the other dichotomy we've not touched on is the word the distinction between sort of syntax and semantics that hasn't come up at all I mean does anyone want to talk a little bit to that well that's the octopus test that I mentioned uh yeah so um look uh it depends on what your theory of semantics is what it is for the machine to know the meanings of the terms I play a like a view called
01:16:00 functional role semantics or conceptual Rose Medics which says that if the roles of the representation is in the machine are the right roles then it will understand the words um and there's a there's a recent paper by Steve banter does he arguing for that view and uh seems to me that they have substantial elements of the right conceptual role so I think there is some amount of understanding of the representation in these machines but if
01:16:31 you forget a little bit about the importance of course the main importance of the content so the semantics of what is being written but in terms of the syntax this is really where the first amazement is is how can these machines under without telling them the rules of syntax can they can write in such a fluent and eloquent way in in a language or even more than one you know that I mean to me was the first the first my first you know uh approach to the
01:17:04 language generators what this one oh my God is writing you know in a very eloquent way then of course if you go and look at the semantics then you then you have things to say you know about the quality of the semantics but the syntax is really much more amazing than the semantics I think you know the way I the way I think about it is is there is an underlying statistical representation of language that that assigns kind of words in their
01:17:34 currencies to like a vector space model like looks like the stars and certain things are likely occurred to next next to other things so when you say I want to eat you know blank some things are probable because they were they're occurred in the training Corpus and some things just rarely or improbable rarely occur so now is that so when you translate language into a statistical model this is where I begin to think okay I I don't is that model sentient is it does it have semantics it is what it
01:18:06 is it's it's a particular mathematical model that that represents language in in a way so so I'm actually much more cautious to not go the next step and to say what you know to to personify and make a metaphor and kind of animate it to say like is it going to do the this kind of stuff I just want to add one more anecdote from history that the the paper by Markov and Mark of the mark the original Mark of chain generator by Mr Markov which was you know published in
01:18:36 either German or Russian and then it made its way in Translation you know which had you earlier you had a great explanation it's like it's a chain you know it looks back it sees a letter and then it says what's the probability it was letter by letter that paper wasn't Pushkin it was an ungenerating pushkin's pros and he by hand created like a simple letter by letter generator that produced and it was also like it's amazingly effective it's the simplest mathematical model it produces like very nonsensical Pushkin but but nevertheless
01:19:09 it was effective right it right away got us to like the point where we are saying like wait a second uh this thing is is producing sentences it also has the same kind of problems he had has difficulty with context obviously because it only looks back one letter right so so these are you know so I think by by trying to not fall into the same metaphorical the same kind of language which we slide into which is intelligence mind you know uh sentience and and just try to like
01:19:39 restate what we mean in other ways like this is a statistical model do we say what do we how do we talk about statistical models I think that helps us kind of move past some of the I think last night you know we saw some beautiful magic tricks like there's some magical tricks here that are in our minds they're they're it's it's the I think they are our failings in a way of incorporating these these techniques into our life it's the last comment and we're gonna have questions from the audience one thing and that just came to mind I haven't thought about this previously is when you think of like the
01:20:10 historical context of Turing test it's kind of an artificial environment right where you you blind yourself to the human in the computer but one thing that's kind of amazing is you know over the last 20 years so much of our interaction is and purely text form right I text friends I type on social media so we're in this world now where we interact with humans in an entirely text-only way so when we see something like gpd3 that I interact with sex only I think that's part of why it feels like it's more human or at least we think of
01:20:40 those terms and ask about it because that's a standard form of interaction now we don't you know if you look at all old school sci-fi it was about physical robots because that's our world but now we live in a world of texting ngpt3 is potentially as real as anything else it's not but that's it's just somehow our world has changed separate from the AI as well yeah foreign we have some time for some questions from the audience would you like to approach the microphone thank you
01:21:13 conversation is very interesting um The Prompt as it's stated in the program here it says that the program can create language that gives the impression that it is thinking and thinking is the thing that I'm sort of want to press the circle to talk a little bit more about and it strikes me that um when computers first arrived we didn't intend to think of them as violating entropy you know you get them to go and then they break and they've you know uh they have a lot of they
01:21:44 require lots of energy and vacuum tubes and all of this but now with the rise of like language learning and the sort of artificial intelligence there's this kind of impression that we have that the computer somehow like us violating the second law of Thermodynamics that is somehow creating centropy and and generating things outside the realm of like the you know the heat death of the Universe I think in what I'm basically saying is that is the problem with computer thinking
01:22:15 basically the same problem we have in understanding our own thinking like if we don't understand what Consciousness is in the first place how are we I mean it seems like there's a really quick move to understand the computer is doing this because it seems to be doing what we do which is make connections we're creative we flourish we do all these things but in the end are we even actually thinking according to to that model when one point doesn't directly address that but kind of tangential is um the the prompt also mentioned
01:22:45 something about does being aware of a code kind of affect its realness and I hate to say it but the fact that okay we don't know exactly the black box of machine learning and deeper neural Nets and all that but we do understand neural networks in the sense that we've designed the algorithms we know what a Transformer is and I hate to say it but the fact that we know exactly what algorithm gpt3 is running not you know the parameters after it's been trained but the the raw algorithm knowing that does take a lot of wind out of the sails it takes away a lot of magic we I can't
01:23:15 look in a human brain and understand it architecturally to the same level that I understand a Transformer so I think part of why it's easier to ascribe Consciousness and thinking and sentence and all these things to organic life is we know much less we don't know everything about neural networks but we know so much that it's really hard to believe it's thinking that it's conscious that it's any of these human type of things or animal things okay the um the idea of Amazement which we brought up so much at the beginning also seems to be a fancy way to talk about
01:23:48 that would be to say oh we observe things with low entropy that surprise us right and we've already learned how to do that and you're right humans in our human intercourse uh among ourselves uh discourse I think it's a better word whoops uh that uh we we we're we're exposed to those sort of flashes of low entropy that our Consciousness can create does that mean that that when computers do it that's another instance of thinking like humans I don't I don't
01:24:19 think so but people may disagree so again we don't really have theories necessarily that make sense of the data and so we are in this experimental phase where we're just um you know one of the frustrations of working on it is you just have to give data point after data point after data point but we can't necessarily say what it all means um you know we had one student who was a Bernie Sanders supporter senior who decided to have gpt3 write a Lullaby by Marx and it did a beautiful job and then
01:24:50 a conversation between Adam Smith and Karl Marx and it did a beautiful job and then and one shot not you know five times and then you know what would Karl Marx say about Bitcoin and it said well he might say this I mean it was all very you know so is that thinking is it conscious of course not you know but it is doing something that we recognize that as difficult in terms of intellect you know and and we don't really have a way of of making sense of that with our current theories we just have to look at
01:25:20 the examples and yeah I think important thing to note here is that even humans right so whatever we say and you know whatever the new knowledge that we seem to create does not necessarily actually become important knowledge but it's always all about looking back right hand side is 2020 so you know that we do increase the entropy I do whatever we say something almost everything is going to be forgotten and it's going to be conserved noise when we look better a couple of years back so that we do increase the
01:25:51 entropy but General and then these machines are same thing right so these models were trained to minimize the entropy based on the data and then what we know is that the the entropy of the trend of the learn distribution has to be greater than equal to the original entropy so it's always going to increase the entropy but then the thing is it all comes down to distillation like process right so we look at all those things and then we pick what our important creative amazing things and they were able to kind of say keep them and then maybe the
01:26:22 important process then this kind of language Generations great thanks yes uh thank you again for a great panel I have a question we talk about language like in the literary sense like putting words together but I would be curious what if gpt3 can formalize and then try to predict mathematical language and what I'm trying to get at is you know about getting theorem right right that like the the mathematical language of
01:26:54 arithmetic is either incomplete or inconsistent I would be curious if if a system like gpt3 can try to do all the combinations that mathematical language can generate all the possible sentences and find out a contradiction in arithmetic which we say could exist but I don't think anybody has fathomed what inconsistency inconsistency lurks within arithmetic I would be curious if if gpt3 can be applied to mathematical
01:27:25 language thank you uh equations and so on but I think the one thing that is interesting is we don't even have to go into the incomplete theorem or anything like that but in mathematics and computer science theory we all we have a very well established theory of the hierarchy of the problems problems that we can solve based on the complexity right so the memory of that is just the container complexity and
01:27:55 then one thing we know is that if it is language models that we build have a very fixed amount of compute that is assigned to each and every input so for instance let's say we're trying to solve traveling sales person problem and we know that it is Olympic complete problem and then we know that d35301 not has only the quadratic complexity with respect to the size of the input let's say graph size then what we know is that the unless unless traveling sales person or the npt.3p unless that happens we know that there will be instances of the traveling salesperson problem that cannot be
01:28:26 solved by this and the gpd3 so I don't think it's about the training or model better getting more data but there are some fundamental or computational limitations that are actually being imposed by our own construction and how to go beyond that is a kind of as a research direction that people are looking into and they ultimately less Universal Transformers but but yes to to just connect to what they said that yes these large language models have been further trained to be used for example for generating code
01:28:58 which is a special kind of text with some rules because of the coding or to generate plans sequences of actions or to generate other structured tests other forms of structural tests and so the way that is done as we mentioned the beginning is that you take this large language model and you further train it for that specific domain whether it's code or plants or other forms of structure text so the not not related to the
01:29:31 computational complexity thing but but to say that yes this this language generators can be used to generate specific forms of language such as code plans and other things but it's important to remember even when you train it to create mathematical language it's still statistical it doesn't know when it's right or wrong so no it's not going to find some contradiction because it doesn't know when it's right or wrong to begin with our students also not know what's happening
01:30:03 but we know whether right but actually failures and logic are is one of the main tells when you're trying to distinguish right right now which is very surprising because we think of computers as highly logical and yet that's precisely what these models fail at yeah I thank you all for this amazing talk um we talked about world models a bit and uh I think it might be interesting to take the view that the successor failure of a of an algorithm is actually nothing to do with the algorithm but rather the
01:30:33 human judgment that deploys in a given context a world model to judge how what a computer has done and we mentioned about the continuing progress of of AI in the future algorithms and I can see actually two vectors one in which based on how we deploy these in actual real world systems we are so used to seeing all this quote-unquote PS that we actually lower our judgment function to say that this is acceptable to us or these we don't have as sophisticated ruled models to judge the outputs of AI
01:31:07 because we're so used to growing up with them um so if you have any comments on that point and the other question I have for all of you is has coming to the earlier question on what does it say about mind that we started with has it changed for you how you understand yourselves as human beings gbt3 is about 20 or 27 correct on two digit
01:31:40 multiplication problems you know 25 times 72 it can't it's very poor uh you know come on that's sort of by you and Sanders that's not a particularly sophisticated kind of issue so you know the the failures are are severe um and as we were discussing earlier I think the key issue is is this a matter of different kind of training bigger
01:32:10 models you know more training um I'm sure that that some Future model will be much better at at these problems but will there still be you know some astonishing failure at some kind of mathematical logical thing but I mean to answer the question I think the answer is yes I mean because you're seeing also this discussion that we always often do these analogies and so whenever we try to
01:32:42 test or even analyze or discuss this large language generators we always all AI General we always think at least I always think in terms of human beings you know we learn from data we learn from examples we learn from Roots we learn we abstract from data to rules so and that thinking about how humans do and reason it strength translated in in our for example in my work in my research
01:33:13 project is translated my understanding of how humans do things is then translated and tries to be adapted into the AI space so like I don't know for example my my recent project is about thinking fast and slow in AI so to take that cognitive theory or how human make decisions by combining The Thinking Fast and the thinking slow and see what it would mean inside by the machine what is the thinking fast it's just machine data-driven approaches in the thinking fast or does it also generate an
01:33:46 emerging thinking slow Behavior or you have to add the thinking slow Behavior because it doesn't emerge from there so in my job I always do this that analog analogy between humans and machines or differences that certainly helped me in recent years well you know to understand better how human Minds works as well it can also say when I confront questions like this so I'm reminded of
01:34:17 the old distinction between kind of functional like is the proof going to be in the outward representations of intelligence or is it going to be in the inward some kind of inside structure that you know there's a long philosophical tradition and thinking about it but I myself am skeptical about these algorithms telling us anything kind of internally a kind of any answering existential questions about the mind or God or Love or Whatever whatever it is but functionally I and
01:34:47 functionally I think the the answer to you know what effect what will these algorithms teach us that's not a speculative question it's a question is how will these algorithms will be integrated into our daily practice and I think that's that's a matter of observation that's a matter of engineering um one example I'll give you is that I guarantee you all of you who teach our students will be using these algorithm so they're already using these algorithms to write fairly mediocre like C plus B minus papers because it's so
01:35:19 easy you can right now go to a website put in a bunch of like really good papers and produce like a someone then sensical but you'll be like oh that's an interesting idea I never thought of that you know B minus now that changes that changes my I mean maybe hopefully not at Colombia but that changes my um practice of teaching that means when I sign papers I can no longer view a paper as well and as this like special insight into my students ability to comprehend
01:35:51 something because I know now that the student is thinking with the computer uh in a hybrid way in a way we've always been doing but now the computer is playing more of a part so now and this is I don't have an answer by the way now I'm thinking okay to be in front of this trade can I give them and I love your the various experiments your lab is doing can I give them papers and say actually explicitly right write them with GPT in some way and then show me
01:36:21 kind of what what is the next student paper format what is it going to look like and I think it's going to be something different post but because because these algorithms are unreasonably effective because they're magical and they're they seem to surprise us in a particular way that means they will transform our practice of teaching in this in this example I think that's excellent yes well I hope my own internal language generator chooses the following word I think this was a staring conversation
01:36:52 and a wonderful one at that very oh we have another question yes oh my goodness now I I hope you will not call the wagon and have me sent to the loony bin for what I'm about to say I'm a jungian psychoanalyst and part of what we do is we we try to learn all the mythologies of the world which is of course impossible but get trained in mythopoietic approaches mythopoeic analogous associative approaches and the way you're talking about what's
01:37:23 in the computers is the same way we approach dreams I'm really so struck by this because a person could have a dream that Egypt was sent over to the Golden Gate Bridge and we would look to see what the unconscious associations are is this a person who's putting Great Value and renewal of life out of an Egyptian system but is drawn to commit suicide and thinks about doing it on the Golden Gate Bridge so then you look at all the
01:37:56 underlying associations some of my fantasy is in some weird way because you are all trained so well in rational thought that the unconscious associative mythopoeic level is getting picked up in some way and so just just think about it but consider it which means look at it from the point of view of the Stars
01:38:28 well thank you again everyone and we will be reconvening at 2 p.m for our talk on the metaverse thanks again thank you thank you thank you