Showing posts with label cognitive science. Show all posts
Showing posts with label cognitive science. Show all posts

Bayes, bias, p-hacking and the Monty Hall problem

TL;DR: Bias is a fundamental concept in the science of agents; bias of a source is fundamentally related to cognitive bias.


You open the news in the morning, and you see the following headline from The Pan-Pan times: 80% of Chimpanzee Party lawmakers have criminal cases filed against them.

"Outrageous!" You cry-- "The Chimpanzee Party has no respect for the law! I certainly shall not be voting for them!" 

But upon subsequent investigation, you find that while this statement is true, in fact, 80% of Bonobo Party lawmakers also have criminal cases filed against them. Or perhaps you find that they have criminal cases filed against them -- in an unrecognized country created by some crackpot on Reddit. And the person responsible for generating these headlines was aware of this fact.

You feel cheated, even though everything you heard was the truth. Even though The Pan-Pan times only gave you information, and you acted as a rational agent when exposed to this information (well, actually you didn't, but let's ignore that for now), you feel like you've been "exploited" somehow, "tricked" even. 

Is this truly possible? Can a Bayes-rational agent truly be fooled by cleverly selecting information to provide?

Let's think about the problem more carefully. 

Our ultimate decision might be governed by the following rule: vote for whichever party you expect to have fewer criminal cases filed against its lawmakers. We may have some prior distribution on the fraction of criminal-accused lawmakers of each party, say $\mathrm{B}(2,3)$ and $\mathrm{B}(3,2)$ for the Chimpanzee and Bonobo parties respectively. Under the prior, the expected fraction of criminal-accused lawmakers are 0.4 and 0.6 respectively, and one would vote for the Chimpanzee Party; under the posterior, the expected fraction of criminal-accused lawmakers are 0.8 and 0.6 respectively, and one would vote for the Bonobo Party. 

Or.

Our ultimate decision might be governed by the following rule: vote for whichever party minimizes the expected sum $X_1+\dots+X_{100}$, where say: $X_1$ is the regulatory burden the party will place upon coming to power, $X_2$ is the tax rate the party will implement upon coming to power, $X_3$ is the number of criminal cases against the party lamakers, $X_4$ is the number of bad words the party's candidates use on TV, $X_5$ is the number of lies the party's candidates say on TV, $X_6$ is the number of dissidents the party will throw in prison upon coming to power, etc.

Let's say, for simplicity, that these are all Bernoulli distributed variables -- further that each $X_i$ (Chimpanzee Party) and $Y_i$ (Bonobo Party) are distributed as $\mathrm{Bernoulli}(0.5)$. Then in this prior distribution, we would be uncertain as to whom to vote for, as both parties have $\mathrm{E}[\sum X_i]=50$ is lower. 

And suppose the Pan-Pan Times tells us: "We looked at $X_{3}$ and it turned out to be 1 for the Chimpanzee Party!" Now, $\mathrm{E}[\sum X_i]=50.5$, while $\mathrm{E}[\sum Y_i]=50$, so we vote for the Bonobo Party. 

But here's the thing though: the precise information you receive isn't $X_3$ is equal to 1 -- it is the Pan-Pan Times's report is "$X_3$ is equal to 1". And that's what you should be conditioning on.

If you were to condition on the Pan-Pan Times's report is "$X_3$ is equal to 1" (call this variable $\Pi$), well, what would your inference look like? You could apply Bayes's theorem, etc. but simply put -- let's say the Pan-Pan Times's report is some generative process that looks at all the $X_i$, and reports one that is equal to 1 (i.e. tosses a hundred coins and reports a heads) -- it can do so in all but $2^{-100}$ of outcomes, so the only information we're given is that that particular outcome (where all $X_i$ are 0) is not the case -- the only information we're given is that the Chimpanzee Party isn't literally perfect -- so that $\mathrm{E}[\sum X_i] = 50/(1-2^{-100})$.

(OK, in this case, the decision is the same -- but for example, suppose our decision was instead "donate some sum of money proportional to the difference in $\mathrm{E}[\sum X_i]-\mathrm{E}[\sum Y_i]$" then the decision would be different.)

To instead condition on just $X_3=1$ -- rather than the full information provided -- is a cognitive bias. As actually implementing Bayes's theorem everywhere is expensive, the mind processes information using heuristics -- one such heuristic is that only some information is selected to be conditioned on, leading to selection bias. Indeed, all such "source biases" are fundamentally manifested as some form of cognitive bias -- the use of negative terms to describe the Chimpanzee Party, for example, is the exploitation of some sort of the association fallacy, repeating the word "Bonobo Party" for hours on screentime is an exploitation of the availability heuristic, etc. 

A perfectly Bayes-rational agent -- that takes into account all information it is exposed to -- is immune to being tricked in this way. But a real agent, which uses heuristics, can be exploited. The idea is that if such biases are systematic, then it can be predicted and avoided cheaply. 

Riddles and mental models

I've been thinking about some assorted cognitive processes over the past week, trying to remove the mysterious/magical feel to cognition. I have no idea if my observations reflect the cognitive science literature, I don't know if what is true of my mind is true of others', and I haven't done any actual serious testing of my guesses, but as long as the ideas described are a possible description of reality, they should work as an acceptable basis for doing AI, which is the main purpose of cognitive science anyway.

Observation: thoughts are wordless

People often describe their thoughts in words, and often insist that their thoughts are in the form of words mentally. This is obviously false, because (1) words need to be rooted to some notion of meaning, and (2) you need to already know what the sentence you're about to think is before you think it, or you won't get into the right grammar, etc. 

So "thoughts", whatever they are, are not words. So what are they? Let's think about some example thoughts one may have, in the form of words:

  • "I'll find carrots tastier than cucumbers, so let me eat it."
  • "Alright, let's focus on this for now."
  • "It's sunny, let me close the curtain."
  • "What are thoughts?"
  • "Saying the words What are thoughts? won't help me progress on the question."
  • "What will help me progress on the question?"
  • "What thoughts am I having right now?"
  • "I need to develop a stronger intuition for this"
  • "Wait, I thought thoughts weren't words, what's this?"
  • "Stupid question, never mind, but I should add that clarification to the sentence above."
  • "Ah, stupid spelling mistake."
I think that they key abstraction here is the idea of mental theories and models. Thoughts are beliefs and reasoning about implications between mental models -- i.e. a very model theoretic concept. This is true for thoughts about how things are, thoughts about decisions to make, introspective thoughts (these are just self-referential sentences), whatever.

One may consider these model theoretic sentences to be generalizations of linguistic sentences -- a sort of language that every conscious being naturally has and vividly imagines, lives in. It is not necessary to talk about how your brain attaches meaning to sentences in these models -- associates them with reality, because your brain lives in these models, that is its reality, it is the source of all meaning. One of these mental models is the very visual picture you see of your surroundings.

Observation: riddle solving.

There are two kinds of riddles -- one, the ancient Greek kind, like "what has four legs at dawn, two legs at noon, three legs at dusk, one foot stuck in its skull, and zero legs after time-traveling back to dawn?" and "what comes first? the chicken or the egg?". These are stupid and pretentious, and have nothing interesting to tell us.

Two, the kind of riddle that people scoff at as childish and uncultured, and is therefore actually somewhat interesting. Some examples:
  • A bus driver was heading down a street in Colorado. He went right past a stop sign without stopping, he turned left where there was a "no left turn" sign, and he went the wrong way on a one-way street. Then he went on the left side of the road past a cop car. Still - he didn't break any traffic laws. Why not?
  • Samuel was out for a walk when it started to rain. He did not have an umbrella and he wasn't wearing a hat. His clothes were soaked, yet not a single hair on his head got wet. How could this happen?
  • A pet shop owner had a parrot with a sign on its cage that said "Parrot repeats everything it hears". Davey bought the parrot and for two weeks he spoke to it and it didn't say a word. He returned the parrot but the shopkeeper said he never lied about the parrot. How can this be?
  • The 22nd and 24th presidents of the United States of America had the same parents, but were not brothers. How can this be possible?
  • What goes place to place yet stays in one place?
  • What is harder to catch the faster you run?
(Source: Riddles.com, a site I frequented as a kid.)

(There's also a third kind riddle, having to do with pattern matching, like What did the ___ tell to the ____? and Why did the _____ _____?)

When you hear about the riddle about Samuel and the rain, your brain immediately infers a particular mental theory from the problem text -- and very quickly ends up at a contradiction (or confusion), something that eliminates all possible models. In a sense, riddles are all about the art of noticing confusion -- about identifying the axiom you subconsciously assumed in your reasoning. You subconsciously assumed that the bus driver was in a bus, that was included in your mental picture, but was not implied by the form of the question itself.

The last two appear different, but are based on the same principle of being able to infer multiple possible theories and evaluate their consequences.

Things being alienly alien

Aliens in science fiction are stupid. They either look like almost exactly like humans in form, or just a hodgepodge of features from earthly animals made with the intention (but not outcome) of appearing truly alien.

The reason for this is that making truly original art is hard. There's no algorithm in the brain that just randomly generates something that has never been seen before, because something "randomly generated" like that will just be noise. It will have to already have learned a distribution to sample from, based on previous knowledge, i.e. it would be an act of transfer.

But we do generate truly new knowledge, not based on observation of the environment or transfer of existing knowledge -- it is not necessarily so that the most important creativity is "truly new" or anything, but it is exciting, and we ought to think about where it comes from.

My hypothesis is that such "truly new art" is a output of some kind of specified algorithm -- indeed, it is not the haphazard exploration and imaginative thinking that lead to truly creative ideas, but algorithms that are cognitively well-defined (i.e. they need not be an actual logical deduction algorithm, but use the standard, honest outputs of basic cognitive processes, e.g. "stick this image to this image here", "pick a point of this image", "find a solution to this problem").

As an illustration, I have attempted to produce such a "truly new/un-inspired" fictional alien species from setting up a basic "axiomatization" (the conditions of the planet on which it forms) and following my brain's emulation of the evolutionary process to see where it leads.

Note that this isn't a scientific prediction (and these aren't actual "axioms")! This is not intended to give you an actual prediction of how an alien species on such a planet will appear, because the algorithm followed is not a truly mathematical model of evolution, it is not actually a logical deductive algorithm, just some cognitive processes in the brain, i.e. taking steps that "feel" right. The point of this is not to make deductions about aliens, it's to understand aspects of human cognition, and get some nice science fiction out of it. "Putting yourself in evolution's shoes" is not scientific, and should not be interpreted as achieving any goal of a scientific process.

I would be interested in seeing what species some other person would come up with from the "axioms" (the part above the "SPOILERS FOLLOW" tag) -- it would give us a clue of how much leeway/arbitrariness is involved in the unit "cognitive processes" mentioned.



Environment overview

(Note: again, the internal explanations here are probably inconsistent or silly, and so are the "deductions" we make from them -- they are not true logical implications or the results of a true evolution algorithm, they are the result of our cognitive algorithms. The point is that these algorithms are still feed-forward, despite not being rational algorithms, they are also not rationalizations. I started afresh from new axioms when I ran into dead ends I didn't like, but that's fine, because I'm not actually faking the forward-feeding steps. Observe, for example, how not everything in the process turns out to be "plot-relevant" in the end.)

The planet has a silicon-rich, glassy crust rich in hydrogen gas and ionic salts. The high atmospheric pressure has resulted in most carbon deposits having been converted into diamond underground. The planet's atmosphere is rich in ammonia gas, which often condenses near the surface at night as a result of the enormous atmospheric pressure -- while there are no permanent lakes, lakes do typically form at the same spots at night. The planet's core has long cooled down, but its proximity to its star causes it to have a very high temperature, heated from the atmosphere down. The high temperature and diurnal variation damages and cracks the upper layers of the glassy crust, the granules becoming coarser as you get to deeper, more insulated layers.

(Assumptions about resulting life forms: they will be predominantly silicon-based, with the use of ammonia as a transport solvent -- they will also require energy, which can be extracted from dissolved substances in the ammonia. We will further assume that )


SPOILERS FOLLOW. I'd be very interested to see what kind of evolutionary path(s) someone who hasn't read the below would come up with from the above prompt -- for reasons mentioned in the introductory passage.


Early life-forms

Reactions begin to take place at night in the cracks in the upper crust, where the ephemeral ammonia lakes seep into. These cracks, or pods, are more "melty", porous, so the first organisms aren't really "cellular" or "multicellular" but rather "neighbour-based"/"distance-based"/"ship of Theseus"  (i.e. you are "basically the same cell" as your neighbour, who is "basically the same cell" as his neighbour, etc., but you never exchange any material with the cell a kilometre from you). 

These life forms (or "life-continuum"), however, quickly die at daytime. But even a life-continuum permits evolution -- you just have a genetic continuum, as mutations influence neighbours (in fact, this means that evolution takes place faster and more haphazardly). Life forms that can revive the next night (i.e. upon re-exposure to liquid ammonia) eventually form. 

These life forms subsequently evolve to be able to be able to apply pressure on specific regions of the surfaces of their pods, allowing them to crack them and travel inwards into the crust. This allows them to move slightly further and further from the hot atmosphere (bringing the ammonia with them), allowing them to survive for slightly longer and longer durations in the day. 

Individualized "cells" or "cracks" result. However, cells that crack further to pursue mixing with other cells are evolutionarily selected, due to the standard advantages of sexual reproduction (sexual reproduction is a built-in feature of life on this planet, and a result of the initial "continuum" nature of the life form). These "branchings" of cracks act as physical family trees. The cells are still connected, and genetic information propagates in all directions across the family tree -- until the older regions of the trees die out, separating the offspring.

Life forms added:
  • Cracks -- simple life forms with a continuum-based notion of identity.
  • Trees -- growing, branching versions of cracks, still with a continuum-based notion of family and reproduction.

Early neural systems

How does a tree reproduce faster? It needs to find -- or be found by -- other trees faster. 

Ideas for finding faster:
  • Vision
  • Hearing sound
  • Memory of maps and family history
  • Asking others
  • Fast motion 
Ideas for being found faster:
  • Displaying bright flashing lights
  • Producing sound 
  • Telling others

Thus develops a primitive neural system -- however, this requires a lot of energy. Thus an advantage forms for regions where some organisms evolve to be particularly adept at boring "ammonia wells" to supply energy and ammonia. These wells are then accessed by both root and borer organisms. Borers have reduced fibrage compared to roots, as they need to be able to efficiently pump fluid across their bodies, and cannot retrain whenever a part of their body is cut off or destroyed.

In the deepest level of abstraction, the cognitive processes attached to "finding" and "being found" are quite different -- the former requires optimization for observation, the second requires optimization for leaving signatures. We call these two "sexes" observers/seekers and markers. Organisms are generally hermaphroditic at this stage (because one area of a tree can move without affecting another), with each "node" (the part that cracks the crust) having a sex that determines its behaviour. Mating is more common between the sexes (because they are more likely to meet), but generally can take place between any two nodes, if they meet by happenstance. Nodes generally refrain from mating within the tree, for standard reasons.

Life forms added:
  • Roots -- hermaphroditic trees with basic cognition 
  • Borers -- hermaphroditic trees with basic cognition and powerful boring abilities

Biodiversity, food chains and systematic cognition

There are now multiple "pathways" for evolution -- e.g. What kind of modes of observing/marking does a creature adopt? Does homosexuality become less prevalent among marker nodes as a result of the low likelihood of two marker nodes meeting? What about among observer nodes? Speciation results, as different efficient equilibria cannot consistently mix to form efficient equilibria. 

As a result, organisms now eat other organisms, when mating is not an option. 

As well-borers have the ability to bore wells, they also have the ability to bore central bubbles to store their minds in, to enable greater and more complex computation. This allows more efficient mate-finding, and their branches become fewer in number, like tentacles (they don't need as much fibrage anymore). 

As a result of having central cognition loci, the bubbleheads can no longer be hermaphroditic. Nonetheless, the bubblehead sexes are not completely dimorphic, as with more advanced cognition necessary to attack and defend from other organisms, each individual organism needs observational and communicational availabilities.

Because genetic and cognitive material are mixed, child bubbleheads can be considered very small backups of each of their parents, i.e. a form of immortality, although memories do fade with time. 

Life forms added:
  • Bubbleheads -- central locus of cognition with emanating boring tentacles for mating, food-finding and other social interaction; not hermaphroditic, but also not fully dimorphic; reproduce via horcruxes
  • General speciation

Colonies

Parasites form, particularly on bubbleheads, as they are great sources of energy. Roots in particular are quite successful in their raids against animals because of their great "fibrage", less concentrated locations that allow them to launch more comprehensive raids on relatively defenseless bubbleheads.

Bubbleheads evolve to rely on their strengths: powerful boring abilities, superior cognition and communication, faster motion and presence of muscles. They band together, forming caves that act as forts, where they can move more freely to evade and kill slow-growing roots with little recurring energy consumption (despite higher upfront energy consumption). Non-bubblehead borers do not form caves, as they are still required to spread roots far and wide.

Note that a single colony may contain multiple bubblehead species. However, the following type of colony proves most efficient: a colony with three species of three specializations:
  • Excavator species, to bore out caves, tunnels, ammonia chambers.
  • Killer species, to gather food, guard excavators, attack other bubbleheads, defend against roots and other bubbleheads.
  • Thinker species, to maintain mental maps and instruct killers and excavators.
Some caves co-operate with each other. Some mate, and some form predator-prey relationships. And more complicated behaviour can arise, e.g. some bubbleheads of the rival cave are eaten, others are mated with, and others are accepted into the colony.

Note how versatile tentacles have gotten -- they hold, eat, mate and communicate.

Cables, therefore, being means of trade, are very important in a co-operative colony, which makes them quite valuable, especially with the lower fibrage. The high cognitive power of bubbleheads means that children take quite a while to form, so it is inefficient to waste tentacles for long periods of time while reproducing. So the tentacles deposit necessary information into a pod during reproduction and move on.

Life forms added:
  • Generic specialization
  • Cave-dwellers -- Bubbleheads that live in caves and reproduce via horcrux-eggs. Individual identities exist amidst a lot of communication/memory-sharing and co-operation. Of note are their cables, highly versatile and important.

Machines of efficiency

Specialization increases, as killers and excavators no longer had to think and thinkers no longer had to bore (the functions were shared, via cables -- the thinkers just completely control the killers and excavators). 

Life forms added:
  • Lichens -- Composite organisms. Bubbles of cognition controlling long cables at the ends of which are nodes that hold, eat, mate, communicate, bore, kill.
Cables, are, however, very expensive and difficult to replace. A particular type of bubblehead gains an advantage over the others. Instead of growing long tentacles for motion and interaction, its cables are short and tractable, and allow locomotion of the entire body. 

This allows them to easily dominate existing colonies by virtue of their speed -- and while they lose their ability to maintain constant communication with their peers, they make up for it by developing more efficient, abstracted form of thinking that doesn't require a constant data stream of information but allows them to make inferences and predictions.

They still do have peers, to help them gang on each other and on simpler cave-dwellers. But their thoughts are more individualized, which promotes competition.

Life forms added:
  • Roamers -- Smart, individualized lichens with short cables that hold, walk, eat, mate, communicate, bore, kill.

Psychics

A Roamer species -- whom we will call Psychs -- gain clearer, finer control over the thoughts they communicate -- however, it is still possible to break these barriers and invade the mind of another Psych. The arts of offense and defense in these areas force the Psychs to plot, make alliances to invade minds and uncover secrets (because someone may have found a great food source they're keeping for themselves -- or may be plotting to invade your mind -- or may be plotting to invade your mind after helping you take down the competition), giving rise to the birth of politics.

Life forms added:
  • Psychs -- a highly intelligent Roamer species. Also psychic: telepathy, rapid teaching, truth verification, mind-reading, memory manipulation, psychic brainwashing, hypnosis, mind-hijacking, mind-killing and mental defense are all serious arts among the Psychs.

Culture, religion and politics

The combination of "individualist" morals and "collectivist" ability to control and understand others creates an ingenious -- albeit violent and ruthless -- society. And as a result of the dangers, a very risk-averse one. 

Psych culture, aka "How do Psychs view the world?":
  • The danger of mind-invasion quickly leads to monogamy (as the same cables control psychic powers and mating); spousal trust highly valued.
  • Psych politics heavily rewards power. States grow, and the only thing standing in the way of complete world domination by a faction or individual is the use of physical (non-psychological) barriers against mind invasion, such as coverings on cables, or just territorial defenses and perimeters.
  • As a result, clan leaders often install themselves as gods in their members' minds, a memory to be passed on and worshipped (to prevent memory from fading). When a clan conquers another clan, the earlier gods are wiped out.
  • High degree of tribalism, as tribe members loyal to their leaders/religion.
  • Heretics who refuse to join a clan or religion are killed to avoid the threat of rebellion. But there are secret heretics who are brilliant at mind defense (everything in Psych society is ultimately a race between mind-offense and mind-defense).
  • Hatred, mistrust and skepticism are traits highly valued in Psych society, because you cannot trust someone who doesn't mistrust people (because they might be hypnotized to harm you).
  • It is important to take your cables very seriously! Psychs develop complex emotions and culture surrounding cables -- they can't just hate and run away from other people's cables, because they serve important purposes, but they need to be vigilant.
  • Psychs do not engage in smalltalk -- talking is very dangerous, and may get you killed -- it should be reserved for important things only!
  • Psychs are curious, of course. But they are also risk-calculating. They wonder about things, but are very reluctant to hold beliefs about the unknown. So they don't have mythology -- just mystery and intense passions towards mysteries they find important. Unfortunately, this limits their creativity, as it's very hard for them to just imagine hypotheses to test, or to imagine interpretations of concepts before seeing if it truly works.
  • Psychs are protective of their children, but view them with suspicion because they are easily brainwashed. As children eventually outgrow brainwashing that are made in their childhood, there is a coming-of-age ritual in which children are brainwashed by the state (generally via their parents).
  • Wars are never publicly-announced, as brainwashing of military by opposing military would make it very easy for 3rd party to swoop in (such wars would be almost entirely a numbers game) -- and treatises to ban mental warfare just don't work (it's possible to fool opponent to believe he isn't being fooled or cheated). Instead elaborate scheming to allow rapid takeovers that kill the king (to allow freeing, re-brainwashing of subjects). Heretics often created this way.

Economics, technology and progress

  • Initial advancements -- use of diamond tools for rapid boring, hunting and gathering replacing their biological lichen-era cable tips; ammonia canals for underground farming of roots, borers and bubbleheads. Writing is not invented, as better forms of communication (human books) exist.
  • Technological advancement driven mostly by heretic Psychs who usurp power. Heretics who refuse to re-brainwash their subjects are quickly killed. Progress is, despite the high intelligence of the Psychs, quite slow.
  • Instead of freedom and capitalism, what causes the Psych industrial revolution is a king who figures out a way to have his subjects be creative for him, rather than for themselves. This is hard -- it's a value alignment problem! Must make sure that whatever the subjects discover, it will never cause their goals and desires to change. 
  • A totalitarian world government results, run by a hereditary absolute monarchy (the king has advisors, of course, but they're concerned with his well-being). Heretics are rapidly and systematically eliminated through ritualized mind-invasions to ensure perfect loyalty. The king himself is regularly invaded to check for impostering.
  • A heretical king may eventually -- hopefully! -- find a way to create a better, more individualistic value alignment. You probably won't have a king screw up and result in a chaotic value alignment, because Psychs are risk-averse.

Overview of neural network architectures

(This article was initially in the form of several disparate articles, which I've compiled into one because they were rather short and trivial on their own, and didn't communicate significant insights.)

Below is a list of some basic "features" of human cognition that come to mind upon a cursory introspection:
  • Data type specific processing: the brain has specific mechanisms to handle visual and audio data, based on hardwired assumptions about how such data must look like.
  • Data streams: The brain does not download a chunk of data and process them in an egalitarian fashion -- we get a continuous stream of data, and learn continually from it. Our brain has a notion of time -- things we saw in the past affect what we see now, and what we see now affects our beliefs about what we saw in the past. In particular, we have memory.
  • Environment interaction: aka "the world is a game". Most of the feedback we get is not in terms of pre-prepared labels, but is feedback from the environment, from experimenting with the environment.
  • Generation: Our brain can come up with new things: artwork, ideas, thoughts, etc. 
  • Making connections: I have often emphasized the importance of transferring "insights" from one academic area to another, etc. (e.g. in mathematics, in engineering) -- but this also occurs at a much more basic level, such as sharing part of a classification algorithm for different scripts.
It is worth noting that these are "orthogonal" architectures -- we want to be able to process streams of images and audio, to respond to a stream of environment feedback, to generate images and streams of audio, language and robotic commands, to apply knowledge learned from previous data streams we processed, etc. In this spirit, I will often use terms like "standard neural network" to refer to a neural network "that doesn't possess the architecture currently being considered", rather than a specific model.



Data type specific processing: Convolutional and recurrent neural networks

When processing visual data, it seems that to simply flatten the image and feed it as a vector is a bit disappointing. I mean, it works -- but remember what I said about the Bayesian prior? The network should a priori understand what the "inherent structure" of a data type is, so that it is more inclined towards more likely models.

The need for convolutional neural networks arises from this need to "bias" our network in favour of learning features that are more likely to be useful -- doing things that a Bayesian prior would do in lieu of actually having some Bayesian inference. More specifically, we want to represent the following prior knowledge about the features we want to learn:
  • "Features are based on local interactions" -- you are likely to be interested in linear combinations of neighbouring points. Further, there is a hierarchial nature to this, in that we are further interested in the interactions between nearby features, etc.
  • "Isotropy" -- the network shouldn't be overfitted for centered characters, etc. Even if it is, the bulk of the network should not be overfitted, i.e. relevant features should be identified throughout the image, and any overfitting that occurs in the last few layers can then be corrected through transfer learning (this is analogous to how humans learn, as well).
The second point corresponds to the idea of parameter-sharing ("basically all" points should undergo the same processing), and the first tells us that the exact kind of processing done should be based on local interactions, i.e. the precise notion of a convolution. Furthermore, the "hierarchial nature" of image processing leads to the notion of either pooling or stride, which brings distant features closer together. 

(Source: https://subscription.packtpub.com/book/game_development/9781789138139/4/ch04lvl1sec31/convolutional-neural-networks)

(Short project ideas: Check the claims above:
  • Garble/permute the rows and columns of the images and see how that affects the training accuracy with convolutional networks (make sure you apply the same permutation to all images!) 
  • Train a CNN on MNIST digits then do transfer learning on randomly off-centered digits with the convolutional layers fixed and check that you are able to get a similar accuracy to before. 
(Link to Colab Notebook where I perform these tests))

(Quick note: There are two forms of pooling, max-pooling, which tests for presence of a feature, and mean-pooling, which tests for sustained presence of a feature. Max pooling may be better if you don't have padding, otherwise it doesn't generally matter which form of pooling you use, as the convolution is capable of spreading a required feature to nearby pixels of a layer.)

In general, the notion that the fundamental operation of the neural network (affine transformations with vectors, convolutions with images) should depend on the "type" of the input data, is an important one. 

Once one understands that convolutions are the "fundamental way" of dealing with images, we should simply write convolutional neural networks in the abstract as any standard neural network, i.e. thinking of these "blocks of convolution kernels, plus bias kernels" as the appropriate generalization of "weights, plus biases".

Convolutional neural network: each block in the network is a vector of images, i.e. an image with channels.
A very important data "format" to care about is that of a data stream. This is something crucial to any AI that has a concept of time. All the data received by humans is in the form of a data stream, so it is easy to see that processing such data is important when attempting to replicate human tasks.

The natural structure on a data stream is given -- much like the natural structure on an image is given by notions of closeness -- by the flow of time. 

Note that simply intuiting the "structure" of the data type doesn't uniquely tell us what the network architecture should be. In fact, the cases of the convolutional and recurrent neural networks aren't even totally analogous -- recurrent neural networks are actually necessary for data streams, because of the input size not being fixed. Nonetheless, principles like parameter sharing are somewhat analogous in each context. 

Recurrent neural network architecture

The "parameter sharing" that goes on here is that the the parameters of each horizontal slice of the network are the same. Of course, here this parameter-sharing is required by the lack of fixed input size. It's also worth noting how this isn't analogous to convolutional networks, importantly: the information about nearby cells is fed into a layer by neurons in the same layer (as opposed to convolutions, which would have diagonal arrows). Nonetheless, the idea that our processing of a data stream is in some sense "consistent" or "uniform" over time somewhat motivates our understanding of this architecture.

Wrapped version of recursive neural network depiction. Input is a data stream.

(In this case, the unwrapped depiction is a better mental model when training the network, as the graph is acyclic, so you can apply backpropagation normally.)

"Bidirectional recurrent network", for updating your knwoledge/memory based on new information. This network has two "memory canals", carrying information forward as well as backward. The wrapping suppresses it in the depiction, but the second memory canal feeds into itself in the opposite direction as the data stream.


Data streams: Recurrent neural networks and Turing-completeness

One way to understand recurrent neural networks -- as we did above -- is as the natural algorithm for processing the specific data format called "data stream".

Of course, this isn't really "all there is", The precise structure of the recurrent neural network seems somewhat arbitrary, and it isn't truly completely determined by saying "it's the natural way to process data streams". Some genuinely new structure is seen here, and we should ask for a corresponding universal approximation theorem for recurrent networks.

The question to ask is: what exactly is the new structure seen in RNNs? How, precisely, is it different from standard feedforward networks?

A standard feed-forward network seeks to simulate functions, right? So the "universal approximation theorem" says that a neural network can approximate any function, up to something. So what does a recurrent neural network seek to simulate? What is a "function on a data stream"?

To answer this, we should talk about what exactly a function in the sense of computer science is. A function cannot really "take in a data stream" in computer science. A data stream that has not yet been fully fixed/captured is not a valid input variable, it's not a valid input data type.

What we're trying to simulate isn't really a function -- it's a program. A function is an example of a program, but a more general computer program can actually look at a data stream as it's running and continually update its output based on it. And the analog of universal approximation for programs is Turing-completeness, which RNNs do possess (as proven in [Siegelmann & Sontag, 1992]).

(A short project you can do to test the claim that RNNs indeed simulate programs: check that an RNN tested for a certain length of data input works reasonably well on inputs of different sizes. Can you do this with a CNN?)

You might get the sense that the RNN architecture we've discussed doesn't really feel the same as the way we process streams of data. It seems too generic, like there are more specific tasks we always do in our minds while processing some stream of audio or video. Like we need a better prior, to tell the network the exact nature of what it means to "mix" past and present features.

With the human mind, we have a very specific notion of memory -- specific actions to add and remove things from our memory. This construct is of importance, while watching a movie, holding a conversation, or reading a sentence. It's not just recalled memory, that is stored somewhere and accessed through searching for keywords in the mind, but actively held short-term memory that is particularly relevant in this context. During any of these activities, the human mind will never adopt any other "mixing" mechanism between knowledge from different points in time, that doesn't go through memory.

This is the idea behind Long Short Term Memory (LSTM) networks. Every recurrent "layer" of an LSTM network actually involves the following three computations:

  1. The forget filter -- based on current input, the network drops of some elements in the "memory channel" (called the cell state in LSTM)
  2. The remember filter -- the network adds some features of the current input to memory at varying intensities. 
  3. The output filter -- the network allows some of the current memory to pass to output.

I call these "filters" and use language like "dropping elements" and "some features", but it is to be noted that these are really about differences in intensity, and involve multiplying by the output of trainable sigmoid layers that decide how much of your memory/features you want to allow to pass through.

Source: Understanding LSTM networks by Christopher Olah

As an exercise, figure out which parts of the network correspond to which of the filters I've mentioned.

(Short project idea: I'm actually not completely sure if I understand whether the bottom channel is needed, i.e. whether its contents must be transfered to the next iteration in the sequence. Experiment with this alternative on standard applications of LSTM and compare the performance.)



Environment interaction: Reinforcement learning

With standard neural networks, feedback is specifically provided in the form of predetermined data and labels that the network is required to predict.

Reinforcement learning can be seen as a generalization of this, where the feedback isn't prepared by something complicated like a human, but instead is the result of dynamic interaction with an environment (or game). The environment typically follows some laws (i.e. the laws of physics, or the laws of a game -- for robotics and game bots respectively), and this automatically generates massive amounts of data, and computes an action's consequences, which act as a generalized notion of data "labels".

A problem is immediately clear with this method: how do you differentiate an environmental response against your network's parameters? This is not a minor technical problem -- differentiation fundamentally requires that you know what happens if you change your parameters a bit.

E.g. suppose you have a network that takes in the current state of a game and outputs a real number between 0 and 1, representing the probability that it tells the agent to "jump". Then you can differentiate each decision with respect to the parameters; however, you cannot differentiate the outcome (win/loss) with respect to the decisions.

The solution to this problem comes from recalling that we are trying to maximize an expected score, so we should be doing some sampling. More formally: we should play around with expectations.

Let $Y$ be the random variable representing the agent's decision (i.e. "jump" or "not") for some given input $x$, with probability function $P(Y|x, \theta)$ where $\theta$ is the network's trainable parameters. Then where $L(x, Y)$ is the environmental loss function, we are interested in minimizing $E(L)$.

\[\begin{align}
  {\nabla _\theta }E\left[ {L(Y)} \right] &= {\nabla _\theta }\sum\limits_Y^{} {P(Y)L(Y)}  \\
   &= \sum\limits_Y^{} {{\nabla _\theta }\left[ {P(Y)} \right]L(Y)}  \\
   &= \sum\limits_Y^{} {P(Y)\frac{{{\nabla _\theta }P(Y)}}{{P(Y)}}L(Y)}  \\
   &= \sum\limits_Y^{} {P(Y){\nabla _\theta }\left[ {\log P(Y)} \right]L(Y)}  \\
   &= {E_Y}\left[ {{\nabla _\theta }\left( {\log P(Y)} \right)L(Y)} \right] \\
\end{align} \]
So the solution is as follows: sample a large number of gameplays. Now pretend that each decision contributed directly to the victory and optimize them -- encourage all the moves in winning gameplays and all the moves in losing gameplays, i.e. update the parameters by the average value of $ {{\nabla _\theta }\left( {\log P(Y)} \right)L(Y)} $ across the sample.

So despite the bizarreness of pretending that every move in a winning play was correct and every move in a losing play was wrong, doing this for a large sample makes incorrect learned features cancel out -- a good move is expected to produce better results when all other moves are averaged out, and a bad move is expected to produce worse results when all other moves are averaged out.

This strategy is known as policy gradients, and is a general technique to deal with non-differentiable feedback.

Image source: Andrej Karpathy

(You may notice that this is incredibly inefficient. Indeed, this article only covers the most basic and superficial elements of cognition -- the human brain is capable of reasoning, and of producing a highly abstracted model of the game in its mind, and of transferring intuition from elsewhere onto the game.)



Generation: Generative (matching and adverserial) neural networks

Equipped with the ability to process data, the obvious next step is to get an AI to produce things -- to get an AI to be creative. To come up with art, compositions, original thoughts and ideas. We'll now describe the most elementary of such neural networks, which we will call Generative Neural Networks, while more complicated ideas would exploit some sort of transfer learning.

It's not at all absurd to expect it to be possible for a neural network to generate images of horses that don't look like any horse it's actually seen -- because humans can do that! If you imagine a horse, it's probably not a horse whose image you've seen before, but it nonetheless possesses the features you've identified as common between horses.

The idea behind a generative neural network can be motivated from the following two statistical notions:

  • The inverse transform method of generating random variables.

Content generated by a mind can be considered to be a random variable in some fancy space. E.g. if we want to get our neural network to produce (28, 28) digit characters, we're training it into a random variable on the space of (28, 28) images whose support is the images we identify as valid digit characters.

The way that computers typically sample random variables is through the "inverse transform method", which is to start with a uniform random sample and apply $F^{-1}$ to your sample where $F$ is the CDF of the random variable $X$ you want to sample. Your result will be a sample of $X$.

Quick explanation of inverse transform method: under the uniform distribution, the probability of getting a value under $u$ is $u$, which under the CDF of $X$ is the probability of getting a value under $F^{-1}(u)$. So you map $u\mapsto F^{-1}(u)$.
So we once again need a function approximator -- to approximate $F^{-1}$.

The inputs to the neural network are randomly generated, typically uniformly

  • A nonlinear generalization of principal component analysis

Think, e.g. of eigenfaces. If you've ever tried to use eigenfaces to generate realistic faces, you'll notice that your results are just terrible. Much of the information in faces is not so linear and nice -- there's no reason to expect it to be. Illumination and angle are pretty much the only properties that can be expected to vary linearly.

I.e. suppose you have some data that varies as follows:


Then a PCA might give you the pink line as your first principal component, but sampling from the pink line gives you a lot of unphysical outputs, those are the areas where your pink line doesn't intersect the data.

But using PCA to generate samples from a distribution can be understood as taking some random inputs, corresponding to the values of each principal component you want to use, and feeding them through a function, the principal component change-of-basis matrix. 

But more generally, replacing this function with something nonlinear allows us to deal with nonlinear models.

OK,so it's clear to us that we need a neural network -- a "generative neural network" -- to construct the inverse CDF. How would one train this network?

Given an initial random guess for the network parameters, what we have is some guessed distribution for "images of horses". And what we really want to do is perturb these parameters to match the distribution of our data.

Well, such an approach is certainly possible -- one could measure some notion of distance from our generated sample distribution and the real distribution and backpropagate this error with each iteration.

Note that we don't actually know the distribution of our data (the red one), so we can't really use something like "the probability of observing this sample given our distribution" as our loss function. Anyway, there are measures of the distance between two distributions that we could use for our error function, such as the maximum mean discrepancy approach, and methods involving moments.

(This approach, generally, is called a Generative Matching Network.)

But being arbitrary is generally disappointing in machine learning, and it's worth asking if there's a way to get the network to learn to discriminate between distributions.

Here's an idea: we could just subjectively tell that the outlier point did not belong to the distribution. We used our human brains. How about rather than defining a discrimination function, we trained a neural network to tell if a given data point could belong to a distribution? Then this neural network would train our generative neural network, and vice versa.

And this makes sense, right? When we learn to draw an object, we're also simultaneously learning to identify one.

And one could imagine showing off a generative network's results and having people guess if they're real or not (alongside actual real images of course) -- and based on whether they thought it was real, we could use it to train the network. These "people" are precisely what a discriminator network is.

In other words, we have two networks: the generator network, which generates random horse faces from random uniform variables, and the discriminator network, which takes the output of the generator network and some actual images, and figures out if the result is real or not.

If the classification is incorrect, the discriminator network is punished, while if it is correct, the generator network is punished.

This is known as a Generative Adverserial Network.

(add: GANN improvements [1], deepfakes, adverserial inputs)



Making connections: transfer learning

Something of crucial importance to human thought is the ability to make connections between ideas, transfering ideas and results from one area to another, either directly or by "abstracting out" the analogies.

There are many "levels" on which this occurs -- examples follow:
  • We don't need 60,000 instances to learn a script -- We can do with 1. It's shocking that we can learn distributional information from a single data point, and suggests that we already have an extraordinarily good prior. And we do -- our prior is continually updated from experience, of course, so this just means we're applying existing knowledge. We already know what features of the script are likely to be important and what can be safely discarded (hint: that accidental wiggle in the straight-ish line is probably noise) away. 
  • Prerequisites exist. So we depend on "applying" existing knowledge in some sense to learn new things. There's a reason why babies typically don't learn algebraic geometry. 
  • Making abstract connections between ideas -- Like recognizing that counting apples and counting sticks is the same task (and frivolous details about what an apple is or what a stick is can be lost for our purposes), recognizing that the algebra of kinematic quantities is the same to some extent as the algebra of quantum states (they both form vector spaces) (see Abstraction in mathematics, Abstraction in engineering), or recognizing the analogy between linear transformations of vector spaces and continuous functions of topological spaces (they're both categories).
There are probably multiple different algorithms in the brain that involve analogy and abstraction, and at least some of them (such as the first example above) have to do with learning.

This is the principle behind transfer learning, the idea that one may use the results of existing encoders for processing in other domains (and perhaps the precise encoder used may be fine-tuned for these other purposes). 

Examples of transfer learning architectures:


Progressive neural networks

This may be a bit of a disappointing solution to the question "How do brains think abstractly?" Our answer seems to depend on our active, hard-coded choice of what layers and encodings to preserve for our next task. Surely our brains do this somewhat "automatically" -- we don't actively tell our brain "Look, you've seen lines, remember? Well, there's lines here."

Well, maybe sometimes we do. It seems there may be some element of a conscious thought (whatever that means on an algorithmic level) in making analogies when it comes to complex intellectual tasks. But we certainly do not undergo any conscious thought when it comes to something like character recognition.

An architecture that may be analogous to our brain's cognition is that of a progressive neural network -- here, layers are algorithmically transferred from a previous task to another, with transferred parameters frozen. The picture below conveniently helps us make the analogy to "lateral thinking".

Source: arXiv:1606.04671

This architecture is capable of (and is about) choosing the sources of information that are most effective for the required task, but what suggests to me that this isn't what our brain does exactly is the parameter wastefulness and increasing level of model complexity (trainable parameters), which doesn't seem to be analogous with our method of learning, which seems to be more or less symmetric between learned ideas.

Multitask learning

A far simpler transfer learning approach that addresses the above concern about symmetry is multi-task learning:
Hard parameter sharing. Source: Sebastian Ruder
Soft parameter sharing. Source: Sebastian Ruder

This does intuitively seem to be the brain's transfer learning algorithm -- even if one of the tasks has previously been trained, the brain seems to be able to retrain the network as needed. And I have often observed the benefits of learning two abstractly related areas of mathematics together, e.g. Hilbert spaces and quantum mechanics, linear algebra and special relativity, topology and probability theory.