Learning — Classical Conditioning, Operant Conditioning, Cognitive Learning, Observational Learning, Rewards and Intrinsic Motivation
How does experience change behaviour? Pavlov's dogs and the Garcia effect, Skinner's schedules of reinforcement and the limits of punishment, Tolman's cognitive maps, Bandura's Bobo doll experiment and self-efficacy, and the overjustification effect in which rewards dampen interest, at university level.
Last reviewed 2026-09-29
What is learning?
In psychology, learning is a relatively lasting change in behaviour or knowledge that comes from experience. Temporary changes, such as those from fatigue or drugs, and changes unrelated to experience, such as maturation, don’t count.
In the early 20th century, Watson (1913) launched behaviourism, arguing that psychology should study observable behaviour rather than invisible things like consciousness. Within this movement, research on learning became psychology’s first sophisticated experimental tradition.
Classical conditioning: learning signals
Pavlov’s discovery
While studying digestion, the Russian physiologist Pavlov (1927) discovered that dogs salivated at the mere sound of the footsteps of the person who fed them, before seeing any food.
| Term | Meaning | Pavlov’s example |
|---|---|---|
| Unconditioned stimulus (US) | A stimulus that triggers a response without learning | Food |
| Unconditioned response (UR) | The innate response to the US | Salivation |
| Conditioned stimulus (CS) | An originally neutral stimulus paired with the US | A bell |
| Conditioned response (CR) | The learned response to the CS alone | Salivating at the bell |
A learned response weakens when the CS is presented repeatedly on its own (extinction) and briefly returns when the CS is presented again after a break (spontaneous recovery). Generalisation, responding to similar stimuli, and discrimination, not responding to different ones, occur as well.
Prediction matters more than pairing
Rescorla (1968) showed that even when a CS and US occur together many times, conditioning is weak if the US also often occurs without the CS. What animals learn is not a simple pairing but the information “does this signal predict that event?” The Rescorla–Wagner model (1972) formalised this as learning that reduces prediction error.
Not all connections are learned equally
Garcia and Koelling’s (1966) rats linked a taste with subsequent nausea after a single experience, but struggled to link lights and sounds with nausea; conversely, they readily linked lights and sounds with electric shock. Organisms are prepared to learn connections that matter for survival more easily — which is why a food that once made you sick can put you off for a long time.
Learning and extinguishing fear
Watson and Rayner (1920) reported that by making a loud noise while showing 11-month-old “Little Albert” a white rat, they made him fear the rat and similar furry objects. By today’s standards the experiment is ethically unacceptable and its methods were sloppy, but the idea that fear can be learned has been supported by later research. Exposure therapy for anxiety disorders builds on extinction learning: repeatedly experiencing a feared stimulus in a safe situation.
Operant conditioning: learning from consequences
The law of effect and the Skinner box
Thorndike (1898) watched cats trapped in puzzle boxes stumble on the latch, escape, and then get out faster and faster — the law of effect: behaviours followed by satisfying outcomes are strengthened. Skinner (1938) systematised this into operant conditioning, the study of how behaviour is shaped by its consequences.
| Something is added | Something is removed | |
|---|---|---|
| Behaviour increases | Positive reinforcement (praise, pocket money) | Negative reinforcement (buckling up stops the warning beep) |
| Behaviour decreases | Positive punishment (a scolding) | Negative punishment (losing game time) |
“Negative reinforcement” is easily confused with punishment, but it means increasing a behaviour by removing something unpleasant.
Schedules of reinforcement
| Schedule | Rule | Behaviour pattern | Example |
|---|---|---|---|
| Fixed ratio | Reward after a set number of responses | Fast responding, brief pause after reward | A free coffee after 10 stamps |
| Variable ratio | Reward after an unpredictable average number | Very fast and steady; most resistant to extinction | Slot machines, loot boxes |
| Fixed interval | Reward for the first response after a set time | Responses bunch up near the deadline | Monthly pay, scheduled exams |
| Variable interval | Reward after unpredictable average times | Slow but steady | Checking for messages |
This table shows why variable-ratio schedules are used in gambling and randomised in-game rewards: when you can’t tell when the next reward will come, it’s hard to stop.
Shaping and the limits of punishment
Through shaping — reinforcing successive approximations of a desired behaviour — animals can be taught complex behaviours. Punishment, by contrast, may stop a behaviour immediately but doesn’t teach what to do instead. A meta-analysis of research on corporal punishment of children found it associated with immediate compliance, but also consistently with negative outcomes such as increased aggression and poorer parent–child relationships (Gershoff, 2002).
Cognitive learning: maps in the head
Some phenomena were hard for behaviourism to explain. In Tolman and Honzik’s (1930) experiment, rats that had wandered a maze without reward, once finally given food, immediately found their way as quickly as rats rewarded from the start. They had been learning a cognitive map of the maze even without reward (latent learning; Tolman, 1948).
Köhler’s (1925) chimpanzees, faced with bananas out of reach, would pause for a while and then suddenly stack boxes or join sticks to get the fruit — not trial and error but insight learning, grasping the structure of the problem all at once.
Observational learning: learning by watching others
In Bandura, Ross and Ross’s (1961) Bobo doll experiment, children who had watched an adult hit and shout at an inflatable doll later imitated the behaviour far more. The study showed that people can learn simply by watching others, without being reinforced themselves.
Bandura (1977) explained observational learning as four processes.
- Attention: you have to notice the model’s behaviour.
- Retention: you have to store what you saw in memory.
- Reproduction: you need the ability to perform the behaviour.
- Motivation: you need a reason to do it, such as having seen the model rewarded.
Bandura also held that self-efficacy — the belief “I can do this” — governs whether people start and persist at a behaviour. Self-efficacy grows from mastery experiences, seeing similar people succeed, encouragement, and how one interprets bodily states.
When rewards undermine interest
Lepper, Greene and Nisbett (1973) promised preschoolers who liked drawing a certificate for drawing, and gave it to them. Weeks later, during free time, these children drew less than those who hadn’t been rewarded. When an external reward is attached to something done for enjoyment, people come to think “I did it for the prize,” and interest drops — the overjustification effect.
A meta-analysis of 128 studies likewise found that expected, tangible rewards contingent on doing a task reduced intrinsic motivation, whereas unexpected rewards and specific praise had little such effect or even helped (Deci, Koestner & Ryan, 1999). Reinforcement is a powerful tool, but it should be used with care for things people already enjoy. Other aspects of motivation continue in the next chapter.
Check questions
Question 1. Someone who felt pain every time they heard the dentist’s drill now tenses up at the sound of the drill alone. What are the US, UR, CS and CR here?
Answer. The US is the pain of treatment, the UR is tension and fear in response to pain, the CS is the sound of the drill, and the CR is the tension produced by the drill sound alone.
Question 2. A child is let off washing the dishes on days they tidy their room. What kind of reinforcement or punishment is this?
Answer. It aims to increase a desired behaviour (tidying) by removing something unpleasant (washing up), so it is negative reinforcement.
Question 3. An app gives users random rewards when they log in. Using schedules of reinforcement, explain why this keeps users hooked.
Answer. It resembles a variable-ratio schedule, where you can’t predict after how many tries a reward will come. This schedule produces high response rates and is the most resistant to extinction, so people don’t easily stop even when rewards dry up for a while.
References
- Watson, J. B. (1913). Psychology as the behaviorist views it. Psychological Review, 20(2), 158–177.
- Pavlov, I. P. (1927). Conditioned Reflexes (G. V. Anrep, Trans.). Oxford University Press.
- Rescorla, R. A. (1968). Probability of shock in the presence and absence of CS in fear conditioning. Journal of Comparative and Physiological Psychology, 66(1), 1–5.
- Rescorla, R. A., & Wagner, A. R. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement. In A. H. Black & W. F. Prokasy (Eds.), Classical Conditioning II (pp. 64–99). Appleton-Century-Crofts.
- Garcia, J., & Koelling, R. A. (1966). Relation of cue to consequence in avoidance learning. Psychonomic Science, 4(1), 123–124.
- Watson, J. B., & Rayner, R. (1920). Conditioned emotional reactions. Journal of Experimental Psychology, 3(1), 1–14.
- Thorndike, E. L. (1898). Animal intelligence: An experimental study of the associative processes in animals. Psychological Review Monograph Supplements, 2(4), i–109.
- Skinner, B. F. (1938). The Behavior of Organisms. Appleton-Century.
- Gershoff, E. T. (2002). Corporal punishment by parents and associated child behaviors and experiences: A meta-analytic and theoretical review. Psychological Bulletin, 128(4), 539–579.
- Tolman, E. C., & Honzik, C. H. (1930). Introduction and removal of reward, and maze performance in rats. University of California Publications in Psychology, 4, 257–275.
- Tolman, E. C. (1948). Cognitive maps in rats and men. Psychological Review, 55(4), 189–208.
- Köhler, W. (1925). The Mentality of Apes. Harcourt, Brace.
- Bandura, A., Ross, D., & Ross, S. A. (1961). Transmission of aggression through imitation of aggressive models. Journal of Abnormal and Social Psychology, 63(3), 575–582.
- Bandura, A. (1977). Social Learning Theory. Prentice-Hall.
- Bandura, A. (1977). Self-efficacy: Toward a unifying theory of behavioral change. Psychological Review, 84(2), 191–215.
- Lepper, M. R., Greene, D., & Nisbett, R. E. (1973). Undermining children’s intrinsic interest with extrinsic reward: A test of the “overjustification” hypothesis. Journal of Personality and Social Psychology, 28(1), 129–137.
- Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627–668.