In mammals, mesencephalic dopamine neurons participate in a number of important cognitive and physiological functions including motivational processes (Wise, 1982; Fibiger and Phillips, 1986; Koob and Bloom, 1988), reward processing (Wise, 1982), working memory (Sawaguchi and Goldman-Rakic, 1991), and conditioned behavior (Schultz, 1992). It is also well known that extreme motor deficits correlate with the loss of midbrain dopamine neurons; however, activity in the substantia nigra and surrounding dopamine nuclei, i.e., areas A8, A9, A10, does not show any systematic relationship with the metrics of various kinds of movements (Delong et al., 1983; Freeman and Bunney, 1987).
Physiological recordings from alert monkeys have shown that midbrain dopamine neurons respond to food and fluid rewards, novel stimuli, conditioned stimuli, and stimuli eliciting behavioral reaction, e.g., eye or arm movements to a target (Romo and Schultz, 1990; Schultz and Romo, 1990; Ljungberg et al., 1992; Schultz, 1992; Schultz et al., 1993). Among a number of findings, these workers have shown that transient responses in these dopamine neurons transfer among significant stimuli during learning. For example, in a naive monkey learning a behavioral task, a significant fraction of these dopamine neurons increase their firing rate to unexpected reward delivery (food or fluid). In these tasks, some sensory stimulus (e.g., a light or sound) is activated so that it consistently predicts the delivery of the reward. After the task has been learned, few cells respond to the delivery of reward
The capacity of these dopamine neurons to represent such predictive temporal relationships and their well described role in reward processing suggest that mesolimbic and mesocortical dopamine projections may carry information related to expectations of future rewarding events. In this paper, we present a brief summary of the physiological data and a theory showing how dopamine neuron output could, in part, deliver information about predictions to their targets in two distinct contexts: (1) during learning, and (2) during ongoing behavioral choice. Under this theory, stimulus-stimulus learning and stimulus-reward learning become different aspects of the same general learning principle. The theory also suggests how information about future events can be represented in ways more subtle than tonic firing during delay periods.
References
Montague, P. R., P. Dayan, and T. J. Sejnowski. 1996. “A framework for mesencephalic dopamine systems based on predictive Hebbian learning.” The Journal of neuroscience 16 (5): 1936–47.

