Sequential decision reinforcement learning training method and system
By designing a two-step sequential decision-making experiment paradigm and building a pigeon sequential decision-making learning and training system, the problem of unclear dynamic representation of behavioral value and successor state in the decision-making process of bird brains is solved, and precise simulation and research on dynamic neural representations of the brain is achieved.
Patent Information
- Application Number
- CN202510145426.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the dynamic neural representation rules of behavioral value and successor state of bird brains in decision-making process are not clear, especially the mechanism of how to dynamically adjust biological learning strategies in sequential decision-making is not clear.
A two-step sequential decision-making experiment paradigm was designed, and a microcontroller with the STM32 microcontroller as the core was built to build a pigeon sequential decision-making learning and training system including a behavioral training box, experimental parameter setting and monitoring module, and a multi-channel neural data detection and recording system. Uncertainty was introduced through state transfer and reward units to simulate the dynamic neural representation of the brain in the decision-making process.
The study of the dynamic neural representation law of the brain in the decision-making process can more accurately simulate the dynamic neural decision-making process of pigeons in sequential decision-making, revealing the dynamic adjustment mechanism of the brain in behavioral value and successor state.
Smart Images

Figure CN120354964A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sequential decision-making reinforcement learning training for birds, and specifically to a sequential decision-making reinforcement learning training method and system thereof. Background Art
[0002] In sequential decision-making, the formation of action value is affected by subsequent states. However, it is still unclear how the brain dynamically represents action value and subsequent states during this process, how subsequent state information promotes the formation of action value, and how the change in the relationship between the two affects the dynamic adjustment of biological learning strategies. Research shows that birds such as pigeons and crows have good behavioral learning abilities and are classic model animals for operant conditioning behavioral experiments.
[0003] Therefore, in order to study and reveal the dynamic neural representation laws of action value and subsequent states in the decision-making process of the brain, a two-step sequential decision-making experimental paradigm was designed using pigeons with good cognitive decision-making abilities as model animals, and a pigeon sequential decision-making learning training system was designed and built according to experimental requirements. Summary of the Invention
[0004] The purpose of the present invention is to provide a sequential decision-making reinforcement learning training method and system thereof to study and reveal the technical problem of the dynamic neural representation laws of action value and subsequent states in the decision-making process of the brain.
[0005] The above invention purpose of the present invention is achieved through the following technical solutions:
[0006] A sequential decision-making reinforcement learning training method and system thereof includes a sequential decision-making reinforcement learning training method and a sequential decision-making experimental training system for implementing this method. The sequential decision-making experimental training system includes a behavior training box, an experimental parameter setting and experimental state monitoring module, and a multi-channel neural data detection and recording system. The communication module among the behavior training box, the experimental parameter setting and experimental state monitoring module, and the multi-channel neural data detection and recording system includes a microcontroller with an STM32 single-chip microcomputer as the core and its peripheral circuits.
[0007] Preferably, the state transition unit and the reward unit respectively correspond to the first step (Step1) and the second step (Step2) of the task.
[0008] Preferably, the sequential decision-making experimental method includes the following experimental process: Step 1: At Step1, two option signs S1 + and S1 - will be presented simultaneously in front of the pigeon. This state is called the initial state (State1, S1) and is marked with red and green squares respectively;
[0009] Step 2: Train the pigeon to choose one of the options within 2 s of option presentation and confirm its choice by pecking the key below the corresponding option; once the pigeon confirms its choice, it will transfer to the second state (State2, S2) with a certain probability T s s′ which is also called the successor state;
[0010] Step 3: The successor state marker S2 + or S2 - will be presented on the screen in front of the pigeon, presented as a blue triangle or a blue circle respectively, and each state marker carries a reward with a different probability; Step 4: The pigeon needs to peck the key below the marker within 2 s of the stimulus presentation to confirm that it notices the current state; finally, the system will present the food reward R with a certain probability P r for 3 s. After the reward ends, it enters a 5-s inter-trial interval (ITI).
[0011] Preferably, in the design of the sequential decision-making experimental paradigm, whether in Step 1 or Step 2, if the pigeon does not make a choice or peck the key within the specified time, the experiment will directly enter the ITI.
[0012] Preferably, in the design of the sequential decision-making experimental paradigm, in order to enable the pigeon to fully learn the state transition relationship and reward information, a certain degree of uncertainty is introduced in both the state transition unit and the reward unit during the experiment design. The state transition probability T s s′ from Step 1 to the second step Step 2 r and the reward probability P
[0013]
[0014] are set as shown in the following formula: where + represents choosing S1 in Step 1 and - represents choosing S1 in Step 1 ; + represents making a pecking action in state S2 in Step 2 and - represents making a pecking action in state S2 in Step 2
[0015] ; R represents the reward information. + and S1 - in Step 1 are respectively
[0016] The red in Step1 (S1 + ) is the preferred option, and the green (S1 - ) is the inferior option. The triangle in Step2 (S2 + ) is the preferred successor state, and the circle (S2 - ) is the inferior successor state. For a state transition with a probability of 0.8, it is called a common state transition, such as selecting S1 + and transitioning to S2 + or selecting S1 - and transitioning to S2 - . Conversely, for a state transition with a probability of 0.2, it is called an uncommon state transition, such as selecting S1 + and transitioning to S2 - or selecting S1 - and transitioning to S2 + . When a reward is obtained, it is marked with R (Reward), and when no reward is obtained, it is marked with NR (NoReward).
[0017] Preferably, the experimental parameter setting and experimental state monitoring module is a man-machine interaction interface used by experimenters to set basic experimental parameters and monitor the current training state of pigeons. This interface communicates bidirectionally with the microcontroller. After the experimenter completes the parameter setting, it is transmitted to the microcontroller, and then the microcontroller transmits the pigeon's behavior data to the experimental parameter setting and experimental state monitoring module. This module has functions such as starting and ending the experiment, switching the food box, and downloading and saving experimental parameters. It can also set information such as the total number of trials, state transition probability, reward probability, stimulus presentation duration, and ITI duration of the experiment. In addition, it can also monitor the pigeon's experimental state in real time and record behavior data, including the current trial number, number of awards, and the pigeon's action selection, successor state, pecking key response time, and whether it has won an award in each trial. Experimenters can judge the pigeon's current learning state based on the above information.
[0018] Preferably, the behavior training box is a cuboid space for the sequential decision-making learning training of pigeons, with a length, width, and height of 55 cm × 50 cm × 60 cm respectively. It is made of acrylic plates with a thickness of 0.6 cm, 5 sides are closed, and the other side is an openable door. There are ventilation holes on the top layer, and a layer of stainless steel wire mesh is provided at the bottom for pigeons to stand on. Inside the training box, there are a status display device, an action detection device, a food reward device, and a video monitoring device. Among them, the status display device is located on the inner wall of the box and is a screen with a length × width of 8 cm × 6 cm; the action detection device is located directly below the position where the stimulus marker is presented and consists of three buttons; the food reward device consists of a stepper motor and a box with a length × width of 4 cm × 3 cm, and is located at the bottom of the box below the buttons; the video monitoring device is installed on the wall directly opposite the status display device and can observe the status of the pigeons. The behavior training box and the microcontroller have two-way communication. The microcontroller transmits experimental parameter information, such as option presentation, presentation duration, and reward information, into the experimental box, and the action detection device in the behavior training box transmits the behavior selection situation of the pigeons to the microcontroller.
[0019] In summary, the beneficial technical effects of the present invention are as follows: The present invention uses pigeons with good cognitive decision-making abilities as model animals, designs a two-step sequential decision-making experimental paradigm, and designs and builds a sequential decision-making learning training system for pigeons according to experimental requirements, which can more accurately simulate the dynamic neural decision-making process of the brain during the decision-making process. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a schematic diagram of the sequential decision-making experimental paradigm and experimental scenario of the present invention;
[0021] Figure 2 is a schematic diagram of the composition of the sequential decision-making learning training system of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0022] The present invention will be further described in detail below with reference to the accompanying drawings.
[0023] Referring to Figure 1 and Figure 2 , a sequential decision-making reinforcement learning training method and system disclosed by the present invention include the design of a sequential decision-making experimental paradigm and the construction of a sequential decision-making experimental training system.
[0024] The sequential decision-making task needs to include two basic units: a state transition unit and a reward unit, corresponding to the first step (Step1) and the second step (Step2) of the task respectively. In addition, when designing the paradigm, the behavioral habits of the model animals need to be considered simultaneously. Therefore, when designing the sequential decision-making experimental paradigm of the present invention, the classic and general two-step sequential decision-making task is referred to, and a sequential decision-making experimental paradigm suitable for pigeons is designed in combination with the characteristics of pigeons. Refer toFigure 1 , in the figure, (a) is the sequential decision-making experimental paradigm for pigeons, (b) is the schematic diagram of the experimental scenario, and (c) is the division of time windows in a complete trial.
[0025] Refer to Figure 1 , the specific experimental procedure is described as follows: Step 1: Two option signs S1 + and S1 - will be presented simultaneously in front of the pigeon. This state is called the initial state (State1, S1), and is marked with a red and a green square respectively. The pigeon is trained to choose one of them within 2 s when the options are presented and confirm its choice by pecking the corresponding button below the option as shown in (b); once the pigeon confirms its choice, it will transfer to the second state (State2, S2) with a certain probability T s s′ , which is also called the successor state.
[0026] Meanwhile, the successor state marker S2 + or S2 - will be presented on the screen in front of the pigeon, presented as a blue triangle or a blue circle respectively. Each state marker carries a reward with a different probability. The pigeon needs to peck the button below the marker within 2 s when the stimulus is presented to confirm that it has noticed the current state; finally, the system will present the food reward R with a certain probability P r . The reward time is 3 s. After the reward ends, it enters the 5-s inter-trial interval (ITI) state. It should be noted that regardless of Step 1 or Step 2, if the pigeon does not make a choice or peck the key within the specified time, the experiment will directly enter the ITI state.
[0027] To enable the pigeon to fully learn the state transition relationship and reward information, certain uncertainties are introduced into both the state transition unit and the reward unit in the experimental design. The state transition probability T s s ' from Step 1 to the second step Step 2 and the reward probability P r are set as shown in the following formula:
[0028]
[0029] In the formula, represents choosing S1 in Step 1 + , represents choosing S1 in Step 1 - , represents making a pecking action in state S2 in Step 2 + , represents state S2 in Step 2 -The pecking action is made under the following conditions, and R represents reward information.
[0030] According to the above preset probability, S1 is selected in Step1 + and S1 - The theoretical expected values (expected value, EV) of
[0031]
[0032] Define the red color (S1 + ) in Step1 of this paradigm as the dominant option, the green color (S1 - ) as the inferior option, the triangle (S2 + ) in Step2 as the dominant successor state, and the circle (S2 - ) as the inferior successor state. For state transitions with a probability of 0.8, they are called common state transitions, such as the case of selecting S1 + transferring to S2 + or selecting S1 - transferring to S2 - Conversely, for state transitions with a probability of 0.2, they are called uncommon state transitions, such as the case of selecting S1 + transferring to S2 - or selecting S1 - transferring to S2 + When obtaining a reward, it is marked with R (Reward), and when not obtaining a reward, it is marked with NR (NoReward).
[0033] Referring to Figure 2 , first of all, the experimental parameter setting and experimental state monitoring module is a human-computer interaction interface for experimenters to set basic experimental parameters and monitor the current training state of pigeons. This interface communicates bidirectionally with the microcontroller. After the experimenter finishes setting the parameters, they are transmitted to the microcontroller, and then the microcontroller transmits the pigeon's behavior data to the experimental parameter setting and experimental state monitoring module. This module has functions such as starting and ending the experiment, switching the food box, and downloading and saving experimental parameters. It can also set information such as the total number of trials in the experiment, state transition probability, reward probability, stimulus presentation duration, and ITI duration. In addition, it can also monitor the experimental state of pigeons in real time and record behavior data, including the current number of trials, the number of awards, as well as information such as the action selection of pigeons in each trial, successor state, pecking response time, and whether they won an award. Experimenters can judge the current learning state of pigeons based on the above information.
[0034] The behavior training box is a cuboid space used for the sequential decision-making learning training of pigeons, with a length, width, and height of 55 cm × 50 cm × 60 cm respectively. It is made of acrylic boards with a thickness of 0.6 cm. Five sides are closed, and the other side is a door that can be opened and closed. There are ventilation holes on the top layer, and a layer of stainless steel wire mesh is provided at the bottom for pigeons to stand on. Inside the training box, there are a status display device, an action detection device, a food reward device, and a video monitoring device. Among them, the status display device is located on the inner wall of the box and is a screen with a length × width of 8 cm × 6 cm; the action detection device is located directly below the position where the stimulus marker is presented and consists of three buttons; the food reward device consists of a stepper motor and a box with a length × width of 4 cm × 3 cm and is located at the bottom of the box below the buttons; the video monitoring device is installed on the wall directly opposite the status display device and can observe the status of the pigeons. The behavior training box and the microcontroller communicate bidirectionally. The microcontroller transmits experimental parameter information into the experimental box, such as option presentation, presentation duration, and reward information, etc. The action detection device in the behavior training box transmits the behavior selection situation of the pigeons to the microcontroller.
[0035] The specific working process is as follows. When the experimental task is started, the light in the experimental box is turned on to prompt the pigeons to enter the experimental state, and the status display screen starts to present a gray background; next, two random red and green square markers are presented simultaneously on both the left and right sides of the screen; the test pigeon needs to select one of the options during the option presentation period and confirm its selection by pecking the button below the marker; when the action detection device detects the button press, the two square markers disappear, and the system will transfer to the subsequent state according to the pigeon's selection. A blue triangle or circle will be presented in the middle position of the status display screen, and the presented pattern depends on the pigeon's selection and the preset probability; similarly, the pigeon needs to peck the button below the pattern within the specified time to confirm; finally, the microcontroller will decide whether to activate the food reward according to the preset probability. When the reward is obtained, the stepper motor located outside the experimental box will drive the food box to be delivered into the experimental box through the reward port, otherwise it will enter the ITI state. It should be noted that if no pecking action is detected during the marker presentation period, the current trial ends and directly enters the ITI state.
[0036] Finally, for the convenience of subsequent behavioral and neural data analysis, the multi-channel neural data detection and recording system simultaneously receives the input of the behavioral event markers of the experimental training box and the input of 32-channel neural signals from the ST and Hp brain regions. When the status monitoring module detects the start of the experimental task, the multi-channel neural data detection and recording system starts to synchronously record neural data and behavioral data. The behavioral data are the time markers of various events during the experiment, including the stimulus presentation time, the pigeon's selection and pecking confirmation time, whether the award is obtained and the reward moment, etc.
[0037] The behavioral data markers in the data acquisition system are shown in the following table.
[0038] Behavioral data marking in data acquisition
[0039]
[0040] The above data is from the sequential decision-making learning training experiment of pigeons. All the pigeon subjects were housed in a cage with a length, width, and height of 80 cm throughout the experiment, maintained a 12-hour day-night cycle each day, and all experimental trainings were carried out during the day of the day-night cycle.
[0041] The present invention uses pigeons with good cognitive decision-making abilities as model animals, designs a two-step sequential decision-making experimental method, and designs and builds a pigeon sequential decision-making learning training system according to the requirements of the experimental method. This method can more accurately simulate the dynamic neural decision-making process of the brain during the decision-making process.
[0042] The embodiments of this specific implementation manner are all preferred embodiments of the present invention, and do not limit the protection scope of the present invention accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A sequential decision-making reinforcement learning training method and system, characterized in that: The method includes the following experimental steps: Step 1: Present two option signs S 1+ and S 1- to a pigeon at the same time, which are marked with red and green boxes respectively, and this state is the initial state (State1, S1); Step 2: Train the pigeon to select one of the options within 2 s of option presentation and confirm its choice by pecking the key below the corresponding option; once the pigeon has confirmed its choice, it will transition to the second state (State2, S2), also known as the successor state, with a certain probability T ss′ transition to the second state (State2, S2), also known as the successor state Step 3: The successor state marker S2 will be presented on the screen in front of the pigeon + or S2 - , presented as a blue triangle or a blue circle respectively, and each state marker carries a reward with a different probability; Step 4: The pigeon needs to peck the key below the marker within 2 s after the stimulus presentation to confirm that it has noticed the current state; finally, the system will present a food reward R with a certain probability P according to the successor state and the pigeon's choice. r Present the food reward R for 3 s. After the reward ends, enter a 5-s inter-trial interval (ITI).
2. The sequential decision-making reinforcement learning training method and system according to claim 1, wherein: In the design of the sequential decision-making experimental paradigm, whether in Step 1 or Step 2, if the pigeon does not make a choice or peck the key within the specified time, the experiment will directly enter the ITI state.
3. The sequential decision-making reinforcement learning training method and system according to claim 1, characterized in that: The state transition probability T from Step1 to the second step Step2 ss′ and the reward probability P r are set as shown in the following formula: In the formula, indicates that S1 is selected in Step1 + , indicates that S1 is selected in Step1 - , indicates the state S2 in Step2 + when the key pecking action is made, indicates the state S2 in Step2 - when the key pecking action is made, and R represents the reward information.
4. A sequential decision-making reinforcement learning training method and system according to claim 1, characterized in that: Select S1 in Step1 + and S1 - The theoretical expected values EV are respectively The red in Step1 (S1 + ) is the dominant option, and the green in Step1 (S1 - ) is the inferior option. The triangle in Step2 (S2 + ) is the dominant successor state, and the circle in Step2 (S2 - ) is the inferior successor state. For a state transition with a probability of 0.8, it is called a Common state transition, such as when choosing S1 + and transitioning to S2 + or choosing S1 - and transitioning to S2 - . On the contrary, for a state transition with a probability of 0.2, it is called an Uncommon state transition, such as when choosing S1 + and transitioning to S2 - or choosing S1 - and transitioning to S2 + . When a reward is obtained, it is marked with R (Reward), and when no reward is obtained, it is marked with NR (NoReward).
5. A sequential decision-making reinforcement learning training method and its system for implementing the sequential decision-making reinforcement learning training method as described in claim 1. The sequential decision-making experimental training system includes a behavior training box, an experimental parameter setting module, an experimental state monitoring module, and a multi-channel neural data detection and recording system. The communication between the behavior training box, the experimental parameter setting module, the experimental state monitoring module, and the multi-channel neural data detection and recording system is realized by a microcontroller with an STM32 single-chip microcomputer as the core and its peripheral circuits.
6. The sequential decision-making reinforcement learning training method and system according to claim 5, characterized in that: The experimental parameter setting module and the experimental state monitoring module are a human-computer interaction interface for setting basic experimental parameters and monitoring the current training state of the pigeon. This interface communicates bidirectionally with the microcontroller. This module can start and end the experiment, switch the food box, and download and save experimental parameters. It can also set the total number of trials, state transition probability, reward probability, stimulus presentation duration, and ITI duration information of the experiment. It can also monitor the experimental state of the pigeon in real time and record behavioral data, including the current number of trials, the number of awards, and the action selection, subsequent state, pecking response time, and whether the pigeon wins the award in each trial.
7. A sequential decision-making reinforcement learning training method and system according to claim 5, characterized in that: The behavior training box is a cuboid space for the sequential decision-making learning training of pigeons, with a length, width, and height of 55 cm × 50 cm × 60 cm respectively. It is made of acrylic plates with a thickness of 0.6 cm. Five sides are closed, and the other side is a door that can be opened and closed. There are ventilation holes on the top layer, and a layer of stainless steel wire mesh is provided at the bottom for the pigeons to stand on. Inside the training box, there are a state display device, an action detection device, a food reward device, and a video monitoring device. Among them, the state display device is located on the inner wall of the box and is a screen with a length × width of 8 cm × 6 cm. The action detection device is located directly below the position where the stimulus marker is presented and consists of three buttons. The food reward device consists of a stepper motor and a box with a length × width of 4 cm × 3 cm and is located at the bottom of the box below the buttons. The video monitoring device is installed on the wall directly opposite the state display device and can observe the state of the pigeon. The behavior training box communicates bidirectionally with the microcontroller. The microcontroller transmits experimental parameter information, such as option presentation, presentation duration, and reward information, into the experimental box. The action detection device in the behavior training box transmits the behavior selection of the pigeon to the microcontroller.