Intelligent agent strengthening control method based on time sequence synchronization pulse memory strategy

By adopting the timing synchronous pulse memory strategy in the enhanced control of agents, the problem of insufficient timing data processing capabilities of existing algorithms in complex POMDP tasks is solved, and more accurate agent control and higher performance performance is achieved.

CN120046649APending Publication Date: 2025-05-27XI AN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510184639.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing pulse reinforcement learning algorithm cannot effectively utilize the time series data processing capabilities in complex partial observable tasks, resulting in inaccurate interpretation of historical information, limiting the performance of agents in complex POMDP tasks.

Method used

The agent enhancement control method based on the timing synchronization pulse memory strategy is adopted to retain the timing information between multiple steps of data through timing synchronization encoding, and a high-performance pulse enhancement network is built through feature enhancement and network memory enhancement.

Benefits of technology

It realizes more accurate agent-enhanced control in complex POMDP tasks, improves the agent's performance ability in some observable environments, and reduces computing burden and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046649A_ABST
    Figure CN120046649A_ABST
Patent Text Reader

Abstract

The invention discloses an agent strengthening control method based on a time sequence synchronization pulse memory strategy, which is used for solving the agent decision-making problem in a part of observable environments. The method comprises the following steps of: 1, constructing a simulation environment, and simulating an information missing scene by randomly covering an observation dimension; 2, a strategy-judgment enhanced network framework is built, a pulse memory strategy network adopts time sequence synchronization pulse coding, real-time observation and memory branch integration historical information are processed through a current branch, and a memory judgment network evaluates a decision value; 3, adopting data obtained by interaction with the task environment to jointly train and strengthen a network framework; and 4, deploying the trained network model to realize agent control, and evaluating a task effect through action signal execution and reward feedback. According to the framework, the biological characteristics of spiking neurons and the time sequence modeling capability of a memory module are fused, the decision-making precision is improved under partial observable conditions, and compared with a traditional method, the framework has higher environmental adaptability and decision-making robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of the application of brain-like computing and reinforcement learning, and particularly relates to an intelligent agent reinforcement control method based on a time-sequential synchronous pulse memory strategy. Background Art

[0002] As a new generation of neural network, the spiking neural network (SNN) has become a research hotspot due to its excellent spatio-temporal information processing ability and biological rationality. Compared with the artificial neural network (ANN), the SNN adopts a pulse-driven communication and computing method, and combined with neuromorphic chips, it can greatly improve energy efficiency, and is very suitable for deep reinforcement learning (DRL) applications that require strategy deployment. In recent years, more and more research has been devoted to combining the SNN with the DRL, aiming to achieve a robust and low-power intelligent agent.

[0003] However, the existing spiking reinforcement learning (SRL) research mostly focuses on the biological rationality under small-scale and fully observable tasks, while ignoring the consideration of practical applications in complex partially observable tasks. At the same time, most of the existing SRL algorithms use the pulse frequency encoding of single-step data to retain the numerical information of each step of data. However, this method severs the temporal relationship between multi-step data, making it impossible to effectively utilize the temporal data processing ability of the SNN in subsequent pulse communication and computing processes, thereby affecting the accurate interpretation of historical information and limiting the performance in partially observable Markov decision process (POMDP) tasks. Therefore, aiming at these shortcomings, constructing a high-performance spiking reinforcement network model that can solve complex POMDP problems is a very important direction. Summary of the Invention

[0004] Aiming at the shortcoming that the pulse encoding of the existing algorithm ignores the temporal information between multi-step data, the purpose of the present invention is to propose an intelligent agent reinforcement control method based on a time-sequential synchronous pulse memory strategy. Different from the existing spiking reinforcement methods, the present invention uses time-sequential synchronous encoding to fully retain the temporal information between multi-step data, and through feature enhancement and network memory enhancement, a higher-performance spiking reinforcement network that can solve complex POMDP problems is realized to achieve more accurate intelligent agent reinforcement control.

[0005] The present invention can be applied to a variety of continuous control intelligent agents, including but not limited to complex partially observable environmental objects such as mobile robots, driverless vehicles, game AIs, etc., and has high practical application value.

[0006] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0007] An intelligent agent reinforcement control method based on a time-sequential synchronous pulse memory strategy, comprising the following steps:

[0008] Step A: Select a simulation environment for continuous control of the agent as the task scenario, and partially or completely mask the observation state received by the agent by adding random interference to form a partially observable environment with random sensor loss or random frame loss for enhanced training and testing;

[0009] Step B: Based on the pulse coding method of time synchronization, a reinforcement learning network framework is constructed. The reinforcement learning network framework consists of a pulse memory strategy network with memory branches and a memory judgment network. The pulse memory strategy network is divided into a current branch and a memory branch and is composed of pulse neurons and uses pulse signals for calculation. The memory judgment network is composed of artificial neurons;

[0010] Step C: Use the data obtained by interacting with the task scenario obtained in step A to jointly train the pulse memory strategy network and the memory judgment network to obtain a network model that converges under the corresponding task scenario;

[0011] Step D: Use the trained network model to test the corresponding task scenario obtained in step A, control the agent through the action signal output by the network model, and use the reward feedback of the scenario to evaluate the task completion effect.

[0012] The specific steps of step B are as follows:

[0013] Step B01: The pulse memory strategy network uses a group encoder with a Gaussian receptive field in the current branch and the memory branch to pulse encode the current observation information and historical information received by the agent respectively; the current branch uses IF neurons to encode the current observation information into pulses through repeated stimulation; the memory branch splices and fuses the historical observation sequence and historical action sequence in the historical information and generates pulses according to the stimulation sequence. The temporal synchronization encoding with enhanced spatiotemporal features is realized by a single stimulation of dynamic neurons with intra-layer connections. The dynamic neurons with intra-layer connections are specifically represented as follows:

[0014]

[0015]

[0016] Among them, t-1 represents the pulses fired by the dynamic neurons connected within the layer at time t-1; W intra and b intra The connection weights and biases of the synapses connecting the dynamic neurons within the layer; represents the current change caused by the stimulation of the intralayer connection at time t; d c Represents the current attenuation coefficient; C t-1 and C trespectively represent the neuron currents at times t-1 and t;

[0017] Step B02: The current branch and the memory branch of the pulse memory strategy network respectively use a non-spiking fully connected layer and a double-output LSTM neuron to further interpret the obtained pulse signals; among them, the non-spiking fully connected layer uses dynamic non-spiking neurons to use the membrane potential at the last moment as pulse decoding for output, and its neuron dynamics are:

[0018]

[0019] V t = d v ·(V t-1 - θ r ) + θ r

[0020]

[0021] V t = V t + V delta

[0022] In the formula represents the pulse output by the previous layer at time t; W N and b N represent the weights and biases of the synaptic connections between the dynamic non-spiking neurons and the previous layer of neurons; d v is the membrane potential decay factor; V delta represents the change in the cell membrane potential; θ r represents the learnable reset threshold; V t and V delta respectively represent the membrane potential and the change amount of the neuron at time t;

[0023] The double-output LSTM neuron uses a double-output dynamic neuron to replace the output gate part in the traditional pulse LSTM neuron, ensuring the transmission of pulse signals within the unit and using the membrane potential as the output, and its neuron dynamics are:

[0024]

[0025] V t = V t-1 ·(1 - O t-1 ) + O t-1 ·θ r ·tanh(V t-1 - V th )

[0026] U t = U t-1 + O t-1 ·θs

[0027]

[0028] U delta = θ v ·V t -θ u ·U t

[0029] V t = V t + V delta

[0030] U t = U t + U delta

[0031] O t = (V t > V th )

[0032] where represents the pulse of the hidden layer input at time t; W D and b D represent the connection weights and biases of the input synapses of the dynamic neurons in this layer; θ v , θ u , θ r , θ s and θ f are all learnable dynamic parameters, representing the membrane potential parameter, hyperpolarization resistance parameter, reset parameter, pulse parameter, and output factor respectively; U t and U delta represent the hyperpolarization resistance term and the change amount of the neuron at time t respectively, and V th represents the membrane potential threshold for neuron spike firing; Using the membrane potential at the last moment for pulse decoding realizes the alignment of the two-branch features, facilitating feature fusion;

[0033] Step B03: After the current branch and the memory branch output the membrane potential at the last moment, the features are concatenated and a fully connected layer is used for feature fusion, which is expressed as:

[0034]

[0035] V combined = FusionFC(V concat )

[0036] where represent the membrane potentials output by the current branch and the memory branch at the last simulation moment T respectively; Concat represents the concatenation operation, FusionFC represents the fusion fully connected operation, and V concat represents the result obtained by concatenation and The obtained membrane potential information vector, V combined represents the membrane potential feature vector obtained by integrating through the fusion of fully connected operations;

[0037] After obtaining the fused features, a group decoder is then used to partition and decode the fused features according to the action dimension to obtain the action and output it, which is specifically expressed as:

[0038]

[0039] where represents the membrane potential output by population i; and represent the weight and bias for decoding the i-th dimensional population respectively; a (i) represents the action output of the i-th dimension;

[0040] Step B04: The spiking neurons in the spiking memory policy network all adopt a scale-varying soft reset mechanism to avoid extreme potentials and enhance memory, which is specifically expressed as follows:

[0041] V t = V t-1 · (1 - O t-1 ) + O t-1 · θ r · tanh(V t-1 - V th )

[0042] where θ r represents the learnable reset threshold, and V th represents the reset threshold; the difference between the potential and the threshold is mapped to the interval [0, 1] using the tanh() function to limit the potential change and retain the historical potential information;

[0043] Step B05: Input the action obtained in step B03, as well as the current observation information and historical information, into the memory critic network for Q-value calculation; the memory critic network is built using artificial neurons, and its network also has a current branch and a memory branch. Among them, the current branch uses a fully connected layer, and the memory branch uses a fully connected layer and an LSTM memory layer. The current branch and the memory branch are fused and decoded using two fully connected layers after feature concatenation, and finally the cumulative reward Q-value corresponding to the state-action expectation of the task scenario is output.

[0044] The specific steps of step C are as follows:

[0045] Step C01: When training the reinforcement learning network framework built in step B, the policy network loss is used for the spiking memory policy network and the memory critic network with memory branches respectively, that is and the discriminator network loss, that is

[0046] Step C02: Use the data obtained by interacting with the task scenario for training, and adopt the spatio-temporal backpropagation algorithm to minimize the loss to optimize the learnable parameters and model parameters of the reinforcement learning network, and obtain a converged network model.

[0047] In step A, the environment is divided into a fully observable condition without masking, a random sensor loss condition where each observation has a 0.1 probability of being masked, and a random frame loss condition where each frame has a 0.2 probability of being completely masked.

[0048] Compared with the prior art, the present invention has the following advantages:

[0049] First, since the temporal synchronous pulse coding used by the present invention at the memory branch can retain temporal information, enhance features through sufficient spatio-temporal fusion, and reduce the information loss caused by coding, it can better adapt to temporal data processing tasks compared with existing pulse coding methods;

[0050] Second, since the temporal synchronous pulse coding used by the present invention at the memory branch represents richer temporal feature information with fewer pulse signals, it reduces the computational burden and power consumption;

[0051] Third, since the present invention adopts an in-layer connected dynamic neuron, a dual-output dynamic neuron, and a scale transformation soft reset mechanism, it enhances the memory ability of the network model and has better performance in partially observable agent sequence decision-making tasks that rely on memory ability;

[0052] Fourth, the scale transformation soft reset mechanism adopted by the present invention limits the potential change to avoid the generation of extreme potentials, and improves the stability during model training. Brief Description of the Drawings

[0053] Figure 1 is a framework diagram of the reinforcement learning network based on the temporal synchronous pulse memory strategy built by the present invention. Detailed Embodiments

[0054] The following details of each step of the present invention will be introduced in detail with reference to the drawings.

[0055] The present invention proposes an intelligent agent reinforcement control method based on a temporal synchronous pulse memory strategy.

[0056] This method mainly includes the following steps:

[0057] Step A: Select a simulation environment for continuous control of the agent as the task scenario. Partially or completely mask the observed state received by the agent by adding random interference to form a partially observable environment with random sensor loss or random frame dropping for reinforcement training and testing.

[0058] Step B: Based on the pulse coding method of time series synchronization, build a reinforcement learning network framework. The reinforcement learning network framework consists of a pulse memory policy network with a memory branch and a memory evaluation network. The pulse memory policy network is divided into a current branch and a memory branch, which are composed of pulse neurons and use pulse signals for operation. The memory evaluation network is composed of artificial neurons. The reinforcement learning network framework is as Figure 1 shown. The overall network framework consists of a pulse memory policy network with a memory branch and a memory evaluator network in time series synchronization. The pulse memory policy network as a whole uses pulse neurons, inputs the observed information and outputs the action control of the agent. The memory evaluation network uses artificial neurons, inputs the observed and action information and outputs the expected cumulative return Q value for network training. The figure shows a detailed display of the population encoder in time series synchronization in the memory branch, which uses a Gaussian receptive field as a whole and outputs a time series synchronized pulse sequence.

[0059] The specific process in the figure and the specific steps of Step B are as follows:

[0060] Step B01: The pulse memory policy network uses population encoders with Gaussian receptive fields in the current branch and the memory branch to perform pulse coding on the current observed information and historical information received by the agent respectively; the current branch uses IF neurons to encode the current observed information into pulses through repeated stimulation; the memory branch splices and fuses the historical observation sequence and historical action sequence in the historical information and performs spatio-temporal feature enhanced time series synchronization coding in the way of single-stimulating dynamic neurons with intra-layer connections. The dynamic neurons with intra-layer connections are specifically represented as follows: The single-stimulating dynamic neurons with intra-layer connections realize spatio-temporal feature enhanced time series synchronization coding. The dynamic neurons with intra-layer connections are specifically represented as follows:

[0061]

[0062] where O t-1 represents the pulse fired by the dynamic neuron with intra-layer connections at time t - 1; W intra and b intra represent the connection weights and biases of the synapses with intra-layer connections of the dynamic neurons with intra-layer connections; represents the current change generated by the intra-layer connection stimulation at time t; d c represents the current decay coefficient; C t-1 and C t represent the neuron currents at times t - 1 and t respectively.

[0063] Step B02: The current branch and the memory branch of the pulse memory strategy network respectively use a non-spiking fully connected layer and a double-output LSTM neuron to further interpret the obtained pulse signals. Among them, the non-spiking fully connected layer uses dynamic non-spiking neurons to output using the membrane potential at the last moment as pulse decoding, and its neuron dynamics are:

[0064]

[0065] V t = d v ·(V t-1 - θ r ) + θ r

[0066]

[0067] V t = V t + V delta

[0068] In the formula represents the pulse output from the previous layer at time t; W N and b N represent the weights and biases of the synaptic connections between the dynamic non-spiking neurons and the previous layer of neurons; d v is the membrane potential decay factor; V delta represents the change in the cell membrane potential; θ r represents the learnable reset threshold; V t and V delta represent the membrane potential and the change amount of the neuron at time t respectively;

[0069] The double-output LSTM neuron uses a double-output dynamic neuron to replace the output gate part in the traditional pulsed LSTM neuron, ensuring the transmission of pulse signals within the unit and using the membrane potential as the output. Its neuron dynamics are:

[0070]

[0071] V t = V t-1 ·(1 - O t-1 ) + O t-1 ·θ r ·tanh(V t-1 - V th )

[0072] U t = U t-1 + O t-1 ·θ s

[0073]

[0074] U delta = θ v ·V t -θ u ·U t

[0075] V t = V t + V delta

[0076] U t = U t + U delta

[0077] O t = (V t > V th )

[0078] Wherein represents the pulse of the hidden layer input at time t; W D and b D represent the connection weights and biases of the input synapses of the dynamic neurons in this layer; θ v , θ u , θ r , θ s and θ f are all learnable dynamic parameters, representing the membrane potential parameter, hyperpolarization resistance parameter, reset parameter, pulse parameter and output factor respectively; U t and U delta represent the hyperpolarization resistance term and the change amount of the neuron at time t respectively, and V th represents the membrane potential threshold for neuron spike firing; Using the membrane potential at the last moment for pulse decoding realizes the alignment of two-branch features, which is convenient for feature fusion;

[0079] Step B03: After the current branch and the memory branch output the membrane potential at the last moment, splice the features and use a fully connected layer for feature fusion, which is expressed as:

[0080]

[0081] V combined = FusionFC(V concat )

[0082] Wherein represent the membrane potentials output by the current branch and the memory branch at the last simulation moment T respectively; Concat represents the splicing operation, FusionFC represents the fusion fully connected operation, and V concat represents the result obtained by splicing and The obtained membrane potential information vector, V combined represents the membrane potential feature vector obtained by integrating through the fusion fully connected operation; after obtaining the fused features, a group decoder is then used to divide and decode the fused features according to the action dimension to obtain the action and output it, specifically expressed as:

[0083]

[0084] where represents the membrane potential output by population i; and respectively represent the weight and bias of the i-th dimensional population decoding; a (i) represents the action output of the i-th dimension.

[0085] Step B04: The spiking neurons in the spiking memory policy network all adopt a scale-varying soft reset mechanism to avoid extreme potentials and enhance memory, specifically expressed as follows:

[0086] V t = V t-1 ·(1 - O t-1 ) + O t-1 ·θ r ·tanh(V t-1 - V th )

[0087] where θ r represents the learnable reset threshold, and V th represents the reset threshold; the difference between the potential and the threshold is mapped to the interval [0, 1] using the tanh() function to limit the potential change and retain the historical potential information;

[0088] Step B05: Input the action obtained in Step B03, as well as the current observation information and historical information, into the memory critic network for Q-value calculation; the memory critic network is built using artificial neurons, and its network also has a current branch and a memory branch. The current branch uses a fully connected layer, and the memory branch uses a fully connected layer and an LSTM memory layer. The current branch and the memory branch are fused and decoded using two fully connected layers after feature concatenation, and finally, the cumulative reward Q-value of the state-action expectation corresponding to the task scenario is output.

[0089] Step C: Use the data obtained by interacting with the task scenario obtained in Step A to jointly train the spiking memory policy network and the memory critic network to obtain a converged network model for the corresponding task scenario.

[0090] The specific steps of Step C are as follows:

[0091] Step C01: When training the reinforcement learning network framework built in Step B, the policy network loss is used respectively, that is and the critic network loss, i.e., acting on two independent Q networks ( and );

[0092] Step C02: Use the data obtained from interacting with the task scenario for training, and adopt the spatio-temporal backpropagation algorithm to minimize the loss to optimize the learnable parameters and model parameters of the reinforcement learning network, and obtain a converged network model.

[0093] Step D: Use the trained network model to test in the corresponding task scenario obtained in Step A, control the agent through the action signal output by the network model, and evaluate the task completion effect using the reward feedback of the scenario.

[0094] On the simulation environment obtained in Step A, train the currently used advanced artificial neural network and spiking neural network methods and the method of the present invention. After the training is completed, use the randomly initialized simulation environment to test different methods. The results are shown in Table 1 below. The larger the corresponding average reward value in the table, the better. Among them, the best result is marked in italic and bold.

[0095] Table 1

[0096]

[0097] As can be seen from Table 1, compared with other methods, the method of the present invention has significantly improved performance under multiple conditions in different environments, indicating the effectiveness of the method of the present invention.

Claims

1. The intelligent agent reinforcement control method based on the timing synchronization pulse memory strategy is characterized in that a pulse encoding method based on timing synchronization is adopted to build a reinforcement learning network framework with a pulse memory strategy network, which specifically includes the following steps: Step A: Select a simulation environment for continuous control of the agent as the task scenario, and partially or completely mask the observation state received by the agent by adding random interference to form a partially observable environment with random sensor loss or random frame loss for enhanced training and testing; Step B: Based on the pulse coding method of time synchronization, a reinforcement learning network framework is constructed. The reinforcement learning network framework consists of a pulse memory strategy network with memory branches and a memory judgment network. The pulse memory strategy network is divided into a current branch and a memory branch and is composed of pulse neurons and uses pulse signals for calculation. The memory judgment network is composed of artificial neurons; Step C: Use the data obtained by interacting with the task scenario obtained in step A to jointly train the pulse memory strategy network and the memory judgment network to obtain a network model that converges under the corresponding task scenario; Step D: Use the trained network model to test the corresponding task scenario obtained in step A, control the agent through the action signal output by the network model, and use the reward feedback of the scenario to evaluate the task completion effect.

2. The intelligent agent enhanced control method based on the timing synchronous pulse memory strategy according to claim 1 is characterized in that: The specific steps of step B are as follows: Step B01: The pulse memory strategy network uses a group encoder with a Gaussian receptive field in the current branch and the memory branch to pulse encode the current observation information and historical information received by the agent respectively; The current branch uses IF neurons to encode the current observation information into pulses through repeated stimulation; the memory branch splices and fuses the historical observation sequence and historical action sequence in the historical information and generates pulses according to the stimulation sequence. The temporal synchronization encoding with enhanced spatiotemporal features is realized by a single stimulation of dynamic neurons with intra-layer connections. The dynamic neurons with intra-layer connections are specifically represented as follows: I intrat =W intra O t-1 +b intra Among them, t-1 represents the pulses fired by the dynamic neurons connected within the layer at time t-1; W intra and b intra The connection weights and biases of the synapses in the dynamic neuron layer that represent the connections within the layer; represents the current change caused by the stimulation of the intralayer connection at time t; d c Represents the current attenuation coefficient; C t-1 and C t represent the neuronal currents at time t-1 and t respectively; Step B02: The current branch and memory branch of the pulse memory strategy network respectively use a fully connected layer without emitting pulses and a dual-output LSTM neuron to further interpret the obtained pulse signal; the fully connected layer without emitting pulses uses dynamic non-pulse neurons to use the membrane potential at the last moment as the pulse decoding output, and its neuron dynamics is: V t =d v ·(V t-1 -θ r )+θ r V t =V t +V delta Where X t o represents the pulse output by the previous layer at time t; W N and b N Represents the weights and bias of the synaptic connection between the dynamic non-spiking neuron and the neurons in the previous layer; d v is the membrane potential attenuation factor; V delta represents the change in neuronal membrane potential; θ r represents the learnable reset threshold; V t and V delta Respectively represent the membrane potential and change of the neuron at time t; The dual-output LSTM neuron uses a dual-output dynamic neuron to replace the output gate part of the traditional pulse LSTM neuron, ensuring the pulse signal transmission within the unit and using the membrane potential as the output. Its neuron dynamics are: V t =V t-1 ·(1-O t-1 )+O t-1 ·θ r ·tanh(V t-1 -V th ) IN t =U t-1 +O t-1 ·θ s U delta =θ v ·V t -θ u ·U t V t =V t +V delta IN t =U t +U delta About t =(V t >V th ) in represents the pulse of hidden layer input at time t; W D and b D Represents the connection weights and biases of the dynamic neuron input synapses in this layer; θ v ,θ u ,θ r ,θ s and θ f are all learnable dynamic parameters, representing membrane potential parameters, hyperpolarization resistance parameters, reset parameters, pulse parameters and output factors respectively; U t and U delta Respectively represent the neuron hyperpolarization resistance term and change at time t, V th The membrane potential threshold representing the neuron pulse emission; using the membrane potential at the last moment as the pulse decoding, the two-branch feature alignment is achieved, which facilitates feature fusion; Step B03: After the current branch and the memory branch output the membrane potential at the last moment, the features are concatenated and the fully connected layer is used for feature fusion, which is expressed as: V combined =FusionFC ( V concat in Respectively represent the membrane potentials of the current branch and the memory branch output at the last simulation time T; Concat represents the concatenation operation, FusionFC represents the fusion full connection operation, and V concat Indicates by splicing and The membrane potential information vector obtained, V combined represents the membrane potential feature vector obtained by integrating the fully connected operation; After obtaining the fused features, the group decoder is used to divide and decode the fused features according to the action dimension to obtain the action and output it, which is specifically expressed as: in represents the membrane potential of the output of group i; and Represent the weight and bias of the i-th dimension group decoding respectively; a (i) represents the action output of the i-th dimension; Step B04: The spike neurons in the spike memory strategy network all use a scale-varying soft reset mechanism to avoid extreme potentials and enhance memory, as shown below: V t =V t-1 ·(1-O t-1 )+O t-1 ·θ r ·tanh(V t-1 -V th ) where θ r represents the learnable reset threshold, V th Represents the reset threshold; the difference between the potential and the threshold is mapped to the interval [0,1] using the tanh() function, thereby limiting the potential change and retaining the historical potential information; Step B05: input the action obtained in step B03, current observation information and historical information into the memory judge network to calculate the Q value; The memory judge network is built with artificial neurons, and its network also has a current branch and a memory branch. The current branch uses a fully connected layer, and the memory branch uses a fully connected layer and an LSTM memory layer. The current branch and the memory branch are fused and decoded using two layers of full connection after feature splicing, and finally the expected cumulative reward Q value of the state action corresponding to the task scenario is output.

3. The intelligent agent enhanced control method based on the timing synchronous pulse memory strategy according to claim 1 is characterized in that: The specific steps of step C are as follows: Step C01: When training the reinforcement learning network framework built in step B, the strategy network loss is used for the pulse memory strategy network with memory branches and the memory judgment network, that is, And the judge network loss, i.e. Step C02: Use the data obtained by interacting with the task scene for training, and use the spatiotemporal back propagation algorithm to minimize the loss to optimize the learnable parameters and model parameters of the reinforcement learning network to obtain a converged network model.

4. The intelligent agent enhanced control method based on the timing synchronous pulse memory strategy according to claim 1 is characterized in that: In step A, the environment is divided into a fully observable condition with no masking, a random sensor loss condition with a 0.1 probability of each observation being masked, and a random frame loss condition with a 0.2 probability of each frame being fully masked.

5. The intelligent agent enhanced control method based on the timing synchronous pulse memory strategy according to claim 1 is characterized in that: The intelligent agents include mobile robots, unmanned vehicles, and game AI.

Citation Information

Cited By

  • Method and system for enhancing body navigation memory based on spiking neurons

    CN121351880A