A multi-agent brain-like tracking control system and method based on reservoir computing
Patent Information
- Application Number
- CN202610714683.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-08-18
AI Technical Summary
[0007]综上,现有多智能体跟踪控制技术在抗干扰性、通信效率、计算资源优化及动态环境适应性等方面存在系统性缺陷,亟需一种兼具高鲁棒性、低通信开销、强环境普适性及理论可证明性的新型控制范式
(1)本发明构建了神经形态非线性动态系统辨识器,利用储层计算的高维动力学特性,能够实时学习并重构包括内部未建模动态与外部干扰在内的系统总扰动。该方法无需依赖精确的系统动力学模型,有效解决了实际工程中模型难以获取或存在显著不确定性的问题,实现了对未知扰动的精准前馈补偿,大幅提升了系统在复杂环境下的控制精度与鲁棒性。
Smart Images

Figure CN122593393A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of neuromorphic computing and intelligent agent control, and relates to a multi-agent neuromorphic tracking control system and method based on reservoir computing. Background Technology
[0002] Multi-agent cooperative control, a core research direction in the field of unmanned systems, has demonstrated enormous application potential in scenarios such as formation flying, swarm operations, and disaster relief in recent years. Existing control schemes mainly employ centralized or distributed architectures, combining wireless communication, sensor networks, and traditional control theory to achieve multi-robot cooperation. However, in practical engineering deployments, these methods generally face the following fundamental challenges: Existing technologies are mostly based on idealized premises of zero-delay and zero-packet-loss communication, precise and known robot dynamics models, and negligible environmental interference. However, in real-world scenarios, the time-varying characteristics of wireless communication lead to unstable data transmission, and the nonlinear friction, saturation characteristics, and sensor noise of robot actuators are difficult to model accurately. Dynamic environmental disturbances (such as sudden changes in wind force and ground friction) further exacerbate system uncertainties, causing theoretical algorithms to degrade sharply in actual operation.
[0003] Traditional distributed control relies heavily on external positioning facilities such as the Global Positioning System (GPS), Ultra-Wideband (UWB), or pre-installed visual markers. In environments where satellite signals are denied, such as indoors, underground, forests, or urban canyons, the system fails because it cannot obtain global location information, which greatly limits the expansion of application scenarios.
[0004] Conventional control methods rely on continuous control signal transmission, requiring agents to frequently exchange large amounts of raw state data. In networks with limited bandwidth, this not only wastes communication resources but also easily leads to network congestion, latency, and packet loss. Especially when the system encounters unknown strong disturbances or malicious attacks, even more data needs to be transmitted to maintain control performance, further exacerbating the communication load and causing a decrease in the stability of the closed-loop system.
[0005] For obstacle avoidance problems involving dynamic obstacles, early model predictive control often simplified them to moving static obstacles, lacking effective prediction of the obstacle's future trajectory, leading to delayed obstacle avoidance decisions. Furthermore, embedding hard constraints such as collision avoidance, communication connection maintenance, and actuator physical limits into the control framework often results in the absence of feasible solutions due to constraint conflicts. Although introducing control barrier functions can ensure safety, it easily disrupts the consistent convergence characteristics of multi-agent systems, and compatibility is difficult to guarantee when multiple constraints coexist.
[0006] While black-box models such as deep learning and reinforcement learning possess strong environmental adaptability, they lack rigorous theoretical proofs of stability and safety, making it difficult to meet the requirements of safety-critical tasks. How to organically combine the adaptive advantages of data-driven methods with the provability of traditional control theory remains a cutting-edge challenge that urgently needs to be overcome.
[0007] In summary, existing multi-agent tracking control technologies suffer from systemic deficiencies in terms of anti-interference capabilities, communication efficiency, computational resource optimization, and adaptability to dynamic environments. There is an urgent need for a novel control paradigm that combines high robustness, low communication overhead, strong environmental universality, and theoretical provability. Distributed control methods based on neuromorphic computing, in particular, offer a novel approach to solving these challenges. Summary of the Invention
[0008] In view of this, the purpose of the present invention is to provide a multi-agent brain-like tracking control system and method based on reservoir computing.
[0009] To achieve the above objectives, the present invention provides the following technical solution: A multi-agent brain-like tracking and control system based on reservoir computing includes a state perception layer, a reservoir computing core layer, a composite control layer, and an execution output layer connected in sequence. The state awareness layer is used to acquire the state of each agent in the multi-agent system and the state information of neighboring agents, and to calculate the tracking error with the navigator. The reservoir computing core layer incorporates a neuromorphic nonlinear dynamic system identifier, which is used to map the tracking error to a high-dimensional feature space to reconstruct the total system disturbance, and generate a dynamic communication trigger signal based on the spatiotemporal prediction information entropy of the reservoir state. The composite control layer is used to generate continuous or discrete neuromorphic sparse control instructions based on the tracking error, the reconstruction result of the total system disturbance, and a preset virtual energy threshold. The execution output layer is used to receive and execute the neuromorphic sparse control instructions, and feed back the physical state after execution to the state perception layer to form a closed-loop control.
[0010] Furthermore, the reservoir computing core layer includes a reservoir network, a readout layer weight training module, and a communication trigger determination module; The reservoir network is an echo state network, whose internal structural weights are randomly generated and fixed, and whose state evolution follows the following discrete-time difference equation:
[0011] in, Let k be the instantaneous activation state of the neuron mapped to the high-dimensional feature space. This represents the total number of reservoir neurons. fork The neural network input vector at time step; The reservoir internal connectivity weight matrix has a spectral radius. ; The input weight matrix; To output the feedback weight matrix; Leakage factor; The readout layer weight training module updates the output weights using a neural modulation reward-punishment synapse training algorithm based on the biological brain dopamine mechanism. ; The communication trigger determination module calculates the trigger energy functional. With dynamic threshold function The relationship determines whether communication is triggered.
[0012] Furthermore, in the communication triggering determination module, the triggering energy functional The calculation formula is:
[0013] in, , These are the weighting coefficients. For intelligent agents i exist t The state vector at time t, For intelligent agents i Last communication time The state vector, Forgetting factor; Dynamic threshold function The calculation formula is:
[0014] in, As the initial threshold, , , For adjustment coefficients, The spatiotemporal prediction information entropy of reservoir state, This is the real-time reconstructed signal for the total system disturbance; Next trigger time satisfy At that time, intelligent agent i Broadcast status information to neighbors.
[0015] Furthermore, the neuromorphic sparse control instructions generated by the composite control layer It is composed of the nominal consistency control law, the neuromorphic compensation term, and the impulse firing gating state term, and its calculation formula is as follows:
[0016] in, The pulse delivery gating state is set. This is a nominal consistency control law. This is a neuromorphic compensation term.
[0017] Furthermore, the pulse delivery gating state The value is determined by the virtual energy storage variable. The virtual energy accumulation variable is determined. The differential expression is:
[0018] in, Leakage coefficient, This is the gain coefficient. This is the total disturbance reconstruction signal output in real time by the reservoir network. For tracking error; When the tracking error does not reach the preset steady-state threshold The system performs continuous control; when the tracking error enters the preset steady-state threshold and When the preset trigger threshold is exceeded, Within the preset time window The internal activation is set to 1 and discrete action pulses are emitted.
[0019] A multi-agent neuromorphic tracking control method based on reservoir computing includes the following steps: S1: Obtain the agents in the multi-agent system through the state-aware layer. i its own position ,speed And the state information of neighboring intelligent agents, calculating the navigator's trajectory. Tracking error ; S2: The tracking error Input the reservoir core layer and reconstruct the total system perturbation using a neuromorphic nonlinear dynamic system identifier. And based on the spatiotemporal prediction information entropy of reservoir state Determine whether the communication mechanism has been triggered; S3: The composite control layer determines the tracking error based on the aforementioned tracking error. The total disturbance of the system and virtual energy accumulation variables Generate brain-like sparse control instructions ; S4: The output layer receives and executes the neuromorphic sparse control instructions. The physical state after execution is fed back to S1 to form a closed-loop control.
[0020] Furthermore, in S2, the neuromorphic nonlinear dynamic system identifier dynamically maps the errors of the low-dimensional physical system to a high-dimensional feature space, reconstructing the total perturbation of the system. The calculation formula is:
[0021] in, for t The readout layer weight matrix at time step [time]. for t The high-dimensional activation field of the reservoir at any given moment This is the gain coefficient. As the attenuation factor, For the Laplace manifold surface curvature smoothing term.
[0022] Furthermore, in S2, the reservoir computing core layer uses a neural modulation reward-punishment synaptic training algorithm based on the biological brain dopamine mechanism to update the readout layer weights. The algorithm aims to maximize the long-term cumulative reward and punishment signal and adjusts the weights through a three-factor synaptic plasticity law. The three factors include presynaptic activity, postsynaptic activity, and the dopamine signal, which represents the evaluation of the global reward and punishment environment.
[0023] Furthermore, in S3, the neuromorphic sparse control instructions nominal consistency control law The calculation formula is:
[0024] in, For the proportional term gain, For the differential term gain, For intelligent agents i The neighborhood group, For intelligent agents i and j Adjacency weights between them To provide navigator access gains, when the intelligent agent i When the navigator can be directly observed ,otherwise .
[0025] Furthermore, in S2, the spatiotemporal prediction information entropy of the reservoir state... The calculation formula is:
[0026] in, This represents the total number of reservoir neurons. For the first kThe activation probability of a neuron is calculated based on the probability distribution of the activation states of high-dimensional neurons within the reservoir.
[0027] The beneficial effects of this invention are as follows: (1) This invention constructs a neuromorphic nonlinear dynamic system identifier, which utilizes the high-dimensional dynamic characteristics of reservoir computing to learn and reconstruct the total system disturbance, including internal unmodeled dynamics and external disturbances, in real time. This method does not rely on an accurate system dynamic model, effectively solving the problem that models are difficult to obtain or have significant uncertainties in practical engineering, and realizing accurate feedforward compensation for unknown disturbances, which greatly improves the control accuracy and robustness of the system in complex environments.
[0028] (2) This invention proposes a reinforcement learning reservoir readout layer weight update rule based on the biological brain dopamine mechanism. This mechanism breaks through the limitation of traditional training methods that only pursue error minimization, and instead aims to maximize the long-term cumulative reward and punishment signal. By introducing a synaptic noise detection term to help the system escape local optima, and by comprehensively considering tracking accuracy, motion smoothness and control energy consumption, the system has the ability to autonomously seek optimization and can continuously optimize performance during use, rather than passively fitting data.
[0029] (3) The environmental complexity is measured by calculating the spatiotemporal predictive information entropy of the activation probability of high-dimensional neurons inside the reservoir in real time, and the communication threshold is dynamically adjusted accordingly. When severe environmental disturbances cause the system state to become chaotic, the threshold decays rapidly to force the agent to broadcast information frequently and quickly reconstruct knowledge sharing; during the stable period, the threshold is adaptively relaxed to significantly reduce redundant communication. This mechanism greatly reduces communication bandwidth and energy consumption while ensuring system stability.
[0030] (4) This invention introduces a brain-like pulse emission gating mechanism into the composite controller. Mimicking the principle that biological nervous systems only emit discrete pulses to maintain posture in steady state, the controller automatically switches to pulse sleep mode when the system tracking error reaches the steady-state threshold. At this time, the control command changes from continuous signal output to discrete action pulses with extremely low duty cycles, effectively avoiding continuous high-load operation of the actuator and significantly reducing the mechanical wear of the actuator and the overall power consumption of the system.
[0031] (5) This invention achieves deep integration of the perception layer and the control layer, and integrates the advantages of brain-like computing throughout the entire control loop. By leveraging the strong representational ability of reservoir computing for nonlinear dynamics, the autonomous evolutionary ability brought by the dopamine mechanism, and the low power consumption characteristics brought by pulse control, it successfully combines the strong adaptability of data-driven methods with the provable stability of traditional control theory, solving the problem of the lack of theoretical guarantees in black box models. This provides a new, biologically rational brain-like control paradigm for the safe and reliable operation of multi-agent systems in complex dynamic environments.
[0032] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein: Figure 1 Here is a diagram of the echo state network structure; Figure 2 A block diagram of a neuromorphic nonlinear dynamic system identifier; Figure 3 This is a flowchart of a neural modulation reward and punishment synaptic readout layer training algorithm based on the biological brain dopamine mechanism. Figure 4 A block diagram of the adaptive topology communication triggering mechanism based on reservoir spatiotemporal prediction entropy; Figure 5 A block diagram is shown for the design of a multi-agent brain-like tracking control algorithm based on reservoir computing. Detailed Implementation
[0034] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0035] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0036] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0037] 1. Overall Architecture Design The overall architecture of this invention aims to achieve deep integration of the perception layer and the control layer, and solves problems such as nonlinear interference, communication redundancy, and excessive power consumption in multi-agent distributed tracking control through neuromorphic computing. The overall system architecture adopts a closed-loop design, which can be divided into three core layers from the perspective of logic and data flow: system input layer, core control layer, and execution output layer.
[0038] The system input layer is mainly responsible for sensing and aggregating the basic spatiotemporal state and environmental information required for the operation of the intelligent agent. It encapsulates the leader's reference trajectory, external unknown interference and intelligent agent state information into a spatiotemporal tensor, which serves as the basic excitation signal for subsequent brain-like core algorithms.
[0039] The core control layer comprises multiple collaborative neuromorphic algorithm modules responsible for information dimensionality reduction and abstraction, high-dimensional mapping, online learning, and decision generation. The consistency error calculation module receives state data from the system input layer, calculates the relative tracking error in the multi-agent network, and provides a benchmark error signal for subsequent identification and control law synthesis. The neuromorphic nonlinear dynamic system identifier abandons traditional precise mathematical modeling, utilizing the high-dimensional dynamic characteristics of reservoir computation to map low-dimensional physical system spatiotemporal signals to a high-dimensional feature space of neuromorphism. A high-dimensional state manifold is generated through complex dynamic iterations within the reservoir, and then a linear readout layer is used to achieve online inverse mapping reconstruction and accurate compensation estimation of unmodeled dynamics and external composite perturbations. The neural modulation reward-punishment synaptic readout layer training module optimizes the readout layer weights of the identifier online. An innovative reinforcement learning algorithm based on the biological brain dopamine mechanism is introduced, moving away from passively pursuing error minimization and constructing a comprehensive neural modulation reward-punishment signal. Under this mechanism, the system can proactively explore optimal synaptic connection states, preventing it from getting trapped in local optima and endowing it with a high degree of autonomous optimization capability and environmental adaptability. The spatiotemporal predictive entropy communication module introduces a brain-like advanced attention mechanism, breaking through the bandwidth and energy bottlenecks caused by continuous broadcasting in traditional distributed systems. It deeply analyzes the degree of chaos in the system under disturbance by calculating the spatiotemporal predictive information entropy of the activation probability of high-dimensional neurons within the reservoir in real time. Combining a dynamic threshold function and a trigger energy functional, the threshold is adaptively relaxed to reduce redundant communication when the environment is stable; when encountering severe disturbances, the threshold decays rapidly to trigger forced communication, quickly reconstructing knowledge sharing on the topology network. The composite controller module is responsible for executing the final control law integrated decision. The control law is composed of three parts: a nominal consistency control law used to stabilize the system and track the trajectory, a neuromorphic online compensation term based on the identifier output, and a brain-like Spiking pulse firing gating state term. In particular, the controller is designed with a state switching mechanism based on a virtual energy threshold. Discrete action pulses are only issued when the system tracking error enters a steady state and the internal virtual energy accumulation exceeds the threshold, thus achieving extremely low power consumption and low mechanical wear.
[0040] Finally, the output layer consists of physical actuators from a multi-agent cluster, responsible for receiving neuromorphic sparse control commands from the core control layer and executing physical movements. The latest physical state generated after execution will serve as state feedback, flowing back to the system input layer, thus forming an end-to-end dynamic closed-loop control system.
[0041] 2. Neuromorphic Nonlinear Dynamic System Identification Algorithm Multi-robot systems based on distributed control strategies lack a central controller; each agent can make local decisions based solely on its own information and that of its neighbors. In the leaderless scenario, the communication topology of the agent cluster is an undirected graph. It means that, among them A set of nodes represents an intelligent agent. The edge set is an information flow. The weighted adjacency matrix represents the communication strength. Consider a set of... A system composed of heterogeneous or homogeneous intelligent agents, the first The dynamic equations of an agent can be described as the following second-order nonlinear system: (1) in These are the position and velocity state vectors, respectively. To control the input, and Given the matrix of the linear part of the system, For internal, unknown nonlinear terms in a system model, which are often difficult to obtain directly, it is more advisable to use function approximation techniques to approximate them. For unknown external disturbances to the system, it is generally assumed that their energy is constrained but their specific form is unknown.
[0042] Assuming that the multi-agent system involved in this invention includes a leader, whose reference trajectory is as follows: To achieve the system's control objective, a distributed tracking controller should be designed. This makes the state of all agents... and Ultimately, it can gradually track the leader. ,Right now: (2) (3) Traditional state observers are prone to phase lag and feedback amplification of high-frequency observation noise when dealing with high-dimensional nonlinear coupled friction or complex aerodynamic disturbances. To maximize the approximation of system nonlinearity, this invention introduces reservoir computational dynamics. The core of the RC system lies in the reservoir, and its mathematical essence is highly consistent with Koopman operator theory. That is, by implicitly elevating a finite-dimensional nonlinear dynamic system to a high-dimensional linear feature space, linear solvability reconstruction of complex nonlinear mappings is achieved. It is a fixed but random recurrent network composed of a large number of sparsely connected neurons, whose internal connection weights are randomly generated and fixed. This random, high-dimensional dynamic system can map the input nonlinear time-series signal to a complex high-dimensional state space, whose state contains historical information of the input. This invention mainly applies the echo state network, one of the reservoir computational paradigms in neuromorphic computing.
[0043] Echoing-state networks achieve their advantages by using a relatively large, sparsely connected library of neurons employing the sigmoid transfer function. The connections within the reservoir are assigned completely randomly, but not unordered; the internal weights of the reservoir are not trained, only the output weights connecting the reservoir and the output neurons are trained, significantly reducing the training load on the neural network. During training, the reservoir state is captured and stored over time. The trained output weights can be integrated into the existing network and used to process new inputs. Its specific structure is as follows... Figure 1 As shown.
[0044] Generally speaking, a discrete-time recurrent neural network can be described as a graph structure with three sets of nodes, i.e. input nodes , Internal network nodes and Output nodes At the time point The activation vector is represented as: , , The weights of the interconnect edges within the reservoir are: It is represented in the form of an adjacency matrix. When This indicates that the node is... To the node There exists a path. Definition Input weights into the neural network, For the internal connection weights of the neural network, Output weights for the neural network. These are the weights fed back to the neural network from the output nodes, where the subscripts indicate the dimension. Here, it is assumed that the reservoir contains... The evolution of the high-dimensional activation field vector of the reservoir, involving a number of neurons, is first presented through the following discrete-time difference equation: (4) in This represents the instantaneous activation state of a neuron mapped to a high-dimensional feature space. The input vector for the neural network; The reservoir internal connectivity weight matrix is typically sparse and randomly generated, and must satisfy the spectral radius requirement. To ensure the echo state attribute; This is the input weight matrix. It is the output feedback weight matrix; The leakage rate represents the rate at which the neuron state is updated, simulating the inertia of a physical system. The output equation, i.e., the linear readout layer, is defined as: (5) in The weight matrix is the only one that needs to be trained, and the input vector and bias constant are usually also connected to the linear readout layer to enhance the expressive power of the neural network.
[0045] To maximize the utilization of reservoir nonlinear dynamics, this invention constructs a neuromorphic nonlinear dynamic system identifier based on a high-dimensional feature manifold projection, defining the total system perturbation as: (6) The core objective of the identifier is to Under the condition that the mathematical analytical expression is completely unknown, the local spatiotemporal input signal sequence is used to approximate and reconstruct the expression in real time. To achieve precise feedforward compensation. The structural block diagram is as follows: Figure 2 As shown.
[0046] This identifier is no longer a simple state mapping, but rather it analytically reconstructs the error dynamics of the low-dimensional physical system by mapping it to a high-dimensional feature space of the neuromorphic system through a nonlinear differential manifold. Assuming that leader information is available, the reservoir input spatiotemporal tensor is defined as follows: It can be in the following form: (7) in This is a spatial adjacency mapping tensor that characterizes the topological connection strength of local multi-agent systems. Furthermore, to adapt to the Lyapunov analysis framework of continuous physical systems, equation (4) can be transformed into a forced nonlinear partial differential continuity equation with long-range spatiotemporal characteristics: (8) in Represents the reservoir dynamic activation field in a high-dimensional characteristic manifold; The diagonal matrix representing the time constant of neuronal membrane potential leakage determines the decay rate of high-dimensional memory manifolds; The truncated remainder term is excited by the higher-order input to absorb nonlinear mutations in extreme cases; The cell membrane potential time constant is the absolute time scale used to regulate the reservoir network's tracking of changes in input signals.
[0047] To prove the inherent stability of the identifier within a rigorous mathematical framework, a local and global Jacobian matrix spectral analysis of its dynamic equations is necessary. Here, the state-space equations of an unforced continuous reservoir at the equilibrium point are analyzed. Expanding on the vicinity, its Jacobian matrix has the following analytical form: (9) Based on the Lyapunov direct method for nonlinear system stability analysis, the positive definite quadratic Lyapunov candidate function is selected as follows: (10) Then calculate its time derivative along the system trajectory: (11) From the above inequality, we can conclude that the derivative of the activation function is strictly bounded to the interval [0, 1]. If suitable scalar multipliers can be selected to make the internal weight matrix... If the spectral radius is strictly constrained, all eigenvalues of the Jacobian matrix will be forcibly stretched to the left half of the complex plane. In this case, regardless of the initial perturbation of the system, as long as the parameters are configured within the feasible region, the high-dimensional characteristic manifold within the reservoir will eventually converge to a bounded forced evolution trajectory, thus preventing numerical divergence.
[0048] When low-dimensional errors and state variables are mapped to a stable high-dimensional reservoir space, the total unknown disturbance of the system... This can be viewed as a linear combination of high-dimensional basis vectors. The readout layer only needs to solve a convex optimization projection problem from high-dimensional to low-dimensional, and its reconstruction formula includes three important parts: linear mapping, long-term memory integration, and smooth manifold curvature correction. (12) in This is the gain coefficient, combined with the attenuation factor. An integral operator with a forgetting mechanism is formed to compensate for constant static deviations or low-frequency disturbances; This is the Laplace manifold surface curvature smoothing term. It is equivalent to an embedded spatial low-pass filter, used to forcibly suppress nonlinear high-frequency noise spikes that may appear in high-dimensional feature reconstruction, preventing high-frequency oscillations from being introduced into the control law.
[0049] 3. Training Algorithm for Neural Modulation Reward and Punishment Synaptic Readout Layer Based on Biological Brain Dopamine Mechanism In training linear readout layers, traditional methods aim to minimize the norm of the output error. This is a passive approach and lacks consideration for long-term cumulative effects and macroscopic performance optimization capabilities. To endow the system with brain-like intelligence that actively explores and comprehensively optimizes, this invention proposes a neural modulation reward-punishment synaptic readout layer training algorithm based on the biological brain dopamine mechanism at the end of the reservoir computing framework. The structural diagram is as follows: Figure 3 As shown, this mechanism, through the three-factor synaptic plasticity principle, deeply couples presynaptic and postsynaptic activities with dopamine signals representing the global reward and punishment environment evaluation. By simulating the synaptic weight adjustment process of biological neurons under positive and negative feedback from the external environment, the reservoir computing system possesses autonomous optimization capabilities. The trained readout layer weights are considered as synaptic connections in the brain; the algorithm no longer merely pursues the minimization of error, but rather the maximization of the long-term accumulated comprehensive reward and punishment signal.
[0050] In the basal ganglia and cortical striatum pathways of the brain, synaptic remodeling depends not only on instantaneous firing matching, but also on generating a transient and decaying biochemically labeled eligibility trace to link causal relationships across time scales. Mathematically, eligibility traces provide a natural mechanism for resolving the time allocation problem between actions and delayed rewards. Now, combining impulse temporal-dependent plasticity, a readout layer is constructed... The first output and reservoir Synaptic connections between neurons The time-varying continuous differential equation: (13) in The natural exponential decay time constant of the qualified trace; Presynaptic activity represents the instantaneous concentration of neuronal input pulses; This is a postsynaptic local error driver, indicating the current bias sensitivity of the execution terminal; Represents the synaptic plasticity window function; For Dirac-delta function, it represents the pulsed impact of discrete spike excitation times on the continuous qualified trace.
[0051] Brain-like mechanisms require the external environment to provide a global evaluation scalar that can comprehensively consider various control indicators. Inspired by the value function and multi-attribute objective optimization in reinforcement learning, we define a comprehensive environmental reward evaluation functional to drive the release of virtual dopamine: (14) The first term is the reward, representing positive dopamine, which encourages the system to track the target trajectory quickly and accurately. The exponential function ensures that the smaller the error, the faster the reward increases, forming a positive incentive gradient. The second term is a soft constraint, penalizing drastic fluctuations in the state, encouraging the system to generate smooth, low-jitter trajectories, and avoiding damage to the mechanical structure or excessive energy consumption due to frequent and drastic acceleration and deceleration. The third term is a direct penalty for control energy consumption, forcing the learning algorithm to use as little control energy as possible while maintaining a certain tracking accuracy, thereby achieving low-power neuromorphic control. Its working principle is that the smaller the system tracking error, the larger the reward; when speed fluctuations are drastic or control energy consumption is too high, the penalty and action cost increase. An increase in signal strength leads to a negative overall signal, triggering a penalty mechanism. For nonlinear activation modulation functions, hyperbolic tangent or activation functions with dead zones are generally used to control the intensity and direction of weight updates: when , At this point, the weight update direction is the same as the gradient descent direction. Similarly, strengthen the current update direction that brings rewards; when , This will inhibit the direction of updates that lead to penalties; All of these are sensitivity adjustment coefficients for the reward and punishment mechanism. This mechanism indicates that the weight training of the readout layer is no longer simply fitting the data, but actively exploring synaptic connection states in a multidimensional space that can maximize long-term cumulative rewards, which significantly helps to enhance the system's robustness and learning autonomy.
[0052] striatum Type II dopamine receptors primarily mediate the direct pathway; stimulation with high concentrations of dopamine induces long-term potentiation (LTP), i.e., positive learning. Type I receptors mediate indirect pathways, releasing inhibition upon decrease in dopamine concentration and inducing long-term inhibition (LTD) to eliminate erroneous pathways. Based on this cortical biological model, the readout layer matrix... Weight elements The learning and updating equations are synthesized into the following formula: (15) in The learning rate is for adaptive annealing; Represents the three-factor core product operation; The noise term for neural synapse detection is essentially a Gaussian random perturbation whose variance is proportional to the error of the current state. This ensures the network's ability to proactively explore the globally optimal solution; The dead-zone nonlinear activation modulation gate function is defined as follows: (16) Among them The dead zone threshold mentioned above can prevent unnecessary continuous weight drift.
[0053] Through the aforementioned neurobiological algorithm reconstruction, this multi-agent control system is a decision-making ecosystem with autonomous environmental cognition, trend prediction, experience accumulation, and long-term planning capabilities.
[0054] 4. Adaptive Topology Communication Triggering Algorithm Based on Reservoir Spatiotemporal Prediction Entropy Since continuous broadcasting of state information in distributed systems consumes significant bandwidth and energy, this invention, inspired by the principle that biological neurons only fire spikes when necessary, designs an adaptive topology communication triggering mechanism based on reservoir spatiotemporal prediction entropy. This mechanism not only considers the errors in the system's physical state but also deeply analyzes the degree of chaos in the high-dimensional states within the reservoir, using this as an intrinsic basis for determining whether communication is necessary. The specific structural diagram is shown below. Figure 4 As shown.
[0055] Here, the normalized probability distribution of the activation vector within the reservoir is defined as: (17) in For the first The activation probability of each neuron; For the first The instantaneous activation value of a neuron; for The sum of the activity levels of all neurons within the reservoir at time t is given. Then the spatiotemporal prediction information entropy of the reservoir state at the current time is: (18) The agent only at time Communicate, next trigger time The following multidimensional topology criteria determine the outcome: (19) The triggering energy functional Designed as follows: (20) Its function is to measure the deviation energy difference between the agent's current state and the previously transmitted state; dynamic threshold function. Designed as follows: (twenty one) Its value is dynamically adjusted according to the entropy value of the reservoir's internal state, among which... This is the real-time output of a neuromorphic nonlinear dynamic system identifier. When the external environment experiences severe disturbances, the activation state of neurons within the reservoir tends to shift from ordered to chaotic, leading to a time-varying rate of the predicted information entropy. The dynamic threshold increases dramatically. At this point, the dynamic threshold decays exponentially, forcing the agent to frequently broadcast high-dimensional features to its neighbors to quickly reconstruct knowledge sharing on the communication topology. During the stable period, the information entropy tends to be constant, and the system operates in normal mode. The dynamic threshold will then adaptively relax.
[0056] 5. State switching control algorithm based on virtual energy threshold Brain-inspired sparse pulse actuation involves cutting off continuous signals when the system reaches a steady state, instead relying on periodic or event-triggered short pulses to maintain posture. Beyond traditional continuous nominal control and observation compensation, to truly align with the event-driven and low-power characteristics of neuromorphic hardware, this invention introduces the brain-inspired Spiking sparse motion execution concept into the composite controller. In biological motion control, when a limb is in a steady state, the brain no longer sends continuous, dense muscle control signals, but only occasionally releases discrete electrical pulses to maintain posture, thereby significantly reducing energy consumption and actuator wear. Based on this principle, the final system control law consists of three parts: a nominal consistency controller... Used for stabilizing systems and tracking trajectories, it is a neuromorphic compensator based on network topology and a cooperative stabilizing term. The interference term is directly canceled out by an online compensation term based on neurodynamics and a pulse firing gating state term. The individual controllers and the composite controller are defined as follows: (twenty two) (twenty three) (twenty four) in , These are the proportional gain and the derivative gain, respectively. Adjacency weight; Define the agent to provide access gain for the leader. Being able to directly observe the virtual leader would make ,otherwise ; The real-time output of the reservoir network is a real-time reconstruction signal of the total system disturbance generated by the reservoir computing module based on the state error sequence; the pulse firing condition is defined as follows: the virtual neuron state space constructed inside the controller has a virtual energy accumulation variable. Its differential expression is: (25) The system generates an impulse moment if and only if the system error is within the stability threshold and the virtual energy accumulation exceeds the trigger threshold. .
[0057] In the composite controller The pulse delivery gating state is set, and When the system is in a phase where the dynamic tracking error is large... At this point, the system adopts continuous composite control; when the error reaches the steady-state threshold, the system will switch to pulse sleep mode, which occurs if and only if the virtual energy accumulation variable in the virtual neuron's state space exceeds the trigger threshold. Within a very short time window The internal activation is set to 1, and discrete action pulses are emitted. This design pattern transforms the continuous output of the actuator into a pulse-driven mode with an extremely low duty cycle, effectively avoiding damage to the physical hardware caused by the impulse function, and can truly achieve end-to-end low power consumption.
[0058] The pseudocode for the entire algorithm flow is as follows:
[0059] Figure 5 A block diagram is shown for the design of a multi-agent brain-like tracking control algorithm based on reservoir computing.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multi-agent brain-like tracking control system based on reservoir computing, characterized in that: It includes a state awareness layer, a reservoir computing core layer, a composite control layer, and an execution output layer connected in sequence; The state awareness layer is used to acquire the state of each agent in the multi-agent system and the state information of neighboring agents, and to calculate the tracking error with the navigator. The reservoir computing core layer incorporates a neuromorphic nonlinear dynamic system identifier, which is used to map the tracking error to a high-dimensional feature space to reconstruct the total system disturbance, and generate a dynamic communication trigger signal based on the spatiotemporal prediction information entropy of the reservoir state. The composite control layer is used to generate continuous or discrete neuromorphic sparse control instructions based on the tracking error, the reconstruction result of the total system disturbance, and a preset virtual energy threshold. The execution output layer is used to receive and execute the neuromorphic sparse control instructions, and feed back the physical state after execution to the state perception layer to form a closed-loop control.
2. The multi-agent brain-like tracking control system based on reservoir computing according to claim 1, characterized in that: The reservoir computing core layer includes a reservoir network, a readout layer weight training module, and a communication trigger determination module. The reservoir network is an echo state network, whose internal structural weights are randomly generated and fixed, and whose state evolution follows the following discrete-time difference equation: in, Let k be the instantaneous activation state of the neuron mapped to the high-dimensional feature space. This represents the total number of reservoir neurons. for k The neural network input vector at time step; The reservoir internal connectivity weight matrix has a spectral radius. ; The input weight matrix; To output the feedback weight matrix; Leakage factor; The readout layer weight training module updates the output weights using a neural modulation reward-punishment synapse training algorithm based on the biological brain dopamine mechanism. ; The communication trigger determination module calculates the trigger energy functional. With dynamic threshold function The relationship determines whether communication is triggered.
3. The multi-agent brain-like tracking control system based on reservoir computing according to claim 2, characterized in that: In the communication trigger determination module, the trigger energy functional is... The calculation formula is: in, , These are the weighting coefficients. For intelligent agents i exist t The state vector at time t, For intelligent agents i Last communication time The state vector, Forgetting factor; Dynamic threshold function The calculation formula is: in, As the initial threshold, , , For adjustment coefficients, The spatiotemporal prediction information entropy of reservoir state, This is the real-time reconstructed signal for the total system disturbance; Next trigger time satisfy At that time, intelligent agent i Broadcast status information to neighbors.
4. The multi-agent brain-like tracking control system based on reservoir computing according to claim 1, characterized in that: The neuromorphic sparse control instructions generated by the composite control layer It is composed of the nominal consistency control law, the neuromorphic compensation term, and the impulse firing gating state term, and its calculation formula is as follows: in, The pulse delivery gating state is set. This is a nominal consistency control law. This is a neuromorphic compensation term.
5. The multi-agent brain-like tracking control system based on reservoir computing according to claim 4, characterized in that: The pulse delivery gating state The value is determined by the virtual energy storage variable. The virtual energy accumulation variable is determined. The differential expression is: in, Leakage coefficient, This is the gain coefficient. This is the total disturbance reconstruction signal output in real time by the reservoir network. For tracking error; When the tracking error does not reach the preset steady-state threshold The system performs continuous control; when the tracking error enters the preset steady-state threshold and When the preset trigger threshold is exceeded, Within the preset time window The internal activation is set to 1 and discrete action pulses are emitted.
6. A multi-agent neuromorphic tracking control method based on reservoir computing, characterized in that: Includes the following steps: S1: Obtain the agents in the multi-agent system through the state-aware layer. i its own position ,speed And the state information of neighboring intelligent agents, calculating the navigator's trajectory. Tracking error ; S2: The tracking error Input the reservoir core layer and reconstruct the total system perturbation using a neuromorphic nonlinear dynamic system identifier. And based on the spatiotemporal prediction information entropy of reservoir state Determine whether the communication mechanism has been triggered; S3: The composite control layer determines the tracking error based on the aforementioned tracking error. The total disturbance of the system and virtual energy accumulation variables Generate brain-like sparse control instructions ; S4: The output layer receives and executes the neuromorphic sparse control instructions. The physical state after execution is fed back to S1 to form a closed-loop control.
7. The multi-agent neuromorphic tracking control method based on reservoir computing according to claim 6, characterized in that: In S2, the neuromorphic nonlinear dynamic system identifier dynamically maps the errors of the low-dimensional physical system to a high-dimensional feature space, reconstructing the total perturbation of the system. The calculation formula is: in, for t The readout layer weight matrix at time step [time]. for t The high-dimensional activation field of the reservoir at any given moment This is the gain coefficient. As the attenuation factor, For the Laplace manifold surface curvature smoothing term.
8. The multi-agent neuromorphic tracking control method based on reservoir computing according to claim 6, characterized in that: In S2, the reservoir computing core layer uses a neural modulation reward-punishment synaptic training algorithm based on the biological brain dopamine mechanism to update the readout layer weights. The algorithm aims to maximize the long-term cumulative reward and punishment signal and adjusts the weights through a three-factor synaptic plasticity law. The three factors include presynaptic activity, postsynaptic activity, and the dopamine signal, which represents the evaluation of the global reward and punishment environment.
9. The multi-agent neuromorphic tracking control method based on reservoir computing according to claim 6, characterized in that: In S3, the neuromorphic sparse control instructions nominal consistency control law The calculation formula is: in, For the proportional term gain, For the differential term gain, For intelligent agents i The neighborhood group, For intelligent agents i and j Adjacency weights between them To provide navigator access gains, when the intelligent agent i When the navigator can be directly observed ,otherwise .
10. The multi-agent neuromorphic tracking control method based on reservoir computing according to claim 6, characterized in that: In S2, the spatiotemporal prediction information entropy of the reservoir state The calculation formula is: in, This represents the total number of reservoir neurons. For the first k The activation probability of a neuron is calculated based on the probability distribution of the activation states of high-dimensional neurons within the reservoir.