Intelligent Control Method and System for Nonlinear Systems Based on Sustainable Deterministic Learning

CN122568993APending Publication Date: 2026-08-14SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0007]针对现有技术中固定结构神经网络覆盖能力有限、多任务持续学习易导致知识遗忘、历史动力学知识难以复用等不足,本发明提供了一种基于可持续确定学习的非线性系统智能控制方法及系统,通过反步法构造误差系统,利用径向基函数神经网络在线逼近未知动力学并动态扩展节点,结合确定学习固化知识并迁移至新任务,同时引入增强节点补偿项在线补偿历史知识退化,解决了多任务轨迹跟踪中的网络自适应扩展、知识迁移与抗遗忘问题,实现了未知非线性系统在多任务场景下的连续学习与高精度控制

Benefits of technology

(1)本发明将持续学习与确定学习相结合,通过反步控制构造误差系统、径向基函数神经网络在线逼近未知动力学、自适应增量节点生成机制动态扩展网络结构,实现了非线性系统动力学知识在多任务轨迹间的持续积累、长期存储与重复利用。相较于传统自适应神经网络控制需要针对不同任务重复在线学习,本发明能够将历史任务中学到的动力学知识以常值权值形式固化并直接用于新任务参数初始化,显著避免了重复学习问题,提高了学习效率并降低了计算资源消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122568993A_ABST
    Figure CN122568993A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent control method and system for nonlinear systems based on sustainable deterministic learning, belonging to the field of intelligent control and deterministic learning technology for nonlinear systems. Addressing the problems of limited fixed network coverage and the tendency for historical knowledge to degrade and be difficult to reuse in existing deterministic learning control methods, this invention constructs state error and a virtual control law based on backstepping; uses the system state and the derivative of the virtual control law as inputs to a radial basis function neural network to approximate the unknown closed-loop dynamic function; dynamically generates a node extension network based on the input activation degree; uses deterministic learning to converge the weights to constant values ​​and stores them for initializing parameters for new tasks; constructs enhanced node compensation terms superimposed on the control law and establishes its adaptive update law to compensate for historical knowledge degradation online. This invention is applicable to high-precision tracking control of systems such as UAVs, robots, and robotic arms under multi-task trajectories, achieving continuous learning, knowledge reuse, and resistance to forgetting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control and deterministic learning technology for nonlinear systems, and in particular to an intelligent control method and system for nonlinear systems based on sustainable deterministic learning. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Complex nonlinear systems are widely found in fields such as aircraft control, robotics systems, industrial automation, and intelligent equipment. Due to the inherent characteristics of these systems, including strong nonlinearity, parameter uncertainty, and external disturbances, traditional model-driven control methods struggle to achieve ideal control performance. Radial basis function neural networks (RBFNNs), with their excellent nonlinear approximation capabilities, are widely used in adaptive neural network control of nonlinear systems. Existing methods typically utilize RBFNNs to approximate the dynamics of unknown systems online, compensating for system uncertainties and ensuring closed-loop stability and tracking performance.

[0004] However, traditional adaptive neural network control methods primarily focus on control performance, neglecting the neural network's dynamic learning and knowledge representation capabilities. In practical applications, neural network parameters need to be repeatedly learned online for different tasks. When the system executes the same or similar trajectories again, it is difficult to directly reuse existing learning results, leading to low learning efficiency and high computational resource consumption. The fundamental reason is that the accurate learning of nonlinear dynamics by neural networks usually depends on continuous excitation conditions, which are difficult to strictly satisfy in actual control processes, thus limiting the stable acquisition, storage, and reuse of dynamic knowledge.

[0005] To address the aforementioned issues, deterministic learning theory was proposed. This theory states that, under continuous excitation conditions, the weights of an RBF neural network can converge to constant values ​​along the system trajectory, thereby achieving accurate learning of local closed-loop dynamics and storing reusable dynamic knowledge as constant neural network parameters. Based on this characteristic, deterministic learning methods have been widely studied in fields such as robot control, fault diagnosis, formation control, and distributed control. Although existing deterministic learning methods can learn and store trajectory-related dynamic knowledge, most studies are still limited to single-trajectory learning scenarios and employ fixed-structure neural networks. When the system state exceeds the coverage range of the preset neural network nodes, its approximation performance is difficult to guarantee, often requiring redesign of the network structure and relearning, making it difficult to adapt to multi-task continuous learning scenarios. Furthermore, existing methods generally lack mechanisms for the continuous accumulation, fusion, and reuse of historical dynamic knowledge, making it difficult to achieve cross-task knowledge transfer and learning acceleration.

[0006] On the other hand, continuous learning aims to enable learning systems to acquire new knowledge continuously during successive tasks, while maintaining and reusing historical knowledge, avoiding starting from scratch every time a new task is encountered. However, in the field of neural network control, combining continuous learning with deterministic learning still faces several key challenges: First, updating neural network parameters during the learning process of a new task can easily destroy existing dynamic knowledge, leading to knowledge forgetting; second, different tasks may have different local dynamic characteristics, making it difficult to uniformly represent them with a fixed network structure; third, how to effectively accumulate, integrate, and reuse dynamic knowledge across multiple tasks to improve the learning efficiency and control performance of new tasks remains a challenge. Summary of the Invention

[0007] To address the shortcomings of existing technologies, such as limited coverage of fixed-structure neural networks, susceptibility to knowledge forgetting during multi-task continuous learning, and difficulty in reusing historical dynamic knowledge, this invention provides an intelligent control method and system for nonlinear systems based on sustainable deterministic learning. It constructs an error system using the backstepping method, utilizes a radial basis function neural network to approximate unknown dynamics online and dynamically expands nodes, combines deterministic learning to solidify knowledge and transfer it to new tasks, and introduces enhanced node compensation terms to compensate for historical knowledge degradation online. This solves the problems of network adaptive expansion, knowledge transfer, and anti-forgetting in multi-task trajectory tracking, enabling continuous learning and high-precision control of unknown nonlinear systems in multi-task scenarios.

[0008] On the one hand, a method for intelligent control of nonlinear systems based on sustainable deterministic learning is provided, including: A nonlinear system dynamics model is established, and the system state error and virtual control law are constructed based on the backstepping control method. Using the system state and the derivative of the virtual control law as inputs to the radial basis function neural network, the unknown closed-loop dynamic function is approximated online. Based on the approximation results and the system state error, the neural network adaptive control law and weight update law are constructed. Calculate the maximum activation level of existing nodes in the neural network. When the maximum activation level is lower than a preset threshold, generate a new node and add it to the neural network. Different task trajectories are learned to make the neural network weights converge to constant values ​​and be saved as dynamic knowledge. The dynamic knowledge is then used to initialize the parameters of new tasks. An enhanced node compensation term is constructed, which is generated by multiplying the Gaussian function value by an adaptive weight after random mapping and nonlinear activation. The compensation term is superimposed on the adaptive control law of the neural network, and an adaptive update law for the compensation term is established to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

[0009] Furthermore, based on the approximation results and the system state error, a neural network adaptive control law and a weight update law are constructed, specifically as follows: The derivatives of the system state and the last layer of virtual control law are used as the input vectors of the radial basis function neural network. The Gaussian function output of the neural network is used to approximate the unknown nonlinear function in the closed-loop dynamics. The approximation result is substituted into the backstepping control law to replace the unknown function term, and combined with the system state error of the last layer, a neural network adaptive control law is constructed. The weight update law is designed based on the system state error and the Gaussian function value.

[0010] Furthermore, the maximum activation level of the existing nodes in the neural network is calculated as follows: Calculate the Gaussian activation function for each existing radial basis function node under the current input vector, and take the maximum value. Preset threshold Where b is the minimum node spacing parameter, d th Set the node coverage threshold distance; when When the time comes, a new node center is generated based on the relative position of the current input vector and the nearest node center, and the distance between the new node and all existing nodes is greater than 2b. Then the new node is added to the neural network.

[0011] Furthermore, deterministic learning is performed on different task trajectories to bring the neural network weights to constant values ​​and store them as dynamic knowledge, specifically: Under periodic reference trajectory excitation, the system satisfies the continuous excitation condition, and the neural network weight estimation error converges exponentially. The weight estimates within the stable time period after convergence are averaged in the time domain to obtain constant weights, which are then stored as the dynamic knowledge corresponding to the task.

[0012] Furthermore, the dynamic knowledge is used to initialize the parameters of the new task, specifically: for the first task, the neural network weights are initialized to zero; for subsequent tasks, the constant weights saved in the previous task are used as the initial weights of the neural network for the current task.

[0013] Furthermore, the enhanced node compensation term specifically includes: ( ); in, For a randomly initialized linear mapping matrix, For bias vectors, To enhance the adaptive weights of nodes, The hyperbolic tangent activation function is used. Let be the Gaussian activation function for the k-th trajectory task, where k represents the k-th trajectory task.

[0014] Furthermore, the weights of the enhanced nodes are updated using the following adaptive law: ; in, To enhance node learning gain, As a robust leakage factor, Let be the system state error of order n.

[0015] On the other hand, a nonlinear system intelligent control system based on sustainable deterministic learning is provided, including: The model building module is configured to: establish a nonlinear system dynamics model and construct the system state error and virtual control law based on the backstepping control method; The neural network learning module is configured to: use the system state and the derivative of the virtual control law as input to the radial basis function neural network, approximate the unknown closed-loop dynamic function online, and construct the neural network adaptive control law and weight update law based on the approximation result and the system state error; The node expansion module is configured to: calculate the maximum activation level of existing nodes in the neural network, and generate new nodes and add them to the neural network when the maximum activation level is lower than a preset threshold; The knowledge storage and transfer module is configured to: perform deterministic learning on different task trajectories, cause the neural network weights to converge to constant values ​​and save them as dynamic knowledge, and use the dynamic knowledge to initialize the parameters of new tasks; The enhancement compensation module is configured to: construct enhancement node compensation terms, which are generated by multiplying the Gaussian function value by an adaptive weight after random mapping and nonlinear activation; superimpose the compensation terms into the neural network adaptive control law, and establish an adaptive update law for the compensation terms to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

[0016] In another aspect, a computer device is also provided, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, when the processor executes the program, it performs the method described in the first aspect.

[0017] In another aspect, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, performs the method described in the first aspect.

[0018] The above technical solution has the following advantages or beneficial effects: (1) This invention combines continuous learning with deterministic learning. By constructing an error system through backstepping control, approximating unknown dynamics online through a radial basis function neural network, and dynamically expanding the network structure through an adaptive incremental node generation mechanism, it achieves continuous accumulation, long-term storage, and reuse of nonlinear system dynamics knowledge across multiple task trajectories. Compared to traditional adaptive neural network control, which requires repeated online learning for different tasks, this invention can solidify the dynamics knowledge learned from historical tasks in the form of constant weights and directly use it for initializing parameters for new tasks. This significantly avoids the problem of repeated learning, improves learning efficiency, and reduces computational resource consumption.

[0019] (2) The adaptive incremental node generation mechanism proposed in this invention can automatically expand the RBF neural network structure according to the system trajectory, thereby overcoming the defect of insufficient coverage of unknown state space regions by fixed structure RBF neural networks and ensuring the approximation accuracy of the network in continuous multi-task scenarios.

[0020] (3) The experience-enhanced node compensation term constructed in this invention generates compensation features through random mapping and nonlinear activation and adjusts its weights online, so that the compensation output can approximate the deviation between historical knowledge and current task knowledge in real time, effectively suppressing the forgetting of historical knowledge caused by parameter updates during continuous learning, and realizing the synergistic integration of historical dynamic knowledge and current task knowledge.

[0021] (4) While ensuring the stability of the closed-loop system, the present invention achieves high-precision trajectory tracking control of unknown nonlinear systems and has good continuous learning and knowledge reuse capabilities.

[0022] (4) This invention is applicable to trajectory tracking and intelligent control scenarios in unmanned aerial vehicle systems, robot systems, electromechanical servo systems and other complex nonlinear dynamic systems, and has strong engineering application value. Attached Figure Description

[0023] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0024] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 2 A schematic diagram of the overall framework of the method in Embodiment 1 of this invention; Figure 3 This refers to the tracking and control performance during the continuous learning process of different task trajectories in Embodiment 1 of the present invention, wherein... Figure 3 (a) in the figure is the tracking control performance diagram of trajectory A. Figure 3 In the diagram, (b) represents the tracking control performance of trajectory B. Figure 3(c) in the figure represents the tracking control performance of trajectory C. Figure 3 (d) in the figure represents the tracking control performance of trajectory D; Figure 4 This is a schematic diagram illustrating the incremental node generation during the continuous learning process of different task trajectories in Embodiment 1 of the present invention, wherein... Figure 4 (a) in the figure shows the node distribution during the learning process of trajectory A. Figure 4 (b) in the figure shows the node distribution during the learning process of trajectory B. Figure 4 (c) is a diagram showing the node distribution during the learning process of trajectory C. Figure 4 (d) is a diagram showing the node distribution during the learning process of trajectory D; Figure 5 This is a convergence graph of RBFNN weights during the continuous learning process of different task trajectories in Embodiment 1 of the present invention, wherein... Figure 5 In the diagram, (a) is the convergence plot of the RBFNN weights under trajectory A. Figure 5 (b) in the figure is the RBFNN weight convergence plot under trajectory B; Figure 5 (c) in the graph is the RBFNN weight convergence plot under trajectory C. Figure 5 (d) in the figure represents the convergence plot of the RBFNN weights under trajectory D. Figure 6 The approximation performance of RBFNN during the continuous learning process of different task trajectories in Embodiment 1 of the present invention is shown below. Figure 6 (a) in the figure is a comparison diagram of the actual dynamics and approximate dynamics of the system under trajectory A. Figure 6 (b) in the figure is a comparison diagram of the actual dynamics and approximate dynamics of the system under trajectory B. Figure 6 (c) in the figure is a comparison diagram of the actual dynamics and approximate dynamics of the system under trajectory C. Figure 6 (d) in the figure is a comparison diagram of the actual dynamics and approximation dynamics of the system under trajectory D; Figure 7 This invention presents a comparison of the approximation performance of RBFNN in the learning process between continuous deterministic learning and traditional deterministic learning in Embodiment 1 of the present invention. Figure 7 (a) in the figure is a comparison of the dynamic approximation error under trajectory B. Figure 7 (b) in the figure is a comparison of the dynamic approximation error under trajectory C; Figure 8 This is a performance evaluation of the continuously deterministic learning control scheme proposed in Embodiment 1 of the present invention, wherein, Figure 8 (a) shows the tracking performance under five different trajectories. Figure 8 (b) in the figure is a comparison of tracking errors under five different trajectories. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. Those skilled in the art should understand that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0026] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0027] Example 1 This embodiment provides an intelligent control method for nonlinear systems based on sustainable deterministic learning. Figure 1 This is an overall flowchart of the method according to Embodiment 1 of the present invention. The method includes the following steps: S101: Establish a dynamic model of the nonlinear system and construct the system state error and virtual control law based on the backstepping control method; S102: Using the derivatives of the system state and the virtual control law as inputs to the radial basis function neural network, the unknown closed-loop dynamic function is approximated online. Based on the approximation results and the system state error, the neural network adaptive control law and weight update law are constructed. S103: Calculate the maximum activation level of existing nodes in the neural network. When the maximum activation level is lower than a preset threshold, generate a new node and add it to the neural network. S104: Perform deterministic learning on different task trajectories to make the neural network weights converge to constant values ​​and save them as dynamic knowledge. Use the dynamic knowledge to initialize the parameters of new tasks. Continuously accumulate, integrate and reuse dynamic knowledge in the process of multi-task trajectory tracking to form a unified multi-task dynamic knowledge representation and realize continuous deterministic learning control. S105: Construct an enhanced node compensation term, which is generated by multiplying the Gaussian function value by an adaptive weight after random mapping and nonlinear activation; superimpose the compensation term into the neural network adaptive control law, and establish an adaptive update law for the compensation term to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

[0028] like Figure 2 As shown, Figure 2This is a schematic diagram of the overall framework of the intelligent control method for nonlinear systems based on sustainable deterministic learning provided in this embodiment of the invention. First, incremental RBF neural network learning is executed sequentially along the task sequence. After each task is completed, the learning is solidified into a constant RBF neural network to achieve layer-by-layer accumulation and unified storage of dynamic knowledge. Then, by reusing historical experience and fusing enhancement nodes, an experience-enhanced deep learning controller is constructed. Finally, the output control quantity is applied to the nonlinear system, and a closed loop is formed by performance evaluation feedback, thereby achieving high-performance tracking control along the sequence trajectory and continuous learning, accumulation, and reuse of dynamic knowledge.

[0029] Specifically, in step S101, a rigorous feedback dynamic model of the nonlinear system to be controlled is first established to lay the foundation for subsequent controller design. This model is specifically expressed as follows: ; ; in, Let be the system state vector. To control the input, Represents an unknown nonlinear dynamic function. This represents an unknown control gain function. satisfy And satisfy ,in, The given positive constant is .

[0030] Based on this model, for the first Each trajectory task constructs the system state error variable: ; ; in,, Represents a virtual control law. Indicates the first Reference trajectory for each task Let be the first-order state variable corresponding to the k-th task.

[0031] Subsequently, a closed-loop error system is constructed based on the backstepping control method. For the first-level error variables, a first virtual control law is constructed: ; in, To control the gain parameter; yes The first-order time differential variable.

[0032] For the Layer error variables, constructing the first layer error variables A virtual control law: ; in, To control the gain parameters, Indicates the first A virtual control law The first-order time differential variable.

[0033] Based on this, the error variables of the last layer (nth order) Its time derivative is: ; Through the design of the backstepping error system and virtual control law, the original nonlinear system is gradually transformed into an error subsystem that can be recursively stabilized, thereby transforming the unknown closed-loop dynamics into a learnable form. This provides a direct foundation for the subsequent design of an online approximation and continuous learning controller using radial basis function neural networks.

[0034] Step S102: Using the derivatives of the system state and the virtual control law as inputs to the radial basis function neural network, the unknown closed-loop dynamic function is approximated online. Based on the approximation results, an adaptive control law containing neural network weight estimation is initially constructed, laying the controller foundation for subsequent deterministic learning and knowledge solidification.

[0035] Step S103: To overcome the limited coverage of traditional fixed-structure RBF neural networks and ensure sufficient nonlinear approximation capability during continuous learning, this invention proposes an adaptive incremental node generation mechanism based on activation function thresholds to achieve dynamic coverage of newly visited state space regions. When the system trajectory exceeds the effective coverage area of ​​existing neural network nodes, new RBF nodes can be automatically generated based on the Gaussian activation function output and node spacing constraints, thereby achieving online expansion of the neural network structure. The proposed node generation mechanism ensures that newly added nodes satisfy local continuous excitation conditions, enabling the neural network to maintain effective approximation capability of newly visited state space regions and avoiding learning performance degradation caused by fixed node structures.

[0036] Specifically, the steps include the following: (1) Initialize the RBF neural network structure.

[0037] Define the center node of RBFNN as ,in The node number is represented by the input vector of the neural network. Set the initial center node to the system's initial input vector, i.e. The corresponding node weights are initialized to zero. Let the input vector be... Dimensions ,parameter This represents half of the minimum allowable spacing between any two RBFNN center nodes.

[0038] (2) Calculate the Gaussian activation function corresponding to the current input vector.

[0039] Based on the current input vector Calculate the Gaussian activation function for each existing node: in, For the first i Gaussian activation functions are represented as follows: ; in, and Let represent the center node and width in the Gaussian activation function, respectively. Then, take... The maximum value of the child elements in the expression is denoted as .

[0040] (3) Construct a node triggering threshold based on the activation function. Determine whether the current input exceeds the effective coverage area of ​​existing nodes, and define the activation threshold as: ; in: ; When satisfied When this occurs, it indicates that the current input vector is outside the coverage area of ​​all existing nodes, thus triggering the new node mechanism.

[0041] (4) Generate a new RBF center node. When the node triggering condition is met, based on the current input vector... and the nearest node center Relative positional relationships are used to generate new node centers. The nearest node center Gaussian activation function To obtain, represented as: ; Then the new node center It can be represented as: ; in, The scaling factor is designed as follows: ; in, The definition is as follows: ; in, and .

[0042] The generated new node satisfies two key conditions: the distance between the new node and the current input vector satisfies: Furthermore, the new node satisfies the minimum interval condition between it and existing nodes: This ensures that newly added nodes can effectively cover the newly accessed area while avoiding excessively dense node distribution.

[0043] According to deterministic learning theory, when RBF nodes satisfy the above-mentioned coverage and spacing conditions, the local regression vector can satisfy some of the continuous excitation conditions, thereby ensuring that the RBF neural network can accurately learn the local closed-loop dynamics near the trajectory.

[0044] Theoretical derivation can prove that for any recursive trajectory input: The proposed incremental node generation mechanism can guarantee that newly added nodes satisfy the following: ;as well as This ensures that the local activation vectors have continuous activation properties and maintains the approximation ability and learning accuracy of the neural network.

[0045] The proof is as follows: (a) Consider when hour:

[0046]

[0047] (b) Consider when hour:

[0048]

[0049] Through the aforementioned adaptive incremental node generation mechanism, this invention can expand the neural network structure online according to the system trajectory, achieve dynamic coverage of unknown state space regions, and provide a structural foundation for subsequent multi-task continuous learning and dynamic knowledge reuse.

[0050] Step S104: Based on the dynamically expanded RBF neural network structure in step S103, step S104 further utilizes deterministic learning theory to achieve accurate learning and storage of unknown closed-loop dynamics, and improves the learning efficiency of new tasks through knowledge transfer mechanism, so as to realize continuous learning, storage and reuse of dynamic knowledge under multi-task trajectory.

[0051] (1) Construct a continuously deterministic learning control law. For the first... For each task trajectory, an online approximation of the unknown closed-loop dynamic function is achieved using a radial basis function neural network. in, For the neural network input vector, For the first RBF basis function vectors of each task trajectory For the ideal weight vector, This is the bounded approximation error.

[0052] Further construct a continuous deterministic learning control law based on an RBF neural network: ; in, For the first The neural network estimated weights for each task. To control the gain parameter.

[0053] (2) Cross-task dynamics knowledge transfer and initialization. To achieve cross-task dynamics knowledge transfer and reuse of historical knowledge, this invention designs a dynamic knowledge initialization mechanism based on continuous learning. During the new task learning phase, the dynamics knowledge retained from historical tasks is used to initialize the current neural network parameters, thereby avoiding the repetitive learning problem in traditional neural network control methods and improving the learning speed and control performance for new tasks. Specifically, when executing the first task, the neural network weights are initialized as follows: ; This means using a zero-initial approach for dynamic learning.

[0054] When executing subsequent tasks When the task is completed, the constant weights saved after the previous task is completed are used as the initial weights for the current task. ; in, Indicates the first The dynamics knowledge matrix saved after each task is completed.

[0055] In this way, the local closed-loop dynamics knowledge learned in historical tasks is directly transferred to new tasks, thereby avoiding the need to relearn from scratch for each task and improving the learning speed and control performance of new tasks.

[0056] (3) Continuous Determination of Adaptive Law Design. In order to simultaneously achieve rapid convergence of the current task dynamics and the preservation of historical knowledge during online learning, the following adaptive law is designed: ; in, To learn the gain matrix, For robust leak items, The knowledge retention factor is designed as follows: ; in, .

[0057] In the above adaptive law, the first term The second item is used for online dynamics learning to achieve the current task. The third term is used to enhance the boundedness and robustness of weights. This method constrains the deviation of current task learning results from historical knowledge, ensuring the preservation and continuous evolution of historical dynamic knowledge and mitigating the problem of knowledge forgetting during continuous learning. Therefore, the proposed method can continuously accumulate, integrate, and reuse dynamic knowledge during multi-task trajectory tracking, thereby forming a unified multi-task dynamic knowledge representation and achieving continuous deterministic learning control.

[0058] This invention establishes a method for analyzing the stability of closed-loop systems based on Lyapunov stability theory, proving that all closed-loop signals are uniformly bounded under the proposed control framework. Simultaneously, by constructing a linear time-varying system model, it proves that the weights of the neural network can converge to constant values ​​under continuous excitation, thereby achieving accurate learning and long-term storage of unknown nonlinear dynamics.

[0059] First, to verify the closed-loop stability of the proposed continuous deterministic learning control method, a Lyapunov function is constructed to perform stability analysis on the system tracking error and neural network weight error.

[0060] Define the system's Lyapunov functions as follows: in, This represents the error in the weight estimation of the neural network.

[0061] Combining error systems, virtual control laws, continuously deterministic learning control laws, and adaptive update laws, for Differentiating the function and applying Young's inequality to the cross terms, we obtain: ; in, The stability coefficient, It is a bounded constant composed of the neural network approximation error, historical knowledge term, and robustness term.

[0062] According to Lyapunov's stability theory: (1) All closed-loop signals in the system are uniformly bounded; (2) Tracking error Bounded convergence; (3) Neural network weight estimation error Uniformly bounded; (4) The proposed continuous deterministic learning control system has semi-globally consistent eventual bounded stability.

[0063] At the same time, due to continuous learning items By introducing this approach, the system can retain knowledge of historical task dynamics while learning the dynamics of the current task, thereby mitigating the knowledge forgetting problem during multi-task continuous learning. Therefore, the proposed method can achieve stable control and continuous dynamic learning during multi-task trajectory tracking.

[0064] To further demonstrate that the neural network weights can converge to stable constant values ​​and to achieve the storage and reuse of dynamic knowledge, this invention further transforms the closed-loop system into a linear time-varying (LTV) system for analysis. First, a coordinate transformation is performed on the error variables: as well as By combining a localized RBF neural network structure, the closed-loop error system can be rewritten as a linear time-varying system that includes state errors and neural network weight errors: in, Includes system tracking error and neural network weight error. For time-varying system matrices, This is the bounded approximation error term. Further construction of a positive definite matrix is ​​needed. And prove that: This demonstrates that the constructed linear time-varying system satisfies uniform exponential stability.

[0065] Since the reference trajectory is periodic, under the proposed control framework, the system state, virtual control law, and RBF neural network regression vector all maintain periodicity. Therefore, the localized RBF neural network regression vector satisfies the partial continuous excitation condition (PEC). Based on deterministic learning theory, under the PEC condition, the neural network weight error converges to near the optimal constant weight. Therefore, the... The unknown closed-loop dynamics corresponding to each task can be represented by the weights of a stable constant neural network: in, Is it with The same small approximation error term. This represents the dynamics knowledge matrix saved after learning is completed.

[0066] Thus, this invention achieves: accurate learning of unknown nonlinear dynamics; stable storage of dynamic knowledge; continuous accumulation and reuse of dynamic knowledge among multiple tasks; and rapid learning and control of new tasks based on historical knowledge.

[0067] In step S105, in order to overcome the problem of degradation and forgetting of historical dynamic knowledge of neural networks during continuous learning, the present invention further designs an experience-enhanced deterministic learning control method (EEDLC) based on enhanced nodes to realize the dynamic fusion of historical knowledge and current task knowledge, thereby suppressing forgetting and improving control performance.

[0068] First, the final dynamic knowledge obtained after multi-task continuous learning is represented as: , in This represents the total number of tasks that have been completed. Based on this, a continuous deterministic learning control law is constructed, which includes compensation terms for augmented nodes: in, Used to access historical dynamics knowledge stored during the continuous learning phase. To indicate the first The dynamics knowledge matrix saved after each task is completed is transposed. Let Gaussian activation function be used for the k-th task. ( To enhance node compensation terms, The first The random initialization linear mapping matrix and bias vector of each task trajectory To enhance the adaptive weights of nodes, The hyperbolic tangent activation function is used. The control law superimposes traditional feedback terms, historical knowledge terms, and enhanced node compensation terms, with the compensation terms specifically used to compensate for the degradation of historical knowledge caused by learning new tasks online.

[0069] Because dynamic differences and knowledge forgetting may exist between different tasks during continuous learning, reinforcement nodes are introduced to compensate for the dynamic deviation between historical knowledge and the current task online, and a knowledge degradation term is defined: This represents the difference between historical dynamics and current task dynamics during continuous learning. Enhancement nodes are used to approximate knowledge degradation terms. in, To ideally enhance node weights, This is the bounded approximation error.

[0070] Further design to enhance the node adaptive update law:

[0071] in, To enhance node learning gain, It is a robust leakage factor.

[0072] Through the above design, the augmentation node can compensate for the degradation of dynamic information during continuous learning online, thereby improving the system's control performance for both historical and unknown tasks. Therefore, the proposed experience-enhanced continuous deterministic learning control method can simultaneously achieve: reuse of historical dynamic knowledge, compensation for forgetting dynamics, enhancement of multi-task control performance, and rapid adaptation to unknown tasks.

[0073] To verify the stability of the proposed continuous deterministic learning control method for augmented nodes, a Lyapunov function incorporating the weight error of the augmented nodes is further constructed: ; in, This represents the error in estimating the weights of the enhanced nodes. Further combining the error system, virtual control law, continuous deterministic learning control law, and enhanced node adaptive update law, we differentiate the Lyapunov function and apply Young's inequality for scaling, resulting in: in, The stability coefficient, It is a bounded constant consisting of the approximation error and the compensation error of the enhanced nodes.

[0074] Based on Lyapunov stability theory, we can further obtain: (1) All closed-loop signals in the system are uniformly bounded; (2) The system tracking error is consistent and eventually bounded; (3) Enhance the consistency and boundedness of node weight errors; (4) Enhanced nodes can stably compensate for knowledge degradation errors in continuous learning.

[0075] Therefore, the proposed experience-enhanced continuous deterministic learning control method can not only achieve continuous accumulation and reuse of dynamic knowledge in multi-task trajectory tracking, but also effectively suppress the problem of knowledge forgetting, thereby improving the long-term continuous learning control performance and system robustness.

[0076] In summary, the experience-enhanced continuous deterministic learning control method proposed in this invention not only enables the continuous accumulation and reuse of dynamic knowledge during multi-task trajectory tracking, but also effectively suppresses knowledge forgetting, thereby improving long-term continuous learning control performance and system robustness. This method enhances the online compensation of dynamic deviations between historical knowledge and the current task by strengthening nodes, achieving the synergistic fusion of historical dynamic knowledge and current task knowledge, while ensuring the stability of the closed-loop system, and has promising engineering application prospects.

[0077] To further verify the effectiveness of the continuous deterministic learning control method proposed in this invention, this embodiment systematically evaluates the tracking performance, network structure expansion, weight convergence characteristics, dynamic approximation accuracy, and knowledge generalization ability in the multi-task trajectory tracking process. In the simulation, four different reference trajectory tasks are considered, denoted as Task A, Task B, Task C, and Task D, and their expressions are as follows: , , and These trajectories are designed as continuous periodic signals to ensure that the resulting regression vectors satisfy the continuous excitation condition. Furthermore, they exhibit diverse dynamic characteristics in terms of frequency and amplitude, facilitating the verification of the effectiveness of the proposed framework in the accumulation and reuse of dynamic knowledge.

[0078] In this process, the parameters of the proposed adaptive incremental node generation mechanism are selected. , and ,in and The main factors determining the node distribution and approximation capability of the RBF neural network are: Based on the Lyapunov-based stability conditions, the controller and adaptive learning parameters are selected as follows: , , , and To ensure system boundedness and good tracking performance, moderate parameter variations primarily affect convergence speed and approximation performance, while closed-loop stability remains intact. The simulation environment is MATLAB / Simulink, with a sampling time of 0.002 seconds. Results are as follows... Figures 3 to 8 As shown.

[0079] like Figure 3 As shown, Figure 3 The figure shows the tracking control performance during the continuous learning process of different task trajectories. It can be seen that during the continuous deterministic learning process, the system state achieved satisfactory tracking performance under each reference trajectory, which indicates that the designed neural network controller can guarantee the closed-loop stability of the system.

[0080] like Figure 4 As shown, Figure 4 A schematic diagram is generated for incremental nodes during the continuous learning process of different task trajectories, where... Figure 4 (a) in the diagram illustrates the generation of incremental nodes during the learning process of trajectory A. Since there are no learned nodes initially, the RBFNN adaptively generates 96 nodes along trajectory A to cover the corresponding state space region. Figure 4(b) shows the node distribution during the learning process of trajectory B. While retaining the nodes already learned from trajectory A, 53 new nodes were incrementally generated only in the uncovered areas, increasing the total number of nodes from 96 to 149. Figure 4 (c) in the diagram illustrates the node distribution during the learning process of trajectory C. Since trajectory C has a significant overlap in state space with the already learned trajectories, only 6 new nodes are needed to complete the dynamic learning of the current trajectory. Figure 4 (d) in the diagram illustrates the node distribution during the trajectory D learning process. Similarly, since most of the state space region has already been covered by existing nodes, only 5 new nodes are added, increasing the total number of nodes from 155 to 160.

[0081] As can be seen, during the continuous learning process, the RBF neural network dynamically generates incremental nodes through an adaptive node generation mechanism to maintain sufficient approximation capability across different task trajectories. Green dots represent nodes generated in historical tasks; black dots represent nodes generated under the current task trajectory; and the solid line represents the trajectory of the neural network input vector under the current task. It can be observed that the newly added nodes are all located in the vicinity of the trajectory, effectively covering the newly visited state space. This verifies the effectiveness of the proposed node generation mechanism, ensuring that the RBF neural network always possesses sufficient local approximation capability within the actual operating region of the system, thereby ensuring that unknown dynamics can be accurately learned and effectively compensated.

[0082] like Figure 5 As shown, Figure 5 The image shows the convergence plot of RBFNN weights during the continuous learning process for different task trajectories. Figure 5 In the diagram, (a) represents the convergence graph of the RBFNN weights under trajectory A. Figure 5 (b) in the figure represents the RBFNN weight convergence graph under trajectory B; Figure 5 In the diagram, (c) represents the convergence graph of the RBFNN weights under trajectory C. Figure 5 In the diagram, (d) represents the convergence graph of the RBFNN weights under trajectory D. It can be seen that, under the condition of continuous excitation, the RBFNN weights converge to a constant value.

[0083] like Figure 6 As shown, Figure 6 To assess the approximation performance of RBFNN during the continuous learning process for different task trajectories, we compared the real system dynamics with the approximation dynamics of RBFNN, verifying that RBFNN can achieve accurate approximation of the dynamics of closed-loop nonlinear systems.

[0084] like Figure 7 As shown, Figure 7The graph compares the approximation performance of RBFNN in the learning process between continuous deterministic learning and traditional deterministic learning, showing the function approximation errors of traditional deterministic learning and continuous deterministic learning for tasks B and C. It can be seen that continuous deterministic learning achieves faster error convergence than traditional deterministic learning, indicating that knowledge reuse improves learning efficiency.

[0085] like Figure 8 As shown, Figure 8 For the performance evaluation of the proposed continuous deterministic learning control scheme, among which, Figure 8 (a) shows the tracking performance under five different trajectories (where 0-120s represents the four learned trajectories). Figure 8 Figure (b) shows a comparison of tracking errors under five different trajectories (compared to Adaptive Neural Network (ANN) control and Traditional Deterministic Learning (DLC) control). The performance evaluation used five trajectories: the 0-120s trajectories were a concatenation of the four learned trajectories, each lasting 30 seconds; the fifth trajectory, 120-150s, was the unlearned trajectory. .like Figure 8 As shown in (a), the system state under the EEDLC scheme can stably and accurately track the test reference trajectory, verifying the effectiveness of the proposed controller. Satisfactory performance was also achieved in the unseen trajectory segment (120-150 s), indicating that the learned knowledge can be effectively generalized to the new trajectory. Figure 8 Figure (b) shows a comparison of the tracking errors of the three controllers. It is evident that ANC has the largest overall error and performs worse than the learning-based strategy. DLC exhibits forgetting during continuous learning, resulting in a significant decrease in accuracy during the earlier learning phase of task A, while performing better during the later learning phase of task D. EEDLC improves its approximation capability by introducing enhancement nodes, exhibiting smaller errors in tasks A, B, and C, and maintaining performance comparable to DLC in task D. Furthermore, EEDLC maintains stable tracking in unseen trajectory segments, further demonstrating its generalization ability.

[0086] Example 2 This embodiment provides an intelligent control system for nonlinear systems based on sustainable deterministic learning, including: The model building module is configured to: establish a nonlinear system dynamics model and construct the system state error and virtual control law based on the backstepping control method; The neural network learning module is configured to: use the derivatives of the system state and the virtual control law as inputs to the radial basis function neural network, approximate the unknown closed-loop dynamic function online, and construct the neural network adaptive control law and weight update law based on the approximation results and system state errors; The node expansion module is configured to: calculate the maximum activation level of existing nodes in the neural network, and generate new nodes and add them to the neural network when the maximum activation level is lower than a preset threshold; The knowledge storage and transfer module is configured to: perform deterministic learning on different task trajectories, cause the neural network weights to converge to constant values ​​and save them as dynamic knowledge, and use the dynamic knowledge to initialize the parameters of new tasks; The enhancement compensation module is configured to: construct enhancement node compensation terms, which are generated by multiplying the Gaussian function values ​​by adaptive weights after random mapping and nonlinear activation; superimpose the compensation terms into the neural network adaptive control law, and establish an adaptive update law for the compensation terms to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

[0087] It should be noted that each module in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation processes are the same, so they will not be repeated here.

[0088] Example 3 This embodiment also provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, when the processor executes the program, it completes the method described in Embodiment 1.

[0089] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0090] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0091] In the implementation process, each step of the above method can be completed by the integrated logic circuits in the processor hardware or by software instructions.

[0092] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0093] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0094] Example 4 This embodiment also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, complete the method described in Embodiment 1.

[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent control method for nonlinear systems based on sustainable deterministic learning, characterized in that, include: A nonlinear system dynamics model is established, and the system state error and virtual control law are constructed based on the backstepping control method. Using the system state and the derivative of the virtual control law as inputs to the radial basis function neural network, the unknown closed-loop dynamic function is approximated online. Based on the approximation results and the system state error, the neural network adaptive control law and weight update law are constructed. Calculate the maximum activation level of existing nodes in the neural network. When the maximum activation level is lower than a preset threshold, generate a new node and add it to the neural network. Different task trajectories are learned to make the neural network weights converge to constant values ​​and be saved as dynamic knowledge. The dynamic knowledge is then used to initialize the parameters of new tasks. An enhanced node compensation term is constructed, which is generated by multiplying the Gaussian function value by an adaptive weight after random mapping and nonlinear activation. The compensation term is superimposed on the adaptive control law of the neural network, and an adaptive update law for the compensation term is established to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

2. The method according to claim 1, characterized in that, Based on the approximation results and the system state error, a neural network adaptive control law and a weight update law are constructed, specifically as follows: The derivatives of the system state and the last layer of virtual control law are used as the input vectors of the radial basis function neural network. The Gaussian function output of the neural network is used to approximate the unknown nonlinear function in the closed-loop dynamics. The approximation result is substituted into the backstepping control law to replace the unknown function term, and combined with the system state error of the last layer, a neural network adaptive control law is constructed. The weight update law is designed based on the system state error and the Gaussian function value.

3. The method according to claim 1, characterized in that, Calculate the maximum activation level of existing nodes in the neural network, specifically as follows: Calculate the Gaussian activation function for each existing radial basis function node under the current input vector, and take the maximum value. ; Preset threshold Where b is the minimum node spacing parameter, d th Set the node coverage threshold distance; when When the time comes, a new node center is generated based on the relative position of the current input vector and the nearest node center, and the distance between the new node and all existing nodes is greater than 2b. Then the new node is added to the neural network.

4. The method according to claim 1, characterized in that, The neural network weights are learned to converge to constant values ​​and stored as dynamic knowledge by performing deterministic learning on different task trajectories. Specifically: Under periodic reference trajectory excitation, the system satisfies the continuous excitation condition, and the neural network weight estimation error converges exponentially. The weight estimates within the stable time period after convergence are averaged in the time domain to obtain constant weights, which are then stored as the dynamic knowledge corresponding to the task.

5. The method according to claim 1, characterized in that, The aforementioned dynamic knowledge is used to initialize the parameters of a new task. Specifically, for the first task, the neural network weights are initialized to zero; for subsequent tasks, the constant weights saved from the previous task are used as the initial weights of the neural network for the current task.

6. The method according to claim 1, characterized in that, The enhanced node compensation item is specifically as follows: ( ); in, For a randomly initialized linear mapping matrix, For bias vectors, To enhance the adaptive weights of nodes, The hyperbolic tangent activation function is used. Let be the Gaussian activation function for the k-th trajectory task, where k represents the k-th trajectory task.

7. The method according to claim 1, characterized in that, The weights of the enhanced nodes are updated using the following adaptive law: ; in, To enhance node learning gain, As a robust leakage factor, Let be the system state error of order n.

8. An intelligent control system for nonlinear systems based on sustainable deterministic learning, characterized in that, include: The model building module is configured to: establish a nonlinear system dynamics model and construct the system state error and virtual control law based on the backstepping control method; The neural network learning module is configured to: use the system state and the derivative of the virtual control law as input to the radial basis function neural network, approximate the unknown closed-loop dynamic function online, and construct the neural network adaptive control law and weight update law based on the approximation result and the system state error; The node expansion module is configured to: calculate the maximum activation level of existing nodes in the neural network, and generate new nodes and add them to the neural network when the maximum activation level is lower than a preset threshold; The knowledge storage and transfer module is configured to: perform deterministic learning on different task trajectories, cause the neural network weights to converge to constant values ​​and save them as dynamic knowledge, and use the dynamic knowledge to initialize the parameters of new tasks; The enhancement compensation module is configured to: construct enhancement node compensation terms, which are generated by multiplying the Gaussian function value by an adaptive weight after random mapping and nonlinear activation; superimpose the compensation terms into the neural network adaptive control law, and establish an adaptive update law for the compensation terms to compensate for the dynamic deviation between historical knowledge and the current task online, thereby realizing multi-task trajectory tracking control.

9. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the intelligent control method for nonlinear systems based on sustainable deterministic learning as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the intelligent control method for nonlinear systems based on sustainable deterministic learning as described in any one of claims 1-7.