A physical information guided selective attention koopman operator learning method and device
By employing the Koopman operator learning method guided by physical information, the problems of lack of physical consistency and blind feature selection in the modeling of complex nonlinear systems are solved, achieving efficient long-term prediction and stable model application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2025-11-17
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies lack physical information guidance mechanisms in modeling complex nonlinear systems, resulting in model predictions that do not conform to physical laws, blind feature selection, insufficient long-term prediction and generalization capabilities, high computational complexity, and difficulty in applying them to real-time control.
We employ a physical information-guided selective attention Koopman operator learning method. By constructing a sparse physical awareness Koopman learning model and combining physical constraints based on acceleration information with a three-timescale training strategy, we automatically select important features to improve the physical consistency and computational efficiency of the model.
It improves the physical consistency and long-term prediction accuracy of the model, enhances the interpretability and generalization ability of the model, reduces computational complexity, and ensures the stability and reliability of the model in complex systems.
Smart Images

Figure CN121455061B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nonlinear system modeling and intelligent control technology, and in particular to a physical information-guided selective attention Koopman operator learning method and apparatus. Background Technology
[0002] Accurate modeling and control of complex nonlinear dynamic systems is one of the core challenges in modern control theory and artificial intelligence. In fields such as robot motion control, autonomous driving, aerospace, and energy systems, accurately capturing the dynamic characteristics of a system is crucial for achieving efficient and stable control. While traditional physics-based modeling methods offer good interpretability and a strong theoretical foundation, they often require extensive expert knowledge and parameter tuning when dealing with complex systems that are high-dimensional, strongly nonlinear, and coupled across multiple scales, and struggle to handle unmodeled dynamic characteristics. In recent years, data-driven modeling methods have offered new approaches to address this challenge. Among these, Koopman operator theory, as a mathematical framework for elevating nonlinear dynamic systems to an infinite-dimensional linear space, has received widespread attention. First proposed by Bernard Koopman in 1931, the Koopman operator's core idea is to describe the system's evolution by acting on observable functions, thereby transforming the nonlinear dynamics in the original state space into linear dynamics in an elevated space. Although modeling techniques for complex nonlinear systems have made some progress in various fields, they still face many limitations when dealing with high-dimensional, strongly nonlinear, and multi-scale coupled systems. Existing technologies can be divided into three main categories based on their modeling principles: physical principle-based modeling methods, pure data-driven modeling methods, and deep Koopman operator modeling methods. Each category has its own specific technical problems and application limitations.
[0003] Physically based modeling methods can provide accurate mathematical descriptions and sound physical meanings. However, they suffer from several fundamental drawbacks: First, for complex multibody systems, such as multi-degree-of-freedom robotic arms, the derivation of their dynamic equations is extremely cumbersome, requiring consideration of numerous coupling effects, frictional characteristics, and unmodeled dynamics, leading to an exponential increase in modeling complexity. Second, this method heavily relies on expert knowledge and extensive parameter tuning, often making it difficult to obtain accurate model parameters for real-world systems with complex features such as flexible joints and nonlinear friction. Finally, when the system exhibits unknown or difficult-to-model dynamic characteristics, physically based methods struggle to adapt to and compensate for these uncertainties, resulting in decreased model accuracy. While purely data-driven modeling methods can automatically learn complex nonlinear relationships from data, avoiding the tedious process of manual modeling, they also have significant technical limitations. First, they generally exhibit "black box" characteristics, lacking interpretability and making it difficult to guarantee that their predictions conform to basic physical laws, especially in areas outside the training data coverage, where the model may produce predictions that contradict common sense. Second, they are prone to overfitting, particularly with limited training data, resulting in severely insufficient generalization ability and poor performance under new working conditions or disturbances. Furthermore, these methods perform poorly in long-term predictions, with prediction errors accumulating and amplifying with increasing time steps, leading to unreliable long-term predictions. Deep Koopman operator modeling overcomes some of the limitations of traditional methods. However, several key issues remain, including difficulty in setting the dimensionality for upscaling, lack of physical constraints, and blind feature selection. This results in a large number of redundant features in the upscaling space that are unrelated to physical laws, increasing computational complexity and potentially interfering with the model's learning process.
[0004] A comprehensive analysis of the limitations of existing technologies reveals the following common shortcomings in solving complex nonlinear system modeling problems: First, the lack of an effective physical information guidance mechanism makes it impossible to ensure that the model prediction results conform to basic physical laws; second, insufficient feature selection and sparsification capabilities make it difficult to automatically identify and retain key features that are important to the dynamic behavior of the system; third, poor performance in long-term prediction and generalization capabilities limits its application in practical control systems; and fourth, high computational complexity and a large number of parameters make it unsuitable for real-time control applications. Summary of the Invention
[0005] To address the technical problem that existing technologies lack physical constraints and effective feature selection mechanisms, leading to decreased accuracy in dynamic modeling's dimensionality-increasing feature selection and further resulting in large state errors in the predicted robotic arm system, this invention provides a physically information-guided selective attention Koopman operator learning method and apparatus. The technical solution is as follows:
[0006] On the one hand, a physically-informed guided selective attention Koopman operator learning method is provided, which is implemented by a physically-informed guided selective attention Koopman operator learning device. The method includes:
[0007] S1. Obtain the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator;
[0008] S2. Construct a deep Koopman model based on the Koopman operator and define the prediction error loss function for the upgraded dimension space; construct a state reconstruction loss function based on the Koopman operator;
[0009] S3. Based on a planar three-free robotic arm system, construct a sparse physical perception Koopman learning model; wherein, the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network.
[0010] S4. Using the acceleration information of the planar three-free manipulator as a physical constraint, construct a physical information loss function; based on the prediction error loss function of the up-dimensional space, the state reconstruction loss function, and the physical information loss function, construct a loss function for a physical information-guided deep Koopman network model.
[0011] S5. Construct the loss function of the attention-like network; based on the loss function of the deep Koopman network model guided by physical information and the loss function of the attention-like network, construct the total loss function of the Koopman learning model; according to the total loss function of the Koopman learning model, train the Koopman learning model using a three-time-scale training strategy to obtain the trained Koopman learning model.
[0012] S6. Obtain the current state and control input of the planar three-free robotic arm system; input the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model, and output the state of the robotic arm system at the next moment.
[0013] On the other hand, a physical information-guided selective attention Koopman operator learning device is provided, which is applied to the physical information-guided selective attention Koopman operator learning method. The device includes:
[0014] The acquisition unit is used to acquire the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator;
[0015] The first building unit is used to construct a deep Koopman model based on the Koopman operator and define the upscaling space prediction error loss function; and to construct a state reconstruction loss function based on the Koopman operator.
[0016] The second building unit is used to construct a sparse physical perception Koopman learning model based on a planar three-free robotic arm system; wherein, the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network.
[0017] The third building unit is used to construct a physical information loss function by using the acceleration information of the planar three-free manipulator as a physical constraint; and to construct the loss function of the physical information-guided deep Koopman network model based on the prediction error loss function of the up-dimensional space, the state reconstruction loss function, and the physical information loss function.
[0018] The training unit is used to construct the loss function of the attention-like network; the loss function of the deep Koopman network model guided by physical information and the loss function of the attention-like network are used to construct the total loss function of the Koopman learning model; based on the total loss function of the Koopman learning model, the Koopman learning model is trained using a three-time-scale training strategy to obtain the trained Koopman learning model.
[0019] The prediction unit is used to obtain the current state and control input of the planar three-free robotic arm system; it inputs the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model and outputs the state of the robotic arm system at the next moment.
[0020] On the other hand, a physical information-guided selective attention Koopman operator learning device is provided, the physical information-guided selective attention Koopman operator learning device comprising: a processor; a memory storing computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, any one of the methods described above for physical information-guided selective attention Koopman operator learning is implemented.
[0021] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods of the physical information-guided selective attention Koopman operator learning method.
[0022] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0023] The sparse physical awareness Koopman learning model proposed in this invention effectively overcomes key problems in nonlinear system modeling, such as lack of physical consistency, blind feature selection, and training instability, by organically integrating physical constraints based on acceleration information, a physical information-guided selective attention mechanism, and a three-timescale coordinated training strategy. This results in significant technological advancements and beneficial effects. This invention improves the physical consistency and long-term prediction accuracy of the model. Existing data-driven methods, including standard deep Koopman models, suffer from a lack of physical constraints, leading to a rapid accumulation of errors in long-term iterative predictions and causing predictions to deviate from the true physical trajectory. This invention innovatively introduces a physical consistency loss function based on acceleration information, forcing the model to follow Newton's laws of mechanics during the learning process. This invention also achieves adaptive sparsity of upgraded features, enhancing the model's interpretability and computational efficiency. Traditional deep Koopman methods require manually setting the upgraded dimension; excessively high dimensions introduce a large number of redundant features, leading to overfitting and increased computational burden. The physical information-guided selective attention mechanism designed in this invention can automatically assess the physical importance of each feature based on its gradient sensitivity to physical constraints, and adaptively assign high weights to important features while suppressing redundant or irrelevant features. This not only makes the model's internal mechanism partially visible, enhancing interpretability, but also reduces the complexity of subsequent linear prediction models and improves computational efficiency by reducing invalid features. This invention improves the model's generalization ability and robustness to higher dimensions. Purely data-driven models are prone to overfitting in high-dimensional spaces, leading to a decline in generalization ability. The attention mechanism in this invention is guided by physical information, ensuring that the model learns universal physical laws rather than merely fitting the surface features of the training data. The feature selection mechanism in this invention effectively suppresses the overfitting risk caused by redundant features, making the model less sensitive to the selection of higher dimensions and possessing stronger robustness and generalization ability.
[0024] This invention ensures stable convergence of a complex multi-objective optimization process. Optimizing data fitting accuracy, physical consistency, and model sparsity is a complex multi-objective optimization problem, which can easily lead to training instability or convergence failure. This invention's unique three-timescale coordinated training strategy effectively decouples conflicts between different optimization objectives by setting different update frequencies for model parameters (fast), physical importance scores (medium), and attention temperature (slow). This strategy ensures the stability of physical importance assessment and guides the attention mechanism smoothly from the initial "exploration" to the later "sparserization," ultimately guaranteeing that the entire complex model can stably and reliably converge to a state with excellent performance across all metrics. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of a physical information-guided selective attention Koopman operator learning method provided in an embodiment of the present invention;
[0027] Figure 2 This is the working principle of a Koopman operator provided in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of the structure of a planar three-degree-of-freedom robotic arm model provided in an embodiment of the present invention;
[0029] Figure 4 This is an overall architecture diagram of a sparse physical perception Koopman learning model provided in an embodiment of the present invention;
[0030] Figure 5 This is a schematic diagram of an upgraded network structure provided in an embodiment of the present invention;
[0031] Figure 6 This is a block diagram of a physical information-guided selective attention Koopman operator learning device provided in an embodiment of the present invention;
[0032] Figure 7 This is a schematic diagram of the structure of a physical information-guided selective attention Koopman operator learning device provided in an embodiment of the present invention. Detailed Implementation
[0033] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0034] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0035] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0036] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0037] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0038] This invention provides a physically-informed guided selective attention Koopman operator learning method. This method can be implemented by a physically-informed guided selective attention Koopman operator learning device, which can be a terminal or a server. Figure 1 The flowchart shown illustrates the physical information-guided selective attention Koopman operator learning method. The processing flow of this method may include the following steps:
[0039] S1. Obtain the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator.
[0040] In one feasible implementation, a Koopman observation function is constructed based on the dynamic equations of a discrete nonlinear system and the Koopman operator; wherein the dynamic equations of the discrete nonlinear system are expressed by the following formula (1):
[0041] (1)
[0042] in, This represents the system state vector at the k-th sampling time, with a system sampling period of . , This represents the control input vector at the k-th sampling time. This represents the nonlinear dynamic function of the system. Among them, the Koopman operator... It is an infinite-dimensional linear operator, which is expressed in the higher-dimensional space by the following formula (2):
[0043] (2)
[0044] in, It is the Koopman observation function, also known as the dimension-upgrading function, used to map the original state vector to a higher-dimensional space. The dimension-upgrading function can be further decomposed into dimension-upgrading functions for the state vector and for the control input vector, expressed by the following formula (3):
[0045] (3)
[0046] In general, to ensure the linearity of the Koopman model, let That is, it can be expressed by the following formula (4):
[0047] (4)
[0048] in, ,and Since the Koopman operator is an infinite-dimensional operator, it is necessary to construct its finite-dimensional approximation, further simplifying equation (2) into the following equation (5):
[0049] (5)
[0050] in, Indicates a state of dimensional ascension. Represents the system matrix in the higher-dimensional space. Let represent the input matrix in the increased-dimensional space. Generally, to facilitate the reconstruction of the original state from the increased-dimensional state, the following formula (6) is defined:
[0051] (6)
[0052] in, The nonlinear part of the updimensional function is represented by the following formula (7):
[0053] (7)
[0054] in, This represents a dimension reduction matrix. Where, for example... Figure 2The diagram illustrates the working principle of a Koopman operator provided in an embodiment of the present invention.
[0055] In one feasible implementation, the collected system state sequence and control input sequence are used to construct a data matrix, which is represented by the following formulas (8)-(10):
[0056] (8)
[0057] (9)
[0058] (10)
[0059] in, The number of sample points is represented by the following formula (11): The finite-dimensional representation of the Koopman operator is obtained by solving a quadratic programming problem.
[0060] (11)
[0061] By processing the above formula (11), the analytical solution is obtained, which is expressed by the following formula (12):
[0062] (12)
[0063] in, express Norm, This indicates the Moore-Penrose pseudo-inverse.
[0064] S2. Construct a deep Koopman model based on the Koopman operator and define the prediction error loss function for the increased dimensionality space; construct a state reconstruction loss function based on the Koopman operator.
[0065] In one feasible implementation, the linear modeling capability of the Koopman operator depends on the quality of the selection of the dimension-upgrading function. However, for complex physical systems, the dimension-upgrading function is often difficult to predefine. Therefore, a multilayer perceptron is used to map the original state vector into a nonlinear dimension-upgrading state. Assuming the network has a hidden layer, the dimension of the input layer is determined by the dimension of the original state space, while the dimension of the output layer is set by the user. The forward propagation process of the network can be represented by the following formula (13):
[0066] (13)
[0067] in, For the first The nonlinear activation function of the hidden layer Indicates the first Layer weight matrix; Indicates the first Layer bias terms.
[0068] The deep Koopman model can simultaneously learn a suitable upscaling function network and its corresponding finite-dimensional approximation of the Koopman operator from the data. The upscaling space prediction error loss function is expressed by the following formula (14):
[0069] (14)
[0070] in, express Norm; The number of training samples is represented by this error loss function, which measures the prediction accuracy of the Koopman model in the increased-dimensional space. Furthermore, to enhance the model's expressive power in the original state space, a state reconstruction loss term is introduced, expressed by the following formula (15):
[0071] (15)
[0072] By jointly minimizing the aforementioned loss, this method can automatically learn an upgraded representation that satisfies the Koopman linearity property without requiring manual setting of basis functions.
[0073] S3. Based on a planar three-free robotic arm system, a sparse physical perception Koopman learning model is constructed; the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network.
[0074] In one feasible implementation, this embodiment of the invention takes a planar degree-of-freedom robotic arm as the research object, and its model is as follows: Figure 3 As shown. The lengths of the robotic arm are respectively... The unit is meters (m), and the masses of the three links are respectively... The unit is kg, and the centroids of the three links are all located at the geometric center of the links. Among them, the planar three-degree-of-freedom manipulator is a typical nonlinear mechanical system, widely used in industrial automation and robotics research. Its dynamic model can be derived from the Euler-Lagrange equations and expressed by the following formula (16):
[0075] (16)
[0076] in, Indicates the angles of the three joints. This represents the angular velocity of the three joints. Represents the input torque of the three joints; matrix It is the inertia matrix of the robotic arm. It is a matrix that includes Coriolis force and centrifugal force; It is the gravity matrix.
[0077] Optionally, the physical information-guided deep Koopman network model includes: an up-dimensional network module and a Koopman operator approximation module;
[0078] One feasible implementation method is, for example Figure 4 The diagram shown is an overall architecture diagram of a sparse physical perception Koopman learning model provided in an embodiment of the present invention; wherein, the dimension-upgrading network module is responsible for processing the original state. Mapping to a higher-dimensional space yields an upgraded state. This module employs a multilayer perceptron structure, progressively extracting high-order nonlinear features of the system state through multiple nonlinear hidden layers. For example... Figure 5 This is a schematic diagram of an upgraded network structure provided in an embodiment of the present invention.
[0079] The upscaling network module employs a multilayer perceptron structure, extracting high-order nonlinear features of the robotic arm system state through multiple nonlinear hidden layers. The Koopman operator approximation module consists of three linear unbiased network layers, corresponding to the system matrix A, input matrix B, and dimensionality reduction matrix C of the deep Koopman network model, respectively. The system matrix A and input matrix B are used for linear prediction in the upscaling space, while the dimensionality reduction matrix C is used to project the state of the upscaling space back into the original state space for state reconstruction.
[0080] In one feasible implementation, the forward propagation process of the network includes: given the system state at the current moment... and control input First, the upgraded state is obtained through the upgraded network. Then, using the system matrix A and the input matrix B, a linear prediction is made in the increased-dimensional space to obtain the predicted value of the increased-dimensional state at the next time step; finally, the increased-dimensional predicted state is projected back into the original state space through the reduced-dimensional matrix C to obtain the predicted value of the state at the next time step.
[0081] S4. Using the acceleration information of the planar three-free manipulator as a physical constraint, construct a physical information loss function; based on the dimensional space prediction error loss function, the state reconstruction loss function, and the physical information loss function, construct a loss function for a physical information-guided deep Koopman network model.
[0082] To train the network model, this embodiment of the invention designs a composite loss function comprising three parts. The first part is the up-dimensional space prediction loss, used to constrain the linear evolution accuracy in the up-dimensional space; the second part is the state reconstruction loss, used to ensure that the model can accurately reconstruct the original state from the up-dimensional space; the third part is the physical constraint loss function, which introduces system acceleration information as a physical constraint to ensure that the model's prediction results strictly follow physical laws.
[0083] Optionally, S4 uses the acceleration information of the planar three-free manipulator as a physical constraint to construct a physical information loss function, including:
[0084] Traditional deep Koopman models focus solely on data fitting, often resulting in a lack of interpretability and insufficient generalization ability. To address these issues, this invention incorporates the robotic arm's acceleration information into the model, introducing a physical constraint loss term to ensure that the model's predictions match actual physical dynamics.
[0085] S41, Assumption To predict the true acceleration of the system at time k using a deep Koopman network model guided by physical information, the acceleration prediction residual is defined; whereby the acceleration prediction residual is expressed by the following formula (17):
[0086] (17)
[0087] in, Indicates the system's first Real-time acceleration can be obtained through sensor measurement or by calculation using the differential method.
[0088] S42. The true acceleration at time k and the predicted true acceleration at time k are calculated using the finite difference method; wherein the predicted true acceleration at time k is expressed by the following formula (18):
[0089] (18)
[0090] in, Indicates the first The actual joint velocity of the time-lapse system; Indicates the first The actual joint velocities of the time-lapse system; among which, Predicted acceleration at time Similarly, the calculation is performed using the finite difference method, and expressed by the following formula (19):
[0091] (19)
[0092] in, The Koopman model of the system represents the first... Prediction of joint velocities at any given moment; The Koopman model of the system represents the first... Prediction of joint velocities at any given moment.
[0093] S43. Based on the actual acceleration and the predicted actual acceleration, the physical information loss function is constructed and expressed by the following formula (20):
[0094] (20)
[0095] Where P represents the number of training samples; Indicates the loss of physical information; This represents the actual acceleration at time i. This represents the predicted actual acceleration at time i.
[0096] Among them, the physical information loss function enhances the physical consistency of the model by minimizing the acceleration prediction error, ensuring that the model prediction results conform to Newton's laws of mechanics.
[0097] Optionally, the loss function of the S4 physical information-guided deep Koopman network model is expressed by the following formula (21):
[0098] (twenty one)
[0099] in, The loss function represents the physical information-guided deep Koopman network model; Represents the state reconstruction loss function; This represents the loss function for prediction in higher-dimensional space. Represents the physical information loss function; The weighting coefficients represent the physical information loss function, used to balance data fitting accuracy and physical consistency.
[0100] In one feasible implementation, a key challenge faced by traditional deep Koopman models lies in how to reasonably set the dimensionality after dimensionality increase. Too low a dimensionality may prevent the effective capture of the system's dynamic features, while too high a dimensionality leads to computational overhead and redundant information, affecting model performance. To address this issue, this embodiment of the invention introduces a physics-information-guided attention-like mechanism in the dimensionality increase module, constructing an upscaling network structure with attachable graph structures. Through this mechanism, the model can automatically filter and highlight key information related to physical laws while suppressing irrelevant features.
[0101] In one feasible implementation, the original state After passing through the upgraded network, the upgraded state is obtained. It consists of three parts, including: the original state The term itself, the constant term "1", and the weighted nonlinear dimension-upgrading part Its upgraded state is represented by the following formula (22):
[0102] (twenty two)
[0103] The constant "1" is added to compensate for the role of constants in the dynamic evolution of the system. The nonlinear dimensionality increase after attention-based weighting is represented by the following formula (23):
[0104] (twenty three)
[0105] in, This represents element-wise multiplication; Represented as the original nonlinear dimensionality-upgrading part; This represents the attention weight vector, which guides the allocation of dimensionality-upgrading features. It can be dynamically adjusted through learning and optimization during training, thereby achieving differentiated weighting for different dimensionality-upgrading features. Its core idea is: for non-linear dimensionality-upgrading features... The network can automatically adjust the weights of each component based on physical information, giving higher weights to features closely related to the dynamic laws of the system, while suppressing the weights of irrelevant or redundant features. Through this mechanism, the model can adaptively focus on physically significant features during dimensionality increase, achieving the prominent expression of key information and the effective reduction of redundant features, thereby improving the physical rationality and computational efficiency of the dimensionality increase representation.
[0106] The formula for calculating the class attention weight vector is expressed by the following formula (24):
[0107] (twenty four)
[0108] in, It is a trainable parameter vector for a layer of an attention network; It is a temperature parameter; This refers to the Sigmoid function, used to ensure the weights... The value of is between . Among them, the temperature parameter . It plays a crucial regulatory role during training, achieving a smooth transition from feature exploration to feature sparsity through dynamic scheduling. The temperature strategy adopted in this embodiment of the invention is expressed by the following formula (25):
[0109] (25)
[0110] in, , For the total number of training rounds, This is the current training epoch. This strategy ensures that the model fully explores the possibilities of all nonlinear features in the early stage of training (high temperature zone), steadily learns importance scores in the middle stage (mild convergence zone), and strengthens the differentiation of weights towards 0 or 1 in the later stage (low temperature zone), thus achieving effective sparsity of the feature space.
[0111] In one feasible implementation, to enable the attention-like mechanism to identify important features closely related to the physical laws of the system, this invention proposes an importance evaluation mechanism based on the physical loss gradient. The core idea of this mechanism is that feature components sensitive to the physical loss function often contain more physical information and should be assigned higher attention-like weights.
[0112] Optionally, the process of constructing the physical guidance loss function includes:
[0113] The original system state is input into the up-dimensional network module, and a multilayer perceptron is used to extract multiple high-order nonlinear feature components of the system state.
[0114] In one feasible implementation, for each nonlinear feature component of the multilayer perceptron output, its physical importance score is defined as the square of the global average gradient norm of that component with respect to the physical loss function. This score aims to quantify the degree of influence of this feature on the physical consistency of the system.
[0115] Based on multiple nonlinear feature components and the physical loss function corresponding to each training sample, the physical importance score is calculated and expressed by the following formula (26):
[0116] (26)
[0117] in, This represents the j-th nonlinear feature component obtained after the i-th sample passes through a multilayer perceptron; Let represent the physical loss corresponding to the i-th training sample; P represents the number of training samples. This represents the physical importance score corresponding to the j-th nonlinear component;
[0118] Formula (26) provides the global sensitivity assessment of each feature component to physical constraints under the current model state. The importance scores of all these individual feature components together constitute the complete importance score vector. .
[0119] In one feasible implementation, considering that the network parameters are continuously updated during training, the importance score also needs to be dynamically adjusted accordingly. If the above formula (26) is recalculated and used directly every time it is needed, it may cause drastic fluctuations due to the randomness of the data batch, which is not conducive to the stable convergence of training. To solve this problem, the embodiments of the present invention adopt an interval smooth update strategy: instead of recalculating every time it is needed, it is done every An update is performed only after a certain number of training epochs. During the update, the latest importance score vector is calculated according to formula (26), denoted as... Using an exponential moving average (EMA) strategy, the latest calculation results are combined with historical scores to obtain the [number]th [score]. The importance score vector obtained after the next update is represented by the following formula (27):
[0120] (27)
[0121] in, Indicates the first The importance score vector obtained after the next update will serve as a benchmark to guide the calculation of attention loss for a period of time in the future; This is the historical score since the last update; To control the smoothing coefficient of historical and current information weights, this design ensures that the adjustment of importance scores can reflect changes in network parameters in a timely manner, and effectively suppress short-term noise through smoothing, thereby providing a stable and physically meaningful guiding signal for the attention mechanism.
[0122] In this embodiment of the invention, based on the physical importance score, a corresponding attention-based loss function is designed to guide the learning of the weights of the attention-based network.
[0123] Optionally, the loss function of the attention-like network includes: physical guidance loss function, entropy regularization loss and contrast enhancement loss;
[0124] The loss function of the attention network is expressed by the following formula (28):
[0125] (28)
[0126] in, Represents the physical guidance loss function; This represents the entropy regularization loss function; This represents the contrast enhancement loss function; The weights represent the entropy regularization loss function. This represents the weights corresponding to the contrast enhancement loss function.
[0127] The physics importance scores are normalized, and a physics guidance loss function is constructed by guiding weight allocation, which is expressed by the following formula (29):
[0128] (29)
[0129] in, The first digit represents the importance score after normalization. One component; This represents the attention weight corresponding to the i-th nonlinear feature component; The dimension of the upgraded state z; The dimension represents the nonlinear feature; n represents the dimension of the original state x. This can be expressed by the following formula (30):
[0130] (30)
[0131] Among them, the physical guidance loss, through the guidance of importance scores, prompts the model to push the weights of high-importance features to 1 and the weights of low-importance features to 0, thus playing a major role in feature selection guidance. The entropy regularization loss, by maximizing the entropy of the weight distribution, promotes the sparsity of the weights, as expressed by the following formula (31):
[0132] (31)
[0133] in, To prevent small constants from becoming numerically unstable, a value of 0 is usually adopted. Power of 1. Entropy regularization loss, by penalizing the uncertainty of weights, encourages weights to move closer to 0 or 1, thereby enhancing the decisiveness and sparsity of feature selection.
[0134] Among them, the contrast enhancement loss further improves the quality of feature selection by enhancing the weight difference between important and unimportant features. The contrast enhancement loss is expressed by the following formula (32):
[0135] (32)
[0136] in, The average weight representing highly important features; The average weight representing features of low importance; To enhance contrast, a marginal parameter is used to delineate the boundary between high and low importance features. The high-importance feature set can be defined as the top 30% of features by physical importance score, and the low-importance feature set can be defined as the bottom 30% of features by score. This percentage is a hyperparameter and can be adjusted according to the specific application. This loss term ensures that the weights of important features are significantly higher than those of unimportant features, strengthening the discriminative power of feature importance and avoiding an overly even weight distribution.
[0137] S5. Construct the loss function of the attention-like network; based on the loss function of the physical information-guided deep Koopman network model and the loss function of the attention-like network, construct the total loss function of the Koopman learning model; according to the total loss function of the Koopman learning model, use a three-time-scale training strategy to train the Koopman learning model to obtain the trained Koopman learning model.
[0138] In one feasible implementation, the total loss function of the Koopman learning model is expressed by the following formula (33):
[0139] (33)
[0140] in, It is the loss function of a deep Koopman network model guided by physical information; This is the loss function of an attention-based network. By minimizing this total loss function, the Koopman learning model can automatically learn sparse attention weights with clear physical meaning, guided by physical information. This effectively suppresses redundant features in the increased-dimensional space, improving the model's prediction accuracy and interpretability. The training algorithm of the Koopman learning model follows a three-timescale coordination strategy.
[0141] The three-timescale training strategy aims to coordinate and optimize multiple objectives in the Koopman learning model, including data fitting accuracy, physical consistency, and model sparsity. In this model, the class attention weight vector... The calculation formula is shown in (24). Wherein is the trainable parameter vector of the attention network layer; T is the temperature parameter. Meanwhile, this embodiment of the invention also proposes an importance evaluation mechanism based on the physical loss gradient, which calculates the physical importance score vector using the following formula (26). This score does not directly participate in the weight vector. The calculation is not performed on the physical guidance loss function used to construct attention-like network layers. As shown in formula (29). In actual model training, the importance score vector It will be through the loss function To guide the trainable parameter vector of the attention network To train the weight vector, a specific distribution is used. The purpose. Among them, the temperature parameter... This plays a crucial regulatory role in the process. On a slow timescale, Scheduling is performed according to formula (25). In the early stages of training, larger... This makes the input of the Sigmoid function in formula (24) more efficient. The overall values are relatively small, at this time The value tends to smooth out (0.5), indicating that the model is in the exploratory stage as training progresses. Gradually decrease, making The value is amplified by m, at which point... The value is pushed to either 0 or 1, thereby achieving sparsity of the feature.
[0142] Optionally, the specific implementation process of S5 includes S51-S57:
[0143] S51. Obtain the training dataset and input it into the Koopman learning model; the training dataset consists of multiple samples representing the current state of the robotic arm system, the current input of the robotic arm system, and the state of the robotic arm system at the next moment; the structure of the training samples is (x k ,y k ,u k ), x k U represents the system state at time k. k y represents the control input of the system at time k. k This indicates that the system is in state x. k Next, by inputting u k The state reached at the next moment satisfies the following relationship: y k =f(x k ,u k Each such set of sample points is a training sample;
[0144] S52. For each round, the temperature parameters are set using a pre-built temperature strategy;
[0145] S53. Based on the temperature parameter, calculate the class attention weight vector according to the Sigmoid function;
[0146] S54. Based on the class attention weight vector, calculate the nonlinear dimensionality increase part after class attention weighting;
[0147] S55. Based on the weighted nonlinear up-dimensional part, the trajectory of the robotic arm is predicted using the Koopman operator;
[0148] S56. Determine if the current round is an integer multiple of the set round interval. If it is, calculate and update the existing physical importance score. If not, skip the physical importance score update step and continue the training process.
[0149] S57. Based on the existing physical importance scores, calculate the total loss function of the model; based on the total loss function, update all network parameters through backpropagation; when the preset maximum number of training rounds is met, stop training and output the trained Koopman learning model.
[0150] S6. Obtain the current state and control input of the planar three-free robotic arm system; input the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model, and output the state of the robotic arm system at the next moment.
[0151] The embodiments of the present invention can be applied not only to robotic arm systems, but also to systems such as generators.
[0152] The sparse physical awareness Koopman learning model proposed in this invention effectively overcomes key problems in nonlinear system modeling, such as lack of physical consistency, blind feature selection, and training instability, by organically integrating physical constraints based on acceleration information, a physical information-guided selective attention mechanism, and a three-timescale coordinated training strategy. This results in significant technological advancements and beneficial effects. This invention improves the physical consistency and long-term prediction accuracy of the model. Existing data-driven methods, including standard deep Koopman models, suffer from a lack of physical constraints, leading to a rapid accumulation of errors in long-term iterative predictions and causing predictions to deviate from the true physical trajectory. This invention innovatively introduces a physical consistency loss function based on acceleration information, forcing the model to follow Newton's laws of mechanics during the learning process. This invention also achieves adaptive sparsity of upgraded features, enhancing the model's interpretability and computational efficiency. Traditional deep Koopman methods require manually setting the upgraded dimension; excessively high dimensions introduce a large number of redundant features, leading to overfitting and increased computational burden. The physical information-guided selective attention mechanism designed in this invention can automatically assess the physical importance of each feature based on its gradient sensitivity to physical constraints, and adaptively assign high weights to important features while suppressing redundant or irrelevant features. This not only makes the model's internal mechanism partially visible, enhancing interpretability, but also reduces the complexity of subsequent linear prediction models and improves computational efficiency by reducing invalid features. This invention improves the model's generalization ability and robustness to higher dimensions. Pure data-driven models are prone to overfitting in high-dimensional spaces, leading to a decline in generalization ability. The attention mechanism in this invention is guided by physical information, ensuring that the model learns universal physical laws rather than merely fitting the surface features of the training data. The feature selection mechanism in this invention effectively suppresses the overfitting risk caused by redundant features, making the model less sensitive to the selection of higher dimensions and possessing stronger robustness and generalization ability. This invention ensures stable convergence of complex multi-objective optimization processes. Optimizing data fitting accuracy, physical consistency, and model sparsity is a complex multi-objective optimization problem that can easily lead to unstable training processes or convergence failures. The invention's unique three-timescale coordinated training strategy effectively decouples conflicts between different optimization objectives by setting different update frequencies for model parameters (fast), physical importance score (medium), and attention temperature (slow). This strategy ensures the stability of physical importance assessment and guides the attention mechanism smoothly from the initial "exploration" to the later "sparseness," ultimately guaranteeing that the entire complex model can stably and reliably converge to an excellent state with all performance indicators being optimal.
[0153] Figure 6This is a block diagram of a physical information-guided selective attention Koopman operator learning device provided in an embodiment of the present invention. This device is used for a physical information-guided selective attention Koopman operator learning method. (Refer to...) Figure 6 The device includes an acquisition unit 610, a first construction unit 620, a second construction unit 630, a third construction unit 640, a training unit 650, and a prediction unit 660. Wherein:
[0154] The acquisition unit 610 is used to acquire the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator;
[0155] The first building unit 620 is used to build a deep Koopman model based on the Koopman operator and define the upscaling space prediction error loss function; and to build a state reconstruction loss function based on the Koopman operator.
[0156] The second building unit 630 is used to build a sparse physical perception Koopman learning model based on a planar three-free robotic arm system; wherein, the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network.
[0157] The third building unit 640 is used to construct a physical information loss function by using the acceleration information of the planar three-free manipulator as a physical constraint; and to construct a loss function for a physical information-guided deep Koopman network model based on the prediction error loss function of the up-dimensional space, the state reconstruction loss function, and the physical information loss function.
[0158] Training unit 650 is used to construct the loss function of an attention-like network; based on the loss function of the physical information-guided deep Koopman network model and the loss function of the attention-like network, the total loss function of the Koopman learning model is constructed; according to the total loss function of the Koopman learning model, the Koopman learning model is trained using a three-time-scale training strategy to obtain a trained Koopman learning model.
[0159] The prediction unit 660 is used to obtain the current state and control input of the planar three-free robotic arm system; input the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model, and output the state of the robotic arm system at the next moment.
[0160] Optionally, the physical information-guided deep Koopman network model includes: an up-dimensional network module and a Koopman operator approximation module;
[0161] Among them, the up-dimensional network module adopts a multilayer perceptron structure, which extracts the high-order nonlinear features of the robotic arm system state through multiple nonlinear hidden layers;
[0162] The Koopman operator approximation module consists of three linear unbiased network layers, corresponding to the system matrix A, input matrix B, and dimensionality reduction matrix C of the deep Koopman network model, respectively. The system matrix A and input matrix B are used for linear prediction in the increased-dimensional space, and the dimensionality reduction matrix C is used to project the state of the increased-dimensional space back to the original state space for state reconstruction.
[0163] Optionally, the step of using the acceleration information of the planar three-free manipulator as a physical constraint to construct a physical information loss function includes:
[0164] Assumption For a deep Koopman network model guided by physical information, the true acceleration of the system at time k is predicted, and the acceleration prediction residual is defined.
[0165] The finite difference method is used to calculate the actual acceleration at time k and the predicted actual acceleration at time k.
[0166] Based on the actual acceleration and the predicted actual acceleration, the physical information loss function is constructed and expressed by the following formula (1):
[0167] (1)
[0168] Where P represents the number of training samples; Indicates the loss of physical information; This represents the actual acceleration at time i. This represents the predicted actual acceleration at time i.
[0169] Optionally, the loss function of the physical information-guided deep Koopman network model is expressed by the following formula (2):
[0170] (2)
[0171] in, The loss function represents the physical information-guided deep Koopman network model; Represents the state reconstruction loss function; This represents the loss function for prediction in higher-dimensional space. Represents the physical information loss function; The weighting coefficients represent the physical information loss function, used to balance data fitting accuracy and physical consistency.
[0172] Optionally, the loss function of the attention-like network includes: physical guidance loss function, entropy regularization loss, and contrast enhancement loss;
[0173] The loss function of the attention network is expressed by the following formula (3):
[0174] (3)
[0175] in, Represents the physical guidance loss function; This represents the entropy regularization loss function; This represents the contrast enhancement loss function; The weights represent the entropy regularization loss function. This represents the weights corresponding to the contrast enhancement loss function.
[0176] Optionally, the process of constructing the physical guidance loss function includes:
[0177] The original system state is input into the up-dimensional network module, and a multilayer perceptron is used to extract multiple high-order nonlinear feature components of the system state.
[0178] Based on multiple nonlinear feature components and the physical loss function corresponding to each training sample, the physical importance score is calculated and expressed by the following formula (4):
[0179] (4)
[0180] in, This represents the j-th nonlinear feature component obtained after the i-th sample passes through a multilayer perceptron; Let represent the physical loss corresponding to the i-th training sample; P represents the number of training samples. This represents the physical importance score corresponding to the j-th nonlinear characteristic component;
[0181] The physics importance scores are normalized, and a physics guidance loss function is constructed by guiding weight allocation, which is expressed by the following formula (5):
[0182] (5)
[0183] in, The first digit represents the importance score after normalization. One component; This represents the attention weight corresponding to the i-th nonlinear feature component; The dimension of the upgraded state z; The dimension represents the nonlinear feature; n represents the dimension of the original state x.
[0184] Optionally, the training unit 650 is used for:
[0185] Obtain the training dataset and input it into the Koopman learning model; wherein, the training dataset consists of multiple samples composed of the current state of the robotic arm system, the current input of the robotic arm system, and the state of the robotic arm system at the next moment;
[0186] For each round, the temperature parameters are set using a pre-built temperature strategy;
[0187] Based on the temperature parameter, the class attention weight vector is calculated according to the Sigmoid function;
[0188] Based on the class attention weight vector, calculate the non-linear dimensionality increase part after class attention weighting;
[0189] Based on the weighted nonlinear up-dimensional component, the trajectory of the robotic arm is predicted using the Koopman operator;
[0190] Determine if the current round is an integer multiple of the set round interval. If it is, calculate and update the existing physical importance score; otherwise, skip the physical importance score update step and continue the training process.
[0191] Based on the existing physical importance scores, calculate the total loss function of the model; based on the total loss function, update all network parameters through backpropagation; when the preset maximum number of training rounds is met, stop training and output the trained Koopman learning model.
[0192] The sparse physical awareness Koopman learning model proposed in this invention effectively overcomes key problems in nonlinear system modeling, such as lack of physical consistency, blind feature selection, and training instability, by organically integrating physical constraints based on acceleration information, a physical information-guided selective attention mechanism, and a three-timescale coordinated training strategy. This results in significant technological advancements and beneficial effects. This invention improves the physical consistency and long-term prediction accuracy of the model. Existing data-driven methods, including standard deep Koopman models, suffer from a lack of physical constraints, leading to a rapid accumulation of errors in long-term iterative predictions and causing predictions to deviate from the true physical trajectory. This invention innovatively introduces a physical consistency loss function based on acceleration information, forcing the model to follow Newton's laws of mechanics during the learning process. This invention also achieves adaptive sparsity of upgraded features, enhancing the model's interpretability and computational efficiency. Traditional deep Koopman methods require manually setting the upgraded dimension; excessively high dimensions introduce a large number of redundant features, leading to overfitting and increased computational burden. The physical information-guided selective attention mechanism designed in this invention can automatically assess the physical importance of each feature based on its gradient sensitivity to physical constraints, and adaptively assign high weights to important features while suppressing redundant or irrelevant features. This not only makes the model's internal mechanism partially visible, enhancing interpretability, but also reduces the complexity of subsequent linear prediction models and improves computational efficiency by reducing invalid features. This invention improves the model's generalization ability and robustness to higher dimensions. Pure data-driven models are prone to overfitting in high-dimensional spaces, leading to a decline in generalization ability. The attention mechanism in this invention is guided by physical information, ensuring that the model learns universal physical laws rather than merely fitting the surface features of the training data. The feature selection mechanism in this invention effectively suppresses the overfitting risk caused by redundant features, making the model less sensitive to the selection of higher dimensions and possessing stronger robustness and generalization ability. This invention ensures stable convergence of complex multi-objective optimization processes. Optimizing data fitting accuracy, physical consistency, and model sparsity is a complex multi-objective optimization problem that can easily lead to unstable training processes or convergence failures. The invention's unique three-timescale coordinated training strategy effectively decouples conflicts between different optimization objectives by setting different update frequencies for model parameters (fast), physical importance score (medium), and attention temperature (slow). This strategy ensures the stability of physical importance assessment and guides the attention mechanism smoothly from the initial "exploration" to the later "sparseness," ultimately guaranteeing that the entire complex model can stably and reliably converge to an excellent state with all performance indicators being optimal.
[0193] Figure 7This is a schematic diagram of the structure of a physical information-guided selective attention Koopman operator learning device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the physical information-guided selective attention Koopman operator learning device can include the above-mentioned Figure 6 The illustrated physical information-guided selective attention Koopman operator learning device 710 may optionally include a first processor 2001.
[0194] Optionally, the physical information-guided selective attention Koopman operator learning device 710 may also include a memory 2002 and a transceiver 2003.
[0195] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0196] The following is combined Figure 7 The components of the Koopman operator learning device 710, which uses physical information-guided selective attention, are described in detail below:
[0197] The first processor 2001 is the control center of the physical information-guided selective attention Koopman operator learning device 710. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0198] Optionally, the first processor 2001 can perform various functions of the physically-guided selective attention Koopman operator learning device 710 by running or executing software programs stored in memory 2002 and calling data stored in memory 2002.
[0199] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 7 CPU0 and CPU1 are shown in the diagram.
[0200] In a specific implementation, as one example, the physically-informed guided selective attention Koopman operator learning device 710 may also include multiple processors, for example... Figure 7 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0201] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0202] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the physically guided selective attention Koopman operator learning device 710. Figure 7 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0203] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0204] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 7 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0205] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected to the interface circuit of the selective attention Koopman operator learning device 710 guided by physical information. Figure 7 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0206] It should be noted that, Figure 7 The structure of the physical information-guided selective attention Koopman operator learning device 710 shown in the figure does not constitute a limitation on the router. Actual physical information-guided selective attention Koopman operator learning devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0207] Furthermore, the technical effects of the physical information-guided selective attention Koopman operator learning device 710 can be referred to the technical effects of the physical information-guided selective attention Koopman operator learning method described in the above method embodiments, and will not be repeated here.
[0208] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.
[0209] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0210] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0211] It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. Additionally, the character " / " in this document generally indicates an "or" relationship between the preceding and following related objects, but it may also represent an "and / or" relationship; please refer to the context for specific interpretation. In this invention, "at least one" refers to one or more, and "more" refers to two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0212] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0213] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0214] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0215] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0216] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0217] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0218] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A physically information-guided selective attention Koopman operator learning method, characterized in that, The method includes: S1. Obtain the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator; S2. Construct a deep Koopman model based on the Koopman operator and define the prediction error loss function for the upgraded dimension space; construct a state reconstruction loss function based on the Koopman operator; S3. Based on a planar three-free robotic arm system, construct a sparse physical perception Koopman learning model; wherein, the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network. S4. Using the acceleration information of the planar three-free manipulator as a physical constraint, construct a physical information loss function; based on the prediction error loss function of the up-dimensional space, the state reconstruction loss function, and the physical information loss function, construct a loss function for a physical information-guided deep Koopman network model. S5. Construct the loss function of the attention-like network; based on the loss function of the deep Koopman network model guided by physical information and the loss function of the attention-like network, construct the total loss function of the Koopman learning model; according to the total loss function of the Koopman learning model, train the Koopman learning model using a three-timescale training strategy to obtain the trained Koopman learning model. S6. Obtain the current state and control input of the planar three-free robotic arm system; input the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model, and output the state of the robotic arm system at the next moment.
2. The physical information-guided selective attention Koopman operator learning method according to claim 1, characterized in that, The physical information-guided deep Koopman network model includes: an upscaling network module and a Koopman operator approximation module; Among them, the up-dimensional network module adopts a multilayer perceptron structure, which extracts the high-order nonlinear features of the robotic arm system state through multiple nonlinear hidden layers; The Koopman operator approximation module consists of three linear unbiased network layers, which correspond to the system matrix A, input matrix B, and dimension reduction matrix C of the deep Koopman network model, respectively. The system matrix A and the input matrix B are used to perform linear prediction in the higher-dimensional space. The dimension reduction matrix C is used to project the state of the increased-dimensional space back to the original state space for state reconstruction.
3. The physical information-guided selective attention Koopman operator learning method according to claim 1, characterized in that, The S4 process uses the acceleration information of the planar three-free robotic arm as a physical constraint to construct a physical information loss function, including: S41, Assumption For a deep Koopman network model guided by physical information, the true acceleration of the system at time k is predicted, and the acceleration prediction residual is defined. S42. Calculate the actual acceleration at time k and the predicted actual acceleration at time k using the finite difference method; S43. Based on the actual acceleration and the predicted actual acceleration, the physical information loss function is constructed and expressed by the following formula (1): (1) Where P represents the number of training samples; Indicates the loss of physical information; This represents the actual acceleration at time i. This represents the predicted actual acceleration at time i.
4. The physical information-guided selective attention Koopman operator learning method according to claim 1, characterized in that, The loss function of the deep Koopman network model guided by the physical information of S4 is expressed by the following formula (2): (2) in, The loss function represents the physical information-guided deep Koopman network model; Represents the state reconstruction loss function; This represents the loss function for prediction in higher-dimensional space. Represents the physical information loss function; The weighting coefficients represent the physical information loss function, used to balance data fitting accuracy and physical consistency.
5. The physical information-guided selective attention Koopman operator learning method according to claim 1, characterized in that, The loss function of the attention-like network includes: physical guidance loss function, entropy regularization loss, and contrast enhancement loss; The loss function of the attention network is expressed by the following formula (3): (3) in, Represents the physical guidance loss function; This represents the entropy regularization loss function; This represents the contrast enhancement loss function; This represents the weights corresponding to the entropy regularization loss function; This represents the weights corresponding to the contrast enhancement loss function.
6. The physical information-guided selective attention Koopman operator learning method according to claim 5, characterized in that, The process of constructing the physical guidance loss function includes: The original system state is input into the up-dimensional network module, and a multilayer perceptron is used to extract multiple high-order nonlinear feature components of the system state. Based on multiple nonlinear feature components and the physical loss function corresponding to each training sample, the physical importance score is calculated and expressed by the following formula (4): (4) in, This represents the j-th nonlinear feature component obtained after the i-th sample passes through a multilayer perceptron; Let represent the physical loss corresponding to the i-th training sample; P represents the number of training samples. This represents the physical importance score corresponding to the j-th nonlinear characteristic component; The physics importance scores are normalized, and a physics guidance loss function is constructed by guiding weight allocation, which is expressed by the following formula (5): (5) in, The first digit represents the importance score after normalization. One component; This represents the attention weight corresponding to the i-th nonlinear feature component; The dimension of the upgraded state z; The dimension represents the nonlinear feature; n represents the dimension of the original state x.
7. The physical information-guided selective attention Koopman operator learning method according to claim 1, characterized in that, S5 trains the Koopman learning model using a three-time-scale training strategy based on the total loss function of the Koopman learning model to obtain a trained Koopman learning model, including: S51. Obtain the training dataset and input it into the Koopman learning model; wherein, the training dataset consists of multiple samples composed of the current state of the robotic arm system, the current input of the robotic arm system, and the state of the robotic arm system at the next moment; S52. For each round, the temperature parameters are set using a pre-built temperature strategy; S53. Based on the temperature parameter, calculate the class attention weight vector according to the Sigmoid function; S54. Based on the class attention weight vector, calculate the nonlinear dimensionality increase part after class attention weighting; S55. Based on the weighted nonlinear up-dimensional part, the trajectory of the robotic arm is predicted using the Koopman operator; S56. Determine if the current round is an integer multiple of the set round interval. If it is, calculate and update the existing physical importance score; otherwise, skip the physical importance score update step and continue the training process. S57. Based on the existing physical importance scores, calculate the total loss function of the model; based on the total loss function, update all network parameters through backpropagation; when the preset maximum number of training rounds is met, stop training and output the trained Koopman learning model.
8. A physical information-guided selective attention Koopman operator learning device, wherein the physical information-guided selective attention Koopman operator learning device is used to implement the physical information-guided selective attention Koopman operator learning method as described in any one of claims 1-7, characterized in that, The device includes: The acquisition unit is used to acquire the state sequence and control input sequence of the planar three-free robotic arm system; construct a data matrix based on the state sequence and control input sequence of the robotic arm system; construct and solve a quadratic programming problem based on the data matrix to obtain the Koopman operator; The first building unit is used to construct a deep Koopman model based on the Koopman operator and define the upscaling space prediction error loss function; and to construct a state reconstruction loss function based on the Koopman operator. The second building unit is used to construct a sparse physical perception Koopman learning model based on a planar three-free robotic arm system; wherein, the Koopman learning model includes a physical information-guided deep Koopman network model and an attention-like network. The third building unit is used to construct a physical information loss function by using the acceleration information of the planar three-free manipulator as a physical constraint; and to construct the loss function of the physical information-guided deep Koopman network model based on the prediction error loss function of the up-dimensional space, the state reconstruction loss function, and the physical information loss function. The training unit is used to construct the loss function of the attention-like network; the loss function of the deep Koopman network model guided by physical information and the loss function of the attention-like network are used to construct the total loss function of the Koopman learning model; based on the total loss function of the Koopman learning model, the Koopman learning model is trained using a three-time-scale training strategy to obtain the trained Koopman learning model. The prediction unit is used to obtain the current state and control input of the planar three-free robotic arm system; it inputs the current state and control input of the planar three-free robotic arm system into the trained Koopman learning model and outputs the state of the robotic arm system at the next moment.
9. A physical information-guided selective attention Koopman operator learning device, characterized in that, The physical information-guided selective attention Koopman operator learning device includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Omni-directional mobile mechanical arm data driving model prediction control method based on Koopman operator
CN112016194A
Soft robot control method, device and equipment based on Koopman operator and medium
CN115213908A