Robot Trajectory Generation Method and System Based on Kolmogorov-Arnold Network
By combining the Kolmogorov-Arnold network and diffusion strategy, an efficient and smooth robot motion trajectory is generated, solving the problems of environmental adaptability and trajectory optimization in the prior art, and achieving efficient motion control in complex environments.
Patent Information
- Application Number
- CN202510588478.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-08
AI Technical Summary
Existing robot trajectory generation methods are difficult to adapt to burst variables in complex dynamic environments. Traditional methods rely on precise dynamic models and lack flexibility. Deep learning-based methods generate trajectories easily jitter and difficult to optimize.
Using a robot trajectory generation method based on the Kolmogorov-Arnold network, combining convolutional neural network and Transformer architecture, smooth trajectories are generated through diffusion strategies, noise is predicted using the KAN policy network and actions are optimized through the inverse diffusion formula.
An efficient and smooth robot motion trajectory is generated, which improves adaptability and robustness in complex environments and significantly improves the trajectory optimization effect.
Smart Images

Figure CN120106146B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot motion control and trajectory optimization, and specifically relates to a robot continuous trajectory generation method based on a deep learning framework, which is particularly suitable for scenarios requiring smooth motion control, such as high-precision industrial robots and service robots. Background Art
[0002] Current robot trajectory generation and planning technologies mainly cover the following two types of methods:
[0003] The first category is the traditional method based on interpolation and optimization. Methods such as cubic spline interpolation and polynomial fitting all use mathematical modeling to generate continuous trajectories. Its advantage is that it can effectively guarantee the smoothness of the trajectory so that the robot can operate stably. However, this type of method has obvious limitations, that is, it is highly dependent on accurate dynamic models. Once in a complex dynamic environment, it is difficult to quickly adapt and adjust in the face of many sudden variables and uncertain factors, thus affecting the trajectory optimization effect.
[0004] The second category is the strategy model based on deep learning, including diffusion strategy models using CNN (convolutional neural network) and Transformer architecture, which generate trajectories based on data-driven. They perform well in complex pattern recognition and feature extraction, and can mine potential rules in massive data, bringing more flexibility to robot trajectory planning. However, its discrete data processing mechanism has drawbacks, which makes the generated trajectory prone to jitter and discontinuity, affecting the smoothness and accuracy of the robot's movements.
[0005] In addition, existing methods such as genetic algorithms and reinforcement learning can improve the adaptability of trajectories to a certain extent, making robots more adaptable in different environments and tasks. However, in high-dimensional space, these algorithms are very likely to fall into the dilemma of local optimality and it is difficult to find the global optimal solution. Moreover, they are not effectively combined with the ability to represent continuous functions, so they lack the fine characterization and efficient optimization of trajectories, which limits the overall development and performance improvement of robot trajectory optimization technology. Summary of the invention
[0006] Purpose of the invention: In view of the shortcomings of the prior art, the present invention provides a robot trajectory generation method and system based on the Kolmogorov-Arnold network, aiming to solve the problems of jitter and redundancy in trajectory generation based on the CNN / Transformer diffusion strategy model and the problem that traditional methods based on interpolation and optimization cannot adapt to complex dynamic environments.
[0007] Technical solution: In order to achieve the above invention objectives, the technical solution of the present invention is as follows:
[0008] In a first aspect, a robot trajectory generation method based on a Kolmogorov-Arnold network includes the following steps:
[0009] Extract observation data and corresponding action data from a data source file storing the robot's motion trajectory. The observation data includes the environmental and self-state information perceived by the robot, and the action data is the specific control instruction executed by the robot at each time step.
[0010] Use the action data as the object to be denoised by a diffusion model, and train a KAN policy network with the observation data as the condition. The KAN policy network is formed by integrating the Kolmogorov-Arnold network into the basic network architecture of a diffusion policy in a modular form, including a KAN policy network based on a convolutional neural network and a KAN policy network based on a Transformer, which is used to predict the noise to be added.
[0011] Select a trained KAN policy network according to the requirements of the on-site task scenario, generate predicted noise based on the actual observation data and action data of the robot, and use the predicted noise to generate the next action through the inverse diffusion formula of the diffusion model.
[0012] Furthermore, the KAN policy network based on a convolutional neural network includes three parts. The first part is a downsampling layer, which generates an embedded input based on the action data and simultaneously uses a skip connection to pass the downsampled features to the third part. The second part is an embedding layer, which, based on the downsampled features and with the observation data as the condition, uses an embedded KAN module to extract intermediate embeddings. The embedded KAN module includes a tokenization layer, a normalization layer, and a linear layer. The third part is an upsampling layer, which generates the final output based on the features of the previous two parts.
[0013] Furthermore, the tokenization layer is used to convert a one-dimensional sequence into a patch embedding representation with overlapping regions. Given an input sequence , where B is the batch size, C in represents the number of input channels, and L in is the sequence length, project it into a higher-dimensional embedding space using a one-dimensional convolutional operation:
[0014] ;
[0015] where, , W is the preset parameter of the convolution, h is the convolution kernel size, C out is the embedding dimension, LN represents layer normalization, Conv represents the convolutional operation, the convolutional operation includes a stride s and padding , and the output shape is , where L patch is determined by the following formula:
[0016] ;
[0017] The processing of the linear layer is described as:
[0018] ;
[0019] where w b and w s are the weights generated by the linear layer, and the function is defined as follows: , , B i (x) represents the basis function, i is the index, and c i represents the weight of each basis function;
[0020] The final output f(x) of the embedded KAN module is:
[0021] ;
[0022] where a and b are the conditional features extracted by the conditional encoder from the global feature observation data and the time step.
[0023] Furthermore, the Transformer-based KAN policy network includes an encoder, a decoder, and a GR-KAN module. The encoder generates a conditional embedding for the combination of the observation data and the time step. The encoded hidden variable and the action data are input into the decoder together. The features output by the decoder are non-linearly transformed through the GR-KAN module to obtain the final result.
[0024] Furthermore, the GR-KAN module includes a linear layer and a grouped KAT. The formula of the whole module is described as:
[0025] ;
[0026] where w1, w2 and b1, b2 are the weights and biases generated by the two linear layers, K is a rational function, and the superscript of K represents different initialization modes of the grouped KAT.
[0027] Assume that i is the index of the current input channel and j is the index of the output channel . The whole input channel is evenly divided into g groups. is the group number, then K is expressed as:
[0028] ;
[0029] where is the weight coefficient and R(x) is a rational function.
[0030] Further, during the training process of the KAN policy network, the goal is to minimize the mean square error between the predicted noise and the true noise.
[0031] Further, the inverse diffusion formula of the diffusion model is expressed as:
[0032] ;
[0033] where is the action at time step, , , are the parameter settings of the noise, is the noise predicted by the network, are the network parameters, is the observed data at time t, is the noise following a specified distribution.
[0034] In a second aspect, a robot trajectory generation system based on a Kolmogorov - Arnold network includes:
[0035] A data processing module, configured to extract observed data and corresponding action data from a data source file storing the robot's motion trajectory. The observed data includes the environmental and self - state information perceived by the robot, and the action data is the specific control instruction executed by the robot at each time step;
[0036] A model training module, configured to use the action data as the object to be denoised by the diffusion model, and train the KAN policy network with the observed data as the condition. The KAN policy network is formed by integrating the Kolmogorov - Arnold network into the basic network architecture of the diffusion policy in a modular form, including a KAN policy network based on a convolutional neural network and a KAN policy network based on a Transformer, and is used to predict the noise to be denoised;
[0037] A trajectory generation module, configured to select a trained KAN policy network according to the requirements of the on - site task scenario, generate predicted noise based on the actual observed data and action data of the robot, and generate the next action using the predicted noise through the inverse diffusion formula of the diffusion model.
[0038] In a third aspect, an electronic device includes: one or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors. When the program is executed by the processor, it implements the steps of the robot trajectory generation method based on the Kolmogorov - Arnold network as described in the first aspect.
[0039] Fourth aspect, a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the robot trajectory generation method based on the Kolmogorov-Arnold network as described in the first aspect are implemented.
[0040] Advantageous effects: The present invention integrates the Kolmogorov-Arnold network (KAN) into the basic network architecture of the diffusion strategy in a modular form. For the convolutional neural network CNN, a novel modular structure Emb-KAN is proposed; for Transformer, the GR-KAN module is introduced and expanded. This can not only generate effective and smooth trajectories, but also significantly improve the overall performance. The effectiveness of the method is verified in simulated and real-world robot control tasks, confirming its practical application value and robustness in different environments. Description of the drawings
[0041] Figure 1 is the KAN Policy-C framework diagram according to the present invention;
[0042] Figure 2 is the structural diagram of the Emb-KAN module according to the present invention;
[0043] Figure 3 is the KAN Policy-T framework diagram according to the present invention;
[0044] Figure 4 is the structural diagram of the GR-KAN module according to the present invention;
[0045] Figure 5 is the schematic diagram of shared parameters according to the present invention;
[0046] Figure 6 is the flowchart of the trajectory generation method according to the present invention;
[0047] Figure 7 is an example of the change in trajectory and efficiency smoothness according to the present invention. Detailed implementation manners
[0048] The technical solution of the present invention will be further described below with reference to the drawings.
[0049] The present invention combines traditional robot trajectory generation methods based on interpolation and optimization, uses the Kolmogorov-Arnold network based on an efficient fitting function (such as a spline function) and combines the diffusion strategy based on CNN / Transformer to integrate the advantages of both, so as to generate an efficient and smooth robot motion trajectory. To have a clearer understanding of the invention purpose, means and beneficial effects, the related technologies will be introduced first.
[0050] Kolmogorov-Arnold Network
[0051] The Kolmogorov-Arnold Network (KAN) is a neural network structure based on the Kolmogorov-Arnold representation theorem. This theorem was proposed by mathematicians Andrey Kolmogorov and Vladimir Arnold in the 1950s to solve the problem of representing multivariate functions. The theorem states that any multivariate continuous function can be represented as a linear combination of a series of univariate functions, in the mathematical form of:
[0052] ,
[0053] where and are univariate continuous functions.
[0054] Compared with traditional neural networks, the key innovations of KAN are as follows: ① Activation function position: The activation function is placed at the edge of the network, rather than at the traditional node position. ② Weight parameter replacement: Each weight parameter is replaced by a learnable univariate function, usually parameterized as a B-spline function, which can flexibly simulate complex non-linear relationships. This design not only endows the network with higher expressive power but also provides significant interpretability advantages.
[0055] Diffusion Policy
[0056] The Diffusion Policy is a data-driven generative model that generates or optimizes data by gradually adding and removing noise. This strategy has achieved remarkable results in the field of robot policy generation. Traditional policy learning methods often have limited distribution expression ability and stability problems when dealing with multi-modal action distributions or uncertainties. The Diffusion Policy introduces a denoising diffusion probability model into robot motion generation, modeling the robot's visual motion policy as a conditional denoising diffusion process. Its reverse diffusion formula is:
[0057] ,
[0058] where is the action at time step, , , are the parameter settings of the noise, is the noise predicted by the network, are the network parameters, is the observation at time t, that is, the current situation observed by the robot, is the noise that follows this distribution. The loss function corresponding to such a diffusion process can be defined as:
[0059] ,
[0060] By learning the gradient of the action score function and performing stochastic Langevin dynamics sampling on this gradient field, this method can represent any normalizable distribution, including multimodal action distributions. At the same time, it has two basic network architectures: convolutional neural network (CNN) and Transformer.
[0061] Diffusion Strategy Based on KAN Network
[0062] The inventor modified the diffusion strategy and incorporated KAN into the basic network architecture of the diffusion strategy in a modular form. For CNN, the present invention created a novel modular structure Emb-KAN; for Transformer, the present invention introduced the GR-KAN module and extended it. The present invention refers to these two structures as KAN Policy-C (KAN strategy based on CNN)) and KAN Policy-T (KAN strategy based on Transformer), collectively referred to as KAN Policy. This invention is a transformation of the basic network architecture of the diffusion strategy at the algorithm level. The specific method will be introduced below:
[0063] (I) KAN Policy-C
[0064] The inventor modified the CNN architecture in the diffusion strategy and divided the main structure into three parts. As Figure 1 shown, the first part is the downsampling layer, which is responsible for generating the embedded input and passing the downsampled features to the third part using skip connections. The downsampling layer includes a series of convolutional blocks and downsampling blocks; the second part is the embedding layer, which is the most critical modified part of this method. Specifically, a new module called Emb-KAN is designed and integrated, with the full name Embedding KAN, which can be called Embedded KAN. This newly developed Emb-KAN module serves as the intermediate embedding layer in the architecture and is the key mechanism for high-dimensional feature extraction. Utilizing the function of Emb-KAN, the entire model can capture more detailed and discriminative representations, thereby generating more efficient actions. The third part is the upsampling layer, which includes a series of convolutional blocks and upsampling blocks. It integrates the feature information of the first two parts and then generates the output through the final convolutional block. The global feature is the global condition generated by the observation and time step, which will be processed by the conditional encoder.
[0065] Referring to Figure 2 , the specific implementation of Emb-KAN is as follows:
[0066] Tokenization is designed to convert a one-dimensional sequence into a patch embedding representation with overlapping regions, which is particularly useful for extracting local features from sequential data while maintaining the continuity between patches. The key operations are described as follows: Given an input sequence , where is the batch size, represents the number of input channels, is the sequence length, and it is projected into a higher-dimensional embedding space using a one-dimensional convolutional operation:
[0067] ,
[0068] where , W is a preset parameter of the convolution, h is the kernel size of the convolution, is the embedding dimension, and LN is layer normalization. Conv represents the convolution operation, and the convolution operation includes a stride and padding to ensure patch overlap. The output shape is , where is determined by the following formula:
[0069] ,
[0070] Then, through layer normalization, the features along the embedding dimension are processed. After completing the patch embedding representation of the features, the main feature processing of this module is performed by the KAN linear layer, and its output can be mathematically described as:
[0071] ,
[0072] where and are the weights generated by the linear layer. is the silu activation function, and the function is defined as follows: . , represents the basis function (spline function), is the index, represents the weight of each basis function, that is, the coefficient. This method constructs an interpolation framework in the input feature space and can extract both linear and nonlinear features. In addition, the conditional encoder is responsible for extracting conditional features a and b from the global features. Specifically, the conditional encoder FiLM (activation function + linear layer + reconstruction) generates an embedding representation, and the embedding representation is reshaped (e.g., by adjusting the dimensions) into a and b.
[0073] The final output f(x) of Emb-KAN has the following structure:
[0074] ,
[0075] So far, all the processing procedures of the Emb-KAN module have been completed.
[0076] (2) KAN Policy-T:
[0077] In this structure, instead of using a spline function as the basis function, a rational function is adopted, and the method of sharing parameters is used, in groups. Refer to Figure 3 , KAN Policy-T adopts a custom GR-KAN module to replace the linear layer at the end of the Transformer architecture. This is an operation similar to an activation function, introducing non-linearity at the end of the model to enable the neural network to learn complex non-linear relationships. KAN Policy-T combines observations and time steps to generate conditional embeddings, which are input as hidden variables after being encoded by the encoder, and together with the action noise are input into the decoder, and the output features pass through GR-KAN to obtain the final result.
[0078] Refer to Figure 4 , GR-KAN introduces Group KAT to achieve complex non-linear transformations, thereby enhancing the representation ability of the network. The formula of the entire module can be described as:
[0079] ,
[0080] Among them, , and , are the weights and biases generated by two linear layers, is a rational function, 's superscript represents different initialization modes of Group KAT. Group KAT shares parameters for input groups and is the main part of the KAN with rational basis functions, responsible for forming features with the KAN style.
[0081] Assume that i is the index of the current input channel , j is the output channel , the entire input channel is evenly divided into g groups, is the group number, so can be expressed as:
[0082] ,
[0083] is the weight coefficient. This means that the input features of the same group will share parameters. As Figure 5As shown, compared with the general KAN on the left, the GR-KAN on the right shares the weights generated by the network in groups.
[0084] R(x) is a rational function and can be expressed as:
[0085] ,
[0086] where m and n are the orders of the polynomials. , are coefficients and also weights. In practice, GR-KAN can be applied to various policy learning models. When integrated into different models, it can also achieve the effect of trajectory optimization.
[0087] Robot Trajectory Generation Method Based on KAN Policy
[0088] The following presents the specific method for the robot trajectory planning and generation based on the proposed KAN Policy of the present invention. Referring to Figure 6 , the method includes the following steps:
[0089] Step 1: Data processing, extracting observations and corresponding actions from the data source file.
[0090] Collect data from the simulation environment (such as MuJoCo) or the log file generated during the operation of a real robot, which includes two parts: observation data and action data. The observation data includes the robot state (position, speed), environmental information (target point, obstacle), and sensor data (image, depth map, etc.). The action data is the corresponding real action sequence, such as the joint angles and end effector coordinates recorded when the robot executes a task.
[0091] Perform the following processing on the data:
[0092] Normalization: Normalize the observation and action data respectively to eliminate the dimension difference.
[0093] Sequence segmentation: Use the sliding window method to cut the data into continuous segments of a fixed length (such as 30 steps as a sequence).
[0094] Package the processed sequences into a PyTorch TensorDataset and load them in batches through a DataLoader to support parallel training.
[0095] Step 2: Use the action as the object to be denoised by the diffusion model, and train the network with the observation as the condition. The network is mainly used to predict the denoising noise.
[0096] The network architecture is as described above. In the robot trajectory generation task based on the diffusion strategy, KAN Policy-C (based on CNN) and KAN Policy-T (based on Transformer) are two different network architecture designs, and the selection needs to be weighed in combination with the task characteristics and data features. Table 1 shows the specific analysis of the two networks.
[0097] Table 1 Network Architecture Differences and Applicable Scenarios
[0098] Characteristic KAN Policy-C (CNN) KAN Policy-T (Transformer) Core module Emb-KAN embedding layer + convolutional down / upsampling GR-KAN module + self-attention mechanism Feature extraction ability Good at capturing local spatial features (such as joint states, obstacle distributions) Good at modeling global temporal dependencies (such as long-sequence action coherence) Computational efficiency Larger number of parameters, lower computational efficiency Smaller number of parameters, higher computational efficiency Data requirements No special requirements for the dataset, there may be requirements for the machine performing the task (such as degrees of freedom) All current experiments have no special requirements for the dataset
[0099] Based on the above analysis, in practical applications, the corresponding network architecture can be selected according to the task scenario. For example, for tasks such as robotic arm grasping static targets and quadruped robot short-distance gait control where action generation depends on local observation features and has weak temporal dependence, KAN Policy-C is selected; for tasks such as humanoid robot continuous obstacle crossing and UAV dynamic path replanning where action generation requires modeling long-range temporal dependence and the observed data contains complex temporal relationships, KAN Policy-T is selected.
[0100] The network training process is as follows:
[0101] Input: Action data , Observation data , Time step ;
[0102] Objective: Minimize the mean square error (MSE) between the predicted noise and the true noise;
[0103] Optimizer: Use AdamW, set the learning rate to 1e-4, and gradient clipping can be added to prevent training instability.
[0104] For KAN Policy-C, the action data generates embedded inputs through the downsampling layer, and at the same time, the downsampled features are passed to the upsampling layer through skip connections; the embedded inputs then extract high-dimensional intermediate embeddings through the Emb-KAN module, using the observation data and the time step k as constraints or preconditions to guide the network learning process. The time step k is converted into a vector through sine / cosine processing and concatenated after the observation data sequence . Multiple Emb-KAN modules can be applied, and the extracted intermediate embeddings are integrated through skip connections. Finally, the upsampling layer integrates the feature information of the first two parts to generate the final predicted noise.
[0105] For KAN Policy-T, the observation data and the time stepk The combined generation conditional embedding, after being encoded by the encoder, is input into the decoder together with the action noise The output features are non-linearly transformed by the GR-KAN module to obtain the prediction result.
[0106] Step 3: Use the predicted noise to generate the current action through the inverse diffusion formula of the diffusion model. The inverse diffusion formula has been described above and will not be elaborated here.
[0107] To verify the performance of the proposed method, the inventors conducted some experiments, and the experimental results are as follows:
[0108] Experiment 1: Verify the performance of the KAN strategy in generating trajectories. As Figure 7 shown, on the intuitive trajectory, the KAN strategy (KP-C) has a smoother action trajectory than the CNN-based diffusion model (DP-C) (the figure contains three important nodes of the task); in terms of completion efficiency, the time to complete a single task is shortened by 10.7%; the curvature drops by nearly six times. The experimental results strongly prove the effectiveness of modularizing and integrating the Kolmogorov-Arnold network (KAN) into the diffusion strategy. This method makes full use of the powerful non-linear expression ability of KAN to generate more efficient and smooth trajectories.
[0109] Experiment 2: Verify the robustness of the KAN strategy to data. The multi-human dataset (Multi-Human) is collected by operators with different skill levels and may show significant differences during the motion planning process. For example, there may be large differences in the trajectory length, and there may also be noise or errors in the robot motion, such as grasping failures. Three tasks, Lift, Can, and Square, are designed. Lift: Use the robotic arm to pick up and lift a small cube; Can: Use the robotic arm to transport a cylinder to a specified area; Square: Use the robotic arm to pick up a collar and put it on a pillar of the correct shape. The KAN strategy of the present invention also achieves significant improvement on the multi-human dataset, as shown in Table 2. Where Succ is the success rate, Time is the task completion time, and Cur is the curvature. It shows that the KAN strategy has strong robustness to changes in the dataset and can improve performance even on relatively poor datasets.
[0110] Table 2 Performance of the KAN strategy on the multi-human dataset
[0111]
[0112] Experiment 3: Verify the applicability of the GR-KAN module to various policy learning models. By integrating the GR-KAN module at the end of various policy learning architectures, significant performance improvements were observed. All methods benefited from the integration of GR-KAN, and significant performance gains were observed on multiple metrics, as shown in Table 3. More importantly, models that initially performed poorly also showed substantial improvements after integrating GR-KAN. This indicates that GR-KAN is not only an effective enhancement module for diffusion policy models but also has the potential to improve the performance of other models, thus establishing its wide application in the policy learning paradigm.
[0113] Table 3 Performance of the GR-KAN module integrated into various policy learning models
[0114]
[0115] Experiment 4: Verify the performance of the KAN policy in different action spaces. The experiments show that in a larger action space, the KAN policy still performs well, as shown in Table 4. Here, Succ is the success rate, Time is the task completion time, and Coverage is the target coverage area. When the maximum number of steps in the action space increases from 300 to 1000 (extending the maximum completion deadline of the task), the method of the present invention can complete the task completely and within a relatively short time. This indicates that the KAN policy (KP-C and KP-T) can still maintain high efficiency and stability in a larger action space and has better task final completion.
[0116] Table 4 Performance of the KAN policy in different action spaces
[0117]
[0118] An embodiment of the present invention also provides a robot trajectory generation system based on a Kolmogorov-Arnold network, including:
[0119] A data processing module, configured to extract observation data and corresponding action data from a data source file storing the robot's motion trajectory. The observation data includes the environment and the robot's own state information perceived by the robot, and the action data is the specific control instruction executed by the robot at each time step;
[0120] A model training module, configured to use the action data as the object to be denoised by a diffusion model and train a KAN policy network with the observation data as the condition. The KAN policy network is formed by integrating the Kolmogorov-Arnold network into the basic network architecture of the diffusion policy in a modular form, including a KAN policy network based on a convolutional neural network and a KAN policy network based on a Transformer, for predicting the noise to be denoised;
[0121] A trajectory generation module, configured to select a trained KAN policy network according to the requirements of the on-site task scenario, generate predicted noise based on the actual observation data and action data of the robot, and generate the next action by using the predicted noise through the inverse diffusion formula of the diffusion model.
[0122] It should be understood that the robot trajectory generation system based on the Kolmogorov-Arnold network in the embodiments of the present invention can implement all the technical solutions in the above method embodiments. The functions of its respective functional modules can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions in the above embodiments, which will not be elaborated here.
[0123] The present invention also provides an electronic device, including: one or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, it implements the steps of the robot trajectory generation method based on the Kolmogorov-Arnold network as described above.
[0124] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the robot trajectory generation method based on the Kolmogorov-Arnold network as described above.
[0125] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a device (system), an electronic device, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0126] The present invention is described with reference to the flowchart of the method according to the embodiments of the present invention. It should be understood that each process in the flowchart and the combination of the processes in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 process or multiple processes.
[0127] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means embodying the function specified in one or more of the procedures Figure 1 for one or more of the procedures.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the procedures Figure 1 for one or more of the procedures.
Claims
1. A robot trajectory generation method based on the Kolmogorov - Arnold network, characterized in that, It includes the following steps: Extract the observation data and the corresponding action data from the data source file storing the robot motion trajectory. The observation data includes the environmental and self-state information perceived by the robot, and the action data is the specific control instruction executed by the robot at each time step; Use the action data as the object to be denoised by the diffusion model, and train the KAN policy network with the observation data as the condition. The KAN policy network is formed by integrating the Kolmogorov-Arnold network into the basic network architecture of the diffusion policy in a modular form, including the KAN policy network based on the convolutional neural network and the KAN policy network based on the Transformer, which is used to predict the noise to be denoised. Among them, the KAN policy network based on the convolutional neural network includes three parts. The first part is the downsampling layer, which generates the embedded input according to the action data, and at the same time uses the skip connection to transfer the downsampled features to the third part. The second part is the embedding layer, which extracts the intermediate embedding using the embedding KAN module with the observation data as the condition according to the downsampled features. The embedding KAN module includes a tokenization layer, a normalization layer, and a linear layer. The tokenization layer performs a convolution operation on the input sequence through preset convolution parameters. The normalization layer normalizes the convolution result of the tokenization layer and the weighted sum of the linear layer. The linear layer performs a weighted sum on the convolution normalization results processed by different functions. The third part is the upsampling layer, which generates the final output according to the features of the previous two parts; The Transformer-based KAN policy network includes an encoder, a decoder, and a GR-KAN module. The encoder generates a conditional embedding for the combination of observation data and time steps. The encoded hidden variables and action data are input into the decoder together. The features output by the decoder are non-linearly transformed through the GR-KAN module to obtain the final result. The calculation formula of the GR-KAN module is ; where w1, w2, b1, and b2 are the weights and biases generated by two linear layers, K is a rational function, the superscript of K represents different initialization modes of grouped KAT, and grouped KAT means sharing parameters for input features in the same group; Select the trained KAN policy network according to the requirements of the on-site task scenario, generate the predicted noise based on the actual observation data and action data of the robot, and use the predicted noise to generate the next action through the inverse diffusion formula of the diffusion model. The processing process of the KAN policy network based on the convolutional neural network is as follows: The action data generates the embedded input through the downsampling layer, and at the same time transfers the downsampled features to the upsampling layer through the skip connection; The embedded input then extracts the high-dimensional intermediate embedding through the embedding KAN module, using the observation data and the time step as the constraints or preconditions to guide the network learning process. The time step is converted into a vector through sine / cosine processing and concatenated after the observation data sequence; Finally, the upsampling layer integrates the feature information of the previous two parts to generate the final predicted noise. The processing process of the KAN policy network based on the Transformer is as follows: Combine the observation data and the time step to generate the conditional embedding, input it into the decoder together with the action noise after being encoded by the encoder, and the output features are non-linearly transformed by the GR-KAN module to obtain the prediction result.
2. The method according to claim 1, wherein The tokenization layer is used to convert a one-dimensional sequence into a patch embedding representation with overlapping regions, given the input sequence , where B is the batch size, C in represents the number of input channels, L in is the sequence length, and it is projected into a higher-dimensional embedding space using a one-dimensional convolutional operation: ; Among them, , W is a preset parameter of convolution, h is the kernel size, and C out is the embedding dimension, LN represents layer normalization, Conv represents the convolution operation, and the convolution operation includes the stride s and padding , and the output shape is , where L patch is determined by the following formula: ; The processing description of the linear layer is: ; Among them, w b and w s are weights generated by a linear layer, and the function is defined as follows: , , B i (x) represents the basis function, i is the index, and c i represents the weight of each basis function; The final output f(x) of the embedding KAN module is: ; Among them, a and b are the conditional features extracted by the conditional encoder from the global feature observation data and the time step.
3. The method according to claim 1, wherein In the GR-KAN module, assume that i is the index of the current input channel and j is the index of the output channel . The entire input channel is evenly divided into g groups, where is the group number. Then K is expressed as: ; where is the weight coefficient, and R(x) is a rational function.
4. The method according to claim 1, characterized in that, During the training process of the KAN policy network, the goal is to minimize the mean square error between the predicted noise and the real noise.
5. The method according to claim 1, characterized in that The inverse diffusion formula of the diffusion model is expressed as: ; Among them, is the action at the time step, , , are the parameter settings of the noise, is the noise predicted by the network, are the network parameters, is the observation data at time t, is the noise following a specified distribution.
6. A robot trajectory generation system based on the Kolmogorov-Arnold network, characterized in that, It includes: A data processing module, configured to extract observation data and corresponding action data from a data source file storing the robot's motion trajectory. The observation data includes the environmental and self-state information perceived by the robot, and the action data is the specific control instructions executed by the robot at each time step; A model training module, configured to use the action data as the object to be denoised by a diffusion model, and use the observation data as a condition to train a KAN policy network. The KAN policy network is formed by integrating the Kolmogorov-Arnold network into the basic network architecture of the diffusion policy in a modular form, and includes a KAN policy network based on a convolutional neural network and a KAN policy network based on a Transformer, which is used to predict the noise to be denoised. Among them, the KAN policy network based on a convolutional neural network includes three parts. The first part is a downsampling layer, which generates an embedded input according to the action data, and at the same time uses a skip connection to transfer the downsampled features to the third part. The second part is an embedding layer, which, according to the downsampled features and with the observation data as a condition, uses an embedded KAN module to extract intermediate embeddings. The embedded KAN module includes a tokenization layer, a normalization layer, and a linear layer. The tokenization layer performs a convolution operation on the input sequence through preset convolution parameters. The normalization layer normalizes the convolution result of the tokenization layer and the weighted sum of the linear layer. The linear layer performs a weighted sum on the convolution normalization results processed by different functions. The third part is an upsampling layer, which generates a final output according to the features of the previous two parts; The Transformer-based KAN policy network includes an encoder, a decoder, and a GR-KAN module. The encoder generates conditional embeddings for the combination of observation data and time steps. The encoded hidden variables and action data are input into the decoder together. The features output by the decoder are non-linearly transformed through the GR-KAN module to obtain the final result. The calculation formula of the GR-KAN module is ; where w1, w2, b1, and b2 are weights and biases generated by two linear layers, K is a rational function, and the superscript of K represents different initialization modes of grouped KAT. Grouped KAT means sharing parameters for input features in the same group; A trajectory generation module, configured to select a trained KAN policy network according to the requirements of the on-site task scenario, generate predicted noise according to the actual observation data and action data of the robot, and generate the next action through the inverse diffusion formula of the diffusion model. The processing process of the KAN policy network based on a convolutional neural network is as follows: the action data generates an embedded input through the downsampling layer, and at the same time transfers the downsampled features to the upsampling layer through a skip connection; the embedded input then extracts high-dimensional intermediate embeddings through the embedded KAN module, and uses the observation data and the time step as constraints or preconditions to guide the network learning process. The time step is converted into a vector through sine / cosine processing and spliced after the observation data sequence; finally, the upsampling layer integrates the feature information of the previous two parts to generate the final predicted noise. The processing process of the KAN policy network based on a Transformer is as follows: the observation data and the time step are combined to generate a conditional embedding, which is encoded by the encoder and then input into the decoder together with the action noise, and the output features are non-linearly transformed by the GR-KAN module to obtain the prediction result.
7. An electronic device, characterized in that, Comprising: One or more processors; A memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors. When the program is executed by the processor, the steps of the robot trajectory generation method based on the Kolmogorov-Arnold network as described in any one of claims 1-5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the robot trajectory generation method based on the Kolmogorov-Arnold network according to any one of claims 1-5.
Citation Information
Patent Citations
MRI image segmentation method based on WKA-Unet
CN119540553A
Satellite network data anomaly detection method based on CNN and RWav-KAN
CN119854795A