Sequence modeling data enhancement method and system based on trajectory branch generation
Through the trajectory branch generation method, trajectory value function and diffusion model are used to generate trajectory branches with higher returns, solving the problem of poor strategy performance caused by the training of sequence modeling methods on suboptimal trajectories, and improving the learning effect of the model and its ability to adapt to complex environments.
Patent Information
- Application Number
- CN202510393610.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-11
AI Technical Summary
When the sequence modeling method is trained on the suboptimal trajectory, the learned strategies are likely to converge to the suboptimal trajectory, resulting in poor strategy performance.
通过轨迹分支生成的方法,利用轨迹价值函数和扩散模型生成通向更高回报的轨迹分支,扩展机器人轨迹数据集,引导模型学习更优策略。
The strategic performance of the sequence modeling method is improved, and by generating more diverse trajectory data, the generalization ability and robustness of the model are enhanced, helping the model make accurate predictions in complex environments.
Smart Images

Figure CN120297440A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep reinforcement learning, and particularly relates to a sequence modeling data enhancement method and system based on trajectory branch generation. Background Art
[0002] For the existing technical solutions of data enhancement methods, there are TATU, DiffStitch, and TS. The core idea of the TATU method is to control the generation of synthetic data by estimating the uncertainty in the trajectory and avoid using unreliable data. In offline reinforcement learning, the model often generates synthetic trajectories through a dynamics model to enhance the training data. However, the generated trajectories may exceed the support range of the original dataset, resulting in model bias. To solve this problem, TATU proposes to truncate the current trajectory when the accumulated uncertainty exceeds a certain threshold during the trajectory generation process to ensure that only reliable data is used. TATU includes three steps: first, training an environment dynamics model, then constructing an ϵ-pessimistic Markov decision process (MDP), and finally generating synthetic trajectories and truncating them according to the uncertainty.
[0003] The DiffStitch method aims to improve the performance of offline reinforcement learning through trajectory stitching. DiffStitch mainly focuses on low-return trajectories and high-return trajectories in the offline dataset. By generating stitched trajectories that connect them, it makes up for the deficiencies in the dataset and helps the reinforcement learning algorithm better learn how to transition from low-return regions to high-return regions. The method first estimates the number of stitching steps required between two trajectories, then generates a state sequence that conforms to the environment dynamics, then pairs corresponding actions and rewards for these state sequences, and finally constructs a complete stitched trajectory. To ensure the quality of the generated trajectories, DiffStitch also evaluates them through a dynamics model to ensure that the generated trajectories conform to the laws of the environment dynamics. In this way, DiffStitch effectively enhances the offline dataset, enabling the reinforcement learning algorithm to obtain more high-quality training samples, thereby improving the effect of its policy learning.
[0004] TS aims to improve the performance of Behavioral Cloning (BC) by enhancing the dataset of offline reinforcement learning. The core idea of TS is to generate new trajectories by stitching high-value states in different trajectories to replace the original low-return trajectories, thereby improving the quality of the entire dataset. Specifically, TS determines the connectable states between two trajectories through the dynamic model of the environment and the state value function, and generates a new action to connect these states. The stitching event occurs when the state in one trajectory jumps to the state in another trajectory, and it will be executed when this jump is expected to improve the future return of the trajectory. Each stitching step ensures the reachability of the state through a learned dynamic model, generates appropriate actions using the inverse dynamic model, and predicts the return of the stitching event through the reward generation model.
[0005] TATU performs trajectory rollout through the forward dynamic model and controls uncertainty through the truncation mechanism. However, its drawback is that the distance of its trajectory rollout is short, and it is difficult to guarantee the quality of long-distance rollout, which limits its application in sequence modeling. DiffStitch performs data augmentation by randomly stitching two trajectories in the dataset, but its disadvantage is that this random selection is difficult to ensure the quality of the stitching segment, which may lead to unreliable generated trajectories. The TS method stitches trajectories by searching for the next better state. Although it can connect to better trajectories, it lacks the ability to generate new trajectories and is overly dependent on the value function and dynamic model, which limits its flexibility to a certain extent.
[0006] In offline reinforcement learning, sequence modeling methods mainly rely on capturing the complex dependencies between state-action-reward sequences. Different from traditional policy optimization methods based on the current state and action, sequence modeling methods regard the reinforcement learning task as a process of dealing with time series problems. In this way, the model not only considers the state and action at the current moment, but also combines historical information such as past states, actions, and rewards to capture long-range temporal dependencies.
[0007] However, sequence modeling methods learn the trajectory sequences in the dataset through supervised learning. When the trajectory sequences in the dataset do not appear sufficiently and there are some sub-optimal trajectories at the same time. When the sequence modeling method is trained on sub-optimal trajectories, the learned policy tends to converge to the sub-optimal trajectories, resulting in poor performance of the learned policy. Summary of the Invention
[0008] The object of the present invention is to overcome the problem that when the sequence modeling method is trained on suboptimal trajectories, the learned policy tends to converge to the suboptimal trajectories, resulting in poor performance of the learned policy. The present invention proposes a sequence modeling data augmentation method and system based on trajectory branch generation, which expands the trajectories in the dataset by generating trajectory branches leading to better trajectories through a diffusion model, so as to prevent the sequence modeling method from converging to the suboptimal trajectories in the dataset, and finally improve the performance of the policy learned by the sequence modeling method.
[0009] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a sequence modeling data augmentation method based on trajectory branch generation, including the following steps: Sampling trajectory segments in the trajectory dataset of the robot; Using a trajectory value function to generate the future return of the predicted trajectory segment for the trajectory segment; Taking the trajectory segment and the future return of the trajectory segment as combined conditions, and generating trajectory branches through a diffusion model; Connecting the trajectory branches with the trajectory segment and the future return of the trajectory segment to obtain an extended trajectory; Filtering the extended trajectory through a branch filter to obtain trajectory branches.
[0010] Further, the trajectory value function takes the Q value of the last state-action pair of the predicted trajectory segment as the future return of the trajectory segment; The trajectory value function is:
[0011]
[0012] Among them, is the expected regression estimation target of the trajectory value function, D is the trajectory dataset of the robot, is the constraint of the trajectory value function, V is the state value, is the future return of the predicted trajectory segment, , , n is the trajectory index of the segment, N represents the total number of trajectories in the trajectory dataset of the robot, is the state-action pair at time t, is the discount factor.
[0013] Further, the diffusion model is constructed based on the EMD formula; The combined conditions of the diffusion model are:
[0014]
[0015]
[0016] Among them, represents the combined condition of the diffusion model, K represents the length of the segment, is the discount factor, is the state at time t, is the action at time t, is the environment at time t, represents the trajectory segment, represents the future return of the trajectory segment, , n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset.
[0017] Furthermore, the loss function of the diffusion model is:
[0018] Among them, is the loss function of the diffusion model, represents the subsequent trajectory segment of in the real trajectory, represents , , n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset.
[0019] Furthermore, the branch filter extends the trajectory by filtering based on the continuity of the return between the trajectory branch and the trajectory segment.
[0020] Furthermore, the filtering formula of the branch filter is:
[0021]
[0022] Among them, is the Q value of the state-action pair at time t, , n is the trajectory index of the segment, N represents the total number of trajectories in the robot's trajectory dataset, is the discount factor, is the reward function of the environment, is the state-action value function, is the value function corresponding to choosing the action in the state , H is the length of the generated trajectory branch, is the threshold hyperparameter.
[0023] In a second aspect, the present invention provides a robot, which adopts the sequence modeling data augmentation method based on trajectory branch generation in an off-line data training data augmentation task scenario.
[0024] In a third aspect, the present invention provides a sequence modeling data augmentation system based on trajectory branch generation, comprising: A sampling module, configured to sample trajectory segments from a trajectory dataset of a robot; A future reward prediction module, configured to generate the future reward of a predicted trajectory segment for a trajectory segment by using a trajectory value function; A trajectory branch generation module, configured to use the trajectory segment and the future reward of the trajectory segment as combination conditions, and generate trajectory branches through a diffusion model; A connected extended trajectory module, configured to connect a trajectory branch with a trajectory segment and the future reward of the trajectory segment to obtain an extended trajectory; A filtered trajectory branch module, configured to filter the extended trajectory through a branch filter to obtain a trajectory branch.
[0025] In a fourth aspect, the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the sequence modeling data augmentation method based on trajectory branch generation is implemented.
[0026] In a fifth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the sequence modeling data augmentation method based on trajectory branch generation is implemented.
[0027] Compared with the prior art, the present invention has the following beneficial technical effects: The sequence modeling data augmentation method based on trajectory branch generation proposed by the present invention expands the trajectories in the trajectory dataset of the robot through trajectory branches, so that the sequence modeling method converges to a sub-optimal trajectory. The branches will guide the trajectory to a path with higher rewards, providing more opportunities for the sequence modeling method to learn strategies that can turn to a better trajectory instead of continuing along the sub-optimal trajectory. In the case of having branches, the agent can deviate from the sub-optimal trajectory by learning actions from the branches, thereby obtaining a better strategy. A diffusion-based trajectory branch generation method (Branch Generate, BG) is proposed to generate branches according to the trajectory segments in the dataset using a diffusion model, and a trajectory value function (Trajectory Value Function, TVF) is used to guide the generation process, so that the generated branches can guide the trajectory to a path with higher rewards. The generated branches are connected to the trajectory segments as an extension of the trajectory. The generated branches provide more opportunities for the Decision Transformer (DT) to branch out from the sub-optimal trajectory and then learn a better strategy. The present invention aims to generate trajectory branches that can guide the trajectory to obtain higher rewards and use the generated branches to expand the trajectories in the dataset. Trajectory branches are generated using a diffusion model based on segments sampled from the trajectories of the robot's trajectory dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure of the present invention in any way. Additionally, the shapes and proportional dimensions of the components in the figures are only schematic and are used to assist in understanding the present invention, rather than specifically defining the shapes and proportional dimensions of the components of the present invention. In the drawings: Figure 1 It is a flowchart of the sequence modeling data augmentation method based on trajectory branch generation of the present invention.
[0029] Figure 2 It is a structural diagram of the sequence modeling data augmentation system based on trajectory branch generation of the present invention.
[0030] Figure 3 It is a diagram of an electronic device of the sequence modeling data augmentation method based on trajectory branch generation of the present invention.
[0031] Figure 4 It is a flowchart of the BG method of the sequence modeling data augmentation method based on trajectory branch generation of the present invention.
[0032] Figure 5 It is an illustration diagram of the effect of the sequence modeling data augmentation method based on trajectory branch generation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] Abbreviations and definitions of key terms: TS: Trajectory Stitching, trajectory stitching algorithm.
[0035] DiffStitch: Trajectory stitching algorithm based on diffusion model.
[0036] TATU: Trajectory stitching data augmentation method based on uncertainty.
[0037] D4RL: Data-driven reinforcement learning dataset.
[0038] DT: Decision Transformr, decision transformer model.
[0039] CQL: Offline reinforcement learning algorithm based on conservative Q-learning.
[0040] IQL: Offline reinforcement learning algorithm based on implicit Q-value learning.
[0041] Embodiment 1 See Figure 1 , a sequence modeling data augmentation method based on trajectory branch generation, comprising the following steps: Sampling trajectory segments in the trajectory dataset of the robot; Generating the future return of the predicted trajectory segment for the trajectory segment using the trajectory value function; Using the trajectory segment and the future return of the trajectory segment as combined conditions to generate a trajectory branch through a diffusion model; Connecting the trajectory branch with the trajectory segment and the future return of the trajectory segment to obtain an extended trajectory; Filtering the extended trajectory through a branch filter to obtain a trajectory branch.
[0042] In this embodiment, by sampling trajectory segments and generating trajectory branches, the sample diversity in the robot's trajectory dataset is effectively increased. This helps the model learn richer features and patterns during training, thereby improving the model's generalization ability. By introducing a trajectory value function to evaluate the future return of trajectory segments and combining it with a diffusion model to generate trajectory branches, simulated data closer to real-world situations can be generated. These high-quality simulated data can be used as training samples to help the model better learn how to predict and interpret complex dynamics in sequence data. Since the generation of trajectory branches is based on combinatorial conditions (trajectory segments and their future returns), the model can exhibit higher robustness when dealing with trajectory data with different characteristics and complexities. This helps the model still make accurate predictions and judgments when facing noise, anomalies, or missing data in practical applications.
[0043] The trajectory value function uses the Q-value of the last state-action pair of the predicted trajectory segment as the future return of the trajectory segment; the trajectory value function is:
[0044]
[0045] Among them, is the expected regression estimation target of the trajectory value function, D is the robot's trajectory dataset, is the constraint of the trajectory value function, V is the state value, is the future return of the predicted trajectory segment, , , n is the trajectory index of the segment, N represents the total number of trajectories in the robot's trajectory dataset, is the state-action pair at time t, is the discount factor.
[0046] The diffusion model is constructed based on the EMD formula; the combinatorial conditions of the diffusion model are:
[0047]
[0048]
[0049] Among them, represents the combinatorial conditions of the diffusion model, K represents the length of the segment, is the discount factor, is the state at time t, is the action at time t, is the environment at time t, represents the trajectory segment, Represents the future return of the trajectory segment, , where n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset.
[0050] The loss function of the diffusion model is:
[0051] Among them, is the loss function of the diffusion model, represents the subsequent trajectory segment in the true trajectory , represents , , where n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset.
[0052] The EMD (Exponential Mean Diameter) distance calculation formula, that is, the exponential mean diameter formula, is a mathematical formula used to calculate the distance between two points in a multi-dimensional space. It has the characteristics of simple calculation, high reliability, and applicability to various distance metric spaces, and has been widely used in the fields of data mining, machine learning, pattern recognition, etc.
[0053] Diffusion Transformer (DiT) is a new type of generative model that combines the Transformer architecture and the diffusion model, and can efficiently capture the dependencies in the data and generate high-quality results.
[0054] The branch filter extends the trajectory by filtering based on the continuity of the rewards between the trajectory branches and the trajectory segments. The filtering formula of the branch filter is:
[0055]
[0056] Among them, is the Q value of the state-action pair at time t, , where n is the trajectory index of the segment, N represents the total number of trajectories in the robot's trajectory dataset, is the discount factor, is the reward function of the environment, is the state-action value function, is the value function corresponding to choosing the action at the state , H is the length of the generated trajectory branch, is the threshold hyperparameter.
[0057] Example 2 SeeFigure 2 , a sequence modeling data augmentation system based on trajectory branch generation, comprising: A sampling module for sampling trajectory segments from the trajectory dataset of the robot; A future return prediction module for generating the future return of the predicted trajectory segment for the trajectory segment using a trajectory value function; A trajectory branch generation module for generating trajectory branches through a diffusion model using the trajectory segment and the future return of the trajectory segment as combined conditions; A connected extended trajectory module for connecting the trajectory branch with the trajectory segment and the future return of the trajectory segment to obtain an extended trajectory; A filtered trajectory branch module for filtering the extended trajectory through a branch filter to obtain a trajectory branch.
[0058] Embodiment III A robot adopts the sequence modeling data augmentation method based on trajectory branch generation described in Embodiment I in an offline data training data augmentation task scenario. The method includes the following steps: Sampling trajectory segments from the trajectory dataset of the robot; Generating the future return of the predicted trajectory segment for the trajectory segment using a trajectory value function; Using the trajectory segment and the future return of the trajectory segment as combined conditions to generate trajectory branches through a diffusion model; Connecting the trajectory branch with the trajectory segment and the future return of the trajectory segment to obtain an extended trajectory; Filtering the extended trajectory through a branch filter to obtain a trajectory branch.
[0059] Embodiment IV See Figure 3 , an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the sequence modeling data augmentation method based on trajectory branch generation is implemented. The method includes the following steps: Sampling trajectory segments from the trajectory dataset of the robot; Generating the future return of the predicted trajectory segment for the trajectory segment using a trajectory value function; Using the trajectory segment and the future return of the trajectory segment as combined conditions to generate trajectory branches through a diffusion model; Connecting the trajectory branch with the trajectory segment and the future return of the trajectory segment to obtain an extended trajectory; Filtering the extended trajectory through a branch filter to obtain a trajectory branch.
[0060] Embodiment V A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the sequence modeling data enhancement method based on trajectory branching is performed. The method includes the following steps: Sample trajectory segments from the trajectory dataset of the robot; Use a trajectory value function for the trajectory segments to generate the future return of the predicted trajectory segments; Use the trajectory segments and the future return of the trajectory segments as combined conditions to generate trajectory branches through a diffusion model; Connect the trajectory branches with the trajectory segments and the future return of the trajectory segments to obtain an extended trajectory; Filter the extended trajectory through a branch filter to obtain trajectory branches.
[0061] Embodiment Six In the sequence modeling data enhancement method based on trajectory branching in this embodiment, trajectory segments are incorporated into the conditions of the diffusion model for generation. At the same time, a trajectory value function (TVF) is used to guide the diffusion model to generate branches that can lead to higher returns. TVF is pre-trained on the dataset and can predict the future return of the trajectory segments used for branch generation. Since the diffusion model is trained on a limited set of trajectories, guiding the generation of returns beyond the training data may reduce the guiding effect. Therefore, TVF predicts the return of branch generation based on the future maximum return and the return of the trajectories in the dataset. The generated branches are then used to extend the trajectories in the dataset. Figure 4 The overall process of the method is shown, and the detailed introduction of each module is as follows: Use a diffusion-based generative model to generate trajectory branches according to the trajectory segments sampled from the trajectory dataset of the robot. To make the generated branches consistent with the sampled segments and make the branches conditional on the return, the sampled segments and their corresponding returns are incorporated into the conditions of the diffusion model.
[0062] When pre-training the diffusion model, this condition can be expressed as:
[0063]
[0064]
[0065] where N represents the total number of trajectories in the trajectory dataset of the robot, n is the trajectory index of the segment, meaning that this segment is sampled from the nth trajectory in the trajectory dataset of the robot; K represents the length of the segment, is the discount factor.
[0066] Use to represent .
[0067] During pre-training, the true return of the sampled segments calculated from the corresponding trajectories in the robot's trajectory dataset is used as the return in the condition. This conditions the generated segments on future returns. Based on this condition, the diffusion model generates subsequent segments.
[0068] The loss function for training the diffusion model is:
[0069] where represents the subsequent trajectory segments in the true trajectory .
[0070] After pre-training, the true return in the condition is replaced with the future return predicted by the trajectory value function (TVF). This guides the diffusion model to generate trajectory branches leading to higher returns. Then, the generated branches are concatenated with the sampled trajectory segments as an extension of the trajectories in the robot's trajectory dataset.
[0071] To guide the generation of branches leading to higher-return trajectories, the pre-trained trajectory value function (TVF) is used to predict the future return of the sampled trajectory segments. Since only the Q-values of the state-action pairs in the dataset need to be estimated, inspired by IQL, a SARSA-style objective is adopted to completely avoid the influence of outlier actions during pre-training. The Q-value of the last state-action pair (s t , a t ) of the predicted trajectory segment is used as the future return of the segment. However, if the future return predicted by TVF is much higher than the return in the dataset, it will lose effectiveness as a guide. To match the return predicted by TVF with the trajectories in the dataset, it is constrained by the maximum return corresponding to (s t , a t ) in the dataset:
[0072] where , .
[0073] When the dataset becomes large, it is difficult to obtain the maximum value of . In practice, it is estimated by expected regression, which leads to the following objective:
[0074] The following discusses the implementation details of the diffusion model and describes the entire training process.
[0075] The diffusion model aims to generate trajectory branches. A diffusion model is constructed based on the EDM formula, which is more effective. The generated branches should be consistent with the environment, which requires the diffusion model to have high accuracy. The diffusion transformer, which performs well in visual tasks, is used to solve this problem. An overview of the method is as Figure 4 shown.
[0076] First, a trajectory segment is randomly sampled from the robot's trajectory dataset. The Trajectory Value Function (TVF) predicts the future return of this segment. Then, the predicted return and the sampled segment are combined as the conditions for the diffusion model, and the diffusion model generates trajectory branches. The branches are connected to the trajectory segment as an extension of the trajectory from which the segment was sampled. To ensure the consistency between the generated branches and the previous trajectory segments, a branch filter is designed. The generated branches are filtered by the continuity of the return between the branches and the trajectory segments. The formula for filtering it is as follows:
[0077]
[0078] The following further illustrates this embodiment in combination with the experimental results: A large number of experiments were conducted on Gym, Maze2d, and Antmaze tasks in the D4RL benchmark (Fu et al., 2020) to verify the effectiveness of BG. In addition, an ablation study was conducted on each module to evaluate the contribution of each component to the overall performance. The generated trajectory branches were also visualized to clearly show the impact of BG.
[0079] BG is evaluated in various domains of the D4RL benchmark, including Maze2d, Antmaze, and MuJoCo tasks. The Maze2d task has three map types: umaze, medium, and large, with gradually increasing scale and complexity. The Antmaze task includes two map types: umaze and medium, each with two task variants: play and diverse. In the MuJoCo task, BG is evaluated in three environments: Halfcheetah, Hopper, and Walker2d. Each environment contains three types of datasets: Medium, Medium-replay, and Medium-expert, where the Medium-expert dataset consists of expert data and sub-optimal data, while the Medium and Medium-replay datasets are collected by interacting the unconverged SAC policy with the environment.
[0080] Compare BG+DT with representative methods from offline reinforcement learning, imitation learning, and sequence modeling. Offline reinforcement learning methods include Conservative Q-Learning (CQL) and Implicit Q-Learning (IQL). For imitation learning, compare BG+DT with Behavior Cloning (BC). Sequence modeling methods include Online Decision Transformer (ODT), Elastic Decision Transformer (EDT), and Q-Learning Decision Transformer (QDT). For fair comparison, evaluate ODT in its offline version. Also compare with the original DT without BG processing. Figure 5 Illustrates the performance of the method proposed in this embodiment.
[0081] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0082] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the specified functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0083] These computer program instructions can also be stored in a computer-readable memory capable of guiding the computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the specified functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0084] These computer program instructions can also be loaded onto the computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide means for realizing the specified functions in Figure 1One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes. Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: it is still possible to modify the specific implementation manners of the present invention or make equivalent substitutions, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the present invention.
Claims
1. A sequence modeling data augmentation method based on trajectory branch generation, characterized in that, It includes the following steps: Sample trajectory segments in the trajectory dataset of the robot; Use the trajectory value function for the trajectory segments to generate the future rewards of the predicted trajectory segments; Take the trajectory segments and the future rewards of the trajectory segments as combined conditions, and generate trajectory branches through the diffusion model; Connect the trajectory branches with the trajectory segments and the future rewards of the trajectory segments to obtain an extended trajectory; Filter the extended trajectory through the branch filter to obtain trajectory branches.
2. The sequence modeling data augmentation method based on trajectory branch generation according to claim 1, wherein The trajectory value function uses the Q value of the last state-action pair of the predicted trajectory segment as the future reward of the trajectory segment; The trajectory value function is: Among them, is the expected regression estimation target of the trajectory value function, D is the trajectory dataset of the robot, is the constraint of the trajectory value function, V is the state value, is the future return of the predicted trajectory segment, , , where n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset, is the state-action pair at time t, is the discount factor.
3. The method for sequence modeling data augmentation based on trajectory branch generation according to claim 2, wherein The diffusion model is constructed based on the EMD formula; The combined condition of the diffusion model is: Among them, represents the combined condition of the diffusion model, K represents the length of the segment, is the discount factor, is the state at time t, is the action at time t, is the environment at time t, represents the trajectory segment, represents the future return of the trajectory segment, , n is the trajectory index of the segment, and N represents the total number of trajectories in the robot's trajectory dataset.
4. The sequence modeling data augmentation method based on trajectory branch generation according to claim 3, characterized in that The loss function of the diffusion model is: Among them, is the loss function of the diffusion model, represents the subsequent trajectory segment in the true trajectory , represents , , where n is the trajectory index of the segment and N represents the total number of trajectories in the robot's trajectory dataset.
5. The sequence modeling data enhancement method based on trajectory branch generation according to claim 4, characterized in that The branch filter filters the extended trajectory through the continuity of the rewards between the trajectory branches and the trajectory segments.
6. The sequence modeling data enhancement method based on trajectory branch generation according to claim 5, wherein The filtering formula of the branch filter is: Among them, is the Q value of the state-action pair at time t , n is the trajectory index of the segment, N represents the total number of trajectories in the robot's trajectory dataset, , γ is the discount factor, is the reward function of the environment, is the state-action value function, is the value function corresponding to choosing action in state , H is the length of the generated trajectory branch, and θ is the threshold hyperparameter. 7. A robot, characterized in that, In the data augmentation task scenario of offline data training, the sequence modeling data augmentation method based on trajectory branches described in any one of claims 1-6 is adopted.
8. A sequence modeling data augmentation system based on trajectory branch generation, characterized in that, It includes: A sampling module for sampling trajectory segments in the trajectory dataset of the robot; A future reward prediction module for using the trajectory value function for the trajectory segments to generate the future rewards of the predicted trajectory segments; A trajectory branch generation module for taking the trajectory segments and the future rewards of the trajectory segments as combined conditions and generating trajectory branches through the diffusion model; A connected extended trajectory module for connecting the trajectory branches with the trajectory segments and the future rewards of the trajectory segments to obtain an extended trajectory; A filtered trajectory branch module for filtering the extended trajectory through the branch filter to obtain trajectory branches.
9. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the sequence modeling data augmentation method based on trajectory branches described in any one of claims 1-7 is implemented.
10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the sequence modeling data augmentation method based on trajectory branches described in any one of claims 1-7 is implemented.