Method and device for generating training data of robot decision model
By dividing the robot task into multiple actions and collecting only decision frames, skipping the simulation of intermediate control frames, efficient and semantically complete training data is generated, solving the problem of high computational overhead in existing technologies and improving the training efficiency and applicability of robot decision models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, the generation of training data for robot decision-making models relies on physical simulation, resulting in huge computational overhead. Especially in large-scale multi-task and multi-scenario training, the cost of simulation computation increases exponentially, becoming a bottleneck that restricts the improvement of system efficiency.
The robot task is divided into multiple actions. Simulation only collects decision frames and skips intermediate control frames. By generating data pairs of start and end frames, a keyframe dataset is constructed to generate training data with complete task semantics.
It significantly reduces data generation time and resource overhead, improves the efficiency and semantic integrity of training data, adapts to various simulation platforms and task scenarios, and supports large-scale offline data generation.
Smart Images

Figure CN122020150A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and apparatus for generating training data for a robot decision-making model. Background Technology
[0002] In existing technologies, the generation of training data for robot decision-making models often relies on physical simulation environments. This simulation method typically requires fully simulating every frame of the robot's actions to ensure continuity and realism in every step from the initial state to the target state. While this method guarantees physical consistency, it also leads to enormous computational overhead. Especially in large-scale, multi-task, multi-scenario training scenarios, the increase in simulation computation costs is often exponential, resulting in lengthy training cycles for robot decision-making models and becoming a key bottleneck restricting system efficiency improvements.
[0003] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section. Summary of the Invention
[0004] This invention provides a method and apparatus for generating training data for a robot decision model, which solves at least some of the technical problems described above in the prior art.
[0005] This specification provides a method for generating training data for a robot decision-making model, including:
[0006] Divide the robot's task into multiple actions;
[0007] Determine the decision frame for the current action and the next action;
[0008] Generate a data pair for the current action based on the decision frame of the current action and the next action;
[0009] Multiple data points are concatenated according to the execution order of the aforementioned actions to generate training data for the task to be executed by the robot decision-making model.
[0010] In one embodiment, the decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
[0011] In one embodiment, the decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
[0012] In one embodiment, dividing the robot's task to be performed into multiple actions includes:
[0013] An action sequence is generated based on the environmental state corresponding to the task to be executed, the preset action library, and the task to be executed.
[0014] The task to be executed is divided into multiple actions according to the action sequence.
[0015] In one embodiment, the training data further includes:
[0016] The initial environmental state in which the robot performs the task to be performed.
[0017] In one embodiment, the training data further includes:
[0018] The state frame following the execution of the last of the multiple actions.
[0019] This specification also provides a device for generating training data for a robot decision-making model, including:
[0020] The task division module is used to divide the robot's tasks into multiple actions;
[0021] The decision frame determination module is used to determine the decision frame for the current action and the next action;
[0022] The data pair generation module is used to generate a data pair for the current action based on the decision frame of the current action and the next action;
[0023] The training data generation module is used to concatenate multiple data according to the execution order of the multiple actions to generate training data for the task to be executed by the robot decision model.
[0024] In one embodiment, the task partitioning module includes:
[0025] An action sequence generation unit is used to generate an action sequence based on the environmental state corresponding to the task to be executed, a preset action library, and the task to be executed.
[0026] A task division unit is used to divide the task to be executed into the multiple actions according to the action sequence.
[0027] This invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described method for generating training data for a robot decision model.
[0028] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating training data for a robot decision model.
[0029] This specification also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method for generating training data for the robot decision model.
[0030] As described above, embodiments of the present invention provide a method and apparatus for generating training data for a robot decision model. The method includes: first, dividing the robot's task to be executed into multiple actions; then, determining the decision frame of the current action and the next action; generating a data pair of the current action based on the decision frame of the current action and the next action; and finally, concatenating the multiple data according to the execution order of the multiple actions to generate training data for the robot decision model's task to be executed.
[0031] The present invention provides a method for generating training data for a robot decision model, which is a frame skipping simulation mechanism that can improve the data acquisition efficiency and computational performance during the training process of a robot task decision model. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0033] Figure 1 This is a flowchart illustrating a method for generating training data for a robot decision model in one embodiment of this specification.
[0034] Figure 2 This is a flowchart illustrating step 100 in yet another embodiment of this specification;
[0035] Figure 3 This is a flowchart illustrating the method for generating training data for a robot decision model in a specific embodiment of this specification.
[0036] Figure 4 This is a mind map illustrating the logical structure of the method for generating training data for the robot decision-making model in specific embodiments of this specification.
[0037] Figure 5 This is a schematic diagram of the structure of a device for generating training data for a robot decision model in one embodiment of this specification;
[0038] Figure 6 This is a schematic diagram of the structure of the task partitioning module 10 in another embodiment of this specification;
[0039] Figure 7This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of this specification; Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0041] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0042] Provide users with corresponding operation entry points, allowing them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0043] Figure 1 This is a flowchart illustrating the method for generating training data for a robot decision-making model in an embodiment of the present invention. Specifically, this method is applied to the server side. In practical implementation, as follows... Figure 1 As shown, the method for generating training data for the robot decision-making model includes steps 100 to 400.
[0044] Step 100: Divide the robot's task into multiple actions;
[0045] Step 200: Determine the decision frame for the current action and the next action;
[0046] Step 300: Generate a data pair for the current action based on the decision frame of the current action and the next action;
[0047] Step 400: Concatenate multiple data sets according to the execution order of the multiple actions to generate training data for the task to be executed by the robot decision model.
[0048] For steps 100 to 400, current research and applications in the field of robotics are gradually shifting from traditional rule-based control methods to methods such as imitation learning and reinforcement learning, which rely on large models. Through imitation learning, robots can learn human or expert operating strategies from demonstration data; while reinforcement learning explores optimal action strategies through interaction with the environment, thereby completing complex tasks. These methods demonstrate excellent versatility and adaptability when facing complex, dynamic, and unstructured task scenarios. However, both imitation learning and reinforcement learning place extremely high demands on high-quality training data, especially requiring a large amount of simulation environment data covering diverse scenarios to enhance the model's generalization ability.
[0049] The applicant argues that existing methods do not fully decouple "task decision-making" from "action execution" during data generation. In traditional simulations, training data acquisition typically requires following an entire action execution path, with each frame needing calculation and recording to ensure the integrity of the task data.
[0050] However, in many robotic tasks, the lower-level motion control is often relatively deterministic. For example, atomic skills such as grasping and moving, after specialized training and optimization, typically achieve a high success rate. During the simulation data acquisition phase, if the decision model and execution module are decoupled, only the lower-level motion control module needs to collect complete motion flow data to ensure the physical consistency and reliability of skill execution. The upper-level task decision model, on the other hand, primarily focuses on "when" and "under what conditions" a particular skill module is invoked. When training the upper-level decision model, it can be assumed that the lower-level control module's motion execution is reliable with a 100% success rate. Thus, knowing only the start frame (initial task state) and the end frame (skill completion state) is sufficient for semantic modeling and logical reasoning of the task.
[0051] Based on this, this application provides a method as described in steps 100 to 400, which decouples "task decision-making" from "action execution," encapsulating lower-level actions into individual skills, which are then invoked by the upper-level decision-making model. In this context, for training the decision-making model, intermediate frames during action execution often do not carry additional decision information and instead create redundant computational burden. Furthermore, this method, by collecting only the start and end frames and skipping the complete simulation of intermediate action frames, can quickly generate semantically complete and decision-making-meaningful data fragments, thereby significantly improving data acquisition efficiency and reducing simulation computational burden. This has significant practical implications and application value for the generation of large-scale task data, especially in the training phase of large models.
[0052] In one embodiment, the decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
[0053] Specifically, based on a unified simulation timeline, the robot task execution process is divided into two categories: decision frames (key semantic frames) and control frames (continuous execution frames). Only decision frames are collected during data generation, skipping the physical simulation and screen rendering of control frames. This achieves a training data acquisition process with complete task semantics, clear causal structure, and extremely low simulation cost.
[0054] Compared with existing methods, this approach significantly reduces redundant acquisition of intermediate continuous action processes, making it more suitable for large-scale data generation and cross-task transfer robot decision model training.
[0055] As described above, embodiments of the present invention provide a method for generating training data for a robot decision model, comprising: first, dividing the robot's task to be executed into multiple actions; then, determining the decision frame of the current action and the next action; generating a data pair of the current action based on the decision frame of the current action and the next action; and finally, concatenating the multiple data according to the execution order of the multiple actions to generate training data for the robot decision model's task to be executed.
[0056] Compared with the prior art, the method for generating training data for a robot decision-making model provided in this application has the following advantages:
[0057] 1. High efficiency: By skipping intermediate frames during action execution, the physical dynamics calculation and rendering of each frame are avoided, which significantly reduces the time and resource overhead of data generation and greatly improves the efficiency of training data generation.
[0058] 2. Semantic integrity: Although the simulation of intermediate frames is omitted, the key decision input-output relationships are preserved by saving the start and end state frames of each action, ensuring the integrity of the causal chain of the task layer and meeting the data semantic requirements of the task layer decision model.
[0059] 3. High scalability: By skipping the specific physical details of action execution, it has good compatibility with the implementation methods of lower-level execution modules and can be adapted to various path planning methods or control strategies of underlying execution models. The method provided in this application can be used on multiple simulation platforms, supports multiple goal-oriented task scenarios, and can be combined with automatic trajectory generators or behavior cloning models for large-scale offline data generation.
[0060] In one embodiment, the decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
[0061] It should be noted that the aforementioned starting state frame of the current action is used to represent the state before the current action is executed, and at the same time, it is the state frame after the previous action of the current action is executed (representing the state after the previous action is completed).
[0062] In one embodiment, see Figure 2 Step 100 includes:
[0063] Step 101: Generate an action sequence based on the environmental state corresponding to the task to be executed, the preset action library, and the task to be executed;
[0064] Step 102: Divide the task to be executed into the multiple actions according to the action sequence.
[0065] In steps 101 and 102, which correspond to the data generation stage, a complete action sequence is generated for the task based on the target task and the environment state using a predefined action library and a task data planning generator. This action sequence defines the logical steps for task execution.
[0066] In one embodiment, the training data includes not only the splicing result of multiple data in step 400, but also the initial environmental state of the robot when performing the task to be performed, i.e., the initial state of the task (the environmental state when the robot is ready to start performing the task).
[0067] In one embodiment, the training data further includes: the state frame after the last action among the plurality of actions is executed, that is, the state frame after the task is finally completed (corresponding to the Done() state).
[0068] As described above, embodiments of the present invention provide a method for generating training data for a robot decision model, comprising: first, dividing the robot's task to be executed into multiple actions; then, determining the decision frame of the current action and the next action; generating a data pair of the current action based on the decision frame of the current action and the next action; and finally, concatenating the multiple data according to the execution order of the multiple actions to generate training data for the robot decision model's task to be executed.
[0069] Compared with the prior art, the method for generating training data for a robot decision-making model provided in this application has the following advantages:
[0070] 1. High efficiency: By skipping intermediate frames during action execution, the physical dynamics calculation and rendering of each frame are avoided, which significantly reduces the time and resource overhead of data generation and greatly improves the efficiency of training data generation.
[0071] 2. Semantic integrity: Although the simulation of intermediate frames is omitted, the key decision input-output relationships are preserved by saving the start and end state frames of each action, ensuring the integrity of the causal chain of the task layer and meeting the data semantic requirements of the task layer decision model.
[0072] 3. High scalability: By skipping the specific physical details of action execution, it has good compatibility with the implementation methods of lower-level execution modules and can be adapted to various path planning methods or control strategies of various underlying execution models. The method provided in this application can be used on multiple simulation platforms, supports multiple goal-oriented task scenarios, and can be combined with automatic trajectory generators or behavior cloning models for large-scale offline data generation.
[0073] Figure 3 This is a flowchart illustrating the method for generating training data for a robot decision-making model according to a specific embodiment of this specification. To further illustrate the solution, the following is combined with... Figure 3 as well as Figure 4 Further description of the method for generating training data for the robot decision-making model in this application:
[0074] S1: Generate action plan.
[0075] During the data generation phase, a complete sequence of actions is first generated for the target task and environmental state using a predefined action library and a task data planning generator. This sequence of actions defines the logical steps for task execution.
[0076] Specifically, suppose that data acquisition is performed for a task T, and the action sequence of this task has n steps, denoted as: ,in This represents the i-th action of the task. The Done() action indicates the end of the task.
[0077] S2: Set the start frame and end frame.
[0078] For a complete sequence of robot action actions, the training objective of the decision model is to learn how to select a reasonable sequence of actions to achieve the task objective. To simplify training and avoid interference from lower-level execution failures with the learning process of higher-level decisions, this application assumes a 100% success rate for the lower-level action execution module during decision model training, thus achieving complete decoupling between the decision model and the action execution model.
[0079] During the simulation data acquisition phase, the simulator performs full-process data acquisition for the entire task. For each action in the action sequence, a unified task execution timeline is established in the simulation environment, dividing the execution process of that action into several consecutive simulation frames. Each simulation frame is either a decision frame or a control frame. Decision frames represent key state transition nodes during task execution, such as the start and end frames of each action. and Control frames represent the continuous control execution state of each action during the lower-level execution process, i.e. .
[0080] When collecting data, only decision frames are retained, while control frames are ignored. That is, the simulation calculation and screen rendering of the low-level control process are skipped, and only the state frame pairs that are semantically complete and causally related to the high-level decision are retained during task execution. It is understandable that constructing the keyframe dataset for training in this way can significantly reduce the computational cost and improve the training efficiency.
[0081] For example, the complete execution process of the action Grasp (bottle) may take up [time / percentage]. arrive A total of 16 frames, of which: This is the initial decision frame for the action; arrive These are intermediate control frames during the action completion process. This is the end decision frame for the current action (and also the start decision frame for the next action). When training the decision model, only the selection logic of high-level actions needs to be learned, without learning the low-level continuous control process. Therefore, this application proposes skipping intermediate control frames during the simulation acquisition phase and only acquiring key decision frames, i.e. and This enables "frame skipping simulation".
[0082] During the simulation data acquisition phase, the entire task execution process is mapped onto the timeline as a series of continuous simulation frames, which can be segmented according to each action and represented as follows: ,in{ }( ) represents all simulation frames during the execution of the i-th action.
[0083] If we follow the traditional approach, a complete simulation would involve each action... Execute the entire process, saving every frame during execution, for a total of +1 frame. However, a large number of intermediate control frames do not participate in decision-making, resulting in serious data waste and an extremely heavy simulation burden.
[0084] Correspondingly, according to the frame skipping simulation method proposed in this application, it is possible to save only n+1 frames of data, that is... .in, Indicates the initial state of the task; Indicates the i-th action The state simulation frame after execution is complete, which is also the next action. Simulation frame before execution.
[0085] S3: Construct the keyframe dataset.
[0086] After data acquisition, the start and end decision frames for each action are paired with the action text and task objective to form a keyframe dataset for training. Each action in the dataset stores one data point, including the task objective, the start state frame, and the action at that step, indicating "given the starting state, the decision model should select this action to ultimately complete the target task." This keyframe dataset preserves the causal relationship of action selection while effectively avoiding redundant acquisition of continuous action processes. It allows training data to be shuffled and processed in batches, greatly improving data utilization and training efficiency, while ensuring the integrity and rationality of task decisions.
[0087] After data collection is complete, each data entry will be structured as follows: That is, for a specific task T, the initial state before each action is executed. This corresponds to the next action that needs to be performed. In particular, the last effective action. Post-execution status , corresponding to the end signal (Done()).
[0088] S4: Piece together multi-step data to generate training data for the robot's decision-making model's task to be executed.
[0089] During the model inference phase, to test the model's multi-step decision-making ability, multiple keyframes in the dataset can be concatenated according to the task sequence to form a complete multi-step task trajectory. This simulates the model's continuous action selection process during execution, verifying the model's multi-step inference ability and task completion performance. This approach ensures task continuity while testing the model's inference performance without increasing simulation overhead, further validating the model's generalization ability and stability in complex tasks.
[0090] Next, the following operations can be performed based on the training data generated in step S4:
[0091] During the training phase, the collected keyframe data is used. The training model learns to complete task T by starting from the current state. To action The mapping relationship.
[0092] During the inference phase, the data is assembled according to the task order: This allows for multi-step reasoning and the complete execution of the task.
[0093] Understandably, this data acquisition method based on frame skipping simulation avoids redundant acquisition of intermediate consecutive frames, thus ensuring the semantic integrity of the task while significantly reducing the simulation computation overhead.
[0094] Example illustration:
[0095] The following example uses the task "Robot, please retrieve the body BB cream from the shelf display in the conference room" to illustrate how to efficiently collect data for robot task decision model training, combined with a robot decision model training data generation method (frame skipping simulation) provided in this application.
[0096] 1. Action plan generation.
[0097] Before starting the simulation, we first determine the sequence of actions required to complete the task, as shown below:
[0098] Task: Robot, retrieve the body BB cream from the shelf display case in the conference room;
[0099] action:
[0100] Navigate to (meeting room);
[0101] Navigate_to (tiered display rack);
[0102] Scan (body BB cream);
[0103] Grasp (body BB cream);
[0104] Navigate to (cafeteria);
[0105] Navigate to (owner);
[0106] Place (body BB cream, owner);
[0107] Done().
[0108] 2. Frame skipping simulation acquisition process.
[0109] In the simulation environment, action planning is performed for the entire task flow. However, when collecting training data, the "frame skipping simulation" mechanism of this invention is used to retain only the following state frame information:
[0110] Initial task state (i.e., the environmental state when the robot is ready to start executing the task);
[0111] The initial state frame of each action (the state before the action is executed).
[0112] The end state frame of each action (the state after the action is completed).
[0113] The final status frame of the task (corresponding to the Done() state);
[0114] The first step is to keep only two states in the “Navigate_to (meeting room)” simulation system: the start state frame of the action (the robot has not yet moved) and the end state frame of the action (the robot has arrived at the meeting room).
[0115] The second step, "Navigate_to(shelf display rack)", the simulation system retains only two states: the start state frame of the action (the robot arrives at the conference room) and the end state frame of the action (the robot arrives at the shelf display rack).
[0116] As can be seen, the starting frame of the second step mentioned above is actually the ending frame of the first step. Therefore, in actual data acquisition, only the starting state of the task and the ending state frame after each step are captured. By capturing these frames sequentially, the dynamic simulation and rendering of consecutive intermediate frames are avoided, greatly improving simulation efficiency.
[0117] 3. Construction of the keyframe dataset.
[0118] For the dataset construction during the training phase, each action corresponds to one data point, which includes the following information: the task objective (the robot asks the user to retrieve the body BB cream from the shelf display in the meeting room); the starting state frame of the action; and the current action.
[0119] Each data point represents "given the task objective and initial state, the action that should be chosen to achieve the objective," and is used to train the decision model to select actions.
[0120] For example, the data stored for the first action of this task is: the task objective (robot, please get me the body BB cream from the shelf display in the conference room); the starting state frame of the action (the robot has not yet moved); and the current action (Navigate_to(conference room)).
[0121] The data stored for the second action is: the task objective (the robot retrieves the body cream from the shelf display in the conference room); the starting state frame of the action (the robot arrives at the conference room); the current action (Navigate_to(shelf display)); and so on, to generate all the data for the task.
[0122] 4. Multi-step data splicing.
[0123] During the model inference phase, the aforementioned keyframe data can be spliced together into a multi-step task trajectory according to the task sequence, which can then be used by the model for multi-step inference testing.
[0124] As described above, the model first receives the initial state and predicts the first step: Navigate_to (meeting room). The model then feeds back the ending state of Navigate_to (meeting room) to predict the next action, and so on, executing sequentially until the entire task trajectory is completed, achieving closed-loop reasoning throughout the entire process. Through the frame skipping mechanism, only keyframe data needs to be stitched together, ensuring task continuity while avoiding redundant physical simulation overhead.
[0125] As described above, a specific embodiment of the present invention provides a method for generating training data for a robot decision model, comprising: first, dividing the robot's task to be executed into multiple actions; then, determining the decision frame of the current action and the next action; generating a data pair of the current action based on the decision frame of the current action and the next action; and finally, concatenating the multiple data according to the execution order of the multiple actions to generate training data for the robot decision model's task to be executed.
[0126] Compared with the prior art, this application has the following characteristics:
[0127] 1. Frame-skipping simulation mechanism: A data generation method for training robot task decision models. By only retaining the start and end state frames of each action during the simulation process and skipping the intermediate frame simulation during the action execution process, data can be generated quickly and efficiently.
[0128] 2. Construction of the keyframe dataset: The keyframe dataset consists of the task objective, the start state frames of the action, and the end state frames, which are used to train the model to select appropriate actions under specific task objectives and start states.
[0129] 3. Multi-step data stitching mechanism: By stitching multiple keyframe data in the order of tasks, a multi-step task trajectory is constructed to support multi-step inference testing of the model.
[0130] 4. Simulation platform adaptability: The method is applicable to various goal-oriented tasks and can be integrated with multiple simulation platforms. It is particularly suitable for training tasks of Transformer or large model architecture with action selection as the core.
[0131] 5. Optimization of computational resources and compatibility of modules: By skipping intermediate frame simulation calculations, simulation time and computational resources can be significantly reduced, while preserving the semantic integrity of the task and adapting to the path planning and control strategies of various underlying execution modules.
[0132] Based on the same inventive concept, this application also provides a device for generating training data for a robot decision model, which can be used to implement the method described in the above embodiments, as shown in the following embodiments. Since the principle of the device for generating training data for a robot decision model is similar to that of the method for generating training data for a robot decision model, the implementation of the device for generating training data for a robot decision model can refer to the implementation of the method for generating training data for a robot decision model, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0133] The embodiments of the present invention provide a specific implementation of a robot decision model training data generation device capable of generating robot decision model training data, which implements a method for generating robot decision model training data. See [link to specific implementation details]. Figure 5 A device for generating training data for a robot decision-making model specifically includes the following components:
[0134] The task division module 10 is used to divide the robot's task to be performed into multiple actions;
[0135] The decision frame determination module 20 is used to determine the decision frame of the current action and the next action;
[0136] The data pair generation module 30 is used to generate a data pair for the current action based on the decision frame of the current action and the next action;
[0137] The training data generation module 40 is used to concatenate multiple data according to the execution order of the multiple actions to generate training data for the task to be executed by the robot decision model.
[0138] In one embodiment, the decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
[0139] In one embodiment, the decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
[0140] In one embodiment, see Figure 6 The task partitioning module 10 includes:
[0141] Action sequence generation unit 10a is used to generate an action sequence based on the environmental state corresponding to the task to be executed, a preset action library and the task to be executed.
[0142] The task division unit 10b is used to divide the task to be executed into the multiple actions according to the action sequence.
[0143] In one embodiment, the training data further includes:
[0144] The initial environmental state in which the robot performs the task to be performed.
[0145] In one embodiment, the training data further includes:
[0146] The state frame following the execution of the last of the multiple actions.
[0147] As described above, a specific embodiment of the present invention provides a device for generating training data for a robot decision model, comprising: a task division module for dividing a robot's task to be executed into multiple actions; a decision frame determination module for determining the decision frame of the current action and the next action; a data pair generation module for generating a data pair of the current action based on the decision frame of the current action and the next action; and a training data generation module for concatenating multiple data according to the execution order of the multiple actions to generate training data for the robot decision model's task to be executed.
[0148] Compared with the prior art, the robot decision model training data generation device provided in this application has the following advantages:
[0149] 1. High efficiency: By skipping intermediate frames during action execution, the physical dynamics calculation and rendering of each frame are avoided, which significantly reduces the time and resource overhead of data generation and greatly improves the efficiency of training data generation.
[0150] 2. Semantic integrity: Although the simulation of intermediate frames is omitted, the key decision input-output relationships are preserved by saving the start and end state frames of each action, ensuring the integrity of the causal chain of the task layer and meeting the data semantic requirements of the task layer decision model.
[0151] 3. High scalability: By skipping the specific physical details of action execution, it has good compatibility with the implementation methods of lower-level execution modules and can be adapted to various path planning methods or control strategies of various underlying execution models. The method provided in this application can be used on multiple simulation platforms, supports multiple goal-oriented task scenarios, and can be combined with automatic trajectory generators or behavior cloning models for large-scale offline data generation.
[0152] Embodiments of this application also provide a specific implementation of an electronic device capable of implementing all steps in the method for generating training data for the robot decision model in the above embodiments, see [link to implementation details]. Figure 7 The electronic devices specifically include the following:
[0153] Processor 1201, memory 1202, communications interface 1203, and bus 1204;
[0154] The processor 1201, memory 1202, and communication interface 1203 communicate with each other via bus 1204; the communication interface 1203 is used to realize information transmission between server-side devices and client-side devices and other related devices.
[0155] The processor 1201 is used to call the computer program in the memory 1202. When the processor executes the computer program, it implements all the steps in the method for generating training data for the robot decision model in the above embodiment. For example, when the processor executes the computer program, it implements the following steps:
[0156] Divide the robot's task into multiple actions;
[0157] Determine the decision frame for the current action and the next action;
[0158] Generate a data pair for the current action based on the decision frame of the current action and the next action;
[0159] Multiple data points are concatenated according to the execution order of the aforementioned actions to generate training data for the task to be executed by the robot decision-making model.
[0160] The decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
[0161] The decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
[0162] The process of dividing the robot's task into multiple actions includes:
[0163] An action sequence is generated based on the environmental state corresponding to the task to be executed, the preset action library, and the task to be executed.
[0164] The task to be executed is divided into multiple actions according to the action sequence.
[0165] The training data also includes:
[0166] The initial environmental state in which the robot performs the task to be performed.
[0167] The training data also includes:
[0168] The state frame following the execution of the last of the multiple actions.
[0169] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps in the method for generating training data for a robot decision model in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the method for generating training data for a robot decision model in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:
[0170] Divide the robot's task into multiple actions;
[0171] Determine the decision frame for the current action and the next action;
[0172] Generate a data pair for the current action based on the decision frame of the current action and the next action;
[0173] Multiple data points are concatenated according to the execution order of the aforementioned actions to generate training data for the task to be executed by the robot decision-making model.
[0174] The decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
[0175] The decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
[0176] The process of dividing the robot's task into multiple actions includes:
[0177] An action sequence is generated based on the environmental state corresponding to the task to be executed, the preset action library, and the task to be executed.
[0178] The task to be executed is divided into multiple actions according to the action sequence.
[0179] The training data also includes:
[0180] The initial environmental state in which the robot performs the task to be performed.
[0181] The training data also includes:
[0182] The state frame following the execution of the last of the multiple actions.
[0183] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, hardware + program embodiments are relatively simple in description because they are fundamentally similar to method embodiments; relevant parts can be referred to the descriptions in the method embodiments.
[0184] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0185] While this application provides method operation steps as shown in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the method can be executed in the order shown in the embodiments or drawings or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0186] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing the embodiments of this specification, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0187] Those skilled in the art will also know that, besides implementing the controller using purely computer-readable program code, the same functions can be achieved by logically programming the method steps, making the controller function as logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers (PLCs), and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the devices within it used to implement various functions can also be considered structures within that hardware component. Alternatively, the devices used to implement various functions can be considered as both software modules implementing the method and structures within a hardware component.
[0188] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0189] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0190] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the embodiments in this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0191] The above description is merely an embodiment of the present specification and is not intended to limit the embodiments of the present specification. For those skilled in the art, various modifications and variations can be made to the embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the embodiments of the present specification should be included within the scope of the claims of the embodiments of the present specification.
Claims
1. A method for generating training data for a robot decision-making model, characterized in that, include: Divide the robot's task into multiple actions; Determine the decision frame for the current action and the next action; Generate a data pair for the current action based on the decision frame of the current action and the next action; Multiple data points are concatenated according to the execution order of the aforementioned actions to generate training data for the task to be executed by the robot decision-making model.
2. The generation method according to claim 1, characterized in that, The decision frame is used to characterize the semantics of the current action, and there is a mapping relationship between the decision frame and the execution result of the current action.
3. The generation method according to claim 2, characterized in that, The decision frame is the starting state frame of the current action and the state frame after the previous action was executed.
4. The generation method according to claim 1, characterized in that, The process of dividing the robot's task into multiple actions includes: An action sequence is generated based on the environmental state corresponding to the task to be executed, the preset action library, and the task to be executed. The task to be executed is divided into multiple actions according to the action sequence.
5. The generation method according to claim 1, characterized in that, The training data also includes: The initial environmental state in which the robot performs the task to be performed.
6. The generation method according to claim 5, characterized in that, The training data also includes: The state frame following the execution of the last of the multiple actions.
7. A device for generating training data for a robot decision-making model, characterized in that, include: The task division module is used to divide the robot's tasks into multiple actions; The decision frame determination module is used to determine the decision frame for the current action and the next action; The data pair generation module is used to generate a data pair for the current action based on the decision frame of the current action and the next action; The training data generation module is used to concatenate multiple data according to the execution order of the multiple actions to generate training data for the task to be executed by the robot decision model.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.