Task instruction analysis and action sequence generation method for spatial embodied intelligence
Patent Information
- Application Number
- CN202611201660.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-18
AI Technical Summary
在复杂在轨作业和深空探测任务中,通信时延、视角遮挡、操作负担以及任务连续性要求都会限制遥操作效率
提升指令解析的规范性与可解释性:通过三级任务需求抽象与InstructionFrame结构化表示,将模糊的自然语言指令转换为字段明确、语义统一的结构化表示,既覆盖空间任务的对象、空间、动作特征,又便于后续任务审查与错误定位。
Smart Images

Figure CN122770001A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of space robotics and embodied intelligence technology, specifically relating to a method for parsing task instructions and generating action sequences for space embodied intelligence. Background Technology
[0002] Space robots are a crucial technological vehicle for expanding human capabilities in orbit. With the development of long-term stays on space stations, lunar research stations, and deep space exploration missions, space missions are no longer limited to one-off launches and short-duration operations. Instead, they are increasingly characterized by long cycles, complex mission chains, frequent equipment maintenance, and intensive human-robot collaboration. In this context, robots need to undertake more repetitive, hazardous, and delicate operational tasks, including in-cabin material transfer, auxiliary operation of experimental devices, module assembly and disassembly, interface alignment, tool retrieval and placement, and emergency response. Compared to terrestrial industrial robots, space robots face more enclosed, resource-constrained, and safety-critical environments. Their control methods cannot rely on long-term, gradual remote operation, nor can they rely entirely on pre-set control scripts to complete all tasks.
[0003] Traditional space robot control methods mainly include pre-programmed execution and teleoperation execution. Pre-programmed execution offers stability advantages when mission boundaries are well-defined and environmental changes are minimal. However, when object positions, task sequences, or target states change, the system often needs to rewrite the mission script. Teleoperation retains the judgment of human operators, but it places high demands on communication links, operator experience, and real-time feedback. In complex on-orbit operations and deep space exploration missions, communication latency, viewpoint obstruction, operational burden, and mission continuity requirements all limit teleoperation efficiency. Therefore, for space robots to truly fulfill their role of assisting astronauts and enhancing mission autonomy, they must possess a certain level of mission understanding, scene perception, and motion planning capabilities.
[0004] In recent years, the development of embodied intelligence and visual-language-action (VLAMA) models has provided new technological pathways for autonomous robot operation. VLAMA attempts to integrate images, language, and robot actions into a unified model, enabling robots to generate continuous control actions based on visual observations and natural language commands. The value of this approach lies in its move away from completely separating language understanding, visual recognition, and motion control into unrelated modules, instead placing them within a single data loop to learn their correspondences. This technological approach has significant application value for space missions. Astronauts or ground control personnel can describe mission objectives using natural language, and the robot can then generate corresponding actions based on visual observations and its own state, thereby reducing the need for gradual human intervention and improving the autonomous execution capability of space operations.
[0005] Following this technological development direction, this research focuses on the key aspects of "task instruction parsing and action sequence generation." Specifically, for a space robot to complete operations based on natural language instructions, it first needs to identify information such as task intent, target object, reference area, and spatial relationships from the instructions; secondly, it needs to convert high-level semantic targets into action semantic representations that the robot can understand; furthermore, it needs to establish a process from structured instructions to action sequence generation and verify the feasibility of this process through test cases in a virtual environment. Therefore, this research, based on a survey of space task processes, constructs a structured instruction parsing method and action semantic representation mode for space tasks, and combines a visual language action model to complete case-level action sequence generation verification, forming a clear, interpretable, and scalable basic technical framework. Summary of the Invention
[0006] In view of this, the present invention provides a method for task instruction parsing and action sequence generation for embodied intelligence in space. It aims to achieve a complete closed loop from instruction parsing to action generation for spatial tasks under small sample conditions by constructing a structured semantic intermediate layer that connects natural language instructions and visual language action models. The overall process of this method is divided into three stages, specifically including: S1. Spatial mission requirements abstraction and structured instruction parsing stage; S2. Spatial task action semantic representation and trajectory stage division stage; S3. Lightweight VLA model adaptation and closed-loop execution verification stage.
[0007] S1. Spatial Mission Requirements Abstraction and Structured Instruction Parsing Stage This phase first systematically reviews four typical tasks within the space station: material transfer, module operation, tool use, and status confirmation. It summarizes the linguistic characteristics of space mission instructions: clear target objects, prominent spatial relationships, rich implicit steps, and strong dependence on execution order. Based on this, space mission requirements are abstracted into a three-level hierarchical structure: object layer, action layer, and execution layer, realizing a progressive advancement from natural language description to computable task representation.
[0008] Furthermore, this stage proposes an InstructionFrame structured instruction representation model, employing a parsing process combining slot filling and template matching to convert natural language instructions into standardized structured instructions. The parsing process consists of four steps: ①Text preprocessing: Normalize the synonyms of Chinese instructions to unify the mapping of the same action intent in different expressions; ② Keyword matching and slot filling: Based on the spatial task entity vocabulary, identify entities such as target objects, reference objects, and target regions and fill in the corresponding fields; ③ Task type determination: Based on the action trigger words, the instructions are classified into task types such as object transfer, interface insertion, button pressing, rotation operation, etc. ④ Action Template Unfolding: Based on the task type, call the corresponding template to automatically unfold the implicit action stages such as grab, aim, and release, and generate action stage prompts.
[0009] The InstructionFrame contains core fields and extended fields: the core fields are task type, target object, target area, and action stage hints, which are the minimum necessary information for task parsing; the extended fields include reference object, spatial relationship, prerequisite steps, and status confirmation conditions, which are used to support the parsing of complex tasks.
[0010] S2. Spatial Task Action Semantic Representation and Trajectory Stage Division This stage constructs a hybrid action semantic representation of "high-level action primitives + low-level continuous parameters", establishing an interpretable bridge between high-level task semantics and low-level continuous control.
[0011] First, based on typical spatial task flows, eight general action primitives are extracted: MoveTo (move to target), Grasp (grab object), Release (release object), Align (attitude alignment), Insert (insert operation), Press (press operation), Rotate (rotate operation), and Verify (state confirmation), covering the core operation units of four typical spatial tasks. Each action primitive is accompanied by three types of continuous parameters: 6-dimensional end-effector pose parameters (x, y, z, roll, pitch, yaw), gripper opening / closing state parameters, and time scale parameters, realizing the mapping from semantic primitives to continuous control variables.
[0012] Secondly, the demonstration trajectory is divided into stages based on the action primitive system: according to the action stage prompts generated by the InstructionFrame, the continuous operation trajectory is divided into stages corresponding to the action primitives according to the semantic boundaries. The start and end frames and primitive labels of each stage are recorded to provide structured annotations for subsequent model training and to provide a staged localization basis for result analysis.
[0013] S3. Lightweight VLA Model Adaptation and Closed-Loop Execution Verification Phase This phase is based on the SmolVLA lightweight visual language action model, combined with a structured semantic layer to complete small-sample adaptation for spatial tasks, and achieves closed-loop verification in the MuJoCo simulation environment.
[0014] At the data organization level, structured instructions, multi-view images, robot states, motion trajectories, and motion stage labels are organized into time-series samples. The structured instructions serialized from the InstructionFrame are used as the formal language input of the model, replacing the original natural language instructions. This reduces the interference of language expression differences on training under small sample conditions and improves the semantic consistency of the samples.
[0015] At the fine-tuning strategy level, a parameter-efficient few-sample adaptation scheme is adopted: the parameters of the visual encoder and most of the language backbone network are frozen, and only the action decoding head, state projection layer, and a small number of adaptation modules are updated, preserving the model's pre-trained visual-language basic capabilities while quickly adapting to the spatial task action distribution. The training process adopts action block supervision, where the model predicts multiple future actions at once, improving the coherence of action execution; the total loss function is composed of a weighted average of the mean square error loss of the robotic arm's continuous actions and the gripper state loss, and is periodically evaluated using a validation set to select the model checkpoint with the lowest loss and the best performance at each stage.
[0016] At the closed-loop execution level, a rolling prediction mechanism is adopted: the model predicts the next H steps of action based on the current image observation, robot state and structured instructions; after the system executes the previous K steps, it re-collects environmental observation and robot state and carries out the next round of prediction. Through continuous visual feedback, the long sequence execution error is corrected, thereby improving the task success rate.
[0017] Compared with the prior art, the present invention has the following beneficial effects: Improve the standardization and interpretability of instruction parsing: Through three-level task requirement abstraction and InstructionFrame structured representation, ambiguous natural language instructions are transformed into structured representations with clear fields and unified semantics, which not only cover the object, space, and action features of spatial tasks, but also facilitate subsequent task review and error localization.
[0018] Constructing a hierarchical action semantic system: A hybrid representation of "action primitives + continuous parameters" is adopted to establish an intermediate layer between high-level task semantics and low-level continuous control, which not only preserves the interpretability of symbolic methods, but also supports the continuous action generation of the learning model.
[0019] Model optimization for small spatial sample scenarios: By using structured instructions to unify semantic input and a highly efficient fine-tuning strategy of freezing some parameters, the VLA model's requirements for spatial task data volume are significantly reduced, and the model adaptation efficiency and stability under small sample conditions are improved.
[0020] A complete and verifiable technical closed loop is formed: a full-process methodological framework is formed from task research, instruction parsing, semantic modeling to model adaptation and simulation verification. The feasibility of the method is verified through simulation tests of two typical space tasks, providing a foundation for subsequent engineering applications. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A schematic diagram of the overall framework of the task instruction parsing and action sequence generation method for spatial embodied intelligence provided by the present invention; Figure 2 A flowchart for the structured parsing of space mission instructions; Figure 3 Flowchart for SmolVLA space mission adaptation and fine-tuning; Figure 4 A schematic diagram of a simulation test case of a plug being inserted into a socket; Figure 5 A schematic diagram of a simulation test case where experimental samples are placed in a box. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0024] This embodiment is based on the MuJoCo physics simulation platform and uses a six-degree-of-freedom robotic arm as the execution vehicle. Two typical spatial operation cases are used to verify the effectiveness of the method of this invention: an interface alignment and insertion task, "inserting a plug into a socket," and a composite operation task with pre-steps, "placing an experimental sample into a box." The specific implementation process is as follows: 1. Abstraction of Space Mission Requirements and Parsing of Structured Instructions First, based on publicly available mission data from the space station, we summarize four typical mission types: material transfer, module operation, tool use, and status confirmation. We clarify the characteristics of mission instructions at the object, space, and action levels, and establish a three-level requirement abstraction framework: object layer, action layer, and execution layer.
[0025] Instruction analysis is performed for each of the two test cases: For the instruction "plug into socket", it is parsed as an insert_operation type task. The InstructionFrame field is: task_type=insert_operation, target_object=plug, reference_object=socket, target_region=socket_port, spatial_relation=into, primitive_hint=MoveTo→Grasp→MoveTo→Align→Insert; For the instruction "Put the experimental sample into the box", it is parsed as a container_placement type task. The InstructionFrame field is: task_type=container_placement, target_object=experimental_sample, container_object=box, precondition_steps=open_box_lid, target_region=inside_box, primitive_hint=MoveTo(handle)→Grasp(handle)→MoveTo(table)→Release(handle)→MoveTo(sample)→Grasp(sample)→MoveTo(inside_box)→Release(sample)→Verify.
[0026] Through equivalent instruction testing, it was verified that instructions for the same task with different expressions can be mapped to consistent structured fields, thus ensuring the semantic stability of language input.
[0027] 2. Action semantic modeling and demonstration data acquisition A simulation environment was built using MuJoCo, configuring a six-DOF robotic arm, an end effector gripper, multi-view cameras, and a task object. Weak gravity conditions were set to simulate a microgravity environment. Successful demonstration trajectories were acquired through manual teleoperation, simultaneously recording multi-view RGB images, robot joint states, end effector pose, gripper states, and motion control parameters.
[0028] A total of 50 valid demonstration trajectories were collected for the case study, which were divided into a 3:1:1 ratio: 30 training tracks, 10 validation tracks, and 10 test tracks. The trajectories were divided into six semantic stages based on action primitives: MoveTo(plug), Grasp(plug), MoveTo(socket), Align(plug,socket), Insert(plug,socket), and Verify(success). Case 2 collected a total of 60 valid demonstration trajectories, which were divided into a 4:1:1 ratio into a training set of 40 trajectories, a validation set of 10 trajectories, and a test set of 10 trajectories. The trajectories were divided into two sub-tasks, namely opening the lid and placing the sample, with a total of nine semantic stages based on the action primitives.
[0029] All trajectory data are labeled with the corresponding action primitive stage, forming a standardized sample of "image-state-structured instruction-action-stage label".
[0030] 3. Small-sample fine-tuning of the SmolVLA model Load the SmolVLA base model and adopt a partially frozen fine-tuning strategy: freeze the main parameters of the visual encoder and the language backbone, and train only the action decoder, state projection layer and a small number of adaptation modules.
[0031] Case 1 training settings: batch size=4, gradient accumulation steps=4, effective batch size=16, learning rate 5e-4, training for 5000 steps, evaluation on the validation set every 500 steps; Case 2 training settings: batch size=8, gradient accumulation steps=2, effective batch size=16, learning rate 5e-4, training for 7000 steps, evaluation on the validation set every 500 steps.
[0032] The training process uses the next 7 steps of action blocks as the supervision signal, and the loss function includes a 6-dimensional robotic arm action loss and a 1-dimensional gripper state loss. Based on the validation set loss curves, 4000 checkpoints were selected for Case 1 and 6000 checkpoints were selected for Case 2 as the final test models.
[0033] 4. Closed-loop simulation testing and result verification A rolling prediction closed-loop execution approach is adopted to conduct task verification on the test set and to count the overall task success rate and the success rate of each stage.
[0034] Case 1 (Plug Insertion into Socket) Test Results: 7 out of 10 test episodes were successful, with a task success rate of 70%; the success rate in the grasping stage was 90%, and the success rates in the alignment and insertion stages were both 70%. Failures were mainly concentrated in the posture alignment stage, reflecting the high requirements of the interface operation for the end-effector pose accuracy.
[0035] Case 2 (Experimental Sample Placement in a Box) Test Results: 6 out of 10 test episodes were successful, with a task success rate of 60%; the success rate was 100% in the opening stage, 80% in sample retrieval, and 60% in sample placement. The results indicate that the method can effectively support complex tasks with pre-processing steps, although the success rate decreases in the latter half of long sequence tasks due to error accumulation.
[0036] The experimental results above verify that the method of the present invention can complete the entire process from natural language instruction parsing, action semantic modeling, model fine-tuning to closed-loop execution under limited data conditions, and can provide technical support for the task understanding and action generation module of the space embodied intelligent system.
[0037] It should be understood that any parts not described in detail in this specification belong to the prior art. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications to certain steps of the method without departing from the scope of protection of the claims of this invention; all such substitutions or modifications fall within the protection scope of this invention, and the scope of protection claimed by this invention should be determined by the appended claims.
Claims
1. A method for parsing task instructions and generating action sequences for spatial embodied intelligence, characterized in that, Includes the following steps: S1. Spatial Mission Requirements Abstraction and Structured Instruction Parsing Stage: Investigate typical mission types within the space station module and summarize the linguistic features of space mission instructions. Abstract mission requirements into a three-level hierarchical structure of object layer, action layer, and execution layer. Based on the InstructionFrame structured representation model, through a parsing process combining slot filling and template matching, convert natural language mission instructions into structured instructions containing mission type, target object, reference object, target region, spatial relationship, and action stage prompts. S2. Semantic Representation of Spatial Task Actions and Trajectory Stage Division: Construct an action primitive system for spatial operation scenarios, and combine end-effector pose parameters, gripper state parameters and time parameters to form a hybrid action semantic representation of "high-level semantic primitives + low-level continuous parameters"; according to the action stage prompts in the structured instructions, divide the continuous operation trajectory into semantic stages of the corresponding action primitives to form an annotable action sequence with stage boundaries. S3. Lightweight VLA Model Adaptation and Closed-Loop Execution Verification Stage: Based on the SmolVLA lightweight visual language action model, structured instructions, multi-view images, robot states, and motion trajectories are organized into time-series training samples. A small-sample fine-tuning strategy with partial network parameter freezing is adopted to complete the spatial task adaptation. Through the action block output mechanism of rolling prediction, closed-loop action execution and task result verification are realized in the MuJoCo virtual simulation environment.
2. The method according to claim 1, characterized in that, In step S1, the three-level hierarchical structure corresponds to the following: the object layer is used to extract the target object, reference object and target region involved in the task, and to clarify the operation entity and spatial reference; the action layer is used to organize the action primitive type and execution order corresponding to the task, and to clarify the phased process of the task; the execution layer is used to map the robot's visual observation, motion state and continuous control parameters, and to connect semantic representation and underlying execution.
3. The method according to claim 1, characterized in that, In step S1, the InstructionFrame structured instruction includes core fields and extended fields. The core fields include task type, target object, target area and action stage prompts. The extended fields include reference object, spatial relationship, prerequisite steps and status confirmation conditions. The parsing process sequentially completes four steps: text preprocessing and synonym normalization, keyword matching and slot filling, task type judgment and action template expansion.
4. The method according to claim 1, characterized in that, In step S2, the action primitive system includes eight basic action primitives: MoveTo, Grasp, Release, Align, Insert, Press, Rotate, and Verify, covering four typical tasks: space material transfer, interface operation, tool use, and status confirmation. Each action primitive corresponds to continuous control parameters, including 6-dimensional end pose parameters, gripper opening and closing state parameters, and time scale parameters.
5. The method according to claim 1, characterized in that, In step S3, the small sample fine-tuning strategy adopts a parameter-efficient adaptation method, freezing the parameters of the visual encoder and most of the language backbone network, and only updating the parameters of the action decoding head, state projection layer and a small number of adaptation modules. The training process uses future multi-step action blocks as supervision signals, and the total loss function is composed of the weighted average error loss of the continuous action of the robotic arm and the gripper state loss. The optimal model checkpoint is selected by the validation set loss and stage performance.
6. The method according to claim 1, characterized in that, In step S3, the closed-loop execution adopts a rolling prediction mechanism: the model receives multi-view images, robot state and structured instructions at the current moment, and predicts the future H-step action blocks; after the system executes the first K steps of the action, it re-collects environmental observations and robot state, and carries out the next round of prediction, correcting the long sequence execution error through closed-loop feedback.
7. A task instruction parsing and action sequence generation system for spatial embodied intelligence, characterized in that, include: ① The task parsing module is used to receive natural language task instructions and output standardized InstructionFrame structured instructions; ② Action semantic modeling module, used to construct the spatial task action primitive system and complete the semantic stage division and annotation of the demonstration trajectory; ③ Model adaptation module, used to call the SmolVLA basic model to complete the spatial task sample organization, small sample fine-tuning and action block prediction; ④ Simulation verification module, used to execute closed-loop action generation in the MuJoCo virtual environment and output the task success rate and stage verification results.