Universal robot operating system and method based on combination of cognition and execution
Through a general robot operating system based on cognition and execution, using advanced cognitive planners and unchanged operable representation modules, combined with vision-language model and closed-loop control, the problem of cognition and execution disconnection of the robot operating system in an open environment is solved, and the operation ability and task success rate are improved.
Patent Information
- Application Number
- CN202510551152.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-29
AI Technical Summary
The existing robot operating systems have problems of disconnection between cognition and execution in dynamic open environments, resulting in limited operational capabilities.
A general robot operating system based on the combination of cognition and execution, including an advanced cognitive planner, an invariable operable representation module and a pre-trained general embodied agent, decomposes tasks into primitive actions, and generates an invariable operable representation using the visual-language model, image segmentation model and target tracking model, and implements action execution in combination with closed-loop control.
Effectively prevent the robot's cognition and execution from being disconnected from the execution of the operation, improve the robot's operation ability in an open environment, and improve the task success rate and robustness.
Smart Images

Figure CN120382503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent robot control, and particularly relates to a general robot operating system and method based on the combination of cognition and execution. Background Art
[0002] With the rapid development of artificial intelligence and robot technology, robots are increasingly widely used in industrial, service, medical, household and other fields. Especially in dynamic and unstructured environments, robots need to have higher autonomous decision-making capabilities and adaptability, which puts higher requirements on the cognitive and execution capabilities of robot operating systems.
[0003] Current robot operation technologies are mainly divided into two categories: declarative skill solutions and procedural skill solutions. Although both have achieved certain effects in specific scenarios, they both have significant defects, resulting in limited performance of robot operations in dynamic open environments. For example, the prior art discloses a declarative skill solution that attempts to integrate a vision-language model into a robot system and uses a large multi-modal model to generate operation instructions for open-domain tasks. These large multi-modal models are good at task understanding and guide the robot to execute tasks by generating executable operation instructions. Although the system integrating the vision-language model performs excellently in understanding, due to the lack of embodied experience, the output needs to be limited to executable actions. This method forces the language model to handle spatio-temporal reasoning without physical intuition, often resulting in unreasonable task planning. For example, in the task of "placing block A on block B", the lack of sufficient spatial understanding (such as shape, height) often leads to the generation of a fatal error action sequence. The prior art also discloses a procedural skill solution based on an embodied model. The embodied model usually adopts a data-driven trajectory fitting method to manipulate objects according to instructions. This method relies on the motion trajectories learned from a large amount of data to achieve object operation capabilities. When facing environmental changes such as fluctuations in lighting conditions, camera pose deviations, and context changes, the performance of the embodied model drops sharply. Therefore, how to avoid the disconnection between the cognition and execution of robot operations and improve the operation capabilities of robots in open environments is an extremely important technical problem to be solved. Summary of the Invention
[0004] To solve the problem in the above prior art that the disconnection between the cognition and execution of robot operations affects the operation capabilities of robots in open environments, the present invention proposes a general robot operating system and method based on the combination of cognition and execution, which effectively prevents the disconnection between the cognition and execution of robot operations and improves the operation capabilities of robots in open environments.
[0005] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0006] A general robot operating system based on the combination of cognition and execution, comprising: a high-level cognitive planner, an invariant operable representation module, and a pre-trained general embodied intelligent agent;
[0007] The high-level cognitive planner includes a high-level cognitive planning model for obtaining an operation object and an instruction, and based on the operation object and the instruction, decomposing the task into a number of primitive actions;
[0008] The invariant operable representation module is used to utilize the high-level cognitive planning model in the high-level cognitive planner to transform the primitive action into an invariant operable representation, and input the invariant operable representation into a preset general embodied intelligent agent;
[0009] The general embodied intelligent agent is used to receive the invariant operable representation and map the invariant operable representation into a robot execution action through closed-loop control to execute the original action.
[0010] Preferably, the high-level cognitive planning model includes a vision-language model, an image segmentation model, and a target tracking model. The high-level cognitive planner takes the operation object and the instruction as the input of the vision-language model, and the vision-language model queries and outputs a number of primitive actions. The mathematical expression of each primitive action is as follows:
[0011] A i ={T i ,obj i ,des i}
[0012] Wherein, A i represents the i-th primitive action, T i represents the action type, obj i represents the operation object name, and des i represents the destination or may not exist.
[0013] Preferably, the transformation of the primitive action into an invariant operable representation includes:
[0014] Using the image segmentation model to segment the object related to the primitive action and output a segmentation result;
[0015] Taking the segmentation result as the input of the vision-language model, and the vision-language model outputs the final selection of the operation object as the third perspective mask M i and the first perspective mask depth D i , and provides a corresponding direction for tasks with directional constraints, and the direction is normalized and incorporated into the constraint C i ;
[0016] Integrate the third - perspective mask M i , the first - perspective mask depth D i , the constraint C i and the acquired visual - sensor data to generate an invariant actionable representation R i containing the action type T i as follows:
[0017] R i ={T i , M i , D i , C i}.
[0018] Preferably, the third - perspective mask M i includes the third - perspective mask of the gripper the third - perspective mask of the object to be operated and the third - perspective mask of the destination
[0019] Preferably, the first - perspective mask depth D i includes the first - perspective depth mask of the gripper the first - perspective depth mask of the object to be operated the first - perspective depth mask of the destination
[0020] Preferably, the closed - loop control includes high - frequency control and low - frequency control.
[0021] Preferably, the high - frequency control is to update the third - perspective mask M i and the first - perspective mask depth D i using the target - tracking model to obtain an updated invariant actionable representation and the mathematical expression is as follows:
[0022]
[0023] where, represents the updated third - perspective mask, represents the updated first - perspective mask depth;
[0024] Preferably, the low - frequency control is to use the object to be operated with a task - related object label, the gripper state, and the primitive action as the input of the vision - language model, and the vision - language model outputs a state judgment result of the current action executed by the robot. Until the state judgment result is successful, execute the next primitive action or end the task.
[0025] Preferably, use reinforcement learning and imitation learning to train the general embodied agent to obtain a trained general embodied agent.
[0026] The present invention also provides a general robot operation method based on the combination of cognition and execution, including:
[0027] S1. Obtain an operation object and an instruction. Based on the operation object and the instruction, use a high-level cognitive planner including a high-level cognitive planning model to decompose the task into several primitive actions;
[0028] S2. Use the high-level cognitive planning model in the high-level cognitive planner to transform the primitive actions into invariant actionable representations;
[0029] S3. Input the invariant actionable representation into a preset general embodied intelligent agent, and the general embodied intelligent agent maps the invariant actionable representation into a robot execution action through closed-loop control to execute the original action.
[0030] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0031] The present invention provides a general robot operation system and method based on the combination of cognition and execution. The system includes a hierarchical intelligent architecture of a high-level cognitive planner, an invariant actionable representation module, and a general embodied intelligent agent. First, the high-level cognitive planner as the "brain" decomposes the task into several primitive actions based on the operation object and the instruction. Then, the invariant actionable representation module, as a symbolic bridge, uses the high-level cognitive planning model in the high-level cognitive planner to transform the primitive actions into invariant actionable representations. Finally, the general embodied intelligent agent maps the invariant actionable representation into a robot execution action to execute the original action, effectively preventing the disconnection between cognition and execution in robot operation through the three-layer intelligent architecture and improving the operation ability of the robot in an open environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It shows a structural block diagram of a general robot operation system based on the combination of cognition and execution proposed in an embodiment of the present invention;
[0033] Figure 2 It shows a working principle diagram of a general robot operation system based on the combination of cognition and execution proposed in an embodiment of the present invention;
[0034] Figure 3 It shows a threshold curve diagram of a training method principle for the general embodied intelligent agent proposed in an embodiment of the present invention;
[0035] Figure 4 It shows a flow block diagram of a general robot operation method based on the combination of cognition and execution proposed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present patent;
[0037] For those skilled in the art, it is understandable that some well-known content descriptions in the drawings may be omitted;
[0038] To facilitate the understanding of this embodiment, first, the prior art information of this embodiment is introduced as follows:
[0039] Regarding robotic manipulation learning, in recent years, robotic manipulation has attracted extensive research attention. A direct approach is to use imitation learning, which is a type of supervised learning that maps observations to actions. Early methods achieved good results by designing effective network architectures, constructing diverse training objectives, and using appropriate representation methods. To deploy policies in various real-world scenarios, several large-scale robotic datasets have been introduced; however, the scale of these datasets still pales in comparison to the diversity of real-world environments. Inspired by the success of large pre-trained models, researchers have started to directly fine-tune vision-language models on robotic data to enhance the generalization ability of robotic models. However, due to the domain differences between robotic data and pre-trained data, the fine-tuned models often suffer from catastrophic forgetting, resulting in a significant decline in cognitive performance. Another approach is to use reinforcement learning for robotic manipulation, which can develop robust policies. However, these methods often encounter difficulties when dealing with complex tasks that require language understanding.
[0040] Regarding the representation methods of robotic manipulation, structured representation aims to address the generalization challenges in robotic manipulation. In robotic manipulation, structured representation is a method of encoding complex environmental or object information into an abstract form with a clear geometric, semantic, or functional structure to support efficient reasoning, planning, and task execution. The core idea is to capture key information through simplified symbolic elements (such as key points, 6D poses, constraints), thereby reducing the complexity of the problem and enhancing the generalization ability. Some methods predict the key frames of the motion process and use them as the original cues for low-level control. Other methods extract interaction trajectories from human videos to utilize diverse data sources, thereby enhancing the generalization ability for a large number of previously unseen objects, but this usually requires task-specific annotations. In addition, some studies utilize the capabilities of vision-language models and foundation models to extract key points and 6D poses as representations and combine them with motion planning for low-level control. However, due to the lack of physical interaction data during pre-training, the plans generated by vision-language models often do not conform to the physics of the real world, thus limiting their effectiveness in practical applications.
[0041] For general robot operations, to enhance the model's generalization ability in different environments, especially in the transfer from simulated to real-world environments, existing methods often adopt domain adaptation and domain randomization techniques. Domain adaptation aims to create additional synthetic images based on existing images. In contrast, domain randomization is more commonly used because it only needs to randomize factors such as material textures, object poses, and camera parameters. Domain randomization is usually applied in a simulated environment, and then the trained policy is transferred to a real robot.
[0042] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0043] Embodiment 1
[0044] See Figure 1 and Figure 2 , this embodiment proposes a general robot operating system based on the combination of cognition and execution, including: a high-level cognitive planner, an invariant operable representation module, and a pre-trained general embodied intelligent agent;
[0045] The high-level cognitive planner includes a high-level cognitive planning model for obtaining an operation object and an instruction, and based on the operation object and the instruction, decomposing the task into a number of primitive actions; the operation object refers to the observed RGB image;
[0046] The invariant operable representation module is used to use the high-level cognitive planning model in the high-level cognitive planner to convert the primitive action into an invariant operable representation, and input the invariant operable representation into a preset general embodied intelligent agent;
[0047] The general embodied intelligent agent is used to receive the invariant operable representation and map the invariant operable representation into a robot execution action through closed-loop control to execute the original action.
[0048] Here, the system includes a hierarchical intelligent architecture of a high-level cognitive planner, an invariant operable representation module, and a general embodied intelligent agent, denoted as the RoBridge framework. First, the high-level cognitive planner as the brain decomposes the task into a number of primitive actions based on the operation object and the instruction, mainly responsible for the high-level cognitive planning of the task. Then, the invariant operable representation module serves as a symbolic bridge, using the high-level cognitive planning model in the high-level cognitive planner to convert the primitive action into an invariant operable representation. Finally, the general embodied intelligent agent uses the invariant operable representation to map into a robot execution action to execute the original action, effectively preventing the disconnection between cognition and execution of robot operations through the three-layer intelligent architecture and improving the operation ability of the robot in an open environment.
[0049] The advanced cognitive planning model includes a vision - language model, an image segmentation model, and an object tracking model. The vision - language model is GPT - 4o,
[0050] GPT - 4o can process multiple input modalities such as text, images, and audio simultaneously and generate natural language or structured outputs. It has functions of task decomposition, semantic understanding, and state judgment. Task decomposition means parsing a natural language instruction like "Put block A on block B" into a structured primitive action sequence like "Grab block A → Move above block B → Place". Semantic understanding means understanding object attributes (shape, position) and task constraints (orientation, spatial relationship) by combining the observed RGB image. State judgment means, in closed - loop control, judging whether the action is executed successfully by analyzing the gripper state and image labels.
[0051] The image segmentation model includes GroundingDINO and SAM. GroundingDINO is a Transformer - based object detection model that can locate and segment corresponding objects from an image according to text descriptions, and output bounding boxes and masks. SAM is a general image segmentation model that can generate high - quality segmentation masks for any object in an image without training.
[0052] The object tracking model is Track - Anything. Track - Anything is an algorithm - based model for object tracking that can continuously track the position and mask of an object across frames and support real - time updates in dynamic scenarios.
[0053] The advanced cognitive planner takes the operation object and instruction as the input of the vision - language model GPT - 4o. The vision - language model GPT - 4o queries and outputs several primitive actions. The mathematical expression of each primitive action is as follows:
[0054] A i ={T i ,obj i ,des i}
[0055] Where A i represents the i - th primitive action, T i represents the action type, obj i represents the name of the operation object, and des i represents the destination or may not exist.
[0056] The conversion of the primitive action into an invariant operable representation includes:
[0057] Segment the objects related to the original action using the image segmentation models GroundingDINO and SAM, and output the segmentation results;
[0058] Use the segmentation results as the input to the vision-language model, and the vision-language model outputs the final selection of the operating object as the third perspective mask M i and the first perspective mask depth D i , and provide corresponding directions for tasks with directional constraints such as opening a drawer or turning a faucet, and the directions are normalized and incorporated into the constraint C i ; Constraint C i includes the pose and movement direction of the end effector;
[0059] Integrate the third perspective mask M i , the first perspective mask depth D i , the constraint C i and the acquired visual sensor data to generate an invariant actionable representation R i containing the action type T i as follows:
[0060] R i ={T i ,M i ,D i ,C i}.
[0061] The invariant actionable representation aims to help RoBridge achieve better domain invariance and reduce the impact of environmental and task variations on the model.
[0062] The third perspective mask M i includes the third perspective mask of the gripper the third perspective mask of the operating object and the third perspective mask of the destination (if any).
[0063] The first perspective mask depth D i includes the first perspective depth mask of the gripper the first perspective depth mask of the operating object the first perspective depth mask of the destination (if any).
[0064] In this embodiment, a hierarchical intelligent architecture RoBridge for general robot operation is proposed. Through a three-layer architecture of a brain: an advanced cognitive planner, a symbolic bridge: an invariant manipulable representation module, and an embodied agent, the paradigm dilemma of the disconnection between cognitive abstraction and physical execution in traditional methods is broken. This embodiment also designs a general embodied agent that can convert the invariant manipulable representation into specific execution actions and maintain excellent performance under various interference conditions. It should also be specifically stated that for the problem of the disconnection between cognitive abstraction and physical execution, the present invention combines a vision-language model and effectively bridges cognition and execution through an invariant manipulable representation, resulting in a significant increase in the task success rate and the ability to effectively execute new tasks. For the poor execution effect in unseen environments, the present invention combines the declarative skills of the vision-language model and the procedural skills of reinforcement learning to improve the operation ability of the robot in diverse environments.
[0065] Embodiment 2
[0066] The general embodied agent maps the invariant manipulable representation to a robot execution action through closed-loop control to execute the original action. In other words, a policy needs to be learned. The original action "reach" usually involves moving the end effector of the robot to a specific target position. Different from tasks such as grasping or placing that involve complex object interaction and decision-making, "reach" can be effectively solved through motion planning. For other original actions, the general embodied agent is trained using reinforcement learning and imitation learning to obtain a trained general embodied agent. The trained general embodied agent can effectively execute actions under variable input conditions while ensuring robustness and consistent performance.
[0067] To maintain the information accuracy and iterativeness of the original action in a dynamic environment, closed-loop control is added. Since each part of the closed-loop control has a different speed, the closed-loop control is divided into high-frequency control and low-frequency control. At each time step, the closed-loop control generates new
[0068] The high-frequency control is to update the third perspective mask M i and the first perspective mask depth D i using the target tracking model to obtain the updated invariant manipulable representation. The mathematical expression is as follows:
[0069]
[0070] where represents the updated third perspective mask, represents the new first perspective mask depth.
[0071] For low-frequency control, we attempt to check and obtain the following three states: success, error, and normal. We use GPT-4o in combination with the gripper state to determine whether the task is successful.
[0072] The low-frequency control takes the manipulated object with task-related object tags, the gripper state, and the primitive action as the input to the vision-language model GPT-4o. The vision-language model GPT-4o outputs the state judgment result of the current action executed by the robot. The state judgment result refers to the judgment of whether the current action is successful. If the state judgment result is successful, the next primitive action will be continued or the task will be ended; otherwise, if the state judgment result is a failure, R will be regenerated i as the input until the state judgment result is successful, and then the next primitive action will be executed or the task will be ended.
[0073] The general embodied agent is trained. The aim is to train a general embodied agent that can achieve high success rates and robustness in diverse scenarios, robot configurations, and tasks to ensure its reliable and consistent performance. To this end, we use multi-stage training to train the high-level cognitive planner, which is denoted as GEA here, as Figure 3 shown. The training process includes:
[0074] First, reinforcement learning (RL) training is carried out. To achieve high performance, the reinforcement learning method is used to train the expert agent π for each specific task e , ensuring that each task can be executed efficiently. In addition, to enhance the robustness of the expert agent, domain randomization of object shape, robot arm position, and camera direction is introduced during the training process.
[0075] Then, imitation learning (IL) training is carried out. Next, the general embodied agent π is trained g . The expert agent is used to generate high-quality data for each task, and the general interaction representation R is extracted from the data as π gInput. In addition to domain randomization used in expert agent training, various domain randomization techniques such as depth distortion, inflation, random offset, and modifying the mask by addition and deletion are further incorporated. These strategies aim to improve the generalization ability and robustness of the trained agent under different environmental conditions. To simulate the situation of noisy depth sensors in the real world, a series of enhancement techniques are used when processing depth images. Specifically, depth distortion is utilized to apply Gaussian offset to simulate perspective changes and sensor noise. In addition, Gaussian blur is introduced to simulate common sensor blur and focusing problems. Moreover, random masks are introduced to create artificial holes in the depth map, simulating occlusions and data loss that often occur in practical applications. Since the masks generated by the model may have inaccuracies and incomplete coverage, we adopt techniques such as random offset and random cropping. These methods are used to mitigate the risk of the agent relying too much on the mask.
[0076] Finally, continuous skill aggregation is performed. Agents often experience cumulative errors in imitation learning. To address this issue, we introduce the iterative optimization strategy of DAgger. Recognizing that online DAgger may lead to instability during each update, while offline DAgger has a slow update speed in a multi-task learning environment, we develop an adaptive sampling mechanism for offline DAgger. This mechanism adjusts the sampling frequency according to task complexity, sampling more samples for more challenging tasks. Initially, all tasks are given equal importance with a unified weight. In each iteration, the policy π g is trained using the current dataset, and tasks are sampled according to their weights. The policy is tested, and the task difficulty is evaluated based on the results, and the task weights are updated accordingly. Failures are recorded, and corrective data is generated using the expert π e and added to the task dataset.
[0077] Example 3
[0078] This example illustrates the effects brought by a general robot operating system based on the combination of cognition and execution proposed in the above example.
[0079] Table 1 Meta-World Benchmark Results
[0080]
[0081] Referring to Table 1, we tested 50 tasks of Metaworld and changed the background, lighting, object color, and camera pose to evaluate the generalization ability. The average success rate is shown in the last column of the table. The general robot operating system proposed in the above embodiment, which combines cognition and execution, achieved the best performance in various tests, with an average success rate of 82.12%, 11.28% higher than the best baseline. This proves the effectiveness and robustness of our system.
[0082] Table 2 shows the experimental results of unseen tasks
[0083]
[0084] Referring to Table 2, the symbol "-" usually indicates that "this method cannot handle this task due to underlying principle limitations". We tested five new tasks on Metaworld that were unrelated to the tasks used in training. RoBridge achieved an average success rate of 75% in these new tasks. This shows that our general robot operating system can effectively execute tasks unseen during training, thus eliminating the need to collect data for each specific task and reducing the data collection cost.
[0085] Table 3 shows the experimental results in the real world
[0086]
[0087] The experimental results in the real world are shown in Table 4 above. Our system achieved an average success rate of 83.3% in four tasks with only 5 pieces of data per task, exceeding Stanford ReKep by 34.1%. The end-to-end method π0 was only 18.3%.
[0088] Table 4 shows the long-term experimental results in the real world
[0089]
[0090] Referring to Table 4, we also demonstrated the execution ability of our system in long-sequence tasks. The average completion length of RoBridge in long-sequence tasks is 3.0. The performance of RoBridge in long-sequence tasks demonstrates its comprehensive planning and execution ability.
[0091] We also observed that RoBridge can correct errors to a certain extent. When the task fails while trying to grasp the second building block, RoBridge can re-plan and successfully complete the task on the second attempt.
[0092] Embodiment 4
[0093] Referring to Figure 4 , this embodiment proposes a general robot operation method that combines cognition and execution, including:
[0094] S1. Obtain an operation object and an instruction. Based on the operation object and the instruction, use a high-level cognitive planner including a high-level cognitive planning model to decompose the task into a number of primitive actions;
[0095] S2. Use the high-level cognitive planning model in the high-level cognitive planner to transform the primitive actions into invariant actionable representations;
[0096] S3. Input the invariant actionable representation into a preset general embodied intelligent agent, and the general embodied intelligent agent maps the invariant actionable representation into a robot execution action through closed-loop control to execute the original action.
[0097] In this embodiment, first, the high-level cognitive planner as the brain decomposes the task into a number of primitive actions based on the operation object and the instruction. Then, the invariant actionable representation module, as a symbolic bridge, uses the high-level cognitive planning model in the high-level cognitive planner to transform the primitive actions into invariant actionable representations. Finally, the general embodied intelligent agent uses the invariant actionable representation to map into a robot execution action to execute the original action, effectively preventing the disconnection between the cognition and execution of the robot operation through a three-layer intelligent architecture, and improving the operation ability of the robot in an open environment.
[0098] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, and are not limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A general robot operating system based on the combination of cognition and execution, characterized in that, Comprising: A high-level cognitive planner, an invariant actionable representation module, and a pre-trained general embodied intelligent agent; The high-level cognitive planner includes a high-level cognitive planning model for obtaining an operation object and an instruction, and based on the operation object and the instruction, decomposing the task into a plurality of primitive actions; The invariant actionable representation module is configured to use the high-level cognitive planning model in the high-level cognitive planner to convert the primitive action into an invariant actionable representation, and input the invariant actionable representation into a preset general embodied intelligent agent; The general embodied intelligent agent is configured to receive the invariant actionable representation and map the invariant actionable representation into a robot execution action through closed-loop control to execute the original action.
2. The general robot operating system based on the combination of cognition and execution according to claim 1, characterized in that The high-level cognitive planning model includes a vision-language model, an image segmentation model, and a target tracking model. The high-level cognitive planner uses the operation object and the instruction as inputs to the vision-language model, and the vision-language model queries and outputs a plurality of primitive actions. The mathematical expression of each primitive action is as follows: A i ={T i ,obj i ,des i} Among them, A i represents the i-th primitive action, T i represents the action type, obj i represents the name of the operation object, des i represents the destination or may not exist.
3. The general robot operating system based on the combination of cognition and execution according to claim 2, characterized in that, The conversion of the primitive action into an invariant actionable representation includes: Using the image segmentation model to segment the object related to the primitive action and output a segmentation result; Use the segmentation result as the input to the vision-language model, and the vision-language model outputs the final selection of the object to be operated as the third perspective mask M i and the first perspective mask depth D i , and provide the corresponding direction for tasks with directional constraints, and the direction is normalized and incorporated into the constraint C i ; Integrate the third - perspective mask M i , the first - perspective mask depth D i , the constraint C i with the acquired visual - sensor data to generate an invariant actionable representation R i that includes the action type T i as follows: R i = {T i , M i , D i , C i}。 4. The general robot operating system based on the combination of cognition and execution according to claim 3, characterized in that, The third - perspective mask M i The third - perspective mask including the gripper The third - perspective mask of the object to be operated and the third - perspective mask of the destination 5. The general robot operating system based on the combination of cognition and execution according to claim 4, characterized in that, The first - perspective mask depth D i The first - perspective depth mask including the gripper The first - perspective depth mask of the object to be operated The first - perspective depth mask of the destination 6. The general robot operating system based on the combination of cognition and execution according to claim 3, characterized in that The closed-loop control includes high-frequency control and low-frequency control.
7. The general robot operating system based on the combination of cognition and execution according to claim 6, characterized in that, The high-frequency control is to use the target tracking model to update the third perspective mask M i and the first perspective mask depth D i to obtain an updated invariant actionable representation The mathematical expression is as follows: Among them, represents the updated third perspective mask, represents the depth of the updated first perspective mask.
8. The general robot operating system based on the combination of cognition and execution according to claim 6, characterized in that The low-frequency control is to use the operation object with the task-related object label, the gripper state, and the primitive action as inputs to the vision-language model, and the vision-language model outputs a state judgment result of the current robot execution action until the state judgment result is successful, and then execute the next primitive action or end the task.
9. The general robot operating system based on the combination of cognition and execution according to any one of claims 1-8, characterized in that Training the general embodied intelligent agent using reinforcement learning and imitation learning to obtain a trained general embodied intelligent agent.
10. A general robot operation method based on the combination of cognition and execution, characterized in that, Comprising: S1. Obtain an operation object and an instruction, and based on the operation object and the instruction, use a high-level cognitive planner including a high-level cognitive planning model to decompose the task into a plurality of primitive actions; S2. Use the high-level cognitive planning model in the high-level cognitive planner to convert the primitive action into an invariant actionable representation; S3. Input the invariant actionable representation into a preset general embodied intelligent agent, and the general embodied intelligent agent maps the invariant actionable representation into a robot execution action through closed-loop control to execute the original action.