An action generation method and system based on ontology self-thinking and discrimination mechanism
By using an action generation method based on ontology self-thinking and discrimination mechanisms, task instructions are decomposed into short-range sub-tasks and multimodal information is fused, which solves the problem of converting abstract instructions into physical paths in robot control strategies, and realizes the accurate execution and reliability improvement of robot actions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROOTCLOUD TECH CO LTD
- Filing Date
- 2026-06-04
- Publication Date
- 2026-07-31
AI Technical Summary
Existing robot control strategies struggle to translate abstract instructions into physical motion paths, leading to a disconnect between task understanding and action execution, resulting in deficiencies in ontology perception and inaccurate action execution.
An action generation method based on ontology self-thinking and discrimination mechanism is adopted. The task planning module decomposes the task instructions into short-term sub-tasks, and the action generation module integrates multimodal information for state thinking and action output. The feasibility of the task and the rationality of the action are analyzed. Combined with a hierarchical decision-making architecture inspired by biological neuroscience, the precise alignment of action planning and mechanical ontology is achieved.
It achieves precise conversion from abstract instructions to physically realizable paths, improving the accuracy and reliability of robot action execution and reducing the risk of execution failure.
Smart Images

Figure CN122480973A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, specifically to a method and system for generating actions based on a self-thinking and discrimination mechanism. Background Technology
[0002] The Vision-Language-Action (VLA) model, as an end-to-end artificial intelligence model that unifies visual perception, natural language understanding, and motion control into a single framework, is widely used in fields such as robotics strategy. It generates mechanical actions that match the mechanical entity by visually observing the real world and understanding the target task instructions. Currently, general robot control strategies face a disconnect between task understanding and action execution due to deficiencies in entity perception, making it difficult to translate abstract instructions into physical motion paths achievable by the physical entity. Summary of the Invention
[0003] This invention provides an action generation method and system based on ontology self-thinking and discrimination mechanism to solve the problem that existing technologies have difficulty in converting abstract instructions into motion paths that can be realized by physical entities.
[0004] In a first aspect, the present invention provides an action generation method based on an ontology self-thinking and discrimination mechanism, applied to an action generation system based on an ontology self-thinking and discrimination mechanism, the system including a task planning module and an action generation module, the method comprising: The received task instructions are broken down into short subtasks using the task planning module; The action generation module receives short-range subtasks, integrates the mechanical body state, action execution history, and environmental changes, performs state thinking and action output, and obtains the body cognitive thinking and generated actions. The action generation module is used to conduct task feasibility analysis and action rationality analysis based on ontology cognition and action generation. If the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, then the action to be executed is output to the robot.
[0005] This invention breaks down task instructions into short-term subtasks through a task planning module, avoiding the gap caused by directly converting abstract instructions into action paths. The action generation module integrates multimodal information, performs state thinking and action output, and uses both feasibility analysis and action rationality analysis to ensure the feasibility of the task and the rationality of the action, achieving accurate conversion from abstract instructions to physically feasible paths.
[0006] In one alternative implementation, the task planning module is constructed as follows: Acquire videos of human actions and robot work, and use video understanding models to annotate the human action videos, robot work videos, and robot work videos with action descriptions, action logic, dynamic principles, and geometric relationships to build a dataset; Fine-tuning and reinforcement learning are performed on a pre-trained multimodal large model on the dataset, and motion-based knowledge is injected to obtain a task planning module. The motion-based knowledge includes kinematic constraints, dynamic principles, and spatial geometric relationships.
[0007] This invention provides training materials for models to learn the laws of physical motion by using videos of human actions and robot work as samples. It fine-tunes and reinforces multimodal large models, injects essential knowledge of motion, and breaks through the limitations of pure semantic understanding.
[0008] In one optional implementation, the action generation module includes a temporal ontology-aware feature fusion module, a Transformer-based state-based language output module, and a Diffusion-based action generation module. The action generation module receives short-range subtasks, fuses the mechanical ontology state, action execution history, and environmental changes, performs state-based thinking and action output, and obtains ontology-based cognitive thinking and generated actions, including: The temporal ontology perception feature fusion module uses a multimodal attention mechanism to uniformly map the mechanical ontology state, action execution history, and environmental changes to obtain fused features; The Transformer-based state-thinking language output module takes fused features and action features based on different inputs of the discrimination task as input, and uses a self-thinking mechanism to output a natural language description that judges the current state. By using a motion generation module based on a Diffusion model, with fused features and linguistic descriptions as conditions, noise is gradually removed from Gaussian noise to generate motion trajectory sequences.
[0009] This invention utilizes a temporal ontology perception feature fusion module to fuse mechanical ontology state, action execution history, and environmental changes, eliminating differences in different modal dimensions. It uses a Transformer-based state thinking language output module to output natural language descriptions, transforming physical states into interpretable semantics. Finally, it uses a Diffusion-based action generation module to generate motion trajectory sequences, outputting actions that conform to physical laws.
[0010] In an optional implementation, after using the temporal ontology perception feature fusion module to uniformly map the mechanical ontology state, action execution history, and environmental changes through a multimodal attention mechanism to obtain fused features, the method further includes: A recursive hidden state embedding mechanism is introduced, which inputs the fusion feature encoding from the previous time step with the visual features, ontology features, and instruction features from the current time step for feature fusion.
[0011] This invention introduces a recursive hidden state embedding mechanism to achieve continuous transmission of temporal information, enhance the correlation of action timing, and improve the smoothness and stability of motion trajectory.
[0012] In one optional implementation, the ontology-based cognitive thinking and generated actions undergo task feasibility analysis and action rationality analysis, including: The generated action, fused features, and trained parameters are jointly processed, and the sigmoid function is used for binary discrimination to output the discrimination symbol. If the discrimination symbol is the first symbol, then the task is determined to be infeasible; If the discrimination symbol is the second symbol, then the task is deemed feasible; If the discrimination symbol is the third symbol, then the action is deemed reasonable; If the discrimination symbol is the fourth symbol, then the action is deemed unreasonable.
[0013] This invention utilizes the Sigmoid function for binary discrimination to determine the feasibility of a task and the rationality of an action, thereby avoiding invalid task execution and risky actions, and thus improving the reliability of decision-making.
[0014] In one alternative implementation, the method further includes: Learning motion logic knowledge using visual data from human operation videos and robotic arm demonstration videos; Utilize macroscopic physical phenomenon data to enhance the model's understanding of physical motion; Basic motion cloning is performed using robotic arm operation data.
[0015] This invention uses visual data from human operation videos and robotic arm demonstration videos to enable models to learn motion logic, laying the foundation for general motion understanding, strengthening the model's cognition of physical motion, and improving the model's understanding of physical laws. Through basic motion cloning, basic motion capabilities can be quickly mastered, taking into account universality, physicality, and proprioceptive adaptability.
[0016] In one alternative implementation, the method further includes: A reward function based on task success rate and physical compliance is constructed, and the discriminator accuracy of the action generation module is optimized through reinforcement learning algorithm.
[0017] This invention enhances the effectiveness of the discrimination results by focusing on the success rate of the task, introduces physical compliance constraints, improves the accuracy of the discriminator in the action generation module, and improves the reliability of the dual discrimination.
[0018] Secondly, this invention provides an action generation system based on an ontology self-thinking and discrimination mechanism, the system comprising: The task planning module is used to break down received task instructions into short subtasks; The action generation module receives short-range subtasks, integrates the mechanical body's state, action execution history, and environmental changes, performs state thinking and action output to obtain the body's cognitive thinking and generated actions; it performs task feasibility analysis and action rationality analysis on the body's cognitive thinking and generated actions; if the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, it outputs the action to be executed by the robot.
[0019] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the action generation method based on ontology self-thinking and discrimination mechanism described in the first aspect or any corresponding embodiment.
[0020] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute the action generation method based on ontology self-thinking and discrimination mechanism described in the first aspect or any corresponding embodiment above. Attached Figure Description
[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0022] Figure 1 This is a flowchart illustrating an action generation method based on an ontology self-thinking and discrimination mechanism according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the task planning module processing flow according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the processing flow of the action generation module according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the action generation process according to an embodiment of the present invention; Figure 5 This is a structural block diagram of an action generation system based on an ontology self-thinking and discrimination mechanism according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0025] When dealing with the collaborative relationship between perception, decision-making, and control, existing VLA models mainly employ the following technical approaches: (1) End-to-End VLA: Visual features, language instructions and action sequences are regarded as a unified sequence modeling problem. Based on the pre-trained VLM (Vision Language Model), the action space is quantized into special "action tokens". The control parameters are predicted directly from pixels and text using a unified Transformer architecture. This path attempts to implicitly cover ontology features through massive data regression. (2) Multi-modal Spatial Perception: This approach attempts to supplement the robot's geometric understanding of the physical world by introducing enhanced modalities such as 3D point clouds, depth information, or tactile / force perception. Instead of relying solely on the semantic information of 2D images, it extracts the 3D structural features of the environment through a point cloud encoder and aligns them with the joint angles and end-effector pose of the robotic arm in a unified three-dimensional coordinate system. This approach aims to provide the robot with a "spatial sense," thereby reducing the geometric deviation between visual understanding and physical execution. (3) World Model-based VLA: It aims to assist decision-making by constructing a virtual physical simulation environment. It not only predicts actions, but also performs self-supervised learning by predicting the visual changes in the next frame after a given action (evolution of the environment state). It attempts to enable robots to understand the causal laws of the physical world and the impact of their own actions on the environment.
[0026] However, the above-mentioned solution failed to explicitly address the alignment issue between ontology perception and task execution in the initial design phase, resulting in the following drawbacks: (1) End-to-end unified modeling: Ontology difference trap - This path implicitly "hard-codes" ontology characteristics into network parameters. Since it cannot distinguish between "action changes caused by the environment" and "action changes caused by structure", it makes it difficult for the model to migrate between different hardware platforms; Training conflict - When training with a mixture of multiple data, a single end-to-end model will receive different action outputs (corresponding to robotic arms with different performance) in the same visual scene, which will cause the model prediction to be averaged or conflict, resulting in "cognitive confusion" and failing to accurately adapt to the physical characteristics of the current entity; (2) Multimodal perception path: Dimensionality curse and sample sparsity - After introducing 3D point cloud and high-frequency proprioception (such as six-axis torque), the dimension of the input space increases exponentially. In general tasks, the massive combination space leads to extremely sparse samples. In order to learn the mapping between point cloud and action, several times more labeled data is needed than pure vision solution, resulting in extremely high training cost and difficulty in convergence; representation drift and modality dominance problem, that is, due to the huge difference in the dimensions and sampling frequency of different sensors, for example, vision is 30Hz, point cloud is 10Hz, and proprioceptive force feedback is 1kHz, the model is prone to "representation drift" during training, that is, the model often overfits to the visual or point cloud features with extremely high information density, while ignoring the small but decisive proprioceptive modality, such as the small compliance adjustment of joints, etc., resulting in the inaccurate action execution in fine operation. (3) Predictive modeling based on world model: The model technology is not yet mature. In complex and unstructured scenarios, the visual state predicted by the model in the next frame often appears as an illusion, and it cannot accurately simulate the dynamic process of interaction between the real ontology and the environment. The difficulty of multimodal alignment is that the visual modality (low frequency, high latency) and the ontology modality (high frequency, low latency) are difficult to perfectly align in terms of timestamps. The model often overfits to strong visual features and ignores the small but key changes in ontology torque, resulting in a deviation between its understanding of "physical laws" and the actual physical performance, making it difficult to support high-precision operations.
[0027] To address the aforementioned shortcomings, this invention provides an action generation method based on ontological self-thinking and discrimination mechanisms. Inspired by the concept of "brain planning and cerebellar fine-tuning" in biological neuroscience, an improved hierarchical cerebellar-cerebellar collaborative decision-making architecture is constructed. Physical common sense constraints for high-level tasks are planned on the brain side, while ontological perception and introspective self-thinking mechanisms are injected on the cerebellar side. Through cerebellar-cerebellar communication, precise alignment between action planning and mechanical ontology is achieved.
[0028] According to an embodiment of the present invention, an action generation method based on an ontology self-thinking and discrimination mechanism is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] This embodiment provides an action generation method based on ontology self-thinking and discrimination mechanism. Figure 1 This is a flowchart of an action generation method based on ontology self-thinking and discrimination mechanism according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Use the task planning module to break down the received task instructions into short subtasks.
[0030] In this embodiment of the invention, a task planning module ("brain") with physical motion cognition is constructed. The task planning module performs environmental perception, receives task instructions, and decomposes the task instructions into a series of logically coherent and physically consistent short sub-tasks. These sub-tasks include not only semantic information, but also expected spatial state transitions. These sub-tasks serve as contextual conditions and are passed to the lower-level action generation module.
[0031] Step S102: Receive short-range subtasks using the action generation module, integrate the mechanical body state, action execution history, and environmental changes, perform state thinking and action output, and obtain the body cognitive thinking and generated actions.
[0032] In this embodiment of the invention, a deep network architecture of "perceptual fusion-dual-head output" is adopted to construct an action generation module ("cerebellum") with multimodal self-thinking ability and ontological perception function. The action generation module integrates visual observation, language commands, and the memory of the robot arm's historical motion sequence. The action generation module receives short-range sub-tasks, maps the feature vectors of the three modalities of the robot's ontological state, action execution history, and environmental changes to a unified latent space, performs feature fusion, performs state thinking to obtain ontological cognitive thinking, and outputs action to obtain the generated action.
[0033] Step S103: Use the action generation module to perform task feasibility analysis and action rationality analysis on ontology cognition thinking and action generation.
[0034] In this embodiment of the invention, an explicit bidirectional communication and alignment mechanism is established between the task planning module ("brain") and the action generation module ("cerebellum"). Specific trigger symbols are introduced. To address the blindness of action generation, a unique dual-discrimination logic is introduced into the action generation module, simulating the pre-rehearsal and error-correction functions of the biological cerebellum. Task feasibility analysis and action rationality analysis are performed on ontological cognitive thinking and action generation, achieving a leap from traditional open-loop execution to closed-loop control of perception-discrimination-feedback-replanning.
[0035] Specifically, bidirectional communication is achieved through downlink instruction alignment and uplink feedback replanning. Downlink instruction alignment refers to the task planning module ("brain") sending low-level subtask instructions based on high-level semantic decomposition to the action generation module ("cerebellum"). Uplink feedback replanning refers to the action generation module ("cerebellum") receiving instructions and then combining the environment, ontological state, and action execution memory encoding features to simultaneously predict actions and cognitively describe ontological state.
[0036] Step S104: If the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, then output the execution action to the robot for execution.
[0037] In this embodiment of the invention, when the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, and both the task feasibility judgment and the action rationality judgment are passed, the execution action is output to the robot for execution.
[0038] The action generation method based on ontology self-thinking and discrimination mechanism provided in this embodiment decomposes task instructions through the task planning module, reducing abstract task instructions to short-term sub-tasks, avoiding the gap caused by directly converting abstract instructions into action paths. The action generation module integrates multimodal information, performs state thinking and action output, and uses dual discrimination of task feasibility analysis and action rationality analysis to ensure the feasibility of the task and the rationality of the action, achieving accurate conversion from abstract instructions to physically realizable paths.
[0039] This embodiment provides an action generation method based on ontology self-thinking and discrimination mechanism, the process of which includes the following steps: Step S201: Use the task planning module to break down the received task instructions into short subtasks.
[0040] Specifically, the task planning module is constructed as follows: Step S2011: Obtain human action videos and robot work videos, and use video understanding models to perform action descriptions, action logic, dynamic principles, and geometric relationship annotations on the human action videos and robot work videos to construct a dataset; In step S2012, the pre-trained multimodal large model is fine-tuned and reinforced on the dataset, and the essential knowledge of motion is injected to obtain the task planning module.
[0041] In this embodiment of the invention, to overcome the limitations of traditional multimodal large models that only focus on semantic logic while ignoring physical rules and motion feasibility, a high-level planning layer is constructed. For example... Figure 2 As shown, based on the pre-trained Visual Large Model (VLM), a large number of human motion videos and robot work videos are collected. Human motion frame data and robotic arm motion frame data are extracted respectively. Existing video understanding models are used to annotate human motion videos and robot work videos with motion descriptions, motion logic, dynamic principles, geometric relationships, etc. Then, the annotation content is corrected and supplemented through manual quality inspection, thereby constructing a video frame-text description dataset.
[0042] A pre-trained multimodal large model is fine-tuned and reinforced using a video frame-text description dataset to inject motion-based knowledge, resulting in a task planning module. This motion-based knowledge includes kinematic constraints, dynamic principles, and spatial geometric relationships. This approach enables the task planning module to not only perform semantic task decomposition upon receiving visual environment information and natural language instructions, but also to make preliminary judgments on the physical logic rationality.
[0043] By using videos of human actions and robots at work as samples, training materials are provided for the model to learn the laws of physical motion. This allows for fine-tuning and reinforcement learning of the multimodal large model, injecting essential knowledge of motion and breaking through the limitations of pure semantic understanding.
[0044] Step S202: Receive short-range subtasks using the action generation module, integrate the mechanical body state, action execution history, and environmental changes, perform state thinking and action output, and obtain the body cognitive thinking and generated actions.
[0045] Specifically, the action generation module includes a temporal ontology-aware feature fusion module, a Transformer-based state-based language output module, and a Diffusion-based action generation module. Step S202 includes: Step S2021: The temporal ontology perception feature fusion module uses a multimodal attention mechanism to uniformly map the mechanical ontology state, action execution history, and environmental changes to obtain fused features. Step S2022: Using the Transformer-based state thinking language output module, with fused features and action features based on different inputs of the discrimination task as input, the module outputs a natural language description that judges the current state using a self-thinking mechanism. Step S2023: Using the motion generation module based on the Diffusion model, with fused features and language description as conditions, noise is gradually removed from Gaussian noise to generate motion trajectory sequences.
[0046] In this embodiment of the invention, a temporal ontology perception feature fusion module is used to receive high-frequency ontology state information (joint angle, torque, end effector speed) from robot joints and sensors in real time. The robot arm's historical action sequence is memorized, and the feature vectors of the three modalities of the robot ontology state, action execution history, and environmental changes are mapped to a unified latent space through a multimodal attention mechanism to ensure that the action generation is based on the accurate perception of the current physical state.
[0047] The Transformer-based state-of-the-art language output module is a decoder based on the Transformer architecture. Its inputs are fused features and action plans or sequences based on different inputs for the judgment task. The output is a natural language description of the current state. This module simulates human introspection, explicitly generating language analysis of the environmental state, ontological constraints, and the current action execution status. For example, it might state that the gripper is fully loaded or the target object is too far away to be reached. This output not only satisfies the rationality judgment of action generation but also serves as a direct basis for task feasibility judgment.
[0048] The action generation module based on the Diffusion model serves as an action strategy generator. It leverages the powerful distribution fitting capability of the Diffusion model, using the language description output by the fusion feature and state thinking language output module as conditions, to gradually remove noise from Gaussian noise and generate diverse, high-fidelity action trajectory sequences. The introduction of the diffusion model enables the cerebellum to process multimodal action distributions and effectively cope with environmental uncertainties.
[0049] By using a temporal ontology perception feature fusion module to fuse mechanical ontology state, action execution history, and environmental changes, the differences between different modal dimensions are eliminated. A Transformer-based state thinking language output module outputs natural language descriptions, transforming physical states into interpretable semantics. A Diffusion-based action generation module generates motion trajectory sequences and outputs actions that conform to physical laws.
[0050] In some alternative implementations, the method further includes: Step S2021a: Input the fused feature encoding from the previous time step together with the visual features, ontology features, and instruction features from the current time step to perform feature fusion.
[0051] In this embodiment of the invention, a recursive hidden state embedding mechanism is introduced, which inputs the fusion feature encoding of the previous time step together with the visual features, ontological features, and instruction features of the current time step for feature fusion, thereby embedding the robotic arm's record of past states, ensuring that the model has continuous memory of the robotic arm's action history, thereby helping the state thinking language output module to output a more reasonable description.
[0052] Step S203: Use the action generation module to perform task feasibility analysis and action rationality analysis on ontology cognition thinking and action generation.
[0053] Specifically, step S203 includes: Step S2031: Perform joint operations on action planning, fused features, and trained parameters, use the Sigmoid function for binary discrimination, and output the discrimination symbol; Step S2032: If the discrimination symbol is the first symbol, then the task is determined to be infeasible; Step S2033: If the discrimination symbol is the second symbol, then the task is deemed feasible; Step S2034: If the discrimination symbol is the third symbol, then the action is deemed reasonable; Step S2035: If the discrimination symbol is the fourth symbol, then the action is deemed unreasonable.
[0054] In this embodiment of the invention, in order to solve the problem of blind action generation, such as Figure 3 , Figure 4 As shown, a unique two-step discrimination logic is introduced into the action generation module ("cerebellum") to simulate the pre-playing and error correction functions of the biological cerebellum.
[0055] First, task feasibility is assessed: the generated actions from the task planning module and the output from the state-based thinking language output module are jointly evaluated. This is then combined with the trained parameters from the state-based thinking language output module to achieve an understanding of the task, environment, and ontology state, enabling joint analysis of task and ontology characteristics. The output is processed through a Sigmoid function, which maps the output value to a probability value between 0 and 1. A binary classification is performed with a threshold of 0.5: values greater than 0.5 are considered feasible, and values less than or equal to 0.5 are considered infeasible. When the task is feasible, the first discrimination symbol "<|CTN|>" is output, and the action generation module proceeds to the next stage of evaluation. When the task is infeasible, the second discrimination symbol "<|UCTN|>" is output. Simultaneously, analysis is performed based on the current state, and an infeasibility consideration is output to help the system determine whether the inability to execute the task is due to an error in action planning or specific ontology characteristics.
[0056] Then, the rationality of the action is judged. After determining the task's feasibility, before generating the action sequence, a judgment process consistent with the task feasibility judgment is executed. The action output from the action generation module is additionally received for action rationality analysis. The generated action sequence and proprioceptive features are input into the state-thinking language output module for information analysis, verifying the smoothness and rationality of the action trajectory. Combining action training analysis with the rationality of the action, the rationality judgment result is also output through the Sigmoid function. The Sigmoid function maps the output value to a probability value between (0, 1), using 0.5 as a threshold for binary classification: values greater than 0.5 are considered rational, and values less than or equal to 0.5 are considered unreasonable. When the action is rational, "|" is output. <ra>The third discrimination symbol, upon receiving this signal, outputs the action sequence to the robotic arm for execution. If the action is unreasonable, it outputs "| <fa>The fourth discrimination symbol prompts the action generation module to regenerate the action.
[0057] By using the Sigmoid function for binary discrimination, the feasibility of a task and the rationality of an action are judged separately, avoiding invalid task execution and risky actions, thereby improving the reliability of decision-making.
[0058] In some alternative implementations, the method further includes: Step Sa involves using visual data from human operation videos and robotic arm demonstration videos to learn motion logic knowledge. Step Sb: Use macroscopic physical phenomenon data to enhance the model's understanding of physical motion; Step Sc involves cloning basic movements using data from the robotic arm's operation.
[0059] In this embodiment of the invention, a phased hybrid training strategy is adopted, using supervised fine-tuning (SFT) to jointly train the model. Human operation videos and demonstration videos of successful task execution by a robotic arm are used as visual data to learn motion logic knowledge from the model. This allows the model to learn the correct order of actions in specific visual scenarios through numerous successful operation demonstrations, thereby establishing a mapping from visual input to action semantics. In this way, the model masters the temporal dependencies and causal logic between actions, ensuring that subsequent task execution follows the steps consistent with the operational attempts, avoiding erroneous behaviors that violate motion logic.
[0060] By leveraging macroscopic physical phenomena data to enhance the model's understanding of physical motion, video clips of fundamental physical laws and their corresponding textual descriptions are used as training samples. This allows the model to learn to infer the implicit physical properties of objects or environments from visual observations and predict the potential physical effects of actions. Through this reinforcement, the model consciously avoids actions that violate physical laws and chooses safe actions that conform to physical expectations, thereby improving the success rate of operations in real physical environments.
[0061] Basic motion cloning is performed using robotic arm operation data. Visual detection and corresponding motion trajectories of successful task executions are used as samples, and the model is trained using imitation learning to master basic motion distributions. Through motion cloning, the model learns a series of reasonable motion pattern machine probability distributions, enabling it to sample actions adapted to the current scene from the learned motion distributions when faced with minor environmental changes.
[0062] By using visual data from human-operated videos and robotic arm demonstration videos, the model learns motion logic, laying the foundation for general motion understanding, strengthening the model's cognition of physical motion, and enhancing the model's understanding of physical laws. Through basic motion cloning, it can quickly master basic motion capabilities, taking into account versatility, physicality, and proprioceptive adaptability.
[0063] In some alternative implementations, the method further includes: Step Sd involves constructing a reward function based on task success rate and physical compliance, and optimizing the discriminator accuracy of the action generation module using a reinforcement learning algorithm.
[0064] In this embodiment of the invention, a phased hybrid training strategy is adopted. Reinforcement Learning (RL) is used to optimize the model. For the self-thinking mechanism, a reward function based on task success rate and physical compliance is constructed. The accuracy of the discriminator is optimized through reinforcement learning algorithms to more accurately predict the consequences of actions and provide correct feedback signals, thereby enhancing robustness in unstructured environments. During the reinforcement learning training phase, the action generation module performs actions in a simulated or real environment, receiving rewards or penalties based on the execution results. If the action successfully completes the task, a positive reward is given; if the action violates physical laws, a negative penalty is given. The reward function is set as a weighted combination of task success rate and physical compliance.
[0065] By continuously interacting and learning from the environment, the reinforcement learning algorithm gradually adjusts the parameters of the state thinking language output module within the action generation module, thereby improving the accuracy of its output natural language descriptions and discrimination symbols, so as to accurately predict the consequences of actions before they are executed.
[0066] By focusing on task success rate, enhancing the effectiveness of the judgment results, introducing physical compliance constraints, improving the accuracy of the discriminator in the action generation module, and improving the reliability of dual judgment.
[0067] Step S204: If the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, then output the execution action to the robot for execution.
[0068] Please see details Figure 1 Step S104 of the illustrated embodiment will not be described again here.
[0069] The action generation method based on ontology self-thinking and discrimination mechanisms provided in this embodiment achieves a closed-loop control of "perception-discrimination-feedback-replanning" between the brain and cerebellum through self-thinking and discrimination mechanisms. Through the introspection mechanism of the cerebellum thinking module, complex physical ontology states are explicitly expressed as semantic features. This enables the model to distinguish between environmental constraints and ontology constraints. When training with mixed multi-machine data, the model no longer experiences conflicts related to predictive meanization, but instead outputs the optimal action adapted to the hardware based on its current ontology self-cognition.
[0070] By introducing a two-step decision-making logic, when instructions are issued, the cerebellum filters the brain's instructions through task feasibility judgment. If the task exceeds the physical boundaries of the entity, an upward feedback replanning is triggered. Before execution, trajectory risks are predicted through action rationality judgment. This closed loop of thinking first, then verifying, and then executing greatly reduces the execution failure rate caused by singularities and overload.
[0071] This embodiment also provides an action generation system based on an ontology self-thinking and discrimination mechanism. This system is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0072] This embodiment provides an action generation system based on ontology self-thinking and discrimination mechanisms, such as Figure 5 As shown, it includes: Task planning module 501 is used to break down received task instructions into short-term subtasks; The motion generation module 502 is used to receive short-range sub-tasks, integrate the mechanical body state, motion execution history, and environmental changes, perform state thinking and motion output, and obtain the body cognitive thinking and generated motion; perform task feasibility analysis and motion rationality analysis on the body cognitive thinking and generated motion; if the task feasibility analysis result is feasible and the motion rationality analysis result is reasonable, then output the execution motion to the robot for execution.
[0073] In some alternative implementations, the task planning module is constructed as follows: Acquire videos of human actions and robot work, and use video understanding models to annotate the human action videos, robot work videos, and robot work videos with action descriptions, action logic, dynamic principles, and geometric relationships to build a dataset; Fine-tuning and reinforcement learning are performed on a pre-trained multimodal large model on the dataset, and motion-based knowledge is injected to obtain a task planning module. The motion-based knowledge includes kinematic constraints, dynamic principles, and spatial geometric relationships.
[0074] In some alternative implementations, the action generation module 502 includes: The temporal ontology perception feature fusion module is used to uniformly map the mechanical ontology state, action execution history, and environmental changes through a multimodal attention mechanism to obtain fused features; The Transformer-based state thinking language output module is used to take fused features and action features based on different inputs of the discrimination task as inputs, and use a self-thinking mechanism to output a natural language description that judges the current state. The action generation module based on the Diffusion model is used to generate action trajectory sequences by gradually removing noise from Gaussian noise, based on fused features and language descriptions.
[0075] In some alternative implementations, the system further includes: The feature fusion unit is used to introduce a recursive hidden state embedding mechanism, which inputs the fused feature encoding from the previous time step with the visual features, ontology features, and instruction features from the current time step for feature fusion.
[0076] In some alternative implementations, the action generation module 502 includes: The output unit is used to perform joint operations on the generated action, fused features, and trained parameters, and then use the Sigmoid function to perform binary discrimination and output the discrimination symbol. The first decision unit is used to determine that the task is not feasible if the discrimination symbol is the first symbol. The second determination unit is used to determine whether the task is feasible if the discrimination symbol is the second symbol. The third determination unit is used to determine whether the action is reasonable if the discrimination symbol is the third symbol. The fourth determination unit is used to determine that the action is unreasonable if the determination symbol is the fourth symbol.
[0077] In some alternative implementations, the system further includes: The motion logic knowledge learning module is used to learn motion logic knowledge using visual data from human operation videos and robotic arm demonstration videos. The enhancement module is used to enhance the model's understanding of physical motion by utilizing data from macroscopic physical phenomena; The motion cloning module is used to clone basic motions using robotic arm operation data.
[0078] In some alternative implementations, the system further includes: The optimization module is used to construct a reward function based on task success rate and physical compliance, and to optimize the discriminator accuracy of the action generation module through reinforcement learning algorithms.
[0079] The action generation system based on ontology self-thinking and discrimination mechanism provided in this embodiment of the invention can execute the action generation method based on ontology self-thinking and discrimination mechanism provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method. Further functional descriptions of the above modules and units are the same as in the corresponding embodiments described above, and will not be repeated here.
[0080] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.
[0081] The following is a detailed reference. Figure 6 This diagram illustrates a suitable structural design for implementing an electronic device according to embodiments of the present invention. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) 602 or a program loaded from memory 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of the electronic device. The processor 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0082] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.
[0083] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a memory 608, or installed from a ROM 602. When the computer program is executed by the processor 601, it performs the functions defined in the action generation method based on ontology self-thinking and discrimination mechanism of the embodiments of the present invention.
[0084] Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0085] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the action generation method based on ontology self-thinking and discrimination mechanisms shown in the above embodiments is implemented.
[0086] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0087] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and all such modifications and variations fall within the scope defined by the appended invention.< / fa> < / ra>
Claims
1. An action generation method based on ontology self-thinking and discrimination mechanism, characterized in that, An action generation system based on ontology self-thinking and discrimination mechanisms is applied, the system including a task planning module and an action generation module, the method including: The received task instructions are broken down into short subtasks using the task planning module; The action generation module receives the short-range sub-tasks, integrates the mechanical body state, action execution history, and environmental changes, performs state thinking and action output, and obtains the body cognitive thinking and generated actions. The action generation module is used to conduct task feasibility analysis and action rationality analysis based on ontology cognition and action generation. If the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, then the action to be executed is output to the robot.
2. The method according to claim 1, characterized in that, Construct the task planning module as follows: Acquire videos of human actions and robot work, and use video understanding models to annotate the human action videos, robot work videos, and robot work videos with action descriptions, action logic, dynamic principles, and geometric relationships to build a dataset; The pre-trained multimodal large model is fine-tuned and reinforced on the dataset, and the essential knowledge of motion is injected to obtain the task planning module. The essential knowledge of motion includes kinematic constraints, dynamic principles, and spatial geometric relationships.
3. The method according to claim 1, characterized in that, The action generation module includes a temporal ontology-aware feature fusion module, a Transformer-based state-based language output module, and a Diffusion-based action generation module. The action generation module receives the short-range subtask, fuses the mechanical ontology state, action execution history, and environmental changes, performs state-based thinking and action output, and obtains ontology-based cognitive thinking and generated actions, including: The temporal ontology perception feature fusion module utilizes a multimodal attention mechanism to uniformly map the mechanical ontology state, action execution history, and environmental changes to obtain fused features. The Transformer-based state-thinking language output module takes fused features and action features based on different inputs of the discrimination task as input, and uses a self-thinking mechanism to output a natural language description that judges the current state. The motion generation module based on the Diffusion model is used to generate motion trajectory sequences by gradually removing noise from Gaussian noise, using fused features and language description as conditions.
4. The method according to claim 3, characterized in that, After using the temporal ontology perception feature fusion module to uniformly map the mechanical ontology state, action execution history, and environmental changes through a multimodal attention mechanism to obtain fused features, the method further includes: A recursive hidden state embedding mechanism is introduced, which inputs the fusion feature encoding from the previous time step with the visual features, ontology features, and instruction features from the current time step for feature fusion.
5. The method according to claim 3, characterized in that, The aforementioned analysis of the ontological cognition and the generation of actions, including task feasibility analysis and action rationality analysis, includes: The generated action, fused features, and trained parameters are jointly processed, and the sigmoid function is used for binary discrimination to output the discrimination symbol. If the discrimination symbol is the first symbol, then the task is determined to be infeasible; If the discrimination symbol is the second symbol, then the task is deemed feasible; If the discrimination symbol is the third symbol, then the action is deemed reasonable; If the discrimination symbol is the fourth symbol, then the action is deemed unreasonable.
6. The method according to claim 1, characterized in that, The method further includes: Learning motion logic knowledge using visual data from human operation videos and robotic arm demonstration videos; Utilize macroscopic physical phenomenon data to enhance the model's understanding of physical motion; Basic motion cloning is performed using robotic arm operation data.
7. The method according to claim 1, characterized in that, The method further includes: A reward function based on task success rate and physical compliance is constructed, and the discriminator accuracy of the action generation module is optimized through reinforcement learning algorithm.
8. An action generation system based on ontology self-thinking and discrimination mechanism, characterized in that, The system includes: The task planning module is used to break down received task instructions into short subtasks; The action generation module is used to receive the short-range sub-task, integrate the mechanical body state, action execution history, and environmental changes, perform state thinking and action output to obtain the body cognitive thinking and generated action; perform task feasibility analysis and action rationality analysis on the body cognitive thinking and generated action; if the task feasibility analysis result is feasible and the action rationality analysis result is reasonable, then output the action to be executed by the robot.
9. An electronic device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the computer instructions to perform the action generation method based on ontology self-thinking and discrimination mechanism as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the action generation method based on ontology self-thinking and discrimination mechanism as described in any one of claims 1 to 7.