Robot, data production method and device thereof, training method, medium and product

By collaborating with multi-layered visual language models and utilizing action constraints and result verification, a generation and structured processing mechanism was established to solve the problem of efficiently and cost-effectively generating high-quality robot thought chain data, thereby improving logical consistency and interpretability.

CN121552332APending Publication Date: 2026-02-24AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511501470.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality robot thought chain data efficiently and cost-effectively. Manual annotation is costly and inefficient, and the quality of data generated by the model itself is uncontrollable, leading to logical fallacies and deficiencies.

Method used

A multi-layer visual language model collaboration mechanism is adopted. The first visual language model generates an initial thought chain, which is verified by action constraint information and results. The second visual language model performs structured processing to ensure data logic consistency and interpretability.

Benefits of technology

It enables the low-cost and high-efficiency generation of high-quality robot thought chain data, improves the logicality, interpretability and scalability of the data, ensures the high quality and consistency of training data, and reduces the reliance on manual quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121552332A_ABST
    Figure CN121552332A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a robot, a data production method and device thereof, a training method, a medium and a product. The method comprises the following steps: receiving image data, action constraint information and reference action decision information corresponding to a robot task scene; inferring at least one action of the robot by using a first visual language model based on the image data and the action constraint information to generate first thinking process information and first action decision information; when the first action decision information is determined to be correct, at least one action of the robot is inferred using a second visual language model based on the image data, the first thinking process information, and the reference action decision information to generate second thinking process information. According to the embodiment of the invention, automatic production from a visual scene to a high-quality thinking chain is realized with low cost and high efficiency through multi-layer cooperation of the visual language model, a restriction mechanism of action constraint input and quality screening of a result verification type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot data annotation technology, and in particular to robots and their data production methods, devices, training methods, media, and products. Background Technology

[0002] With the rapid development of artificial intelligence and robotics, robots possessing autonomous cognition and reasoning abilities have gradually become a research hotspot. Among them, "Chain-of-Thought" (CoT) technology is considered an important means to improve robots' reasoning ability, task interpretability, and decision-making transparency. By enabling robots to generate a series of logical reasoning steps before performing a task, CoT technology allows robots to think like humans, thereby improving their accuracy and robustness in complex environments. However, training robot models with high-quality CoT capabilities requires obtaining large-scale, high-quality Chain-of-Thought datasets.

[0003] Based on this, embodiments of this application provide robots and their data production methods, apparatuses, training methods, media, and products to improve related technologies. Summary of the Invention

[0004] The purpose of this application is to provide a robot and its operating method, device, and storage medium, which realizes automated production from visual scenes to high-quality thought chains in a low-cost and high-efficiency manner.

[0005] The objective of this application embodiment is achieved using the following technical solutions: In a first aspect, embodiments of this application provide a robot data production method, the method comprising: receiving image data, motion constraint information, and reference motion decision information corresponding to a robot task scenario; based on the image data and the motion constraint information, using a first visual language model to infer at least one action of the robot to generate first thought process information and first motion decision information; and, if the first motion decision information is determined to be correct, using a second visual language model to output structured second thought process information based on the image data, the first thought process information, and the reference motion decision information.

[0006] In some embodiments, the action constraint information includes a set of executable actions; and / or, the second visual language model is configured to use the reference action decision information as the final action decision information.

[0007] In some embodiments, the step of reasoning about at least one action of the robot using a first visual language model based on the image data and the action constraint information to generate first thought process information and first action decision information includes: reasoning about at least one action of the robot using the first visual language model based on the image data and a first prompt text to generate the first thought process information and the first action decision information; wherein the first prompt text is constructed based on the action constraint information and is used to guide the first visual language model to generate the first thought process information and the first action decision information.

[0008] In some embodiments, the process of determining whether the first action decision information is correct includes: matching the first action decision information and the reference action decision information based on a specified matching algorithm; if the matching is successful, determining that the first action decision information is correct; or, if the matching fails, determining that the first action decision information is incorrect.

[0009] In some embodiments, the step of outputting structured second thinking process information using a second visual language model based on the image data, the first thinking process information, and the reference action decision information includes: outputting structured second thinking process information using a second visual language model based on the image data and a second prompt text; wherein the second prompt text is constructed based on the first thinking process information, the reference action decision information, and output template information, and is used to guide the second visual language model to generate the second thinking process information by using the reference action decision information as the final action decision information, referring to the logic of the first thinking process information, and according to the specified structure of the output template information.

[0010] In some embodiments, the output template information includes a first guidance prompt text corresponding to the following information: scene analysis information, analysis of the current task status information, information on the next action to be taken, and information on reflection and correction actions.

[0011] In some embodiments, the output template information further includes a second guiding prompt text corresponding to the reference action decision information; the second thinking process information is represented by a first tag, the reference action decision information is represented by a second tag, and the first tag includes separately set scenario analysis information, analysis of the current task status information, information on the next action to be taken, and information on reflection and correction actions.

[0012] Secondly, embodiments of this application provide a robot data production apparatus, the apparatus comprising: a data receiving module, configured to receive image data, motion constraint information, and reference motion decision information corresponding to a robot task scenario; a first reasoning module, configured to reason about at least one action of the robot based on the image data and the motion constraint information using a first visual language model, to generate first thought process information and first motion decision information; and a second reasoning module, configured to output structured second thought process information based on the image data, the first thought process information, and the reference motion decision information, using a second visual language model, if the first motion decision information is deemed correct.

[0013] Thirdly, embodiments of this application provide a robot model training method, the method comprising: receiving image data, motion constraint information, and reference motion decision information corresponding to a robot task scenario; based on the image data and the motion constraint information, using a first visual language model to infer at least one action of the robot to generate first thought process information and first motion decision information; if the first motion decision information is determined to be correct, using a second visual language model to output structured second thought process information based on the image data, the first thought process information, and the reference motion decision information; and training a specified robot model based on the second thought process information and the reference motion decision information to update at least one model parameter of the robot model.

[0014] Fourthly, embodiments of this application provide a robot, which stores a robot model, and the robot model is trained using the above-described training method.

[0015] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above methods.

[0016] Sixthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above methods.

[0017] This application provides a robot and its data production method, apparatus, training method, medium, and product. The method includes: receiving image data, motion constraint information, and reference motion decision information corresponding to a robot task scenario; based on the image data and motion constraint information, using a first visual language model to infer at least one action of the robot to generate first thought process information and first motion decision information; and, if the first motion decision information is determined to be correct, using a second visual language model to output structured second thought process information based on the image data, the first thought process information, and the reference motion decision information.

[0018] This application's embodiments achieve low-cost, high-efficiency automated production of high-quality thought chains from visual scenes through multi-layered collaboration of visual language models, a constraint mechanism for action-constrained input, and result-verification-based quality screening. This solution not only improves upon the high cost of manual annotation and the uncontrollable quality of model-generated data, but also ensures that the generated data possesses high logic, interpretability, and scalability, providing high-quality foundational data for robot cognitive and reasoning training. The aforementioned embodiments also verify the correctness of the first action decision information generated by the first visual language model to infer and verify the rationality of its thought process. This result-oriented verification method cleverly transforms a complex and ambiguous process evaluation problem into a simple and clear result matching problem, thereby achieving automated and efficient screening of thought chain data quality. Furthermore, only high-quality thought chains corresponding to verified model answers can enter the final processing stage to generate second thought process information. This model ensures that every piece of data ultimately produced by the pipeline has been "tested," possessing high logic and reliability, and fundamentally guarantees the overall high quality level of the robot model training dataset, effectively suppressing data pollution caused by model "illusions" and logical errors. Attached Figure Description

[0019] The embodiments of this application are further described below with reference to the accompanying drawings and specific implementation details.

[0020] Figure 1 This is a flowchart illustrating a robot data production method provided in an embodiment of this application.

[0021] Figure 2 This is a schematic diagram of the structure of a robot data production device provided in an embodiment of this application.

[0022] Figure 3 This is a schematic diagram of a robot data production process provided in an embodiment of this application.

[0023] Figure 4 This is an image data provided in the embodiments of this application.

[0024] Figure 5 This is a flowchart illustrating a robot model training method provided in an embodiment of this application.

[0025] Figure 6 This is a structural block diagram of a robot provided in an embodiment of this application.

[0026] Figure 7 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the embodiments of this application.

[0028] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0029] Chain-of-Thought (CoT) technology empowers robots to understand abstract human instructions and autonomously plan their actions. The core idea of ​​CoT is to allow the model to generate a series of intermediate, logically coherent reasoning steps before outputting the final answer. When a robot applies CoT, it is no longer a "black box" that directly outputs actions after receiving instructions, but can analyze, reason, and plan like a human. For example, faced with the instruction "make a cup of coffee," a robot with CoT capabilities will first generate a thought process internally: "1. Check the coffee machine status to ensure there is water and coffee beans; 2. Find a clean cup; 3. Place the cup under the coffee machine outlet; 4. Start the coffee machine; 5. Wait for the coffee to be made." This ability not only greatly improves the success rate and robustness of robots in performing complex tasks, but more importantly, its decision-making process becomes transparent and explainable, facilitating human supervision and adjustment.

[0030] However, a major obstacle to training robot models with high-quality CoT capabilities is the scarcity of high-quality training data. Currently, there are two main ways to obtain robot CoT data.

[0031] The first approach is manual coding and annotation. This is currently the gold standard for ensuring data quality. Human experts (such as robotics engineers and experienced users) manually write detailed thought processes, decision-making logic, and final actions for specific task scenarios. However, this method has significant drawbacks. First, it is costly. It requires a large investment of manpower with specialized knowledge, resulting in very high time and financial costs. Second, it is inefficient. The manual coding process is time-consuming and labor-intensive, making it unsuitable for training deep learning models that often require millions of samples. Furthermore, it suffers from poor consistency. Different annotators have varying experience, styles, and levels of focus, leading to difficulties in standardizing data quality and format, introducing noise into model training.

[0032] The second approach is simple model self-generation. This involves using existing large models to directly generate CoT data through simple prompts. For example, inputting an image containing the task scenario into the model and asking "What should I do?". While this method is fast and low-cost, the quality of the generated data often fails to meet the stringent requirements of robot training. The main problems include: First, logical fallacies and factual errors. The model may "imagine" steps that contradict common sense about the physical world or the robot's actual capabilities. Second, thought shortcuts. The model tends to generate the simplest and most direct thought processes, lacking consideration for anomalies and exploration of multiple possibilities, resulting in insufficient CoT depth. Furthermore, there is a lack of quality assurance. Verifying the rationality of the generated CoT process requires extensive manual review and quality control, which again leads to the old problems of high cost and low efficiency.

[0033] See Figure 1 , Figure 1 This is a flowchart illustrating a robot data production method provided in an embodiment of this application.

[0034] In order to automate and efficiently produce large-scale, high-quality robot thought chain data, this application provides a robot data production method, which includes steps S101 to S103.

[0035] Step S101: Receive image data, motion constraint information and reference motion decision information corresponding to the robot task scenario.

[0036] Step S102: Based on the image data and the action constraint information, use the first visual language model to infer at least one action of the robot to generate first thinking process information and first action decision information.

[0037] Step S103: If the first action decision information is determined to be correct, based on the image data, the first thought process information and the reference action decision information, a structured second thought process information is output using a second visual language model.

[0038] Visual-Language Models (VLMs) are a class of deep learning models with powerful visual understanding and language generation capabilities, capable of processing both visual information (images, video frames, etc.) and linguistic information (text, instructions, etc.) simultaneously.

[0039] In this embodiment, the first visual language model (VLM-1) can be used in the "generation" stage. Its function is to infer the robot's next or subsequent actions based on the image data and motion constraint information of the robot's task scene, and output the original thought chain text (i.e., the first thought process information) corresponding to the inference process. The second visual language model (VLM-2) can be used in the "processing and reinforcement" stage. Its function is to use the verified first thought process information as a reference to output the structured thought chain text (i.e., the second thought process information) corresponding to the inference process.

[0040] In some embodiments, the second visual language model may be configured to use the reference action decision information as the final action decision information.

[0041] Through the collaborative mechanism of two models (i.e., the first visual language model and the second visual language model), a hierarchical thinking modeling structure of "primary generation and advanced correction" is formed, which ensures that the data logic is consistent, the semantics are reasonable, and the training is usable.

[0042] The above embodiments do not limit the first visual language model and the second visual language model, which can be, for example, Gemini-2.5-Pro ​​or a model with equivalent capabilities.

[0043] Image data is used to characterize the task scenario in which the robot is located and serves as the perceptual input for model reasoning. This image data can be acquired by the robot's vision system, environmental cameras, or simulation systems. Its role is to provide visual cues such as the position, state, and relationships of objects in the environment, as well as to provide contextual semantics for the visual language model, supporting the model in generating action reasoning that conforms to physical common sense.

[0044] Action constraint information is used to limit the action inference space of a model and is a semantic or structured input information. By introducing action constraint information, it is possible to reduce the generation of illusory actions that exceed physical or task boundaries, thereby improving the executability and physical consistency of the generated thought chain.

[0045] In some embodiments, the action constraint information may include a set of executable actions. The set of executable actions is, for example, a predefined action knowledge base containing all possible correct actions for the scenario. The set of executable actions may include multiple SubTask Action Options (which can be considered as executable actions), representing several reasonable operations that the robot can perform in the current task scenario, such as "grabbing an object," "moving to a specified location," or "placing an object."

[0046] The reference action decision information serves as a reference for the correct answer, and can also be called the ground truth answer. The source of the reference action decision information is, for example, the aforementioned action knowledge base. In this embodiment, the reference action decision information is not directly provided to the first visual language model, but is only used in the verification stage to determine whether the first action decision information output by the first visual language model is correct, thus serving as a "result verification signal." For example, in a scenario of repairing a bicycle, the reference action decision information might be "pick up the wrench on the ground." Of course, this reference action decision information is not visible to the first visual language model at this step.

[0047] The first thought process information, such as inference text (e.g., unstructured text) autonomously generated by the first visual language model based on scene images and action constraint information, reflects the model's logical deduction and decision-making process for the current task. The first thought process information can be regarded as the model's "original thought chain" (original CoT), which is used for subsequent quality screening and enhanced generation.

[0048] The first action decision information is, for example, the final action conclusion derived by the first visual language model during the thinking process. That is, the model predicts the robot's next action, and its role is equivalent to the model's output "Model-Generated Answer." As an example, the correctness of this reasoning result can be verified by matching the first action decision information with reference action decision information. When the first action decision information is correct, the corresponding first thought process information can be considered logically sound, thus proceeding to the next stage of generation. When the first action decision information is incorrect, the corresponding first thought process information can be considered logically flawed, thus preventing it from proceeding to the next stage. Therefore, only first thought process information with correct results and logical soundness is used in the generation step of the second thought process information, resulting in high-quality and logically rigorous second thought process information.

[0049] The second thought process information, such as that generated by the second visual language model, can be seen as a structured, high-quality thought chain formed by further processing the first thought process information after it has been verified.

[0050] The control logic of the above method can be viewed as including the following three stages: Stage 1, Generation Stage. The first visual language model freely generates thought chains and predictive actions, obtaining first thought process information and first action decision information. Stage 2, Verification Stage. The correctness of the first action decision information is determined. Stage 3, Reinforcement Stage. For the first action decision information whose result is deemed correct, the second visual language model logically expands and rewrites its corresponding first thought process information, and outputs it in a structured manner. In this way, by using the correctness of the result as the criterion for judging the rationality of the process, automated quality control is achieved by inferring the rationality of thinking from the result-oriented approach. This mechanism reduces reliance on manual quality inspection, ensures data consistency and logical reliability, and significantly improves production efficiency.

[0051] This application proposes an innovative method for the automated generation and annotation of robot thought chain data. Through multi-layered collaboration of visual language models, a constraint mechanism for action-based input, and result-verification-based automatic quality screening, it achieves low-cost and high-efficiency automated production of high-quality thought chains from visual scenes. Since manual annotation of the first thought process information is unnecessary, labor costs are significantly reduced. Collaboration based on multi-layered visual language models rapidly generates high-quality second thought process information, improving the production efficiency of high-quality thought chain data. This solution not only addresses the problems of high manual annotation costs and uncontrollable quality of model-generated data, but also ensures that the generated data possesses high logic, interpretability, and scalability, providing high-quality foundational data for robot cognitive and reasoning training. Unlike data methods that rely on manual writing or single-model self-generation, this method automatically generates diverse thought chain data through multi-model collaboration, significantly improving the diversity and logical richness of samples, providing more generalizable foundational data for robot model training. This method does not limit reasoning steps and paths during data generation, allowing models to autonomously generate multiple thought processes with open-ended prompts, thus more closely resembling the reasoning logic of complex real-world tasks and effectively overcoming the limitations of traditional single-path generation methods.

[0052] The above embodiments also provide an automated CoT quality verification mechanism based on "result alignment". Specifically, the method designs an automated CoT quality inspection process that no longer attempts to directly evaluate the merits of the thought process itself. Instead, it verifies the correctness of the thought process by inferring and verifying the rationality of the thought process through the correctness of the "model answer" (i.e., the first action decision information) generated by the first visual language model. If the result is correct, the thought process is adopted. This "result-oriented" verification method cleverly transforms a complex and ambiguous process evaluation problem into a simple and clear result matching problem, thereby achieving automated and efficient screening of CoT data quality.

[0053] The above embodiments also construct a "Generate-and-Verify" data production closed loop. The first-stage model (i.e., the first visual language model) is responsible for freely "generating" the CoT (i.e., the first thought process information) and the model's answer (i.e., the first action decision information). The second stage is responsible for rigorously "verifying" the correctness of the model's answer. Only high-quality CoTs corresponding to verified model answers can enter the final processing stage to generate the second thought process information. This model ensures that every piece of data ultimately produced by the pipeline (e.g., including the second thought process information and reference action decision information) has been "tested" (i.e., automatically verified), possessing high logicality and reliability. This mechanism guarantees the overall high quality of the robot model training dataset and effectively suppresses data pollution caused by model "illusions" and logical errors.

[0054] The above embodiments do not limit the input method of action constraint information. In actual input, action constraint information can be input into the first visual language model in any of the following ways: First input method: explicit input. The action constraint information is input as an independent parameter to the first visual language model. Second input method: implicit embedding. The action constraint information is embedded in the Prompt text in natural language form, guiding the model to select from a specified set of actions.

[0055] To ensure that the model generates a reasoning process that conforms to both semantic logic and task constraints, allowing the model to generate completely freely may result in unrestrained imaginary actions; conversely, excessive restrictions would sacrifice scene diversity. Therefore, this embodiment of the application inputs the task scene image and the prompt text containing action constraint information (i.e., the first prompt text) into the first visual language model (VLM-1). When generating reasoning, the model retains open-ended thinking while being semantically guided by the action set, thus achieving a balance between freedom and constraints.

[0056] In some embodiments, the step of reasoning about at least one action of the robot using a first visual language model based on the image data and the action constraint information to generate first thought process information and first action decision information may include: reasoning about at least one action of the robot using the first visual language model based on the image data and a first prompt text to generate the first thought process information and the first action decision information; wherein the first prompt text is constructed based on the action constraint information and is used to guide the first visual language model to generate the first thought process information and the first action decision information.

[0057] The first prompt text serves as guiding language input for the first visual language model to generate thought chains and action decisions, controlling the model's reasoning direction. As an example, the first prompt text can be constructed based on action constraint information, such as using natural language templates, to stimulate the model to autonomously engage in "observation-analysis-decision" thinking without directly providing the correct answer. As another example, the first prompt text can include descriptions of the task objective, operational context, and set of executable actions, thus maintaining the flexibility of open-ended reasoning while ensuring the output conforms to task constraint logic during thought chain generation. For example, for a "cleaning the desktop" task, the first prompt text could be constructed as: "Please determine the current state of the robot's task scenario based on the following diagram, and consider the reasonable action to be performed next. Executable actions include: grabbing trash, moving to the recycling area, and wiping the desktop. Please explain your thought process and give your final decision." The first prompt text plays two roles in model input. First, semantic constraint: By including a set of executable actions in the prompt, the model is constrained by the task space during reasoning. Second, thought guidance: Through open-ended question expression, it guides the model to generate text with reasoning logic, rather than simple action classification results.

[0058] In other words, the above embodiments can construct a first prompt text based on action constraint information and input it along with image data into a first visual language model. As an example, the first visual language model first extracts visual features from the input image data to identify key objects and environmental layouts in the scene. Subsequently, the first visual language model uses a multimodal semantic fusion mechanism to jointly encode the extracted visual features and the semantic information of the first prompt text, establishing a logical connection between vision and language in a unified semantic space. Based on this fused representation, the first visual language model performs semantic reasoning, generating textual information including thought processes (i.e., first thought process information) and the specific task decision result corresponding to this thought process (i.e., first action decision information).

[0059] The advantages of this approach are threefold. First, by introducing initial prompt text based on action constraint information into the input of the first visual language model, the model's inference space is reasonably constrained, reducing "illusionary actions" that do not conform to physical rules or task objectives during free generation. This significantly improves the logical consistency and executability of thought chain generation. Second, leveraging the multimodal semantic fusion capability of the first visual language model enables deep semantic understanding of scene images. Third, guided by open-ended prompt text, the model can generate diverse thought paths, reducing single, templated outputs and enhancing the diversity and generalization ability of robot training data. This allows the model to better adapt to complex and varied real-world task environments when subsequently used for robot inference model training. Finally, the above embodiment achieves a balance between open generation and task constraints, preserving the inference capabilities of the large model while ensuring the correctness and reliability of the output results. This provides an efficient, controllable, and scalable technical approach for the automated production of robot thought chain data, laying a solid foundation for subsequent action verification and multi-layered logic reinforcement.

[0060] In the research on automated production of robot thought chain data, it is necessary to accurately determine the rationality of action decisions autonomously generated by the model without relying on manual review, thereby selecting high-quality thought chain samples. This is because, in actual testing, it has been found that even with high-performance visual language models, models are still prone to "illusionary actions" or logical deviations in complex task scenarios. For example, in a cleaning task, the model may generate illogical actions such as "putting the trash can on the table"; in an assembly task, the model may skip necessary steps and directly give incorrect conclusions. If manual quality inspection is still relied upon, it is not only inefficient but also difficult to achieve a closed loop in large-scale data production. Therefore, it is necessary to design an automated correctness judgment mechanism that can "self-examine" its own thought chain output like a human. To this end, the embodiments of this application do not directly evaluate the text quality of the model's thinking process, but instead start from the result and indirectly infer the rationality of the thought chain by judging whether the action decisions output by the model are correct. In practical applications, one approach is to match the first action decision information output by the first visual language model with reference action decision information. Another approach is to employ task-context-based group relative strategy optimization, evaluating the merits of each answer by assessing its relative performance among a set of answers. In some possible implementations, multi-dimensional verification signals can be incorporated, such as using action execution simulation modules to verify action feasibility, or combining multi-model voting mechanisms (e.g., consistency scoring between different VLMs) to improve decision robustness.

[0061] To achieve automated screening with high confidence at low computational cost, in some embodiments, the process of determining whether the first action decision information is correct may include: matching the first action decision information and the reference action decision information based on a specified matching algorithm; if the matching is successful, determining that the first action decision information is correct; or, if the matching fails, determining that the first action decision information is incorrect.

[0062] The above method provides an answer matching filtering and CoT quality automated verification approach, which functions as an automated "quality inspector." It compares the answer generated by the first visual language model (i.e., the first action decision information) with the preset real answer (i.e., the reference action decision information) to determine whether the thought chain (i.e., the first thinking process information) generated by the first visual language model is reasonable and worth retaining.

[0063] The above-mentioned matching algorithm is not limited; it can be a simple string matching algorithm or a semantic matching algorithm that considers synonyms and near-synonyms. If the match is successful, it means that the first visual language model, through its internal thought process (i.e., the "original CoT"), has ultimately arrived at a correct action decision. Based on the principle that "if the result is correct, the process is likely reasonable," this "original CoT" is judged to be high-quality and reliable. This CoT and its corresponding "true answer" will be approved and proceed to the next stage. If the match fails, it means that the first visual language model's thought process may have biases or errors, leading it to give an incorrect action decision. In practical applications, this "original CoT" and "true answer" can be discarded and not used for subsequent production. This step uses automated answer comparison to replace the costly and time-consuming manual CoT quality review process, which is conducive to producing high-quality data at low cost and high efficiency.

[0064] In the research and development of automated production of robot thought chain data, even though high-quality CoT samples can be selected from the first thought process information generated by the first visual language model (VLM-1), these samples generally suffer from problems such as inconsistent format, unclear logical hierarchy, and insufficient reasoning explanation. As a result, even if the first action decision information corresponding to the first thought process information is judged to be correct, the text structure of the first thought process information is still disorganized, making it unsuitable for subsequent use as standardized training samples. Furthermore, directly using this type of "unstructured" data for robot reasoning model training may cause the model to capture irrelevant noise, logical jumps, or non-standard expressions during the learning phase, thereby reducing training convergence efficiency and generalization performance. Therefore, it is necessary to automatically strengthen and structure the validated thought chains while maintaining the diversity of autonomous thinking in the model, thereby generating high-quality data that can be directly used for training.

[0065] To address this issue, this application employs a second visual language model (VLM-2) to "re-reason" and "structure rewrite" the validated thought process chain, thereby evolving from raw logic to normative logic. VLM-2 does not require re-determining actions; instead, it uses "reference action decision information" as an anchor point and follows the logical flow of the "first thought process information," guiding the model to generate a second thought process information that is logically more rigorous and structurally more regular through output templates. Thus, from VLM-1 to VLM-2, a hierarchical thought modeling mechanism of "generation-verification-reinforcement" is formed, with the former responsible for generating diversity and the latter responsible for uniformity and logical quality control.

[0066] This application does not limit the method by which the second visual language model generates second thought process information, and such methods may include, but are not limited to, the following: First, a template-driven approach. By setting a unified output template, the model is guided to generate second thought process information text that conforms to a standard logical framework (such as a four-segment structure of "scene analysis - task status - decision logic - reflection and correction"), ensuring the comparability and interpretability of different samples. By explicitly requiring the model to perform "self-check" or "reflection" operations in the prompt text, it is guided to automatically identify potential errors and propose correction strategies during the generation process, forming advanced thought chain data with self-supervised characteristics. Second, a logical mapping approach. By extracting and mapping key logical nodes (such as "reason," "goal," "risk," etc.) from the first thought process information to designated slots, the model is guided to perform logical expansion and refinement, improving the reasoning depth and interpretability of CoT. Third, a multimodal semantic correction approach. By combining image data features and reference action decision information, the consistency between visual semantics and behavioral semantics in the generated text is strengthened, improving the model's ability to prevent "semantic drift" (i.e., reasoning logic deviating from visual facts).

[0067] In some embodiments, the step of outputting structured second thinking process information using a second visual language model based on the image data, the first thinking process information, and the reference action decision information may include: outputting structured second thinking process information using a second visual language model based on the image data and a second prompt text; wherein the second prompt text is constructed based on the first thinking process information, the reference action decision information, and output template information, and is used to guide the second visual language model to generate the second thinking process information by using the reference action decision information as the final action decision information, referring to the logic of the first thinking process information, and according to the specified structure of the output template information.

[0068] The above embodiments do not limit the structure of the output template information. In some embodiments, the output template information may include a first guidance prompt text corresponding to the following information: scene analysis information, analysis of the current task status information, information on the next action to be taken, and information on reflection and correction actions.

[0069] In some embodiments, the output template information may further include a second guidance prompt text corresponding to the reference action decision information. In other embodiments, the output template information may only include the first guidance prompt text, without including the second guidance prompt text; this application does not limit this.

[0070] The above embodiments do not limit the representation of the first and second guidance prompt texts; for example, they can be represented using tags. The first guidance prompt text can be represented using the "think" tag, and the second guidance prompt text can be represented using the "answer" tag.

[0071] The above embodiments do not limit the representation of the second thinking process information and the reference action decision information. In some embodiments, the second thinking process information can be represented by a first label (e.g., the think label), and the reference action decision information can be represented by a second label (e.g., the answer label).

[0072] The above embodiments do not limit the specific structure of the first tag. In some embodiments, the first tag may include one or more of the following: separately set scene analysis information, analysis of the current task status information, information on the next action to be taken, and reflection and correction action information. As an example, the first tag may include separately set scene analysis information, analysis of the current task status information, information on the next action to be taken, and reflection and correction action information.

[0073] The Second Visual Language Model (VLM-2) is a deep learning model with multimodal semantic understanding and text generation capabilities, used for logical reconstruction and semantic enhancement of validated thought processes. Unlike the First Visual Language Model (VLM-1), which primarily performs "free generation," VLM-2 undertakes "structured processing and logical refinement." VLM-2 receives multimodal inputs including image data, first thought process information, reference action decision information, and output template information, enabling the model to output textual information that conforms to a specified logical format, is complete in content, and possesses reflective depth. For example, VLM-2 can obtain visual feature information of the task scene through a visual feature extraction module; use the first thought process information and reference action decision information as semantic guidance inputs, and unify the modeling of visual feature information and textual semantic information through a cross-modal fusion layer; finally, the language generation module outputs structured text that conforms to template constraints. This structured text is the "second thought process information," whose content logic references the reasoning of the first thought process information, but is more rigorous, clearer in structure, and more semantically controllable. By introducing two collaborative models (VLM-1 and VLM-2), a "generation-verification-reinforcement" mechanism is formed. VLM-1 is responsible for generating diverse inference samples, while VLM-2 is responsible for logical reconstruction and unified standards, ensuring that the final data has high consistency and training availability.

[0074] The second prompt text serves as a guiding input to control the output behavior of the second visual language model, unifying "outcome guidance, structural constraints, and semantic reference" into a single input template. During the construction of the second prompt text, the first thought process information acts as a semantic reference, providing the original reasoning logic generated by VLM-1, enabling VLM-2 to understand the task context and thought chain structure. Reference action decision information serves as an outcome guide, specifying the final correct action direction and ensuring that the model's output thought results align with the target decision. Output template information acts as a format constraint, limiting the structured expression of the generated text (such as paragraph headings, logical order, and reflection levels). By integrating these three types of information in the second prompt text, fine-grained control over the model's generation behavior can be achieved. Specifically, the second prompt text guides the model to "refer to the first thought logic, but with the correct action as the endpoint," and requires the model to output according to the template format, ensuring that the final generated second thought process information retains the diversity of thought processes while possessing a unified and interpretable structure.

[0075] In practical implementation, the second prompt text can be constructed in the form of a natural language template, for example: "Please generate a systematic analysis of this task based on the following: (1) Reference thinking process: {first thinking process information}; (2) Final correct action: {reference action decision information}; please elaborate in detail according to the specified format of {output template information}." This natural language template gives the VLM-2 reasoning process a clear goal orientation and output specification, thereby ensuring that different data samples have a unified expression style and logical hierarchy.

[0076] Output template information is a type of format control information used to constrain the structure of generated text, ensuring consistency in logical hierarchy, semantic content, and paragraph format. Output template information can be represented, for example, as parsed structured instructions, as follows.

[0077] <think> 1. **Analyze the scene:** [Please describe your observations of the current environment in detail here, including the position and state of relevant objects.] 2. **Identify the current state of the task:** [Describe the overall goal and current progress of the task here.] 3. **Determine the next logical step:** [Please explain in detail why '{true answer}' is the most logical and highest priority next step, and state its purpose and expected result.] 4. **Reflection and Correction:** [Please reflect here. For example, what potential risks are involved in this choice? Are there any alternatives? Why were other options ultimately ruled out? How can we ensure this action is executed accurately?] <answer> # [Please enter your '{true answer}' here] < / answer> < / think> Output template information can constrain the semantic order of the model's output, aligning its logical chain with the human reasoning process. Through explicit instructional settings of templates, it is possible to maintain the uniformity, readability, and parsability of the output format when generating large-scale automatic data, facilitating modular data retrieval and feature annotation in subsequent training systems.

[0078] The second thought process information, for example, is a structured, high-quality thought chain text generated by the second visual language model, used to reinforce the unprocessed raw reasoning output. Essentially, it is a "thought process record" processed through logical reinforcement and semantic standardization, and may include features such as: scene analysis information (corresponding to the visual semantic layer); analysis of the current task state information (corresponding to the task cognition layer); information on the next action to be taken (corresponding to the reasoning layer); and information on reflection and correction of actions (corresponding to the metacognitive layer). In terms of control logic, the generation of the second thought process information relies on the goal-oriented approach of "reference action decision information." When outputting, the second visual language model takes the correct action as the final conclusion, refers to the logical path of the first thought process information, and completes expansion and reconstruction according to the template. This mechanism not only ensures the correctness and rationality of the reasoning chain but also makes the output systematic and reflective, providing crucial data support for training robot models capable of autonomous reasoning and self-correction. Therefore, the second thought process information is not merely a linguistic output result but also high-confidence data filtered through a triple mechanism of "correctness verification + logical reinforcement + template standardization."

[0079] In the above embodiments, when generating the second thought process information, a second visual language model is invoked for reasoning based on image data and the second prompt text. The second prompt text is constructed by comprehensively utilizing the first thought process information, reference action decision information, and preset or specified output template information. This allows the model to reference verified logical paths while using the reference action decision information as the final action target, outputting high-quality thought chain text with clear logical hierarchy, standardized language expression, and reflective qualities according to a templated structure. This design achieves an automatic conversion from "raw thought chain" to "structured reinforced thought chain," enabling the model to generate a large amount of highly consistent, highly logical, and directly trainable robot thought chain data with low manual cost, realizing the generation and encapsulation of structured and reflective thinking.

[0080] The above embodiments can automatically generate higher-order thinking elements including "reflection and correction." In the final generation stage, through the design of structured prompt text, the "reflection and correction" component is stably and automatically introduced into the CoT data. This generates data that includes self-examination, risk assessment, and alternative solution analysis. This means that robots trained with this data will not only be passive executors of tasks but are more likely to become proactive "thinkers" with preliminary critical thinking and risk awareness. This has immeasurable value in improving the safety, adaptability, and task success rate of robots in the complex, dynamic, and uncertain real world.

[0081] Furthermore, the above embodiments implement a highly structured data schema and its optimization for training. The labeled think-answer structure clearly decouples and structures the robot's implicit, complex thinking process (the four parts within the think label) from the explicit, concise final decision (the answer label). This data schema is extremely friendly to model training, allowing developers to train AI models in a more refined and modular way. For example, data from the "scene analysis" section can be used separately to specifically improve the robot's environmental perception and understanding capabilities, or data from the "reflection and correction" section can be used to specifically train its risk control and decision optimization modules, greatly facilitating the "partial optimization" and "precise upgrading" of the robot's cognitive abilities.

[0082] See Figure 2 , Figure 2 This is a schematic diagram of the structure of a robot data production device provided in an embodiment of this application.

[0083] This application also provides a robot data production device, which includes a data receiving module, a first inference module, and a second inference module.

[0084] The data receiving module is used to receive image data, motion constraint information, and reference motion decision information corresponding to the robot's task scenario.

[0085] The first reasoning module is used to reason about at least one action of the robot based on the image data and the action constraint information using a first visual language model, so as to generate first thinking process information and first action decision information.

[0086] The second reasoning module is used to output structured second thinking process information using a second visual language model, based on the image data, the first thinking process information, and the reference action decision information, when the first action decision information is determined to be correct.

[0087] In the above embodiments, the first reasoning module can be called the initial thought chain and model answer generation module, and the second reasoning module can be called the open-ended answer and open-ended thinking generation module (Open-Answer+Open Thinking).

[0088] In some embodiments, the apparatus may further include an option match filter, which performs the following steps: determining whether the first action decision information is correct.

[0089] See Figure 3 , Figure 3 This is a schematic diagram of a robot data production process provided in an embodiment of this application.

[0090] First, scenario-based thought process and initial answer generation. The goal of this stage is to enable the First Visual Language Model (VLM-1) to autonomously think and make an initial action decision based on a given visual scenario. The input to the First Visual Language Model may include, for example, image data and an open-ended guiding prompt (i.e., the first prompt text). The image data may be, for example, an image representing the robot's current task scenario. The first prompt text may contain, for example, a set of executable actions (i.e., a set of "SubTask Action Options"), which can be provided by an action knowledge base. The action knowledge base may also provide a pre-defined "Ground Truth Answer" (i.e., reference action decision information) associated with the scenario, which is, for example, selected from the set of executable actions. As an example, for a scenario of repairing a bicycle, the ground truth answer might be "pick up the wrench on the ground." This ground truth answer is not visible to the model at this step. The first prompt text encourages VLM-1 to generate both the thought process and the final action decision simultaneously. The output of VLM-1 may be, for example, a data pair containing the "raw CoT" and the "model answer." The “original CoT” (i.e., first thought process information) can be an unstructured text. The “model-generated answer” (i.e., first action decision information) can be an action decision autonomously generated by VLM-1.

[0091] Secondly, the answer matching filter and automated CoT quality verification act as an automated "quality inspector." By comparing the model-generated answers with preset real answers, it determines whether the thought chain (CoT) generated by the model is reasonable and worth retaining. For example, the "model answer," the "real answer," and the corresponding "original CoT" are input into the option matching filter. The option matching filter performs a matching operation, determining whether the "model answer" matches the "real answer." If the match is successful, the corresponding "original CoT" is deemed high-quality and reliable. This "original CoT" and its corresponding real answer are verified and proceed to the next processing stage. If the match fails, this "original CoT" and its corresponding "real answer" are discarded and not used in subsequent production. The output of the option matching filter may include, for example, a verified, high-quality "original CoT" and its corresponding "real answer." In practical applications, this combination can be called "anchor data" (Think-Answer).

[0092] Next, the generation and encapsulation of structured, reflective thinking occurs. The goal of this stage is to deepen and format the validated "anchor data" from the previous step, producing high-quality, structurally sound data samples that can be directly used for model training. This can be achieved through a second visual language model (VLM-2) for generating open-ended answers and open thinking. The input to the second visual language model can, for example, include the "anchor data" from the second step, namely the validated "original CoT" and "true answer." In practical applications, the "original CoT" and "true answer" can be embedded as key information into a carefully designed second prompt text that forces the output of a specific structure. This specific structure can be defined, for example, by output template information. Then, the image data and the constructed second prompt text are input into VLM-2. The second prompt text can, for example, require the model to perform scene analysis based on the image data and, with the "true answer" as the final decision, refer to the logic of the "original CoT" to generate a richer, more logically rigorous, and more formatted (outputting according to the specified structure in the output template information) complete thought process.

[0093] See Figure 4 , Figure 4 This is an image data provided in the embodiments of this application.

[0094] In a specific application scenario, the input Figure 4 The image data shown can be used to obtain the following output.

[0095] <think> 1. Scene Analysis: The image shows a robotic arm holding a square slice of brown bread. The arm is positioned directly above a purple plate on a white table. Also on the table are another slice of bread, a container with lettuce, and another container with a red clip. A second robotic arm is visible on the right.

[0096] 2. Identify the current task state: The overall task is "make a sandwich". The robot has successfully completed the first sub-task, which is picking up a slice of bread. It is currently holding the bread.

[0097] 3. Determine the next logical step: In the process of making a sandwich, after picking up the first slice of bread, the next step is to place it down on a work surface, such as a plate, to serve as a base for the other ingredients. The robot is already holding the bread over the plate, so putting it down is the immediate and logical next action.

[0098] 4. Reflection and Correction: The robot cannot pick up another item while already holding the bread. Placing the bread anywhere other than the plate would be inefficient or incorrect for the task. Therefore, placing the bread on the plate is the correct progression.

[0099] < / think> <answer> Place the held bread slice into the plate on the countertop.

[0100] < / answer> in," <think> ……< / think> The "" section represents the second thought process information, labeled with a first tag. This first tag includes separately defined scene analysis information (corresponding to "1. Analyze the scene"), current task state analysis information (corresponding to "2. Identify the current state of the task"), next action information (corresponding to "3. Determine the next logical step"), and reflection and correction action information (corresponding to "4. Reflection and Correction"). <answer> ……< / answer> The "" section contains reference action decision information represented by the second label.

[0101] See Figure 5 , Figure 5 This is a flowchart illustrating a robot model training method provided in an embodiment of this application.

[0102] This application also provides a robot model training method, which includes steps S201 to S204.

[0103] Step S201: Receive image data, motion constraint information and reference motion decision information corresponding to the robot task scenario.

[0104] Step S202: Based on the image data and the action constraint information, use a first visual language model to infer at least one action of the robot to generate first thought process information and first action decision information.

[0105] Step S203: If the first action decision information is determined to be correct, based on the image data, the first thought process information and the reference action decision information, a structured second thought process information is output using a second visual language model.

[0106] Step S204: Based on the second thinking process information and the reference action decision information, train the specified robot model to update at least one model parameter of the robot model.

[0107] See Figure 6 , Figure 6 This is a structural block diagram of a robot provided in an embodiment of this application.

[0108] This application also provides a robot, which stores a robot model, and the robot model is trained using any of the above-described training methods.

[0109] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the above methods.

[0110] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of any of the above methods.

[0111] The computer program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the computer program product of the embodiments of this application is not limited thereto, and the computer program product may be any combination of one or more computer-readable media.

[0112] See Figure 7 , Figure 7 This is a structural block diagram of a computer device provided in an embodiment of this application.

[0113] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above methods.

[0114] The embodiments of this application do not limit the computer device, which may be, for example, a local computer device, a cloud computer device, a distributed computer device, etc.

[0115] The computer device may include: a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected through internal connection paths.

[0116] The memory 110 is used to store computer programs, which in some implementations may include code for implementing the methods of the embodiments of this application.

[0117] The processor 120 executes the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information, and output operation results and other data. In some implementations, when the solutions of the embodiments of this application are implemented by software or firmware, the computer program used to implement the solutions of the embodiments of this application can be stored in the processor 120 and executed by the processor 120.

[0118] The memory 110 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, any memory of these and other suitable types. As an example, the memory 110 includes random access memory (RAM), cache memory, and read-only memory (ROM). The memory 110 stores a computer program that can be executed by processor 120, causing processor 120 to implement the steps of any of the methods described above.

[0119] The processor 120 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 120 can be any conventional processor.

[0120] In implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware in the processor 120 or by instructions in software form. The method disclosed in the embodiments of this application can be directly implemented by the hardware processor, or by a combination of hardware and software modules in the processor 120. The software modules can be located in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are not provided here.

[0121] In some implementations, in addition to the hardware units described above, computer devices may also include software modules, such as operating systems, basic input / output systems (BIOS), and application software.

[0122] An operating system is used to manage the hardware and / or software resources of a computer device; it is the kernel and foundation of the computer. The operating system handles fundamental tasks such as managing and configuring memory, determining the priority of system resource allocation, controlling input and output devices, operating the network, and managing the file system. To facilitate user operation, most operating systems provide a user interface for interaction with the system.

[0123] The BIOS is used to perform hardware initialization during the power-on boot phase and to provide runtime services for the operating system and applications. In some implementations, the BIOS can also monitor and display processor temperature and execute temperature protection strategies.

[0124] Application software, also known as an application program, can be understood as software written for a specific user application purpose, and is one of the main categories of computer software. For example, application software can be a program used to achieve purposes such as power control and temperature management.

[0125] It is understood that the specific examples in this application are only intended to help those skilled in the art better understand the implementation of this application, and are not intended to limit the scope of protection of this application.

[0126] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application.

[0127] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and this application does not limit them.

[0128] Unless otherwise stated, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0129] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the embodiments described above can be referred to the corresponding processes in other embodiments, and will not be repeated here.

[0131] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0132] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the technical solution in this application, depending on actual needs.

[0133] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0134] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, essentially, or the part that contributes to related technologies, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0135] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating robot data, characterized in that, The method includes: Receive image data, motion constraint information, and reference motion decision information corresponding to the robot task scenario; Based on the image data and the action constraint information, at least one action of the robot is inferred using a first visual language model to generate first thought process information and first action decision information. If the first action decision information is determined to be correct, a structured second thought process information is output using a second visual language model based on the image data, the first thought process information, and the reference action decision information.

2. The robot data production method according to claim 1, characterized in that, The action constraint information includes a set of executable actions; and / or, the second visual language model is configured to use the reference action decision information as the final action decision information.

3. The robot data production method according to claim 1 or 2, characterized in that, Based on the image data and the action constraint information, the robot uses a first visual language model to infer at least one action of the robot to generate first thought process information and first action decision information, including: Based on the image data and the first prompt text, the first visual language model is used to infer at least one action of the robot to generate the first thinking process information and the first action decision information. The first prompt text is constructed based on the action constraint information and is used to guide the first visual language model to generate the first thought process information and the first action decision information.

4. The robot data production method according to claim 1, characterized in that, The process of determining whether the first action decision information is correct includes: The first action decision information and the reference action decision information are matched based on the specified matching algorithm. If the match is successful, the first action decision information is considered correct; or, if the match fails, the first action decision information is considered incorrect.

5. The robot data production method according to claim 1, characterized in that, The step of outputting structured second thought process information using a second visual language model based on the image data, the first thought process information, and the reference action decision information includes: Based on the image data and the second prompt text, the second visual language model is used to output structured second thought process information; The second prompt text is constructed based on the first thought process information, the reference action decision information, and the output template information, and is used to guide the second visual language model to generate the second thought process information by using the reference action decision information as the final action decision information, referring to the logic of the first thought process information, and according to the specified structure of the output template information.

6. The robot data production method according to claim 5, characterized in that, The output template information includes a first guidance prompt text corresponding to the following information: scene analysis information, analysis of the current task status information, information on the next action to be taken, and information on reflection and correction actions.

7. The robot data production method according to claim 6, characterized in that, The output template information also includes a second guidance prompt text corresponding to the reference action decision information; The second thinking process information is represented by a first label, and the reference action decision information is represented by a second label. Furthermore, the first label includes separately set scenario analysis information, analysis of the current task status information, information on the next action to be taken, and information on reflection and correction actions.

8. A robot data production device, characterized in that, The device includes: The data receiving module is used to receive image data, motion constraint information, and reference motion decision information corresponding to the robot task scenario; The first reasoning module is used to reason about at least one action of the robot based on the image data and the action constraint information using a first visual language model, so as to generate first thinking process information and first action decision information. The second reasoning module is used to output structured second thinking process information using a second visual language model, based on the image data, the first thinking process information, and the reference action decision information, when the first action decision information is determined to be correct.

9. A robot model training method, characterized in that, The method includes: Receive image data, motion constraint information, and reference motion decision information corresponding to the robot task scenario; Based on the image data and the action constraint information, at least one action of the robot is inferred using a first visual language model to generate first thought process information and first action decision information. If the first action decision information is determined to be correct, based on the image data, the first thought process information, and the reference action decision information, a second visual language model is used to output structured second thought process information. Based on the second thought process information and the reference action decision information, a specified robot model is trained to update at least one model parameter of the robot model.

10. A robot, characterized in that, The robot stores a robot model, which is trained using the method described in claim 9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7, or the steps of the method according to claim 9.

12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7, or the steps of the method according to claim 9.