A robot task generation method, system and computer device
Patent Information
- Application Number
- CN202610162644.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-02-04
AI Technical Summary
[0004]本公开实施例提供了一种机器人任务生成方法、系统及计算机设备,用以解决现有机器人训练数据分布不均匀的问题
[0019]本公开实施例的有益效果包括:
Smart Images

Figure CN121973215B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of embodied intelligence, and more particularly to a method, system, and computer device for generating robot tasks. Background Technology
[0002] With the development of embodied intelligence, the generalization ability of Vision-Language-Action (VLA) models highly depends on large-scale, high-quality training data. Currently, data construction is mainly based on manually designed and teleoperated paths. This approach is physically and human-intensive, and limited by the cognitive biases of human designers, tasks are mostly confined to simple primitives such as "grasping" and "placing". This results in a severe long-tail distribution of the dataset, lacking complex interaction data such as long-range planning and hand-eye coordination, which limits the model's generalization ability in novel scenarios and cannot solve the problem of uneven distribution of robot training data.
[0003] In summary, there is an urgent need for a robot task generation method that can automatically generate training data with a balanced distribution to address the shortcomings of existing technologies. Summary of the Invention
[0004] This disclosure provides a robot task generation method, system, and computer device to solve the problem of uneven distribution of existing robot training data.
[0005] In view of the above problems, firstly, embodiments of this disclosure provide a robot task generation method, including: Determine a task space tuple to describe the robot's tasks; the task space tuple includes multiple task description items. Based on the historical task data corresponding to the description data of the preset task description item in the task space tuple, data sampling is performed on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item. The task description data set is provided as a constraint to the proposal generator to generate a target task proposal.
[0006] In conjunction with the first aspect, in one possible implementation, the step of sampling data from historical task data corresponding to preset task description items in the task space tuple, to obtain a task description data set containing at least the sampled data corresponding to the preset task description items, includes: Based on the usage frequency of the description data corresponding to the preset task description item in the task space tuple in historical task data, data sampling is performed on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item; the usage frequency satisfies: lower than a preset frequency.
[0007] In conjunction with the first aspect, in one possible implementation, the preset task description item includes: a set of scene categories representing robot usage scenarios, a task object library representing robot interaction objects, and a task skill library representing robot interaction methods; Based on the frequency of use of the description data corresponding to the preset task description items in the task space tuple in historical task data, data sampling is performed on the preset task description items to obtain a task description data group that at least contains the sampled data corresponding to the preset task description items, including: Based on the frequency of the robot performing tasks in each scene category in the scene category set, scene categories with a usage frequency lower than a preset first threshold are identified as target scene categories; Based on the semantics of the target scene category, task objects logically associated with the target scene category are selected from the task object library to obtain a task object set; Based on the semantics of the target scene category, task skills logically associated with the target scene category are selected from the task skill library to obtain a task skill set; Based on the historical interaction frequency between the robot and the task object, task objects whose usage frequency is within the first preset low frequency order range are selected from the task object set to obtain a candidate task object set. Based on the robot’s historical usage frequency of task skills, task skills whose usage frequency falls within the second preset low frequency order range are selected from the task skill set to obtain a candidate task skill set. The target scene category, the candidate task object set, and the candidate task skill set are determined as the task description data group.
[0008] In conjunction with the first aspect, in one possible implementation, providing the task description data set as constraints to the proposal generator to generate a target task proposal includes: The generative model is invoked to determine prompt words based on the task description data set, and an initial task proposal is generated to instruct the target robot to interact. The initial task proposal is reviewed using the review model to obtain the review results; The reflective optimizer is invoked to determine whether the initial task proposal needs to be revised based on the retrieved historical interaction records, the review results, and the initial task proposal. If corrections are needed, a task proposal correction instruction is generated to correct the initial task proposal and obtain the target task proposal.
[0009] In conjunction with the first aspect, in one possible implementation, providing the task description data set as constraints to the proposal generator to generate a target task proposal includes: The generative model is invoked to determine prompt words based on the task description data set, and an initial task proposal is generated to instruct the target robot to interact. The initial task proposal is used as the current task proposal, and the following proposal review steps are performed: The review model is invoked to review the current task proposal and the review results are obtained. The reflective optimizer is invoked to determine whether the current task proposal needs to be revised based on the retrieved historical interaction records, current review results, and current task proposal. If the current task proposal needs to be modified, generate a task proposal modification instruction to modify the current task proposal and obtain the modified task proposal. The revised task proposal is used as the new current task proposal. The process returns to the proposal review step until it is determined that the current task proposal does not need to be revised, thus obtaining the target task proposal.
[0010] In conjunction with the first aspect, in one possible implementation, the review model includes at least one of the following evaluators: a physical feasibility evaluator, a novelty evaluator, and a constraint consistency evaluator; The review model is invoked in the following manner to review the proposals for the review tasks, and the review results are obtained, including: The evaluators are invoked respectively to review the task proposals to be reviewed from the corresponding evaluation dimensions, and the corresponding review results are obtained respectively; the task proposals to be reviewed include the initial task proposals and / or the current task proposals.
[0011] In conjunction with the first aspect, in one possible implementation, the evaluation dimensions of the physical feasibility assessor include: kinematic range, physical interaction logic, and interaction capability requirements; When the evaluator includes a physical feasibility evaluator, the step of respectively invoking the evaluator to review the proposal to be reviewed from the corresponding evaluation dimensions, and obtaining the corresponding review results, includes: The physical feasibility assessor is invoked to review the proposed task from the corresponding assessment dimensions, and the review results from the physical feasibility assessor are obtained, including: The physical feasibility assessor is invoked to analyze the kinematic range of the robot in the task proposal to be reviewed, and to examine whether the kinematic range conflicts with the kinematic constraints of the target robot. Analyze the physical interaction logic of the robot in the task proposal to be reviewed, and review whether the physical interaction logic is reasonable; Analyze the robot's interaction capability requirements in the proposed task to be reviewed, and examine whether the interaction capability requirements conflict with the hardware capability constraints of the target robot; Based on the review results of kinematic range, physical interaction logic, and interaction capability requirements, the review results of the physical feasibility assessor are integrated.
[0012] In conjunction with the first aspect, in one possible implementation, the evaluation dimensions of the novelty evaluator include: complexity and redundancy; When the evaluator includes a novelty evaluator, the step of respectively invoking the evaluator to examine the proposal to be examined from the corresponding evaluation dimensions and obtaining the corresponding examination results includes: The novelty evaluator is invoked to examine the proposal to be examined from the evaluation dimensions corresponding to the novelty evaluator, and the examination results of the novelty evaluator are obtained, including: The novelty evaluator is invoked to analyze the robot's interaction process in the task proposal to be reviewed, and the complexity of the task proposal to be reviewed is examined according to a predetermined complexity indicator standard; wherein, the complexity indicator standard includes at least one of the following: whether the task proposal includes long-range planning, whether the task proposal includes tool use, or whether the task proposal includes deformable object manipulation. The redundancy of the initial task proposal is reviewed by searching the preset task dataset based on the robot's task logic in the task proposal to be reviewed. Based on the review results of complexity and redundancy, the review results of the novelty evaluator are integrated.
[0013] In conjunction with the first aspect, in one possible implementation, the evaluation dimensions of the constraint consistency evaluator include: content authenticity and scenario logic adaptability; When the evaluator includes a constraint consistency evaluator, the step of respectively invoking the evaluator to review the proposal to be reviewed from the corresponding evaluation dimensions, and obtaining the corresponding review results, includes: Invoke the constraint consistency evaluator, review the proposal to be reviewed from the evaluation dimensions corresponding to the constraint consistency evaluator, and obtain the review results of the constraint consistency evaluator, including: Invoke the constraint consistency evaluator to check whether the task description items involved in the task proposal to be reviewed are included in the task description data group, and review the authenticity of the content of the task proposal to be reviewed. Invoke the constraint consistency evaluator to detect the logical relationships between the task description items involved in the task proposal to be reviewed, and review the scenario logic adaptability of the task proposal to be reviewed. Based on the review results of content authenticity and scenario logic adaptability, the review results of the constraint consistency evaluator are integrated.
[0014] In conjunction with the first aspect, in one possible implementation, the step of sampling data for the preset task description item based on historical task data corresponding to the description data in the task space tuple includes: Based on the historical count information of the category to which the description data of the preset task description item in the task space tuple belongs, data sampling is performed on the preset task description item; The method further includes: Modify the corresponding counter value for the description data category corresponding to the preset task description item involved in the target task proposal.
[0015] In conjunction with the first aspect, in one possible implementation, the historical interaction record includes: a pair of task statements and feedback statements in a historical task proposal; Also includes: Upon obtaining the runtime feedback data corresponding to the target task proposal, the feedback statements in the runtime feedback data are matched with the task statements in the corresponding target task proposal and stored in the historical interaction record.
[0016] In conjunction with the first aspect, in one possible implementation, it further includes: Obtain the image data related to the task description data group; The task description data set is provided as constraints to the proposal generator to generate a target task proposal, including: The task description data set, along with the image data, is provided to the proposal generator to generate a target task proposal.
[0017] A second aspect of this disclosure provides a robot task generation system, comprising: The determination module is used to determine the task space tuple used to describe the robot's task; The task space tuple includes multiple task description items; The sampling module is used to sample the preset task description item based on historical task data corresponding to the description data of the preset task description item in the task space tuple, so as to obtain a task description data group that contains at least the sampled data corresponding to the preset task description item. The generation module is used to provide the task description data set as constraints to the proposal generator to generate a target task proposal.
[0018] A third aspect of this disclosure provides a computer device, characterized in that it includes a processor, a memory, and a bus; The memory stores machine-readable instructions that can be executed by the processor; When the computer device is running, the processor and the memory communicate via a bus; When the machine-readable instructions are executed by the processor, they perform the steps of a robot task generation method as described in the first aspect or in any possible embodiment of the first aspect.
[0019] The beneficial effects of the embodiments disclosed herein include: This disclosure provides a robot task generation method, system, and computer device. The method includes: determining a task space tuple for describing robot tasks; the task space tuple includes multiple task description items; based on historical task data corresponding to the description data of preset task description items in the task space tuple, sampling the preset task description items to obtain a task description data set containing at least the sampled data corresponding to the preset task description items; and providing the task description data set as a constraint to a proposal generator to generate a target task proposal. The method provided in this disclosure predefines a task space tuple for robot task proposals and samples the preset task description items according to actual needs. For example, if training is to be performed on rare scenarios, data sampling can be performed on preset task description items with low usage frequency in historical task data; if training is to be performed on high-risk tasks, data sampling can be performed on preset task description items with a high failure rate in historical task data. Through this data sampling mechanism, the proposal generator is forced to generate corresponding target task proposals based on the constraints of the task description data set. The task proposals generated by this method can explore uncovered task scenarios according to actual needs, solving the problem of uneven distribution of existing training data. Attached Figure Description
[0020] Figure 1 A flowchart illustrating the robot task generation method provided in this embodiment of the disclosure; Figure 2 This is one of the structural schematic diagrams of the robot task generation system provided in the embodiments of this disclosure; Figure 3 This is a second schematic diagram of the structure of the robot task generation system provided in the embodiments of this disclosure; Figure 4 This is the third schematic diagram of the structure of the robot task generation system provided in the embodiments of this disclosure; Figure 5 Fourth schematic diagram of the structure of the robot task generation system provided in the embodiments of this disclosure; Figure 6 This is the fifth schematic diagram of the structure of the robot task generation system provided in the embodiments of this disclosure. Detailed Implementation
[0021] This disclosure provides a robot task generation method, system, and computer device. Preferred embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit this disclosure. Furthermore, the embodiments and features described in this application can be combined with each other unless otherwise specified.
[0022] This disclosure provides a robot task generation method, such as... Figure 1 As shown, it can be implemented as follows: S101. Determine the task space tuple used to describe the robot's task; The task space tuple includes multiple task description items; S102. Based on the historical task data corresponding to the description data of the preset task description item in the task space tuple, perform data sampling on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item. S103. Provide the task description data set as a constraint to the proposal generator to generate a target task proposal.
[0023] In this embodiment of the disclosure, the task space tuple can be a structured collection of elements required for the robot to perform a task. The task space tuple may include multiple task description items. Each task description item is the smallest descriptive unit of the elements required for the robot task within the task space tuple, corresponding to the definition of specific attributes of the robot task from different dimensions. Each type of task description item can correspond to specific elements in the task execution process (e.g., interaction objects, interaction behaviors, and spatial constraints). By combining different types of task description items, a specific task model can be constructed.
[0024] Historical task data refers to the collection of descriptive data generated by the robot system within each work cycle for each task description item. Specifically, it can include the task description data values corresponding to each task description item (for example, if the task description item includes a scene category, then historical task data can include specific historical scene category values, such as: hospital), and can also include the usage frequency, time characteristics, quality performance, and correlations of each task description item. Usage frequency measures the popularity of each task description item in task proposals; time characteristics record the timestamps of activation of each task description item, used to measure the active period of each task description item in task proposals; quality performance records the success and failure ratios of each task description item during execution, used to measure the execution risk of each task description item; correlations record the frequency of simultaneous occurrence of different task description items, used to measure the coupling relationship between task description items.
[0025] Based on records in historical task data, data sampling can be performed on pre-defined task description items in the task space tuples of historical task data, according to the needs of robot training. The sampled task description items and their corresponding sampled data are then grouped into task description data sets to generate candidate task proposals. For example, if training is to be performed on rare scenarios, task description items with low usage frequency in historical task data can be sampled to form task description data sets; if training is to be performed on high-risk tasks, task description items with a high failure rate in historical task data can be sampled to form task description data sets. By sampling specific task description items, this disclosure can force the robot to explore a specific space (e.g., the long-tail space) for task proposals, thus avoiding the problem of insufficient coverage of task proposals due to cognitive biases in related methods.
[0026] It should be noted that the initial value of the description data for each task description item in the historical task data of this disclosure can be 0, and the value of the description data is updated according to the generated task proposal. Alternatively, relevant description data based on existing robot operations can be imported into this historical task data as the initial value of the historical task data.
[0027] The obtained task description data set can be provided as constraints to the proposal generator, which then generates a target task proposal with a preset format based on each task description item in the task description data set. In this disclosure, a task proposal can refer to an executable scheme containing robot action logic and execution parameters. After obtaining the target task proposal, it can be executed in a simulation environment or a real robot environment to train the robot. Generating task proposals based on the above constraints ensures that the generated task proposals are within a feasible and implementable range, avoiding the generation of illusory objects and operational instructions that violate the laws of physics, thus improving the quality of the generated task proposals. It also avoids the manual verification steps required in existing technologies, thereby increasing the efficiency of task generation.
[0028] It should be noted that the method provided in this disclosure can be implemented on a large model-driven business platform, which can be deployed on computer devices with computing capabilities. Following the steps given in this disclosure, based on a predefined task space tuple, software and hardware resources are invoked to ultimately output the target task proposal.
[0029] In one possible implementation, the task space tuple can be defined as... .in, This indicates the configuration type of the target robot, including but not limited to single-arm robots, dual-arm collaborative robots, and mobile manipulation composite robots; It represents a set of scene categories, such as predefined typical environments including home, office, education, laboratory, kitchen, industry, retail and medical, etc. This represents a collection of task object assets containing physical attributes, such as a collection of different task objects. This represents the robot's operational skill library, which contains a set of atomic operations for task skills such as "grabbing", "rotating", "pouring water", and "wiping". It is a task description in natural language.
[0030] In another embodiment provided in this disclosure, the step S102 above, "based on the historical task data corresponding to the description data of the preset task description item in the task space tuple, performing data sampling on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item," can be implemented as follows: Step 1: Based on the usage frequency of the description data corresponding to the preset task description item in the task space tuple in historical task data, perform data sampling on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item; the usage frequency satisfies: lower than a preset frequency.
[0031] In this embodiment of the disclosure, in response to the problems of insufficient data diversity and severe long-tail distribution in robot operation data collection in existing embodied AI research, the frequency of use of the description data corresponding to the preset task description item in historical task data can be considered according to actual needs. For preset task description items with a usage frequency lower than the preset frequency, it can be considered that the relevant data is scarce. By sampling the data of the preset task description item, the corresponding task description data can be increased, thereby increasing the diversity of robot operation data collection.
[0032] Furthermore, the preset frequencies corresponding to different preset task description items can be different, and can be set according to the actual situation during implementation; there are no restrictions here.
[0033] In another embodiment provided in this disclosure, the preset task description item includes: a set of scene categories representing robot usage scenarios, a task object library representing robot interaction objects, and a task skill library representing robot interaction methods. Step one above, "based on the frequency of use of the description data corresponding to the preset task description item in the task space tuple in historical task data, perform data sampling on the preset task description item to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item," can be implemented as follows: Step 1: Based on the frequency of the robot performing tasks in each scene category in the scene category set, determine the scene categories whose usage frequency is lower than a preset first threshold as the target scene categories; Step 2: Based on the semantics of the target scene category, select the task objects logically associated with the target scene category from the task object library to obtain the task object set; Step 3: Based on the semantics of the target scene category, select the task skills that are logically associated with the target scene category from the task skill library to obtain the task skill set; Step 4: Based on the historical interaction frequency between the robot and the task object, select the task objects whose usage frequency is within the first preset low frequency order range from the task object set to obtain the candidate task object set. Step 5: Based on the robot's historical usage frequency of task skills, select task skills whose usage frequency falls within the second preset low-frequency order range from the task skill set to obtain a candidate task skill set; Step 6: Determine the target scene category, the candidate task object set, and the candidate task skill set as the task description data group.
[0034] In this embodiment, the preset task description items may include: a set of scene categories, a task object library, and a task skill library. The set of scene categories refers to the set of environmental contexts for robot usage scenarios. Each scene category can be presented as a semantic tag, such as home scenario, office scenario, education scenario, laboratory scenario, kitchen scenario, industrial scenario, retail scenario, and medical scenario, etc. Each scene category implicitly contains specific rules and corresponding behavioral logic. The task object library refers to the set of specific physical objects that the robot needs to interact with, manipulate, or perceive. These objects are the operational objects in the robot's task proposal. In the task object library, each task object can be a tag representing the object's name, which can correspond to the task object's physical attributes in the real environment (e.g., visual information, geometric information, mass information, coefficient of friction, and center of gravity, etc.). The task skill library refers to the set of capabilities possessed by the robot, which are the actions the robot performs to perform tasks, such as grasping, selecting, pouring water, and wiping. The tags corresponding to each task skill may include: a semantic description of the task skill (e.g., the action type and target type of the skill), constraints on the task skill (e.g., the prerequisites for implementing the skill), and the level of the task skill (e.g., atomic skills and combined skills).
[0035] In practical applications, the preset task description item can also include the robot configuration type (e.g., single-arm robot, dual-arm collaborative robot, or mobile operation composite robot) to define the implementing entity of the task proposal.
[0036] Specifically, to address the issue of insufficient coverage of long-tail scenarios in existing training data, we can first determine the scene categories with usage frequencies below a first threshold based on the usage frequency of descriptive data for the scene category set in historical task data. That is, from the scene category set, we select one or more scene categories with low usage frequencies and record them as the target scene categories generated for this task proposal.
[0037] Based on the selected target scene category, task objects and task skills that conform to the logic of the target scene category can be selected from the task skill library and task object library, respectively, based on the semantics of the target scene category. Here, the semantic-based selection method can employ semantic similarity calculation based on vector space, mapping the text descriptions of scene categories, task objects, and task skills to a vector space and determining similarity through mathematical distance; alternatively, a large language model can be invoked, using the semantics of the scene category as a basis, with the lists of task object libraries and task skill libraries as optional sets, and structured prompts to allow the large language model to infer task objects and task skills that appear in the target scene category and conform to real-world logic.
[0038] After selecting the task objects and task skills, the selected task objects and task skills can be combined to obtain the task object set and task skill set, respectively. Similarly, to address the issues of insufficient coverage and lack of diversity in long-tail scenarios in existing training data, one or more task objects within the task object set can be selected from the low-frequency side based on corresponding records in historical task data, sorted by usage frequency, to serve as the candidate task object set for the task proposal generation process. Likewise, one or more task skills within the task skill set can be selected from the low-frequency side based on corresponding records in historical task data, sorted by usage frequency, to serve as the candidate task skill set for the task proposal generation process.
[0039] After selecting the target scenario category, candidate task object set, and candidate task skill set, the above items can be combined to form a task description data set as constraints for task proposal generation. It should be noted that this task description data set may also include constraints on the configuration type of the robot that performs the task.
[0040] In one possible implementation, the target scene category can be defined as... That is, select the scene category The least frequently used scene category is selected as the target scene category. The resulting task description data set can be defined as follows: ;in, This is the set of candidate task objects. For candidate task skill sets.
[0041] In another embodiment provided in this disclosure, step S103 above, "providing the task description data set as a constraint condition to the proposal generator to generate a target task proposal," can be implemented as follows: Step 1: Invoke the generation model, determine prompt words based on the task description data set, and generate an initial task proposal to instruct the target robot to interact. Step 2: Invoke the review model to review the initial task proposal and obtain the review results; Step 3: Invoke the reflective optimizer to determine whether the initial task proposal needs to be revised based on the retrieved historical interaction records, the review results, and the initial task proposal; Step 4: If corrections are needed, generate a task proposal correction instruction to correct the initial task proposal and obtain the target task proposal.
[0042] In this embodiment of the disclosure, the generative model can refer to a large language model and / or a visual language model. When the generative model is a large language model, prompts written based on a task description data set can be input into the large language model, which will then generate an initial task proposal conforming to the prompt definition format.
[0043] Specifically, the prompt can include: role definition, constraint injection, and format specification. Role definition defines the large language model as a task design expert for a specific robot configuration type; constraint injection restricts the task description items involved in the task proposal to the content of the task description data group during the generation of the task proposal by the large language model; format specification provides the large model with a specific format for the task proposal output, ensuring that the proposal text output is in a specific framework (e.g., JavaScript Object Notation (JSON) framework), and that all predefined fields within the framework are effectively filled.
[0044] The output framework can include: task name, natural language instructions, object names and position parameters, skill sequence, scene layout, and task context description. The task name is a concise and standardized identifier for the task proposal, summarizing its specific content. Natural language instructions are specific commands given to the target robot in natural language to convey the task's intent. Object names and position parameters are the semantic labels of the task objects involved in the task proposal and their specific coordinates in physical space; these objects can be selected from the aforementioned set of candidate task objects. The skill sequence refers to the task skills the target robot needs to perform in the task proposal, and the sequence of execution; these skills can be selected from the aforementioned set of candidate task skills. The scene layout is a structured description of the spatial relationships between all objects related to the task proposal within the target scene category, defining the relative relationships between objects in the target scene category. The task context description includes the robot's state during task proposal execution, the task's preconditions, constraints, or historical information.
[0045] It should be noted that if the generated task proposal involves movement or requires spatial awareness, the generation model can be a visual language model. A set of multiple images that are prepared in advance and show the actual situation of the target scene category can be input into the visual language model, which will then ground the images into text descriptions to help the generation model understand the actual situation of the target scene category.
[0046] After the initial task proposal is generated, a review model can be invoked to review it, ensuring that the generated initial task proposal can be executed normally by the robot and has training value. Specifically, the review model can be implemented as a system composed of multiple evaluation agents. Based on the content of the initial task proposal, each evaluation agent reviews the initial task proposal from its own perspective and generates a corresponding review result in a preset language form (e.g., natural language).
[0047] These review results can be output to the reflective optimizer. Combined with the historical interaction records of the initial task proposal retrieved in advance, it can be determined whether the initial proposal needs to be modified. If the initial proposal needs to be modified, a task proposal modification instruction is generated to modify the initial task proposal, and finally the target task proposal is obtained.
[0048] The aforementioned evaluation agent can be implemented as a large model based on prompt words, which are defined to examine the initial task proposal from different perspectives. The aforementioned reflective optimizer can be implemented as a large model with retrieval capabilities, retrieving historical interaction records related to the initial task proposal through a pre-defined memory module. Combining this with the prompt words defined for the reflective optimizer, the large model can analyze whether the initial task proposal needs modification.
[0049] If corrections are needed, a task proposal correction instruction can be generated to instruct the generation model to revise the initial task proposal and obtain the target task proposal. Corrections may include: correcting incorrect object names in the initial proposal, adjusting the spatial layout to avoid collisions, or adding intermediate steps to comply with physical constraints.
[0050] If no correction is needed, the target task proposal can be added to the robot's task dataset, and relevant counts of descriptive data such as task objects, task skills, and scene categories from the relevant historical task data can be added. The task dataset here can refer to storage space containing task proposals executed by the target robot and task proposals expected to be executed. Furthermore, for executed task proposals, corresponding execution feedback describing the execution status can also be stored. The aforementioned task dataset can be stored in a preset memory module and can be read or queried as needed during the execution of the method provided in this disclosure.
[0051] In one possible implementation, the obtained target task proposal can be defined as .in, The modified algorithm used by the aforementioned reflective optimizer; For the initial mission proposal; The review results output by the physical feasibility assessor; The examination results output by the novelty evaluator; The review results output by the constraint consistency evaluator; This refers to the retrieved historical interaction records.
[0052] It should be noted that historical interaction records can be stored in a memory module, and each historical interaction record can be associated with the corresponding task statement of the task proposal as a key-value pair. It is stored in the form of . Among them, Semantic vector embedding for task statements. This includes specific implementation details, reasons for failure, and suggested corrective actions. When a search is required, it should be based on the initial task proposal. The embedding vector is used to retrieve the one or more key-value pairs with the highest similarity in the memory module. And extract the historical interaction records from them.
[0053] Using the above method, when generating a new task proposal, historical experience can be retrieved by using Retrieval-Augmented Generation (RAG) technology to avoid repeating past mistakes and achieve self-evolution of generation quality.
[0054] In another embodiment provided in this disclosure, step S103 above, "providing the task description data set as a constraint condition to the proposal generator to generate a target task proposal," can be implemented as follows: Step 1: Invoke the generation model, determine prompt words based on the task description data set, and generate an initial task proposal to instruct the target robot to interact. Step 2: Treat the initial task proposal as the current task proposal and perform the following proposal review steps: Step 1: Use the review model to review the current task proposal and obtain the review results; Step 2: Invoke the reflective optimizer to determine whether the current task proposal needs to be revised based on the retrieved historical interaction records, current review results, and current task proposal; Step 3: If the current task proposal needs to be modified, generate a task proposal modification command to modify the current task proposal and obtain the modified task proposal. Step 4: Use the revised task proposal as the new current task proposal, return to the proposal review step, and continue until it is determined that the current task proposal does not need to be revised, thus obtaining the target task proposal.
[0055] In this embodiment of the disclosure, the generation model, the review model, and the reflective optimizer can also be set as a closed-loop structure. After the generation model corrects the initial task proposal according to the task proposal correction instruction, the review model reviews the newly generated task proposal after correction, and the reflective optimizer determines whether further correction is needed.
[0056] By repeating the above process, a target task proposal that requires no modification can be obtained, which serves as a task proposal to instruct the target robot to perform the exchange.
[0057] In another embodiment provided in this disclosure, the examination model includes at least one of the following evaluators: a physical feasibility evaluator, a novelty evaluator, and a constraint conformity evaluator; The review model is invoked to review proposals for review tasks in the following manner to obtain review results, including the following steps: Step 1: Call the evaluator respectively to review the task proposals to be reviewed from the corresponding evaluation dimensions and obtain the corresponding review results respectively; the task proposals to be reviewed include the initial task proposals and / or the current task proposals.
[0058] In this embodiment of the disclosure, the physical feasibility evaluator, the novelty evaluator, and the constraint consistency evaluator are all evaluation agents driven by a large model. By using predefined prompts, a corresponding review angle can be defined for each evaluation agent.
[0059] When a task proposal to be reviewed is input into the review model, each evaluator can review the proposal individually from the dimensions defined by that evaluator and derive the corresponding review results. It should be noted that the task to be reviewed here can refer to the initial task proposal or the current task proposal during the revision process, depending on the different implementation methods described above.
[0060] In yet another embodiment provided in this disclosure, the evaluation dimensions of the physical feasibility assessor include: kinematic range, physical interaction logic, and interaction capability requirements; When the evaluator includes a physical feasibility evaluator, step 1 above, "calling the evaluator respectively to review the proposal to be reviewed from the corresponding evaluation dimensions and obtaining the corresponding review results respectively," can be implemented as follows: Step 1: Call the physical feasibility assessor, review the proposal to be reviewed from the assessment dimensions corresponding to the physical feasibility assessor, and obtain the review results of the physical feasibility assessor; Step one above can be implemented as follows: Step 1: Call the physical feasibility evaluator to analyze the kinematic range of the robot in the task proposal to be reviewed, and check whether the kinematic range conflicts with the kinematic constraints of the target robot. Step 2: Analyze the physical interaction logic of the robot in the task proposal to be reviewed, and review whether the physical interaction logic is reasonable; Step 3: Analyze the robot's interaction capability requirements in the task proposal to be reviewed, and examine whether the interaction capability requirements conflict with the hardware capability constraints of the target robot. Step 4: Based on the review results of the kinematic range, physical interaction logic, and interaction capability requirements, integrate them to obtain the review results of the physical feasibility assessor.
[0061] In this embodiment of the disclosure, the physical feasibility evaluator can be defined by prompts as an evaluative agent capable of reviewing the kinematic and dynamic rationality of a task proposal. The prompts for this physical feasibility evaluator may include the following components: kinematic verification, interaction logic analysis, safety and stability assessment, and feedback information.
[0062] Among them, kinematic verification can be used to determine whether the configuration type of the robot in the task proposal has the physical capability to achieve the task skills in the task proposal (e.g., robot accessibility, robot load capacity, and whether there is interference between the robot's actuators (e.g., whether the workspaces of the robotic arms overlap and whether there are singularities, etc.)). In other words, it examines whether the robot's actual capabilities can reach the kinematic range required in the task proposal. Interaction logic analysis can be used to verify whether the task skills used by the robot in the task proposal to the task object conform to the laws of physics (for example, a single-arm robot cannot disassemble a suspended object and requires a gripper or multiple robotic arms to cooperate), that is, to examine the rationality of the robot's physical interaction logic in the task proposal.
[0063] Safety and stability assessments can be used to determine whether a target robot can safely and stably complete a proposed task. Specifically, this can be achieved by verifying whether the target robot's hardware interaction capabilities meet the requirements of the task proposal (e.g., the requirements for precise force control and operational synchronization). It can also be used to verify whether the robot's movement during the task proposal poses a risk of collision with objects in the environment, thereby determining the safety of the task proposal.
[0064] The above information can be summarized and feedback obtained by the physical feasibility assessor. This feedback can be a summary of whether the task proposal is feasible from a kinematic and dynamic perspective, and the reasons why. For example: not feasible; the contact area between cylindrical objects is too small to achieve stable stacking.
[0065] Then, the physical feasibility assessor can integrate the above content into a review result in a specific format based on the constraints of the prompt words, and hand the review result over to the reflective optimizer to perform subsequent steps.
[0066] In another embodiment provided in this disclosure, the evaluation dimensions of the novelty evaluator include: complexity and redundancy; When the evaluator includes a novelty evaluator, step 1 above, "calling the evaluator respectively to examine the proposal to be examined from the corresponding evaluation dimensions and obtaining the corresponding examination results respectively," can be implemented as follows: Step 1: Invoke the novelty evaluator, examine the proposal to be examined from the evaluation dimensions corresponding to the novelty evaluator, and obtain the examination results of the novelty evaluator. This can be implemented as follows: Step 1: Call the novelty evaluator to analyze the robot's interaction process in the task proposal to be reviewed, and review the complexity of the task proposal to be reviewed according to the predetermined complexity indicator criteria; wherein, the complexity indicator criteria include at least one of the following: whether the task proposal includes long-range planning, whether the task proposal includes tool use, or whether the task proposal includes deformable object manipulation. Step 2: Based on the robot's task logic in the task proposal to be reviewed, search the preset task dataset to review the redundancy of the initial task proposal; Step 3: Based on the review results of complexity and redundancy, integrate the review results of the novelty evaluator.
[0067] In this embodiment of the disclosure, the novelty evaluator can be defined by prompts as an evaluation agent capable of examining the diversity of the dataset of the task proposal itself. The prompts for this novelty evaluator may include the following components: complexity analysis, redundancy checks, and feedback information.
[0068] Complexity analysis is used to evaluate the richness of the task proposal's interactions, such as whether it involves the use of task objects, whether it involves the manipulation of deformable task objects, and the overall length of the task proposal.
[0069] Redundancy checks are used to analyze whether the logic of the task proposal is a simple, repetitive design or a routine operation of the robot. This can be achieved by using a large model driving the novelty evaluator to search a pre-defined task dataset containing historical task proposals for the robot, thereby determining whether the current task proposal is a redundant proposal with repetitive logic.
[0070] The above information can be summarized by the novelty evaluator to obtain feedback. This feedback can be based on the above summary, indicating whether the diversity of the task proposal is sufficiently rich, and the reasons for such richness. For example: high novelty; the proposal introduces complex rotational primitives and rare contact mechanics mechanisms.
[0071] Furthermore, the novelty evaluator can integrate the above content into a specific format of review results based on the constraints of the prompt words, and then hand the review results over to the reflective optimizer for subsequent steps.
[0072] In yet another embodiment provided in this disclosure, the evaluation dimensions of the constraint consistency evaluator include: content authenticity and scenario logic adaptability; When the evaluator includes a constraint consistency evaluator, step 1 above, "calling the evaluator respectively to review the proposal to be reviewed from the corresponding evaluation dimension and obtaining the corresponding review results respectively," can be implemented as follows: Step 1: Invoke the constraint consistency evaluator, review the proposal to be reviewed from the evaluation dimensions corresponding to the constraint consistency evaluator, and obtain the review results of the constraint consistency evaluator. This can be implemented as follows: Step 1: Call the constraint consistency evaluator to check whether the task description items involved in the task proposal to be reviewed are included in the task description data group, and review the authenticity of the content of the task proposal to be reviewed. Step 2: Call the constraint consistency evaluator to detect the logical relationships between the task description items involved in the task proposal to be reviewed, and review the scenario logic adaptability of the task proposal to be reviewed. Step 3: Based on the review results of the content authenticity and scenario logic adaptability, integrate them to obtain the review results of the constraint consistency evaluator.
[0073] In this embodiment of the disclosure, the constraint consistency evaluator can be defined by prompt words as an evaluative agent capable of reviewing the content boundaries of task proposals. The prompt words for this constraint consistency evaluator may include the following components: illusion check, scene logic check, and feedback information.
[0074] The illusion check is used to check whether the task description items involved in the task proposal are included in the task description data group through string matching, and to identify the task description items and fictitious objects that exist in the task description data group, so as to ensure the authenticity of the generated task proposal.
[0075] Scene logic checks are used to ensure that the generated task proposals conform to the logic of the scene category (e.g., welding operations do not occur in a bedroom scene) to review the suitability of the task proposals for the scene logic of the target scene category.
[0076] The above content can be evaluated by the constraint consistency evaluator to obtain feedback information. This feedback information can be a summary of the above content, indicating whether there are problems with the content boundaries of the task proposal, and the corresponding reasons. For example: consistency error; the object "wrench" was detected as not being in the current task description data group, and it is recommended to replace it with "screwdriver".
[0077] In another embodiment provided in this disclosure, the step S102 above, "sampling data for the preset task description item based on historical task data corresponding to the description data of the preset task description item in the task space tuple," can be implemented as follows: Step 1: Based on the historical count information of the category to which the description data of the preset task description item in the task space tuple belongs, perform data sampling on the preset task description item; Robot task generation methods also include: Modify the corresponding counter value for the description data category corresponding to the preset task description item involved in the target task proposal.
[0078] In this embodiment of the disclosure, if the target task proposal passes the review, the task proposal can be added to the robot's task dataset. Furthermore, based on the task description items involved in the target task proposal, the corresponding counter value for the description data category in the historical task data is changed.
[0079] For example, if a task description item is mentioned in the target task proposal, the counter value for the number of times that task description item is used can be incremented by one, and the frequency of use of that task description item in the same dimension of task description items can be changed. The time when the target task proposal is added to the task dataset can also be recorded as the timestamp of the activation of the task description items involved, so as to characterize the temporal characteristics of each task description item. After the target robot executes the task proposal or executes the task proposal in the simulation space, the counter value for the number of successes or failures can be incremented for each task description item according to the success or failure of the execution, so as to characterize the execution risk of each task description item. According to the number of times different task description items appear in the task proposal at the same time, the correlation between task description items can be counted to measure the coupling relationship between task description items.
[0080] In yet another embodiment provided in this disclosure, the historical interaction record includes: a pair of task statements and feedback statements in a historical task proposal; Robot task generation methods also include: Upon obtaining the runtime feedback data corresponding to the target task proposal, the feedback statements in the runtime feedback data are matched with the task statements in the corresponding target task proposal and stored in the historical interaction record.
[0081] In this embodiment, the target robot can execute a task proposal in a real or simulated space. Experts can provide corresponding natural language feedback based on the execution of the task proposal. Then, a large language model can be invoked to summarize the expert's feedback and map the content of the feedback to different task statements in the domain task proposal, obtaining different correspondences. This forms multiple memory key-value pairs, which are stored in a pre-defined memory module as historical interaction records for the reflective optimizer to query.
[0082] It should be noted that the aforementioned historical interaction records can be combined with the task dataset. The memory module can store different task statements and corresponding feedback statements of executed task proposals in the form of memory key-value pairs.
[0083] In another embodiment provided in this disclosure, the robot task generation method further includes: Obtain the image data related to the task description data group; The task description data set is provided as constraints to the proposal generator to generate a target task proposal, including: The task description data set, along with the image data, is provided to the proposal generator to generate a target task proposal.
[0084] In this disclosure, if the generated task proposal involves movement or requires spatial awareness, the generation model in the proposal generator can be a visual language model or other large model with multimodal capabilities.
[0085] Pre-prepared image data related to each task description item in the task description data set can be provided as constraints to the generative model. The generative model then generates corresponding task proposals based on this image data.
[0086] For example, image data for scene categories can be a collection of multiple images showing the actual situation of the target scene category. This data is input into a visual language model, which then grounds the images into text descriptions to help the generation model understand the actual situation of the target scene category. Image data for task objects can be multi-angle images of the task object to help the task proposal better guide the robot to interact with the task object. Image data for task skills can be images of the continuous operation process of the task skill.
[0087] The following is an example illustration. Figure 2 This is a schematic diagram of the structure of a robot task generation system provided in an embodiment of the present disclosure, based on... Figure 2 The provided task generation system offers a robot task generation method based on diversity-driven and self-reflective mechanisms. Its execution flow mainly includes system initialization and definition, diversity context sampling, initial task proposal generation, multi-dimensional self-reflective evaluation, memory-enhanced task optimization, and closed-loop feedback update. Specifically, it includes: Step S100: System initialization and task space definition: Before task generation begins, the robot task space tuple is first formally defined, as described above. Initialize historical statistics (i.e., historical task data), including scene counters, object counters, and skill counters, all with initial values set to zero.
[0088] Step S200: Diversity-driven hierarchical sampling: To address the long-tail effect of data distribution, this step employs a hierarchical LFU strategy to construct the context constraints for task generation: Scene-level sampling: The system queries historical statistics in real time and selects the scenario category with the lowest current global usage frequency. As the background for this generation, the formula is as follows:
[0089] in, A scene counter is used to represent the training data. This ensures a balanced distribution of training data across different environments, such as expanding from common kitchen scenes to rare medical scenes.
[0090] Context-Aware Asset Sampling: Semantic filtering: based on selected scenarios First, the semantic similarity between each object in the object library and the scene is calculated, and a subset that conforms to the logic of the scene is filtered out (for example, excluding unreasonable objects such as "shampoo" in the "industrial" scene).
[0091] LFU sampling: In the filtered subset and skill subset, the LFU strategy is applied again to prioritize sampling the k objects with the lowest historical usage to form a candidate object set. and candidate skill sets The final output is the generated strong constraints: .
[0092] Step S300: Initial task proposal generation: Input the constraints obtained in step S200 into the proposal generator.
[0093] Generative model: Employs either a Large Language Model (LLM) or a Visual Language Model (VLM). When motion manipulation or spatial awareness is required, the VLM receives scene images as auxiliary input to achieve visual grounding.
[0094] Prompt word engineering: The generator receives input containing "role definition (robotics expert)" and "constraint injection (must be used)". , The structure of the "elements in the document" and the "format specification" (Prompt).
[0095] Output: Generate a standardized JSON format initial task. It includes the following fields: task name, language instruction, object name and position parameters, skill sequence, scene layout, and task context description.
[0096] Step S400: Multidimensional Self-Reflection Assessment To ensure the generated The system is executable and has training value. It launches three specialized evaluation agents in parallel to review and generate critical feedback in natural language.
[0097] Physical feasibility assessor ( ): Kinematics test: Check whether the task exceeds the robot's kinematic range (e.g., in a two-arm collaborative task, whether the workspaces of the two arms overlap or whether there are singularities).
[0098] Physical logic verification: Evaluate the rationality of the interaction (e.g., "A single arm cannot unscrew a suspended bottle cap without the assistance of a clamp").
[0099] Synchronization and Force Control: Check for precise force control or millisecond-level synchronization requirements that exceed the hardware's capabilities.
[0100] Output feedback For example, "Not feasible: The contact area between cylindrical objects is too small to allow for stable stacking."
[0101] Novelty evaluator ( ): Complexity analysis: Evaluate whether the task involves long-range planning, tool use, or manipulation of deformable objects, rather than a simple "grab-place" task.
[0102] Redundancy check: Determine whether the task logic is highly repetitive with the existing dataset.
[0103] Output feedback For example, “High novelty: It introduces complex rotational primitives and rare contact mechanics mechanisms.”
[0104] Constraint Consistency Evaluator ): Hallucination detection: Employs a rigorous string matching algorithm to verify whether the objects and skills referenced in the task description strictly exist in the sample set. and middle.
[0105] Scene logic verification: Ensure that the task description does not violate common sense in the scene (e.g., do not perform "welding" in the "bedroom").
[0106] Output feedback For example, "Consistency error: Object 'wrench' was not detected in the current candidate list". In the middle, it is recommended to replace it with 'screwdriver'.
[0107] Step S500: Memory Enhancement and Task Refinement This step uses a self-reflection refiner to refine the task and generate the final task. .
[0108] Retrieval Enhancement (RAG): The system maintains a long-term memory module M to store past human-computer interaction feedback. Each memory entry exists in the form of a key-value pair (K, V), where K is the semantic vector embedding of the task context, and V is the specific reason for failure or a suggested correction (e.g., "Drawer handle operation should be forced to use the left arm due to visual obstruction"). System calculation Given the embedding vectors, retrieve the top-k heuristic rules with the highest similarity from M. .
[0109] Comprehensive Correction: The reflective optimizer combines the following three types of information for reasoning: Original mission proposal ; Critique of the Three Evaluators ; Historical experience retrieved .
[0110] By optimizing the algorithm Generate the revised task ,Right now:
[0111] This process can fix incorrect object names, adjust spatial layouts to avoid collisions, or add intermediate steps to conform to physical constraints.
[0112] Step S600: Closed-loop update and human-machine loopback: Data import: If Once the final validity check is passed, it is officially added to the task dataset, and the counts of related scenes, objects and skills in the historical statistics are increased accordingly, thereby dynamically affecting the LFU sampling weights for the next time.
[0113] Memory Evolution: When the generated task is executed in a real robot or simulation environment, if a failure occurs or human intervention is required (Human-in-the-Loop), the system uses LLM to summarize the natural language feedback provided by experts, forms a new (K,V) pair and stores it in the memory module M, so as to realize the continuous accumulation and evolution of the system's cognition of the physical world.
[0114] Compared with the prior art, the technical solution of the present invention has the following significant advantages: First, it significantly improves the diversity and coverage of training data. Existing methods (e.g., manually designed task proposals or generation through large language models) often suffer from cognitive biases, resulting in common objects in the dataset accounting for the majority of the tasks. This invention forces the system to explore the long-tail space by employing a hierarchical least frequently used sampling strategy (LFU). Experimental data show that the number of unique task objects generated by this system exceeds that generated by human experts and large models (e.g., GPT-4o), and it improves task skill coverage, achieving decentralization and equalization of data distribution.
[0115] Secondly, it significantly enhances the physical feasibility and accuracy of the generated tasks. Addressing the "illusion" problem commonly found in large generative models, the multi-dimensional self-reflective mechanism of this invention (i.e., a combination of an evaluator and a reflective optimizer) acts as a "gatekeeper." Experiments show that the tasks generated by this invention outperform those generated directly using large models (e.g., Gemini 2.5-Pro) in terms of physical feasibility, logical validity, and object consistency, ensuring that the generated task proposals can be used for robot training without extensive manual cleaning.
[0116] Third, it endows the model with excellent zero-shot generalization ability. Thanks to the high diversity and physical realism of the generated task proposals, if the dataset generated by this invention is used to pre-train the Vision-Language-Action (VLA) model, it can demonstrate excellent robustness when facing unseen real-world challenges. In generalization tests that include changes in lighting, background interference, and changes in new objects and instructions, the model pre-trained on the dataset of this invention has a significantly higher average success rate than the baseline model (for example, in lighting change scenarios, the success rate is about 3.5 times that of the Gemini2.5-Pro baseline), demonstrating the core value of this invention in building general embodied agents.
[0117] 4. Possesses experience-based continuous self-evolution capability: The invention features a specially designed memory module that enables the system to learn from historical interaction records. This not only resolves errors in the current task proposal but also, through accumulated physical knowledge (such as grasping strategies for specific objects and handling occlusion), allows the quality and success rate of generated task proposals to spiral upward over time. This is a characteristic that traditional open-loop generation methods do not possess.
[0118] Furthermore, there is currently no single alternative that can simultaneously address the issues of insufficient diversity and poor physical feasibility in existing task proposals. While it is possible to directly generate tasks using general-purpose large models with higher parameter counts (e.g., GPT-4o or Gemini 2.5Pro), this only partially improves the coherence of the task proposal text logic and cannot fundamentally solve the problem of physical infeasibility caused by the lack of robot ontology perception, nor can it proactively balance the data distribution. In addition, using pure random sampling instead of the LFU strategy will result in the generation of a large number of semantically incoherent and meaningless tasks (such as "cook fish in the bedroom"), reducing data validity.
[0119] The specific analysis is as follows: One approach is to directly utilize general-purpose large models (e.g., Large Language Models (LLMs) or Vision Language Models (VLMs)) for generation. This involves using current general-purpose base models (such as GPT-4o and Gemini 2.5Pro) and carefully designed prompts, directly inputting a scene definition and a list of objects, requiring the model to generate a robot task proposal. While general-purpose large models possess extremely strong language understanding and generation capabilities, directly applying them to the robotics field has the following insurmountable drawbacks.
[0120] One drawback is a severe asset illusion; the general-purpose large model lacks "embodied perception" of the robot's actual working environment. Experiments show that even when provided with a list of objects, the model tends to generate objects or properties that do not exist in the actual environment, causing the generated task proposals to be unenforceable in the physical world.
[0121] The second drawback is the distribution bias and long-tail problem. General-purpose large models are limited by the distribution of pre-training data and exhibit a strong scene bias. For example, the tasks generated by general-purpose large models are concentrated in home, kitchen and office scenarios, and rarely involve industrial, medical or laboratory scenarios; at the same time, the generated skills are highly concentrated in common "grab-place" operations, making it difficult to cover long-tail and complex combined skills.
[0122] The third defect is the lack of physical constraints. The general large model cannot perceive the robot's kinematic limitations (such as the difference between single-arm / dual-arm configurations and the range of workspace), and often generates tasks that exceed the robot's dynamic capabilities or violate physical laws (such as requiring a single-arm robot to complete a cover removal operation that requires the cooperation of two arms).
[0123] Secondly, the method involves manual design and data collection by human experts, who design the task flow and demonstrate its remote operation. This is currently the main source of robot datasets (such as OpenX-Embodiment). While manual design ensures high physical feasibility, it suffers from severe scalability bottlenecks and cognitive biases. This method also has the following insurmountable drawbacks.
[0124] One drawback is that cognitive bias leads to a lack of diversity. Human designers tend to design simple, repetitive atomic tasks (such as "picking up an apple"), often neglecting complex, multi-stage tasks. Data shows that in manually designed datasets, simple task skills account for a high proportion, and the use of task objects exhibits an extreme long-tail distribution, with a few common objects making up the vast majority of the samples.
[0125] The second drawback is the high cost and lack of scalability. The cost of collecting robot data in the physical world is far higher than that of text or images. Relying entirely on manual design cannot meet the demand for massive and diverse data for VLA model pre-training.
[0126] In summary, existing technologies are prone to falling into illusion traps, are hampered by scalability bottlenecks, or lack task logic. The method provided in this disclosure, through its unique closed-loop system of "diversity-driven sampling (i.e., the data sampling method adopted in this disclosure) – multi-dimensional self-reflection mechanism (i.e., the combination of the aforementioned review model and reflective optimizer) – human-machine loop long-term memory (i.e., the memory module and corresponding historical interaction record retrieval method provided in this disclosure)," is currently the only solution capable of simultaneously achieving high semantic logic, rigorous physical feasibility, and global statistical distribution balance in the automated task proposal generation process. Therefore, there is no existing solution that can completely replace the technical effects of this invention.
[0127] This disclosure also provides a robot task generation system, such as Figure 3 As shown, it includes: Module 201 is used to determine the task space tuple used to describe the robot task; The task space tuple includes multiple task description items; The sampling module 202 is used to sample the preset task description items based on historical task data corresponding to the description data of the preset task description items in the task space tuple, so as to obtain a task description data group that contains at least the sampled data corresponding to the preset task description items. The generation module 203 is used to provide the task description data group as constraints to the proposal generator to generate a target task proposal.
[0128] In another embodiment provided in this disclosure, the sampling module 202 is used to sample the preset task description item based on the usage frequency of the description data corresponding to the preset task description item in the task space tuple in historical task data, to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item; the usage frequency satisfies: lower than a preset frequency.
[0129] In another embodiment provided in this disclosure, the sampling module 202 is configured to: determine the scene categories whose usage frequency is lower than a preset first threshold as target scene categories based on the frequency of robot task execution in each scene category in the scene category set; select task objects logically associated with the target scene category from the task object library based on the semantics of the target scene category to obtain a task object set; select task skills logically associated with the target scene category from the task skill library based on the semantics of the target scene category to obtain a task skill set; filter task objects whose usage frequency is within a first preset low frequency order range from the task object set based on the historical interaction frequency between the robot and the task objects to obtain a candidate task object set; filter task skills whose usage frequency is within a second preset low frequency order range from the task skill set based on the historical usage frequency of the robot for the task skills to obtain a candidate task skill set; and determine the target scene category, the candidate task object set, and the candidate task skill set as the task description data group; the preset task description item includes: a scene category set representing the robot's usage scenario, a task object library representing the robot's interaction objects, and a task skill library representing the robot's interaction methods.
[0130] In another embodiment provided in this disclosure, the generation module 203 is configured to call a generation model to determine prompt words based on the task description data set and generate an initial task proposal for instructing the target robot to interact; call a review model to review the initial task proposal and obtain a review result; call a reflective optimizer to determine whether the initial task proposal needs to be corrected based on the retrieved historical interaction records, the review result, and the initial task proposal; if correction is required, generate a task proposal correction instruction to correct the initial task proposal and obtain the target task proposal.
[0131] In another embodiment provided in this disclosure, the generation module 203 is used to call the generation model, determine prompt words based on the task description data set, and generate an initial task proposal to instruct the target robot to interact; the initial task proposal is used as the current task proposal, and the following proposal review steps are performed: The review model is invoked to review the current task proposal and obtain the review result. The reflective optimizer is invoked to determine whether the current task proposal needs to be revised based on the retrieved historical interaction records, the current review result, and the current task proposal. If the current task proposal needs to be revised, a task proposal revision instruction is generated to revise the current task proposal and obtain the revised task proposal. The revised task proposal is used as the new current task proposal, and the process returns to execute the proposal review steps until it is determined that the current task proposal does not need to be revised, thus obtaining the target task proposal.
[0132] In another embodiment provided in this disclosure, the examination model includes at least one of the following evaluators: a physical feasibility evaluator, a novelty evaluator, and a constraint conformity evaluator; The generation module 203 is used to call the evaluator respectively to review the task proposals to be reviewed from the corresponding evaluation dimensions and obtain the corresponding review results respectively; the task proposals to be reviewed include the initial task proposals and / or the current task proposals.
[0133] In yet another embodiment provided in this disclosure, the evaluation dimensions of the physical feasibility assessor include: kinematic range, physical interaction logic, and interaction capability requirements; The generation module 203 is used to call the physical feasibility assessor, review the proposal to be reviewed from the assessment dimension corresponding to the physical feasibility assessor, and obtain the review result of the physical feasibility assessor. The generation module 203 is used to call the physical feasibility assessor to analyze the kinematic range of the robot in the task proposal to be reviewed, and to review whether the kinematic range conflicts with the kinematic constraints of the target robot; to analyze the physical interaction logic of the robot in the task proposal to be reviewed, and to review whether the physical interaction logic is reasonable; to analyze the interaction capability requirements of the robot in the task proposal to be reviewed, and to review whether the interaction capability requirements conflict with the hardware capability constraints of the target robot; and to integrate the review results of the physical feasibility assessor based on the review results of the kinematic range, physical interaction logic, and interaction capability requirements.
[0134] In another embodiment provided in this disclosure, the evaluation dimensions of the novelty evaluator include: complexity and redundancy; The generation module 203 is used to call the novelty evaluator, examine the proposal to be examined from the evaluation dimension corresponding to the novelty evaluator, and obtain the examination result of the novelty evaluator. The generation module 203 is used to invoke the novelty evaluator to analyze the robot's interaction process in the task proposal to be reviewed, and to review the complexity of the task proposal to be reviewed according to a predetermined complexity indicator standard; wherein, the complexity indicator standard includes at least one of the following: whether the task proposal includes long-range planning, whether the task proposal includes tool use, or whether the task proposal includes deformable object operation; to search in a preset task dataset according to the robot's task logic in the task proposal to be reviewed, and to review the redundancy of the initial task proposal; and to integrate the review results of the novelty evaluator based on the review results of complexity and redundancy.
[0135] In yet another embodiment provided in this disclosure, the evaluation dimensions of the constraint consistency evaluator include: content authenticity and scenario logic adaptability; The generation module 203 is used to call the constraint consistency evaluator, review the proposal to be reviewed from the evaluation dimension corresponding to the constraint consistency evaluator, and obtain the review result of the constraint consistency evaluator. The generation module 203 is used to call the constraint consistency evaluator to detect whether the task description items involved in the task proposal to be reviewed are included in the task description data group, and to review the authenticity of the content of the task proposal to be reviewed; to call the constraint consistency evaluator to detect the logical relationship between the task description items involved in the task proposal to be reviewed, and to review the scenario logic adaptability of the task proposal to be reviewed; and to integrate the review results of the constraint consistency evaluator based on the review results of the content authenticity and scenario logic adaptability.
[0136] In another embodiment provided in this disclosure, the sampling module 202 is used to sample data of the preset task description item based on the historical count information of the category to which the description data of the preset task description item in the task space tuple belongs; Robot task generation system, such as Figure 4 As shown, it also includes: The counting module 204 is used to modify the corresponding counter value for the description data category corresponding to the preset task description item involved in the target task proposal.
[0137] In yet another embodiment provided in this disclosure, the historical interaction record includes: a pair of task statements and feedback statements in a historical task proposal; Robot task generation system, such as Figure 5 As shown, it also includes: The storage module 205 is used to, when obtaining the running feedback data corresponding to the target task proposal, match the feedback statement in the running feedback data with the task statement in the corresponding target task proposal and store it in the historical interaction record.
[0138] In yet another embodiment provided in this disclosure, the robot task generation system, such as Figure 6 As shown, it also includes: Image data processing module 206 is used to acquire image data related to the task description data group; and to provide the task description data group as a constraint condition to the proposal generator to generate a target task proposal. The image data processing module 206 is used to provide the task description data set as a constraint and the image data to the proposal generator to generate a target task proposal.
[0139] This disclosure also provides a computer device, including a processor, a memory, and a bus; The memory stores machine-readable instructions that can be executed by the processor; When the computer device is running, the processor and the memory communicate via a bus; When the machine-readable instructions are executed by the processor, they perform the steps of a robot task generation method as described in any of the above embodiments.
[0140] Through the above description of the embodiments, those skilled in the art can clearly understand that the embodiments of this disclosure can be implemented in hardware or by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0141] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes in the drawings are not necessarily essential for implementing this disclosure.
[0142] Those skilled in the art will understand that the modules in the system of the embodiments can be distributed in the system of the embodiments as described in the embodiments, or they can be located in one or more systems different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.
[0143] The sequence numbers of the embodiments disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0144] Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from its spirit and scope. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.
Claims
1. A method of robot task generation, the method comprising: include: Determine the task space tuples used to describe the robot's tasks; The task space tuple includes multiple task description items; Based on the frequency of use of the description data corresponding to the preset task description item in the task space tuple in the historical task data, the preset task description item is sampled to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item. The task description data set is provided as a constraint to the proposal generator to generate a target task proposal.
2. The method as described in claim 1, characterized in that, The usage frequency meets the requirement of being lower than a preset frequency.
3. The method as described in claim 1, characterized in that, The preset task description items include: a set of scenario categories representing robot usage scenarios, a task object library representing robot interaction objects, and a task skill library representing robot interaction methods; Based on the frequency of use of the description data corresponding to the preset task description items in the task space tuple in historical task data, data sampling is performed on the preset task description items to obtain a task description data group that at least contains the sampled data corresponding to the preset task description items, including: Based on the frequency of the robot performing tasks in each scene category in the scene category set, scene categories with a usage frequency lower than a preset first threshold are identified as target scene categories; Based on the semantics of the target scene category, task objects logically associated with the target scene category are selected from the task object library to obtain a task object set; Based on the semantics of the target scene category, task skills logically associated with the target scene category are selected from the task skill library to obtain a task skill set; Based on the historical interaction frequency between the robot and the task object, task objects whose usage frequency is within the first preset low frequency order range are selected from the task object set to obtain a candidate task object set. Based on the robot’s historical usage frequency of task skills, task skills whose usage frequency falls within the second preset low frequency order range are selected from the task skill set to obtain a candidate task skill set. The target scene category, the candidate task object set, and the candidate task skill set are determined as the task description data group.
4. The method as described in claim 1, characterized in that, The step of providing the task description data set as constraints to the proposal generator to generate a target task proposal includes: The generative model is invoked to determine prompt words based on the task description data set, and an initial task proposal is generated to instruct the target robot to interact. The initial task proposal is reviewed using the review model to obtain the review results; The reflective optimizer is invoked to determine whether the initial task proposal needs to be revised based on the retrieved historical interaction records, the review results, and the initial task proposal. If corrections are needed, a task proposal correction instruction is generated to correct the initial task proposal and obtain the target task proposal.
5. The method as described in claim 1, characterized in that, The step of providing the task description data set as constraints to the proposal generator to generate a target task proposal includes: The generative model is invoked to determine prompt words based on the task description data set, and an initial task proposal is generated to instruct the target robot to interact. The initial task proposal is used as the current task proposal, and the following proposal review steps are performed: The review model is invoked to review the current task proposal and the review results are obtained. The reflective optimizer is invoked to determine whether the current task proposal needs to be revised based on the retrieved historical interaction records, current review results, and current task proposal. If the current task proposal needs to be modified, generate a task proposal modification instruction to modify the current task proposal and obtain the modified task proposal. The revised task proposal is used as the new current task proposal. The process returns to the proposal review step until it is determined that the current task proposal does not need to be revised, thus obtaining the target task proposal.
6. The method as described in claim 4 or 5, characterized in that, The review model includes at least one of the following evaluators: physical feasibility evaluator, novelty evaluator, and constraint conformity evaluator; The review model is invoked in the following manner to review the proposals for the review tasks, and the review results are obtained, including: The evaluators are invoked respectively to review the task proposals to be reviewed from the corresponding evaluation dimensions, and the corresponding review results are obtained respectively; the task proposals to be reviewed include initial task proposals or current task proposals.
7. The method as described in claim 6, characterized in that, The evaluation dimensions of the physical feasibility assessor include: kinematic range, physical interaction logic, and interaction capability requirements; When the evaluator includes a physical feasibility evaluator, the step of respectively invoking the evaluator to review the proposal to be reviewed from the corresponding evaluation dimensions, and obtaining the corresponding review results, includes: The physical feasibility assessor is invoked to review the proposed task from the corresponding assessment dimensions, and the review results from the physical feasibility assessor are obtained, including: The physical feasibility assessor is invoked to analyze the kinematic range of the robot in the task proposal to be reviewed, and to examine whether the kinematic range conflicts with the kinematic constraints of the target robot. Analyze the physical interaction logic of the robot in the task proposal to be reviewed, and review whether the physical interaction logic is reasonable; Analyze the robot's interaction capability requirements in the proposed task to be reviewed, and examine whether the interaction capability requirements conflict with the hardware capability constraints of the target robot; Based on the review results of kinematic range, physical interaction logic, and interaction capability requirements, the review results of the physical feasibility assessor are integrated.
8. The method as described in claim 6, characterized in that, The evaluation dimensions of the novelty evaluator include: complexity and redundancy; When the evaluator includes a novelty evaluator, the step of respectively invoking the evaluator to examine the proposal to be examined from the corresponding evaluation dimensions and obtaining the corresponding examination results includes: The novelty evaluator is invoked to examine the proposal to be examined from the evaluation dimensions corresponding to the novelty evaluator, and the examination results of the novelty evaluator are obtained, including: The novelty evaluator is invoked to analyze the robot's interaction process in the task proposal to be reviewed, and the complexity of the task proposal to be reviewed is examined according to a predetermined complexity indicator standard; wherein, the complexity indicator standard includes at least one of the following: whether the task proposal includes long-range planning, whether the task proposal includes tool use, or whether the task proposal includes deformable object manipulation. The redundancy of the task proposals to be reviewed is examined by searching the preset task dataset based on the robot's task logic in the proposals to be reviewed. Based on the review results of complexity and redundancy, the review results of the novelty evaluator are integrated.
9. The method as described in claim 6, characterized in that, The evaluation dimensions of the constraint consistency evaluator include: content authenticity and scenario logic adaptability; When the evaluator includes a constraint consistency evaluator, the step of respectively invoking the evaluator to review the proposal to be reviewed from the corresponding evaluation dimensions, and obtaining the corresponding review results, includes: The constraint consistency evaluator is invoked to review the proposal to be reviewed from the evaluation dimensions corresponding to the constraint consistency evaluator, and the review results of the constraint consistency evaluator are obtained, including: Invoke the constraint consistency evaluator to check whether the task description items involved in the task proposal to be reviewed are included in the task description data group, and review the authenticity of the content of the task proposal to be reviewed. Invoke the constraint consistency evaluator to detect the logical relationships between the task description items involved in the task proposal to be reviewed, and review the scenario logic adaptability of the task proposal to be reviewed. Based on the review results of content authenticity and scenario logic adaptability, the review results of the constraint consistency evaluator are integrated.
10. The method as described in claim 1, characterized in that, Also includes: Based on the historical count information of the category to which the description data of the preset task description item in the task space tuple belongs, data sampling is performed on the preset task description item; Modify the corresponding counter value for the description data category corresponding to the preset task description item involved in the target task proposal.
11. The method as described in claim 4 or 5, characterized in that, The historical interaction record includes: task statements and feedback statement pairs in historical task proposals; Also includes: Upon obtaining the runtime feedback data corresponding to the target task proposal, the feedback statements in the runtime feedback data are matched with the task statements in the corresponding target task proposal and stored in the historical interaction record.
12. The method as described in claim 1, characterized in that, Also includes: Obtain the image data related to the task description data group; The task description data set is provided as constraints to the proposal generator to generate a target task proposal, including: The task description data set, along with the image data, is provided to the proposal generator to generate a target task proposal.
13. A robot task generation system, characterized in that, include: The determination module is used to determine the task space tuple used to describe the robot's task; The task space tuple includes multiple task description items; The sampling module is used to sample the preset task description item based on the frequency of use of the description data corresponding to the preset task description item in the task space tuple in historical task data, and to obtain a task description data group that at least contains the sampled data corresponding to the preset task description item. The generation module is used to provide the task description data set as constraints to the proposal generator to generate a target task proposal.
14. A computer device, characterized in that, Includes processor, memory, and bus; The memory stores machine-readable instructions that can be executed by the processor; When the computer device is running, the processor and the memory communicate via a bus; When the machine-readable instructions are executed by the processor, they perform the steps of a robot task generation method as described in any one of claims 1-12.
Citation Information
Patent Citations
Robot simulation task generation method, medium and equipment
CN121009708A