Action correction method and device, equipment and medium
By acquiring task descriptions and observation data, and utilizing motion prediction, evaluation, adjustment, and correction strategies to generate refined target motion sequences, the problem of low processing efficiency in robot self-correction technology is solved, achieving efficient motion correction and environmental adaptability.
Patent Information
- Application Number
- CN202511047420.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-14
AI Technical Summary
Existing robots suffer from semantic reflection and action execution gaps in self-correction technology, lack of fine action correction capabilities, and contradictions between data dependence and real-time requirements, resulting in low processing efficiency, insufficient action correction accuracy, and inadequate dynamic environment adaptability and lifelong learning capabilities.
By acquiring task descriptions and observation data, a motion prediction strategy is used to generate initial action instructions, a motion evaluation strategy is used to evaluate the action execution results, an action adjustment strategy is used to adjust failed actions, and an action correction strategy is used to generate a target action sequence, thus establishing a bridge between high-level semantic understanding and low-level action execution.
It improves the robot's autonomous decision-making and continuous optimization capabilities in dynamic and complex environments, significantly enhances processing efficiency and motion correction accuracy, and adapts to a variety of tasks and scenarios.
Smart Images

Figure CN120941378A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a motion correction method, apparatus, device, and medium. Background Technology
[0002] Currently, existing robots in the financial and medical fields suffer from several shortcomings in self-correction technology. These include a gap between semantic reflection and action execution, a lack of fine-grained action correction capabilities, and a conflict between data dependence and real-time requirements. These shortcomings lead to difficulties in handling untrained tasks, poor generalization in long-term operations, and insufficient accuracy in action correction. Furthermore, they exhibit significant weaknesses in adaptability to dynamic environments and lifelong learning capabilities, resulting in low processing efficiency.
[0003] Therefore, existing motion correction methods suffer from low processing efficiency. Summary of the Invention
[0004] This invention provides a motion correction method, apparatus, device, and medium, aiming to solve the problem of low processing efficiency in existing motion correction methods.
[0005] To address the aforementioned problems, in a first aspect, embodiments of the present invention provide a motion correction method applied to a robot, comprising:
[0006] Obtain task descriptions and observation data;
[0007] Initial action instructions are obtained by predicting the task description and the observation data based on a motion prediction strategy.
[0008] After each motion instruction in the initial motion instruction is executed, an evaluation result is obtained by evaluating each motion instruction according to the motion evaluation strategy.
[0009] When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain an intermediate action instruction;
[0010] The intermediate action instructions are corrected based on the action correction strategy to obtain the target action sequence.
[0011] Secondly, embodiments of this application provide a motion correction device applied to a robot, comprising:
[0012] The acquisition unit is used to acquire task descriptions and observation data;
[0013] The prediction unit is used to perform prediction processing on the task description and the observation data based on the motion prediction strategy to obtain the initial action command;
[0014] An evaluation unit is used to evaluate each motion instruction according to a motion evaluation strategy after each motion instruction in the initial motion instruction is executed, and obtain an evaluation result.
[0015] An adjustment unit is used to adjust each action instruction based on an action adjustment strategy to obtain an intermediate action instruction when the evaluation result is an execution failure.
[0016] The correction unit is used to correct the intermediate action instructions based on the action correction strategy to obtain the target action sequence.
[0017] Thirdly, embodiments of this application provide a computer device, the computer device including a memory and a processor connected to the memory; the memory is used to store a computer program, and the processor is used to run the computer program stored in the memory to perform the method described in the first aspect above.
[0018] Fourthly, embodiments of this application provide a storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, implement the method described in the first aspect above.
[0019] This invention provides a motion correction method, apparatus, device, and medium for robots. The method includes: acquiring task description and observation data; performing predictive processing on the task description and observation data based on a motion prediction strategy to obtain initial motion instructions; after each motion instruction in the initial motion instructions is executed, evaluating each motion instruction according to a motion evaluation strategy to obtain an evaluation result; when the evaluation result is an execution failure, adjusting each motion instruction based on a motion adjustment strategy to obtain intermediate motion instructions; and correcting the intermediate motion instructions based on a motion correction strategy to obtain a target motion sequence. Therefore, this invention obtains a refined target motion sequence by predicting, evaluating, adjusting, and correcting task description and observation data, thereby improving autonomous decision-making and continuous optimization capabilities in dynamic and complex environments, and thus improving processing efficiency. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A schematic flowchart of the motion correction method provided in an embodiment of the present invention;
[0022] Figure 2 This is a schematic block diagram of a motion correction device provided in an embodiment of the present invention;
[0023] Figure 3 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0027] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating the motion correction method provided in an embodiment of the present invention. Figure 1 As shown, this embodiment of the invention provides a motion correction method applied to a robot, which includes the following steps S110-S150.
[0029] S110. Obtain task description and observation data.
[0030] In this embodiment, the task description clearly tells the robot the task to be performed, such as "If you were a robot, how would you make coffee?", providing target guidance for the robot's subsequent operations; the observation data is relevant information obtained by the robot through its own sensors, camera equipment, etc., such as the surrounding environment and its own state, which may include data such as the position of objects around the robotic arm and the angle of its own joints, and is the basis for subsequent action decisions.
[0031] The application scenarios of this invention can be in the financial, medical, and industrial fields; for example, it can be used for business services in the financial field, for health care services in the medical field, and for manufacturing in the industrial field.
[0032] As can be seen from the above embodiments, obtaining task descriptions and observation data, and then performing subsequent targeted processing based on these data, improves the accuracy of subsequent data and processing efficiency.
[0033] S120. Based on the motion prediction strategy, the task description and the observation data are processed to obtain the initial action command.
[0034] In this embodiment, after obtaining the task description and the observation data, the initial action command is obtained by performing prediction processing on the task description and the observation data based on a motion prediction strategy.
[0035] In one embodiment, the step of performing predictive processing on the task description and the observation data based on a motion prediction strategy to obtain initial action instructions includes:
[0036] Obtain preset sample data;
[0037] The initial action instruction is generated by processing the demonstration data, the task description, and the observation data.
[0038] In this embodiment, the initial action instructions can be quickly generated using the demonstration data, the task description, and the observation data. The demonstration data represents the knowledge reserve of expert demonstration data. For example, for a coffee-making task, initial action instructions such as "open the gripper and move the robotic arm upwards" can be given.
[0039] As demonstrated by the above embodiments, the initial action command can be quickly generated and processed using the demonstration data, the task description, and the observation data. Therefore, by using a motion prediction strategy to predict and process the task description and observation data to obtain the initial action command, precise control over the task description and observation data is achieved, improving data accuracy and effectively reducing execution failures caused by unreasonable action planning, thereby improving processing efficiency.
[0040] S130. After each action instruction in the initial action instruction is executed, each action instruction is evaluated according to the motion evaluation strategy to obtain the evaluation result.
[0041] In this embodiment, each action instruction in the initial action instruction is executed sequentially, and after each action instruction in the initial action instruction is executed, the evaluation result is obtained by evaluating each action instruction according to the motion evaluation strategy.
[0042] In one embodiment, after each motion instruction in the initial motion command is executed, the evaluation process for each motion instruction is performed according to a motion evaluation strategy to obtain an evaluation result, including:
[0043] Execute each action instruction in the initial action instructions, and use each action instruction as the current action instruction;
[0044] Obtain the current scene, and perform semantic parsing on the current scene to obtain the current description;
[0045] The evaluation result is obtained by comparing the current description and the current action instruction.
[0046] In this embodiment, each action instruction in the initial action instruction is executed sequentially, and each action instruction is used as the current action instruction; the current scene can be obtained from the camera device, sensor, etc., and the current scene can be semantically parsed to obtain the current description; the current description and the current action instruction are compared to obtain the evaluation result.
[0047] In one embodiment, the step of comparing the current description and the current action instruction to obtain the evaluation result includes:
[0048] If the current description is inconsistent with the current action instruction, the evaluation result is execution failure;
[0049] If the current description is consistent with the current action instruction, then the evaluation result is successful execution.
[0050] In this embodiment, the evaluation result includes two cases: successful execution and failed execution. When the current description is inconsistent with the current action instruction, the evaluation result is a failed execution; when the current description is consistent with the current action instruction, the evaluation result is a successful execution.
[0051] Through the above embodiments, it can be seen that each action instruction in the initial action instructions is executed, and each action instruction is used as the current action instruction; the current scene is obtained, and semantic parsing of the current scene is performed to obtain the current description; the current description and the current action instruction are compared to obtain the evaluation result; if the current description and the current action instruction are inconsistent, the evaluation result is execution failure; if the current description and the current action instruction are consistent, the evaluation result is execution success. Therefore, after each action instruction in the initial action instructions is executed, an evaluation result is obtained by evaluating each action instruction according to the motion evaluation strategy, and the accuracy of the executed action instructions is evaluated, thereby improving processing efficiency.
[0052] S140. When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain intermediate action instructions.
[0053] In this embodiment, when the evaluation result is determined to be an execution failure, the intermediate action instruction is obtained by adjusting each action instruction based on the action adjustment strategy.
[0054] In one embodiment, when the evaluation result is execution failure, adjusting each action instruction based on the action adjustment strategy to obtain intermediate action instructions includes:
[0055] If the evaluation result is execution failure, then the action object corresponding to each action instruction is identified and processed to obtain the current problem;
[0056] Obtain the preset data chain;
[0057] The intermediate action instruction is obtained by optimizing the current problem based on the preset data chain.
[0058] In this embodiment, when the evaluation result is execution failure, the current problem can be obtained by identifying the action object corresponding to each action instruction using camera equipment, sensors, etc. For example, if the action instruction is "insert the coffee pot into the coffee machine," and the evaluation result is execution failure, the current problem obtained by identifying the action object corresponding to the action instruction might be "the coffee pot is not fully inserted into the coffee machine," where the action object is the coffee pot and the coffee machine. Simultaneously, the optimization processing of the current problem based on the preset data chain to obtain the intermediate action instruction specifically includes: analyzing and processing the current problem to obtain the action target, and using the preset data chain and the action target to generate the intermediate action instruction. For example, if the current problem is "the coffee pot is not fully inserted into the coffee machine," the action target could be "the coffee pot needs to be fully inserted into the coffee machine," and the intermediate action instruction could be a coarse-grained action instruction such as "move the arm backward" or "adjust the grip angle." The preset data chain is a thought chain data combining online human intervention, offline annotation, and expert demonstration data.
[0059] Furthermore, after each action instruction in the initial action instruction is executed, and after evaluating each action instruction according to the motion evaluation strategy to obtain an evaluation result, the method further includes: when the evaluation result is successful, continuing to execute the next action instruction in the initial action instruction.
[0060] As can be seen from the above embodiments, if the evaluation result is execution failure, the action object corresponding to each action instruction is identified and processed to obtain the current problem; a preset data chain is obtained; and the current problem is optimized based on the preset data chain to obtain the intermediate action instruction. Therefore, by adjusting each action instruction based on the action adjustment strategy to obtain the intermediate action instruction when the evaluation result is execution failure, an intermediate layer of "action instruction" is introduced, breaking the disconnect between semantic reflection and action execution in traditional robot self-correction technology, and innovatively building a bridge between high-level semantic understanding and low-level action execution, thereby improving processing efficiency.
[0061] S150. Based on the action correction strategy, the intermediate action instructions are corrected to obtain the target action sequence.
[0062] In this embodiment, after obtaining the intermediate action instruction, the intermediate action instruction can be corrected based on the action correction strategy to obtain the target action sequence.
[0063] In one embodiment, the step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence includes:
[0064] Acquire motion codebook, current visual data, and time step data;
[0065] Motion features are obtained by conversion processing based on the motion codebook and the intermediate motion instructions;
[0066] Visual features are obtained by extracting and processing the current visual information.
[0067] An action embedding vector is obtained by performing preset processing based on the motion features, visual features, and time step data;
[0068] The target action sequence is obtained by calculating and processing the action embedding vector using a preset set of algorithms.
[0069] In this embodiment, after acquiring the motion codebook, the current visual data, and the time step data, the motion features are obtained by conversion processing based on the motion codebook and the intermediate action instructions; the visual features are obtained by extraction processing based on the current visual information; the action embedding vector is obtained by preset processing based on the motion features, the visual features, and the time step data; and the target action sequence is obtained by calculating and processing the action embedding vector using a preset algorithm set. The motion codebook stores a series of basic information and patterns related to robot actions, providing a reference for subsequent action generation; the current visual data is image data acquired using cameras, sensors, etc., and may include, for example, the position and color of a coffee pot; the time step data can be a preset time step; and the preset algorithm set may include linear transformations, convolution operations, MSE loss functions, etc. Therefore, through the above embodiment, it has good adaptability to out-of-distribution scenes such as color interference and random changes in object position.
[0070] The process of calculating and processing the motion embedding vector using a preset algorithm set to obtain the target motion sequence specifically includes: progressively calculating and adjusting the motion embedding vector using linear transformation and convolution operations to generate a fine joint motion sequence; calculating the joint motion sequence using the MSE loss function to minimize the error between the predicted and actual actions (MSE loss), achieving fine motion adjustment at the 0.1 mm level to obtain the target motion sequence. For example, the target motion sequence can be a precise joint motion sequence such as "closing the gripper and moving the robotic arm to the right"; wherein, the preset algorithm set may include linear transformation, convolution operations, and the MSE loss function, etc. Therefore, high-frequency and precise motion generation is achieved with low computational cost based on the motion codebook, avoiding the resource waste of traditional reinforcement learning that requires a large amount of trial and error training; calculating and processing the motion embedding vector using the preset algorithm set to obtain the target motion sequence ensures that the generated actions meet the task requirements and environmental conditions, and finally outputs the specific motion sequence.
[0071] In one embodiment, after correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, the method further includes:
[0072] Execute the target action sequence.
[0073] In this embodiment, the robot is controlled to execute each action of the target action sequence sequentially at a frequency of 20 times per second. After the target action sequence is completed, the robot continues to execute the next action instruction in the initial action instruction. After the next action instruction is executed, each action instruction is evaluated according to the motion evaluation strategy to obtain an evaluation result. When the evaluation result is an execution failure, each action instruction is adjusted according to the motion adjustment strategy to obtain an intermediate action instruction. The intermediate action instruction is corrected according to the motion correction strategy to obtain the target action sequence.
[0074] After the robot successfully completes its task, the optimized interaction trajectory is collected, including observed environmental information and fine-grained movements. This successful experience is fed back to the motion prediction module for fine-tuning and optimization. This allows the robot to generate action commands more efficiently and accurately based on previous successes when performing similar or related tasks in the future, continuously improving its task execution capabilities and thus enhancing processing efficiency.
[0075] This invention framework integrates a dual-process approach of action adjustment and motion correction strategies, achieving complementary advantages. The action adjustment strategy focuses on the conversion and optimization of semantics into actions, while the motion correction strategy is responsible for generating fine-grained actions. Their parallel operation enables the robot to detect and correct complex failure cases that are difficult to handle with traditional methods, significantly improving self-correction performance and execution accuracy in complex tasks. Furthermore, this framework demonstrates good versatility across various robot tasks and scenarios, effectively enabling autonomous learning and action optimization in tasks ranging from home service and industrial assembly to complex operations. Moreover, it maintains a high task success rate even when faced with changing backgrounds and target object position shifts, significantly enhancing the robot's autonomous decision-making and continuous optimization capabilities in dynamic and complex environments. This lays the foundation for the deployment of general-purpose robots in scenarios such as home service and industrial production.
[0076] In summary, this embodiment of the invention acquires task description and observation data; performs prediction processing on the task description and observation data based on a motion prediction strategy to obtain initial action instructions; after each action instruction in the initial action instructions is executed, each action instruction is evaluated according to a motion evaluation strategy to obtain an evaluation result; when the evaluation result is an execution failure, each action instruction is adjusted based on a motion adjustment strategy to obtain intermediate action instructions; and the intermediate action instructions are corrected based on a motion correction strategy to obtain a target action sequence. Therefore, this embodiment of the invention obtains a target action sequence by predicting, evaluating, adjusting, and correcting task description and observation data, and by introducing an intermediate "action instruction" layer, innovatively builds a bridge between high-level semantic understanding and low-level action execution, thereby improving autonomous decision-making and continuous optimization capabilities in dynamic and complex environments, and thus improving processing efficiency.
[0077] Figure 2 This is a schematic block diagram of a motion correction device provided in an embodiment of the present invention. Figure 2 As shown, this embodiment of the invention provides a motion correction device 700 that implements the method described above, applied to a robot. For details, please refer to... Figure 2 The motion correction device 700 includes:
[0078] Acquisition unit 701 is used to acquire task descriptions and observation data;
[0079] Prediction unit 702 is used to perform prediction processing on the task description and the observation data based on a motion prediction strategy to obtain initial action instructions;
[0080] The evaluation unit 703 is used to evaluate each motion instruction according to the motion evaluation strategy after each motion instruction in the initial motion instruction is executed to obtain an evaluation result.
[0081] The adjustment unit 704 is used to adjust each action instruction based on the action adjustment strategy to obtain an intermediate action instruction when the evaluation result is an execution failure.
[0082] The correction unit 705 is used to correct the intermediate action instructions based on the action correction strategy to obtain the target action sequence.
[0083] In some embodiments, when the correction unit 705 performs the processing step of correcting the intermediate motion command based on the motion correction strategy to obtain the target motion sequence, it is specifically used for:
[0084] Acquire motion codebook, current visual data, and time step data;
[0085] Motion features are obtained by conversion processing based on the motion codebook and the intermediate motion instructions;
[0086] Visual features are obtained by extracting and processing the current visual information.
[0087] An action embedding vector is obtained by performing preset processing based on the motion features, visual features, and time step data;
[0088] The target action sequence is obtained by calculating and processing the action embedding vector using a preset set of algorithms.
[0089] In some embodiments, when the adjustment unit 704 performs the step of adjusting each action instruction based on the action adjustment strategy to obtain an intermediate action instruction when the evaluation result is an execution failure, it is specifically used for:
[0090] If the evaluation result is execution failure, then the action object corresponding to each action instruction is identified and processed to obtain the current problem;
[0091] Obtain the preset data chain;
[0092] The intermediate action instruction is obtained by optimizing the current problem based on the preset data chain.
[0093] In some embodiments, when the evaluation unit 703 performs the processing step of evaluating each motion instruction according to a motion evaluation strategy to obtain an evaluation result after executing each motion instruction in the initial motion instruction, it is specifically used for:
[0094] Execute each action instruction in the initial action instructions, and use each action instruction as the current action instruction;
[0095] Obtain the current scene, and perform semantic parsing on the current scene to obtain the current description;
[0096] The evaluation result is obtained by comparing the current description and the current action instruction.
[0097] In some embodiments, when the evaluation unit 703 performs the processing step of comparing the current description and the current action instruction to obtain the evaluation result, it is specifically used for:
[0098] If the current description is inconsistent with the current action instruction, the evaluation result is execution failure;
[0099] If the current description is consistent with the current action instruction, then the evaluation result is successful execution.
[0100] In some embodiments, when the prediction unit 702 performs the step of predicting the task description and the observation data based on the motion prediction strategy to obtain the initial action command, it is specifically used for:
[0101] Obtain preset sample data;
[0102] The initial action instruction is generated by processing the demonstration data, the task description, and the observation data.
[0103] In some embodiments, after performing the processing step of correcting the intermediate motion command based on the motion correction strategy to obtain the target motion sequence, the correction unit 705 is further configured to:
[0104] Execute the target action sequence.
[0105] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned device can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0106] The above-described device can be implemented as a computer program, and the computer program can be implemented in, for example... Figure 3 It runs on the computer device shown.
[0107] Please see Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 800 can be a terminal or a server. The terminal can be an electronic device with communication functions. The server can be a standalone server or a server cluster composed of multiple servers.
[0108] See Figure 3 The electronic device 800 includes a processor 802, a memory, and a network interface 805 connected via a system bus 801. The memory may include a non-volatile storage medium 803 and internal memory 804.
[0109] The non-volatile storage medium 803 may store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions that, when executed, cause the processor 802 to perform an action correction method.
[0110] The processor 802 provides computing and control capabilities to support the operation of the entire electronic device 800.
[0111] The internal memory 804 provides an environment for the execution of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can perform an action correction method.
[0112] This network interface 805 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device 800 to which the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0113] The processor 802 is used to run a computer program 8032 stored in the memory to perform the following steps:
[0114] Obtain task descriptions and observation data;
[0115] Initial action instructions are obtained by predicting the task description and the observation data based on a motion prediction strategy.
[0116] After each motion instruction in the initial motion instruction is executed, an evaluation result is obtained by evaluating each motion instruction according to the motion evaluation strategy.
[0117] When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain an intermediate action instruction;
[0118] The intermediate action instructions are corrected based on the action correction strategy to obtain the target action sequence.
[0119] In some embodiments, when implementing the processing step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, the processor 802 is specifically used for:
[0120] Acquire motion codebook, current visual data, and time step data;
[0121] Motion features are obtained by conversion processing based on the motion codebook and the intermediate motion instructions;
[0122] Visual features are obtained by extracting and processing the current visual information.
[0123] An action embedding vector is obtained by performing preset processing based on the motion features, visual features, and time step data;
[0124] The target action sequence is obtained by calculating and processing the action embedding vector using a preset set of algorithms.
[0125] In some embodiments, when implementing the processing step of adjusting each action instruction based on an action adjustment strategy to obtain intermediate action instructions when the evaluation result is an execution failure, the processor 802 is specifically used for:
[0126] If the evaluation result is execution failure, then the action object corresponding to each action instruction is identified and processed to obtain the current problem;
[0127] Obtain the preset data chain;
[0128] The intermediate action instruction is obtained by optimizing the current problem based on the preset data chain.
[0129] In some embodiments, when the processor 802 performs the processing step of evaluating each action instruction according to a motion evaluation strategy to obtain an evaluation result after executing each action instruction in the initial action instruction, the specific steps are as follows:
[0130] Execute each action instruction in the initial action instructions, and use each action instruction as the current action instruction;
[0131] Obtain the current scene, and perform semantic parsing on the current scene to obtain the current description;
[0132] The evaluation result is obtained by comparing the current description and the current action instruction.
[0133] In some embodiments, when implementing the processing step of comparing the current description and the current action instruction to obtain the evaluation result, the processor 802 is specifically configured to:
[0134] If the current description is inconsistent with the current action instruction, the evaluation result is execution failure;
[0135] If the current description is consistent with the current action instruction, then the evaluation result is successful execution.
[0136] In some embodiments, when implementing the processing step of predicting the task description and the observation data based on a motion prediction strategy to obtain initial action instructions, the processor 802 is specifically used for:
[0137] Obtain preset sample data;
[0138] The initial action instruction is generated by processing the demonstration data, the task description, and the observation data.
[0139] In some embodiments, after implementing the processing step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, the processor 802 is further configured to:
[0140] Execute the target action sequence.
[0141] It should be understood that, in this embodiment of the invention, the processor 802 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0142] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0143] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the following steps:
[0144] Obtain task descriptions and observation data;
[0145] Initial action instructions are obtained by predicting the task description and the observation data based on a motion prediction strategy.
[0146] After each motion instruction in the initial motion instruction is executed, an evaluation result is obtained by evaluating each motion instruction according to the motion evaluation strategy.
[0147] When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain an intermediate action instruction;
[0148] The intermediate action instructions are corrected based on the action correction strategy to obtain the target action sequence.
[0149] In one embodiment, when the processor executes the program instructions to implement the processing step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, it is specifically used for:
[0150] Acquire motion codebook, current visual data, and time step data;
[0151] Motion features are obtained by conversion processing based on the motion codebook and the intermediate motion instructions;
[0152] Visual features are obtained by extracting and processing the current visual information.
[0153] An action embedding vector is obtained by performing preset processing based on the motion features, visual features, and time step data;
[0154] The target action sequence is obtained by calculating and processing the action embedding vector using a preset set of algorithms.
[0155] In one embodiment, when the processor executes the program instructions to implement the processing step of adjusting each action instruction based on an action adjustment strategy to obtain intermediate action instructions when the evaluation result is an execution failure, the processor is specifically used for:
[0156] If the evaluation result is execution failure, then the action object corresponding to each action instruction is identified and processed to obtain the current problem;
[0157] Obtain the preset data chain;
[0158] The intermediate action instruction is obtained by optimizing the current problem based on the preset data chain.
[0159] In one embodiment, when the processor executes the program instructions to implement the process of evaluating each action instruction in the initial action instructions and then evaluates each action instruction according to a motion evaluation strategy to obtain an evaluation result, the specific steps are as follows:
[0160] Execute each action instruction in the initial action instructions, and use each action instruction as the current action instruction;
[0161] Obtain the current scene, and perform semantic parsing on the current scene to obtain the current description;
[0162] The evaluation result is obtained by comparing the current description and the current action instruction.
[0163] In one embodiment, when the processor executes the program instructions to implement the processing step of comparing the current description and the current action instruction to obtain the evaluation result, it is specifically used for:
[0164] If the current description is inconsistent with the current action instruction, the evaluation result is execution failure;
[0165] If the current description is consistent with the current action instruction, then the evaluation result is successful execution.
[0166] In one embodiment, when the processor executes the program instructions to implement the processing step of predicting the task description and the observation data based on the motion prediction strategy to obtain the initial action instruction, it is specifically used for:
[0167] Obtain preset sample data;
[0168] The initial action instruction is generated by processing the demonstration data, the task description, and the observation data.
[0169] In one embodiment, after the processor executes the program instructions to implement the processing step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, it is further configured to:
[0170] Execute the target action sequence.
[0171] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0172] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0173] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0174] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0175] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0176] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The software tools, models, or components appearing in the embodiments of the present invention are merely illustrative examples and do not represent actual use.
Claims
1. A motion correction method applied to a robot, characterized in that, include: Obtain task descriptions and observation data; Initial action instructions are obtained by predicting the task description and the observation data based on a motion prediction strategy. After each motion instruction in the initial motion instruction is executed, an evaluation result is obtained by evaluating each motion instruction according to the motion evaluation strategy. When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain an intermediate action instruction; The intermediate action instructions are corrected based on the action correction strategy to obtain the target action sequence.
2. The method according to claim 1, characterized in that, The step of correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence includes: Acquire motion codebook, current visual data, and time step data; Motion features are obtained by conversion processing based on the motion codebook and the intermediate motion instructions; Visual features are obtained by extracting and processing the current visual information. An action embedding vector is obtained by performing preset processing based on the motion features, visual features, and time step data; The target action sequence is obtained by calculating and processing the action embedding vector using a preset set of algorithms.
3. The method according to claim 1, characterized in that, When the evaluation result is execution failure, each action instruction is adjusted based on the action adjustment strategy to obtain intermediate action instructions, including: If the evaluation result is execution failure, then the action object corresponding to each action instruction is identified and processed to obtain the current problem; Obtain the preset data chain; The intermediate action instruction is obtained by optimizing the current problem based on the preset data chain.
4. The method according to claim 1, characterized in that, After each motion instruction in the initial motion command is executed, an evaluation result is obtained by evaluating each motion instruction according to a motion evaluation strategy, including: Execute each action instruction in the initial action instructions, and use each action instruction as the current action instruction; Obtain the current scene, and perform semantic parsing on the current scene to obtain the current description; The evaluation result is obtained by comparing the current description and the current action instruction.
5. The method according to claim 4, characterized in that, The step of comparing the current description and the current action instruction to obtain the evaluation result includes: If the current description is inconsistent with the current action instruction, the evaluation result is execution failure; If the current description is consistent with the current action instruction, then the evaluation result is successful execution.
6. The method according to claim 1, characterized in that, The process of predicting and processing the task description and the observation data based on a motion prediction strategy to obtain initial action instructions includes: Obtain preset sample data; The initial action instruction is generated by processing the demonstration data, the task description, and the observation data.
7. The method according to claim 1, characterized in that, After correcting the intermediate action instructions based on the action correction strategy to obtain the target action sequence, the method further includes: Execute the target action sequence.
8. A motion correction device, applied to a robot, characterized in that, include: The acquisition unit is used to acquire task descriptions and observation data; The prediction unit is used to perform prediction processing on the task description and the observation data based on the motion prediction strategy to obtain the initial action command; An evaluation unit is used to evaluate each motion instruction according to a motion evaluation strategy after each motion instruction in the initial motion instruction is executed, and obtain an evaluation result. An adjustment unit is used to adjust each action instruction based on an action adjustment strategy to obtain an intermediate action instruction when the evaluation result is an execution failure. The correction unit is used to correct the intermediate action instructions based on the action correction strategy to obtain the target action sequence.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-7.