Multi-object rearrangement model training method, multi-object rearrangement method, equipment and medium

By using a multi-stage training framework guided by vision and language, a multi-object rearrangement model is generated, which solves the problem that humanoid robots cannot achieve continuous rearrangement of multiple objects. This enables long-term continuous rearrangement tasks of the multi-object rearrangement model, improving the success rate and adaptability of the task.

CN122065874APending Publication Date: 2026-05-19BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
Filing Date
2026-03-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Current humanoid robots only support single-object rearrangement and cannot achieve continuous rearrangement of multiple objects, which limits their capabilities.

Method used

A multi-stage training framework guided by vision and language is adopted. By generating a multi-object rearrangement model, using student models and pre-trained teacher models, and combining viewpoint images and natural language instructions, multiple action instructions are generated to control the simulated robot to perform multi-object rearrangement tasks. The model performance is improved through supervised training.

Benefits of technology

It realizes long-term continuous rearrangement tasks for multi-object rearrangement models, improves physical stability, task success rate and generalization ability, and can adapt to unknown environments, unknown objects and unknown forms of natural language instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065874A_ABST
    Figure CN122065874A_ABST
Patent Text Reader

Abstract

The invention provides a multi-object rearrangement model training method, a multi-object rearrangement method, equipment and a medium, and the model training method comprises the steps: respectively employing a student model and a first teacher model according to a first visual angle image of a simulation robot and a first natural language instruction of a first multi-object rearrangement task, generating a first action instruction and a second action instruction of the first single object rearrangement task, obtaining a second visual angle image of the simulation robot after the first single object rearrangement task is executed, and according to the second visual angle image and the first natural language instruction, respectively adopting a student model and a second teacher model to rearrange the first single object. And generating a third action instruction and a fourth action instruction of a non-first single object rearrangement task, performing supervised training on the student model according to the first action instruction, the second action instruction, the third action instruction and the fourth action instruction, and generating a multi-object rearrangement model. And continuously finishing multi-object rearrangement by adopting a multi-object rearrangement model guided by a visual language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics technology, and more specifically, to a method for training a multi-object rearrangement model, a multi-object rearrangement method, an apparatus, and a medium. Background Technology

[0002] Humanoid robots, also known as androids, are robots designed to mimic human appearance and behavior, especially those with human-like physiques. Humanoid robots have made significant progress in the areas of motion imitation, environmental interaction, and object manipulation.

[0003] In related technologies, humanoid robots are often used to rearrange single objects. Object rearrangement refers to the process by which robots, through observation and manipulation, rearrange scattered and disordered objects in the environment according to certain target rules or spatial relationships.

[0004] However, the above methods only support single-object rearrangement and cannot achieve continuous rearrangement of multiple objects, which has certain limitations for the task. Summary of the Invention

[0005] In view of this, embodiments of this application provide a training method, a method, a device, and a medium for multi-object rearrangement models, to solve the problem that existing methods only support single-object operations and cannot achieve continuous rearrangement of multiple objects, thus having certain task limitations.

[0006] In a first aspect, embodiments of this application provide a method for training a multi-object rearrangement model, including: Based on the first-view image of the simulated robot and the first natural language instruction of the first multi-object rearrangement task, the first action instruction and the second action instruction of the first single-object rearrangement task in the first multi-object rearrangement task are generated by using the student model and the pre-trained first teacher model, respectively. The first action instruction is used to control the simulated robot to execute the first single-object rearrangement task in the first multi-object rearrangement task. Acquire a second-view image of the simulated robot after the first single-object rearrangement task in the first multi-object rearrangement task has been completed; Based on the second perspective image and the first natural language instruction, the student model and the pre-trained second teacher model are used respectively to generate the third action instruction and the fourth action instruction for the non-first single object rearrangement task in the first multi-object rearrangement task. The third action instruction is used to control the simulation robot to perform the non-first single object rearrangement task in the first multi-object rearrangement task. The student model is trained under supervision based on the first action instruction, the second action instruction, the third action instruction, and the fourth action instruction to generate a multi-object rearrangement model.

[0007] In an optional implementation, the method further includes: Based on the second natural language instruction for the single object rearrangement task and the viewpoint image corresponding to the single object rearrangement task, the preset rearrangement model is trained for the single object rearrangement task to obtain the first teacher model. Based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task, the preset rearrangement model is trained on the multi-object rearrangement task to obtain the second teacher model.

[0008] In an optional implementation, the step of training a preset rearrangement model to perform a single-object rearrangement task based on the second natural language instruction for the single-object rearrangement task and the viewpoint image corresponding to the single-object rearrangement task, to obtain the first teacher model, includes: Using the preset rearrangement model, a third action instruction is generated based on the second natural language instruction and the viewpoint image corresponding to the single object rearrangement task. The third action instruction is used to control the simulation robot to perform the single object rearrangement task. Obtain the first reward parameter when the simulated robot performs the single-object rearrangement task; Based on the first reward parameter, the preset rearrangement model is trained to obtain the first teacher model.

[0009] In an optional implementation, the method further includes: The task completion status information of the simulated robot after performing the single object rearrangement task is obtained, wherein the task completion status information includes: the distance between the simulated robot's torso and the single object corresponding to the single object rearrangement task, the distance between the simulated robot's hand and the single object, and the distance between the single object and the rearrangement termination position. Based on the task completion status information, obtain the second reward parameter; The step of training the preset rearrangement model according to the first reward parameter to obtain the first teacher model includes: The preset rearrangement model is trained based on the first reward parameter and the second reward parameter to obtain the first teacher model.

[0010] In an optional implementation, the step of training the preset rearrangement model for the multi-object rearrangement task based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task to obtain the second teacher model includes: Based on the third natural language instruction and the third-view image of the simulated robot, the first teacher model is used to generate a fourth action instruction; the fourth action instruction is used to control the simulated robot to perform the first single-object rearrangement task in the second multi-object rearrangement task. Acquire a fourth-view image of the simulated robot after it has performed the first single-object rearrangement task in the second multi-object rearrangement task; Based on the third natural language instruction and the fourth perspective image, the fifth action instruction is generated using the preset rearrangement model. The fifth action instruction is used to control the simulation robot to perform a non-first single-object rearrangement task in the second multi-object rearrangement task. Obtain the third reward parameter when the simulated robot performs the second multi-object rearrangement task, which is not the first single-object rearrangement task; Based on the third reward parameter, the preset rearrangement model is trained to obtain the second teacher model.

[0011] In an optional implementation, the step of generating the first action instruction and the second action instruction for the first single-object rearrangement task in the first multi-object rearrangement task, based on the first-view image of the simulated robot and the first natural language instruction for the first multi-object rearrangement task, using a student model and a pre-trained first teacher model respectively, includes: Based on the first perspective image, the first natural language instruction, and the first ontology history data, the first action instruction and the second action instruction are generated using the student model and the first teacher model, respectively. The first ontology history data is the ontological perception data of the simulated robot before the acquisition time corresponding to the first perspective image. The step involves generating third and fourth action instructions for non-first single-object rearrangement tasks in the first multi-object rearrangement task, based on the second-view image and the first natural language instruction, using the student model and the pre-trained second teacher model respectively, including: Based on the second perspective image, the first natural language instruction, and the second ontology history data, the third action instruction and the fourth action instruction are generated using the student model and the second teacher model, respectively. The second ontology history data is the ontological perception data of the simulated robot before the acquisition time corresponding to the second perspective image.

[0012] Secondly, embodiments of this application also provide a method for rearranging multiple objects based on a humanoid robot, including: Acquire robot-perspective images of a humanoid robot and natural speech commands for a multi-object rearrangement task; Based on the robot's perspective image and the natural speech command of the target multi-object rearrangement task, a multi-object rearrangement model is used to generate the target action command of the humanoid robot. The multi-object rearrangement model is a model trained using the method described in any of the first aspects. According to the target action command, the humanoid robot is controlled to perform the target multi-object rearrangement task.

[0013] In an optional implementation, the step of generating the target action command for the humanoid robot using a multi-object rearrangement model based on the robot's perspective image and the natural language speech command for the target multi-object rearrangement task includes: Based on the robot's perspective image, the natural language command for the target multi-object rearrangement task, and the target's historical data, the target action command is generated using the multi-object rearrangement model. The target's historical data refers to the humanoid robot's proprioceptive data prior to the acquisition time corresponding to the robot's perspective image.

[0014] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and the processor executes the machine-readable instructions to perform the method described in any of the first aspects.

[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the method described in any of the first aspects.

[0016] This application provides a training method, apparatus, and medium for a multi-object rearrangement model. The training method includes: generating first and second action instructions for a first single-object rearrangement task using a student model and a pre-trained first teacher model, based on a first-view image of a simulated robot and a first natural language instruction for a first multi-object rearrangement task; acquiring a second-view image of the simulated robot after the first single-object rearrangement task is completed; generating third and fourth action instructions for tasks other than the first single-object rearrangement task using the student model and a pre-trained second teacher model, based on the second-view image and the first natural language instruction; and supervising the student model based on the first, second, third, and fourth action instructions to generate a multi-object rearrangement model. Thus, the visually guided multi-object rearrangement model continuously completes multi-object rearrangements. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 1 ; Figure 2 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 2 ; Figure 3 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 3 ; Figure 4 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 4 ; Figure 5 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 5 ; Figure 6 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 6 ; Figure 7 A schematic diagram illustrating the training process of the multi-object rearrangement model provided in this application embodiment; Figure 8 A flowchart illustrating the multi-object rearrangement method based on a humanoid robot provided in this application embodiment; Figure 9 This is a schematic diagram of the structure of the multi-object rearrangement model training device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structure of a multi-object rearrangement device based on a humanoid robot provided in an embodiment of this application; Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0020] To address the limitation of current humanoid robots that only support single-object rearrangement and cannot perform continuous rearrangement of multiple objects, thus restricting their task capabilities, this application provides a multi-object rearrangement model based on a vision- and language-guided multi-stage training framework. This model enables long-term continuous multi-object rearrangement tasks, improving physical stability, task success rate, and generalization ability. Generalization ability refers to the model's adaptability to unknown environments, unknown objects, and unknown forms of natural language commands.

[0021] Figure 1 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 1 In this embodiment, the execution entity can be a model training device, such as a computer device.

[0022] like Figure 1 As shown, the method may include: S101. Based on the first-view image of the simulated robot and the first natural language instruction of the first multi-object rearrangement task, the first action instruction and the second action instruction of the first single-object rearrangement task in the first multi-object rearrangement task are generated using the student model and the pre-trained first teacher model, respectively.

[0023] A humanoid robot is a physically simulated humanoid robot. A first-person perspective image of a humanoid robot is an image taken from the robot's own viewpoint within a simulated environment. Specifically, the humanoid robot's head is equipped with a camera to capture this first-person perspective image, which includes visual information about the environment and objects.

[0024] The first multi-object rearrangement task is the multi-object rearrangement task used when training the multi-object rearrangement model. The first multi-object rearrangement task includes multiple consecutively rearranged single-object tasks, such as two consecutively rearranged single-object tasks.

[0025] Multiple consecutive single-object rearrangement tasks are divided into two task phases. The first task phase is used to execute the first single-object rearrangement task, and the second task phase is used to execute subsequent single-object rearrangement tasks. For the first single-object rearrangement task, the simulated robot's initial posture is usually a standard posture, such as a standard upright position. For subsequent single-object rearrangement tasks, the simulated robot is often in a non-standard posture after completing the first single-object rearrangement task, such as bending over, or even turning its back to the subsequent single object.

[0026] The number of first-order multi-object rearrangement tasks is multiple. Assuming there are 350 first-order multi-object rearrangement tasks (covering different simulation environments such as bedrooms, kitchens, living rooms, and warehouses), and each first-order multi-object rearrangement task includes two consecutive single-object rearrangement tasks, there are a total of 700 single-object tasks. Two consecutive single-object rearrangement tasks could be, for example, moving a laptop from the bed to the desk on the right side of the room, and moving a chair from the left side of the desk to the front of the desk. These two tasks are consecutive; that is, after moving the laptop from the bed to the desk on the right side of the room, the task of moving the chair from the left side of the desk to the front of the desk is executed consecutively. During this process, before moving the laptop from the bed to the desk on the right side of the room, the simulated robot's posture is ideal; after moving the laptop from the bed to the desk on the right side of the room, the simulated robot's posture is non-ideal.

[0027] The first natural language instruction is used to instruct each individual object rearrangement task included in the first multi-object rearrangement task. The first natural language instruction is a text instruction for each individual object rearrangement task included in the first multi-object rearrangement task, such as "move the laptop from the bed to the desk on the right side of the room" or "move the chair from the left side of the desk to the front of the desk".

[0028] The first-person perspective image includes the first single object (e.g., a laptop), the starting position (e.g., a bed), the ending position (e.g., a desk), and obstacles in the first single-object rearrangement task within the multi-object rearrangement task. By acquiring this first-person perspective image, the simulated robot can perceive the spatial position of the first single object to be rearranged within the simulated environment, thereby establishing its spatial relationships with the first single object, the starting position, the ending position, and obstacles. This allows it to autonomously navigate around obstacles to perform the single-object rearrangement task.

[0029] The first natural language instruction for the first single-object rearrangement task in the first multi-object rearrangement task is determined from the first natural language instruction (e.g., "move the laptop from the bed to the desk on the right side of the room"). Based on the first-view image and the natural language instruction for the first single-object rearrangement task, a student model is used to parse the natural language instruction for the first single-object rearrangement task and extract the visual features of the first-view image to generate the first action instruction for the first single-object rearrangement task. Based on the first-view image and the natural language instruction for the first single-object rearrangement task, a pre-trained first teacher model is used to parse the natural language instruction for the first single-object rearrangement task and extract the visual features of the first-view image to generate the second action instruction for the first single-object rearrangement task.

[0030] The first action command is used to control the simulated robot to perform the first single-object rearrangement task in the first multi-object rearrangement task. If the simulated robot is in the first task stage, the second action command is the action supervision signal of the first action command, provided by the first teacher model.

[0031] S102. Obtain the second-view image of the simulated robot after the first single-object rearrangement task is completed in the first multi-object rearrangement task.

[0032] According to the first action command, the simulation robot is controlled to perform the first single object rearrangement task in the first multi-object rearrangement task, and a second-view image of the simulation robot after the first single object rearrangement task is completed is acquired. The second-view image is the first-person view image of the simulation robot in the simulation environment scene after the simulation robot performs the first single object rearrangement task. It refers to the image collected from the robot's own perspective and is used to indicate what the simulation robot sees after performing the first single object rearrangement task.

[0033] The second-view image includes the non-first single object (such as the second single object) in the first multi-object rearrangement task, the rearrangement start position (such as the left side of the desk), the rearrangement end position (directly in front of the desk), and obstacles. Thus, by acquiring the second-view image, the simulated robot can perceive the spatial positions of the second single object to be rearranged, the rearrangement start position (such as the left side of the desk), the rearrangement end position (directly in front of the desk), and obstacles within the simulated environment. This allows it to establish spatial relationships with the second single object, the target position, and the obstacles, enabling it to autonomously bypass obstacles to perform the single-object rearrangement task.

[0034] S103. Based on the second-view image and the first natural language instruction, the third and fourth action instructions for the non-first single-object rearrangement task in the first multi-object rearrangement task are generated using the student model and the pre-trained second teacher model, respectively.

[0035] From the first natural language instruction, determine the natural language instruction for the non-first single-object rearrangement task in the first multi-object rearrangement task (e.g., "move the chair from the left side of the desk to the front of the desk"). Based on the second-view image and the natural language instruction for the non-first single-object rearrangement task, use a student model to parse the natural language instruction for the non-first single-object rearrangement task and extract the visual features of the second-view image to generate the third action instruction for the non-first single-object rearrangement task. Based on the second-view image and the natural language instruction for the non-first single-object rearrangement task, use a pre-trained second teacher model to parse the natural language instruction for the non-first single-object rearrangement task and extract the visual features of the second-view image to generate the fourth action instruction for the non-first single-object rearrangement task.

[0036] The third action instruction is used to control the simulated robot to perform a single-object rearrangement task that is not the first one in the first multi-object rearrangement task. If the simulated robot is in the second task phase, the fourth action instruction is the action supervision signal of the third action instruction, which is provided by the second teacher model.

[0037] It should be noted that the motion commands mentioned in this solution may include joint control actions, which control the robot's joint movements to drive the robot to perform single object rearrangement tasks.

[0038] S104. Based on the first action instruction, the second action instruction, the third action instruction, and the fourth action instruction, supervise the training of the student model to generate a multi-object rearrangement model.

[0039] The first loss is calculated based on the first action instruction and the second action instruction, the second loss is calculated based on the third action instruction and the fourth action instruction, and the total loss is calculated based on the first loss and the second loss. The student model is then trained under supervision based on the total loss, and the trained student model is used as a multi-object rearrangement model.

[0040] In some embodiments, the total loss can be a weighted sum of the first loss and the second loss, with a weight value of, for example, 0.5. The specific weight can be selected according to the actual situation, and this embodiment does not impose any particular limitation on it.

[0041] It should be noted that by distilling the dual-teacher model into a student model, such as a Visual-Language-Action (VLA) model, and employing an imitation learning algorithm (such as the DAgger algorithm), the dual-teacher model is distilled into a VLA model that only receives first-person view images (such as RGB images) and natural language instructions. Here, dual-teacher distillation is an augmentation technique of knowledge distillation, which uses two complementary teacher models to guide the training of a single student model to achieve better performance.

[0042] It should be noted that if there are three single-object rearrangement tasks, a third-view image of the simulated robot after the second single-object rearrangement task is completed can be obtained, and the natural language instructions for the third single-object rearrangement task can be determined from the first natural language instructions. Based on the third-view image and the natural language instructions for the third single-object rearrangement task, the fifth and sixth action instructions for the third single-object rearrangement task are generated using the student model and the second teacher model, respectively. Then, the student model is supervised and trained based on the first, second, third, fourth, fifth, and sixth action instructions to generate a multi-object rearrangement model.

[0043] Of course, if there are more single-object rearrangement tasks, they can all be executed in a loop according to the above process to supervise the training of the student model.

[0044] In this embodiment, to address the complex operational requirements of multi-object rearrangement, a multi-object rearrangement model based on a vision- and language-guided multi-stage training framework is provided. This model supports continuous responses to first natural language commands and seamless task transitions, enabling long-term continuous multi-object rearrangement tasks and improving physical stability, task success rate, and generalization ability. Furthermore, it requires no specific task training and can directly respond to natural language commands and robot-perspective images, adapting to different scenarios and achieving zero-shot adaptation.

[0045] Figure 2 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 2 ,like Figure 2 As shown, in an optional implementation, the method may further include: S201. Based on the second natural language instruction of the single object rearrangement task and the viewpoint image corresponding to the single object rearrangement task, train the preset rearrangement model for the single object rearrangement task to obtain the first teacher model.

[0046] Among them, the single object rearrangement task is the single object rearrangement task used to train the first teacher model. There are multiple single object rearrangement tasks. For the single object rearrangement task, the initial posture of the simulated robot is usually a standard posture, such as a standard upright state.

[0047] The single-object rearrangement task is used to instruct the rearrangement of a single object. The second natural language instruction can be a text instruction for the single-object rearrangement task, such as "move the box on the right side of the ground to the left side of the right shelf", "lift another box on the ground to the right side of the left shelf", "move the chair to the right side of the entryway table", and "move the green plant from the left side of the TV cabinet to the center of the coffee table".

[0048] The perspective image corresponding to the single object rearrangement task is the perspective image of the simulated humanoid robot under the single object rearrangement task. It refers to the first-person perspective image of the simulated robot in the simulated environment scene, that is, the image collected from the robot's own perspective. The perspective image includes the single object (such as a box) corresponding to the single object rearrangement task, the rearrangement start position (such as the box on the right side of the ground), the rearrangement end position (such as the left side of the right shelf), and obstacles, etc.

[0049] Based on the second natural language instruction for the single object rearrangement task and the corresponding viewpoint image, the preset rearrangement model is trained for the single object rearrangement task, and the trained preset rearrangement model is used as the first teacher model.

[0050] It should be noted that the preset rearrangement model can be, for example, HumanVLA. HumanVLA is an end-to-end Vision-Language-Action (VLA) model specifically designed for physical humanoid robots. Its core feature is that the robot can complete complex object rearrangements using only visual and natural language commands.

[0051] S202. Based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task, train the preset rearrangement model for the multi-object rearrangement task to obtain the second teacher model.

[0052] The second multi-object rearrangement task is used to train the second teacher model. It comprises multiple consecutively rearranged single-object tasks, divided into two phases. The first phase executes the first single-object rearrangement task, and the second phase executes subsequent single-object rearrangement tasks. For the first single-object rearrangement task, the simulated robot's initial posture is typically a standard posture, such as a standard upright position. For subsequent single-object rearrangement tasks, the simulated robot often assumes a non-standard posture after completing the first task, such as bending over or even having its back to the non-first single object.

[0053] The third natural language instruction is used to instruct each individual object rearrangement task contained in the second multi-object rearrangement task. The third natural language instruction is the text instruction of each individual object rearrangement task contained in the second multi-object rearrangement task.

[0054] The perspective images corresponding to the second multi-object rearrangement task include the perspective images corresponding to each single-object rearrangement task within the second multi-object rearrangement task. That is, the first-person perspective images of the simulated robot in the simulation environment scene. The perspective images corresponding to the single-object rearrangement task are the perspective images of the simulated humanoid robot under the single-object rearrangement task. Among them, the perspective images corresponding to the single-object rearrangement task include the single object (such as a box) corresponding to the single-object rearrangement task, the rearrangement start position (the starting position and end position of placing the single object, such as the box on the right side of the ground, the left side of the right shelf), and obstacles, etc.

[0055] Based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task, the preset rearrangement model is trained for the multi-object rearrangement task, and the trained preset rearrangement model is used as the second teacher model.

[0056] Figure 3 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 3 ,like Figure 3 As shown, in an optional implementation, step S201 above, which trains a preset rearrangement model for the single-object rearrangement task based on the second natural language instruction for the single-object rearrangement task and the viewpoint image corresponding to the single-object rearrangement task, to obtain a first teacher model, may include: S301. Using a preset rearrangement model, a third action instruction is generated based on the second natural language instruction and the viewpoint image corresponding to the single object rearrangement task.

[0057] The second natural language instruction and the viewpoint image corresponding to the single object rearrangement task are used as model input. The second natural language instruction is parsed using a preset rearrangement model, and the visual features of the viewpoint image corresponding to the single object rearrangement task are extracted to generate the third action instruction. The third action instruction is used to control the simulated robot to perform the single object rearrangement task.

[0058] S302. Obtain the first reward parameter when the simulated robot performs a single object rearrangement task.

[0059] The simulation robot is controlled to perform a single object rearrangement task according to the third action instruction, and the first reward parameter is obtained when performing the single object rearrangement task. The first reward parameter can be the sum of the task reward and the style reward. The task reward is used to indicate the degree of task completion of the simulation robot when performing the single object rearrangement task, and the style reward is used to indicate the degree of humanization of the simulation robot when performing the single object rearrangement task.

[0060] In some embodiments, when the simulated robot performs a single object rearrangement task, the task reward is calculated using a preset calculation formula based on the distance between the simulated robot and the corresponding single object before grasping, whether the single object is accurately grasped, whether the single object is accurately moved to the rearrangement termination position after grasping, and whether the single object is accurately released when it is moved to the rearrangement termination position.

[0061] The task reward is determined by the following factors: the distance between the simulated robot and the corresponding single object before grasping; whether the single object was successfully grasped; whether the single object was stably moved to the rearrangement termination position after grasping; and whether the single object was stably released when it was moved to the rearrangement termination position. The closer the distance between the simulated robot and the corresponding single object before grasping, the larger the corresponding sub-reward. The more accurately the single object was grasped, the larger the corresponding sub-reward. The more accurately the single object was moved to the rearrangement termination position after grasping, the larger the corresponding sub-reward. The more accurately the single object was released, the larger the corresponding sub-reward. The task reward is obtained by weighted summing the multiple sub-rewards.

[0062] In some embodiments, when the simulated robot performs a single object rearrangement task, an Adversarial Motion Prior (AMP) discriminator is used to evaluate the human-likeness of the third action instruction used by the simulated robot in performing the single object rearrangement task. The higher the evaluation score, the more human-like it is, and the greater the style reward.

[0063] The total weight of task rewards can be 0.9, and the weight of style rewards can be 0.1. The first reward parameter is obtained by weighting and summing the task rewards and style rewards.

[0064] S303. Based on the first reward parameter, train the preset rearrangement model to obtain the first teacher model.

[0065] Based on the first reward parameter, a reinforcement learning strategy is adopted to train the pre-set rearrangement model, and the rearrangement model obtained when the pre-set iteration stopping condition is reached is used as the first teacher model.

[0066] In this embodiment, model training is performed by combining task rewards and style rewards, guiding the model to accurately execute single object rearrangement and human-like actions, enabling the first teacher model to master stable object handling and safe release capabilities, and the action instructions output by the first teacher model have both task completion and human-like characteristics.

[0067] Figure 4 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 4 ,like Figure 4 As shown, in an optional implementation, the method may further include: S401. Obtain the task completion status information after the simulation robot performs a single object rearrangement task.

[0068] The simulation robot is controlled to perform a single object rearrangement task according to the third action command, and the task completion status information after the simulation robot performs the single object rearrangement task is obtained. The task completion status information includes: the distance between the simulation robot's torso and the single object corresponding to the single object rearrangement task, the distance between the simulation robot's hand and the single object, and the distance between the single object and the rearrangement termination position (i.e., the target position where the single object is placed).

[0069] In other words, after the simulated robot completes the single-object rearrangement task, the distance between the simulated robot's torso and the single object, the distance between the simulated robot's hand and the single object, and the distance between the single object and the rearrangement termination position are obtained.

[0070] S402. Obtain the second reward parameter based on the task completion status information.

[0071] The second reward parameter is calculated using a preset calculation formula based on the distance between the robot's torso and the single object in the single object rearrangement task, the distance between the robot's hand and the single object, and the distance between the single object and the rearrangement termination position.

[0072] In some embodiments, an exponentially increasing reward function rrobot2object=1 is used based on the distance between the simulated robot's torso and a single object. exp( Calculate sub-reward parameter 1 using 0.5×drobot2object, where rrobot2object is sub-reward parameter 1 and drobot2object is the distance between the simulated robot's torso and a single object.

[0073] In some embodiments, an exponentially increasing reward function rhand2object=1 is used based on the distance between the simulated robot's hand and a single object. exp( Calculate sub-reward parameter 2 using 0.5 × dhand2object), where rhand2object is sub-reward parameter 2, and dhand2object is the distance between the robot's hand and a single object. Specifically, dhand2object = min(dlefthand2object, drighthand2object), where dlefthand2object is the distance between the robot's left hand and a single object, and drighthand2object is the distance between the robot's right hand and a single object.

[0074] In some embodiments, if the distance between a single object and the rearrangement termination position is greater than or equal to a preset distance threshold, it indicates that the single object has not been moved to the vicinity of the rearrangement termination position. In this case, no reward is given regardless of the location of the simulated robot's torso and hand, and sub-reward parameter 1 and sub-reward parameter 2 are set to zero. If the distance between a single object and the rearrangement termination position is less than a preset distance threshold, it indicates that the single object has been moved to the vicinity of the rearrangement termination position. In this case, a second reward parameter is calculated based on sub-reward parameter 1 and sub-reward parameter 2. The second reward parameter can be a weighted sum of sub-reward parameter 1 and sub-reward parameter 2, for example, it can be expressed as: 0.5×rrobot2object+0.5×rhand2object.

[0075] It should be noted that setting an upper limit threshold for the second reward parameter guides the robot to maintain a distance of approximately 1 meter between its torso and a single object, and approximately 0.5 meters between its hand and the object. These two specific values ​​are chosen to balance task requirements: the robot needs to retreat a sufficient distance to avoid interfering with the placed object, but not too far to prevent it from falling or colliding with other unrelated objects. Therefore, when the distance exceeds the aforementioned threshold, the second reward parameter is fixed at 1, indicating that no further incentive is needed for continued retreat. Furthermore, to prevent premature release or behavioral degradation, if the object is not moved close enough to the rearrangement termination position, the second reward parameter will be set to zero, ensuring that the robot only performs the release and retreat actions after successfully placing the object.

[0076] Step S303 above, which involves training the preset rearrangement model based on the first reward parameter to obtain the first teacher model, may include: S403. Based on the first reward parameter and the second reward parameter, train the preset rearrangement model to obtain the first teacher model.

[0077] Based on the first reward parameter and the second reward parameter, a reinforcement learning strategy is adopted to train the preset rearrangement model, and the rearrangement model obtained when the preset iteration stopping condition is reached is used as the first teacher model.

[0078] In this embodiment, the first teacher model is trained on placement, release, and retreat by combining the second reward parameter. This ensures that after placing a single object, the simulated robot's torso retreats to a safe distance of about 1 meter, and its hand is about 0.5 meters away from the single object. This focuses on the distance between the robot's torso and the single body, and between the hand and the single object. The reward is triggered only after the object is successfully placed, avoiding premature retreat. Instead, the robot automatically retreats after operation to avoid collisions with already placed objects, resulting in a smooth task transition. In other words, by designing a task switching mechanism, three conditions are detected: the speed of the single object (nearly stationary, indicating successful placement at the rearrangement termination position), the distance between the object and the target (distance between the single object and the rearrangement termination position), and the distance between the robot and objects (including the distance between the torso and the single object, and the distance between the hand and the single object). The switching command is automatically triggered, dynamically switching from the first teacher module to the second teacher model to provide supervision.

[0079] Figure 5 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 5 ,like Figure 5 As shown, in an optional implementation, step S202 above, training a preset rearrangement model for the multi-object rearrangement task based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task to obtain a second teacher model, may include: S501. Based on the third natural language instruction and the third-view image of the simulated robot, the fourth action instruction is generated using the first teacher model.

[0080] Among them, the perspective images corresponding to the second multi-object rearrangement task include the third-person perspective images of the simulated robot. The third-person perspective images are the first-person perspective images of the simulated robot in the simulated environment scene, which refer to the images collected from the robot's own perspective.

[0081] The third-person perspective image includes the first single object (such as a laptop) corresponding to the first single object rearrangement task in the second multi-object rearrangement task, the rearrangement start position (such as a bed), the rearrangement end position (desk), and obstacles.

[0082] The natural language instructions for the first single-object rearrangement task in the second multi-object rearrangement task are determined from the third natural language instructions. Based on the third-view image and the natural language instructions for this first single-object rearrangement task, a pre-trained first teacher model is used to parse the natural language instructions for this first single-object rearrangement task and extract the visual features of the third-view image to generate a fourth action instruction. This fourth action instruction is used to control the simulated robot to perform the first single-object rearrangement task in the second multi-object rearrangement task.

[0083] S502. Obtain the fourth-view image after the simulated robot performs the first single-object rearrangement task in the second multi-object rearrangement task.

[0084] According to the fourth action command, the simulation robot is controlled to perform the first single object rearrangement task in the second multi-object rearrangement task, and the fourth-view image after the simulation robot performs the first single object rearrangement task is obtained. The fourth-view image is the first-person view image of the simulation robot in the simulation environment scene after the simulation robot performs the first single object rearrangement task, which is used to indicate what the simulation robot sees after performing the first single object rearrangement task.

[0085] The fourth-view image includes the non-first single object in the second multi-object rearrangement task, the rearrangement start position, the rearrangement end position, and obstacles.

[0086] S503. Based on the third natural language instruction and the fourth perspective image, a preset rearrangement model is used to generate the fifth action instruction.

[0087] The natural language instructions for the non-first single-object rearrangement task in the second multi-object rearrangement task are determined from the third natural language instructions. Based on the natural language instructions for the non-first single-object rearrangement task and the fourth-view image, the natural language instructions for the non-first single-object rearrangement task are parsed using a preset rearrangement model, and the visual features of the fourth-view image are extracted to generate the fifth action instructions. The fifth action instructions are used to control the simulated robot to perform the non-first single-object rearrangement task in the second multi-object rearrangement task.

[0088] S504. Obtain the third reward parameter when the simulated robot performs the second multi-object rearrangement task, which is not the first single-object rearrangement task.

[0089] According to the fifth action instruction, the simulation robot is controlled to perform a non-first single-object rearrangement task in the second multi-object rearrangement task, and the third reward parameter is obtained when the simulation robot performs the non-first single-object rearrangement task.

[0090] The third reward parameter may include: standing up reward, movement reward, stability reward, obstacle avoidance reward, orientation reward, and target approach reward.

[0091] The stand-up reward encourages the emulator to stand up from various initial postures; the movement reward rewards the emulator to move toward the non-first single object; the stability reward encourages the emulator to move toward the non-first single object in a stable and coordinated posture; the obstacle avoidance reward encourages the emulator to automatically avoid obstacles; the orientation reward encourages the emulator to face the non-first single object; and the target approach reward encourages the emulator to reduce the distance between itself and the non-first single object.

[0092] Based on the standing information, movement information, stability information, obstacle avoidance information, orientation information, and target proximity information of the simulated robot when performing this non-first single-object rearrangement task, the third reward parameter is calculated using a preset calculation formula.

[0093] S505. Based on the third reward parameter, train the preset rearrangement model to obtain the second teacher model.

[0094] Based on the third reward parameter, a reinforcement learning strategy is adopted to train the preset rearrangement model, and the trained preset rearrangement model is used as the second teacher model.

[0095] In this embodiment, the first teacher model, after training, completes the rearrangement, release, and backward movement of the first single object. Using this non-standard posture as the initial posture, the second teacher model is trained. During the movement, the model can perceive and respond to the surrounding environment, autonomously avoiding potential obstacles without relying on an explicitly designed navigation module. Furthermore, even if the initial posture is non-standard, the model can control the simulated robot to turn around and approach non-first single objects, achieving a smooth transition and reliable control from a non-standard posture. At the same time, when obstacles are present, it can smoothly detour without relying on an explicit motion planning module, proving that obstacle avoidance behavior can be autonomously learned through the training process. Therefore, the training of the second teacher model significantly improves the overall success rate of the task.

[0096] Figure 6 A flowchart illustrating the multi-object rearrangement model training method provided in this application embodiment. Figure 6 ,like Figure 6 As shown, in an optional implementation, step S101 above, which generates the first action instruction and the second action instruction for the first single-object rearrangement task in the first multi-object rearrangement task based on the first-view image of the simulated robot and the first natural language instruction for the first multi-object rearrangement task, using a student model and a pre-trained first teacher model respectively, may include: S601. Based on the first-view image, the first natural language instruction, and the first ontology historical data, the first action instruction and the second action instruction are generated using the student model and the first teacher model, respectively.

[0097] The first body history data is the body perception data of the simulated robot before the acquisition time corresponding to the first-view image.

[0098] Proprioceptive data refers to the proprioceptive data collected by the simulated robot through its internal sensors, which may include joint positions, joint velocities, joint torques, and centers of mass. In other words, historical proprioceptive data is collected from historical time prior to the acquisition time corresponding to the first-view image. The natural language command for the first single-object rearrangement task in the first multi-object rearrangement task is determined from the first natural language command. Based on the first-view image, the natural language command for the first single-object rearrangement task, and the historical proprioceptive data, a student model is used to generate the first action command. Similarly, based on the first-view image, the natural language command for the first single-object rearrangement task, and the historical proprioceptive data, a teacher model is used to generate the second action command.

[0099] Step S103 above, based on the second-view image and the first natural language instruction, uses the student model and the pre-trained second teacher model respectively to generate the third and fourth action instructions for the non-first single-object rearrangement task in the first multi-object rearrangement task, which may include: S602. Based on the second-view image, the first natural language instruction, and the second ontology historical data, the third action instruction and the fourth action instruction are generated using the student model and the second teacher model, respectively.

[0100] The second body history data is the body perception data of the simulated robot before the acquisition time corresponding to the second-view image.

[0101] Historical data of the second ontology is collected at a historical time prior to the acquisition time corresponding to the second-view image. Natural language instructions for non-first single-object rearrangement tasks in the first multi-object rearrangement task are determined from the first natural language instructions. Based on the second-view image, the natural language instructions for the non-first single-object rearrangement task, and the historical data of the second ontology, a student model is used to generate a third action instruction. Based on the second-view image, the natural language instructions for the non-first single-object rearrangement task, and the historical data of the second ontology, a second teacher model is used to generate a fourth action instruction.

[0102] In this embodiment, compared to relying solely on perspective images and natural language commands, by introducing the ontological historical data of the simulated robot, the model has the ability to continuously perceive the state of the simulated robot itself, which significantly improves the physical consistency of action generation (such as avoiding large-scale upper limb grasping when supporting on one leg) and task continuity (such as naturally connecting backward and turning postures after completing the previous action).

[0103] Figure 7 This is a schematic diagram illustrating the training process of the multi-object rearrangement model provided in the embodiments of this application, as shown below. Figure 7As shown, Stage 1: Single-object rearrangement training is performed on the preset rearrangement model to obtain the first teacher model. Stage 2: Placement and release and backtracking training is performed on the first teacher model. Stage 3: Multi-object rearrangement training is performed on the preset rearrangement model to obtain the second teacher model. Stage 4: The dual-teacher model is distilled into a VLA student model.

[0104] For Phase 4, a task switching mechanism is triggered by a state discriminator. When the state discriminator determines that the simulated robot is performing its first single-object task, it switches to the first teacher model to provide motion supervision signals. When the state discriminator determines that the simulated robot is performing a task other than the first single-object task, it switches to the second teacher model to provide motion supervision signals. This dynamic teacher model selection mechanism ensures that the VLA student model receives supervision and guidance matching the task phase throughout the training process. The inputs to the VLA student model include natural language commands, first-person perspective images, and ontology history data.

[0105] The multi-object rearrangement model obtained after model training can complete multiple consecutive rearrangements of a single object using first-person perspective images and natural language commands. The following describes a multi-object rearrangement method provided in this embodiment.

[0106] Figure 8 This is a flowchart illustrating a multi-object rearrangement method based on a humanoid robot provided in an embodiment of this application. The executing entity in this embodiment can be a humanoid robot.

[0107] like Figure 8 As shown, the method may include: S701, acquire robot-view images of the humanoid robot and natural voice commands for the target multi-object rearrangement task.

[0108] Among them, the robot's perspective image is the first-person perspective image of the humanoid robot in the actual environment scene. It refers to the image collected from the perspective of the humanoid robot itself and is used to indicate what the humanoid robot sees.

[0109] The target multi-object rearrangement task comprises multiple consecutive rearrangement tasks of individual objects. The robot's perspective image includes: the individual object corresponding to each rearrangement task, the starting position of the rearrangement, the ending position of the rearrangement, and obstacles. Thus, by acquiring the robot's perspective image, the humanoid robot can perceive the spatial position of each individual object to be rearranged in the actual environment, thereby establishing its spatial relationship with each individual object, the starting position of the rearrangement, the ending position of the rearrangement, and obstacles, enabling it to autonomously avoid obstacles and perform the individual object rearrangement task.

[0110] The natural language instructions for the target multi-object rearrangement task are used to instruct each individual object rearrangement task contained in the target multi-object rearrangement task. These natural language instructions are text instructions for each individual object rearrangement task contained in the target multi-object rearrangement task, and their specific forms are similar to the first natural language instructions mentioned above, so no further examples will be given here.

[0111] S702. Based on the robot's perspective image and the natural speech instructions for the target multi-object rearrangement task, a multi-object rearrangement model is used to generate the target action instructions for the humanoid robot.

[0112] The multi-object rearrangement model is a model trained using the multi-object rearrangement model training method described above.

[0113] Based on the robot's perspective image and the natural language instructions for the target multi-object rearrangement task, a multi-object rearrangement model is used to extract features from the robot's perspective image to obtain visual features. Semantic parsing is then performed on the natural language instructions for the target multi-object rearrangement task to generate target action instructions for the humanoid robot. These target action instructions are used to control the humanoid robot to perform the object rearrangement task in the target multi-object rearrangement task.

[0114] In an optional implementation, a multi-object rearrangement model is used to generate target action commands based on robot-view images, natural speech commands for the target multi-object rearrangement task, and target ontological historical data. The target ontological historical data refers to the humanoid robot's proprioceptive data prior to the acquisition time corresponding to the robot-view images.

[0115] Specifically, historical data of the target body is collected before the acquisition time corresponding to the robot's perspective image. Based on the robot's perspective image, the natural language commands for the target multi-object rearrangement task, and the historical data of the target body, a multi-object rearrangement model is used to generate target action commands. In this way, by introducing the historical data of the target body of the simulated robot, the model has the ability to continuously perceive the state of the humanoid robot itself, significantly improving the physical consistency and task coherence of action generation.

[0116] S703. Control the humanoid robot to perform the target multi-object rearrangement task according to the target action command.

[0117] Among them, the target action instructions include the joint control actions of the humanoid robot when performing the object rearrangement task in the target multi-object rearrangement task. According to the target action instructions, the humanoid robot can be controlled to perform the target multi-object rearrangement task.

[0118] In this embodiment, by using continuous multi-step target natural instructions and robot perspective images, semantic parsing is performed through a multi-object rearrangement model to extract visual features. Combined with operational experience learned through distillation, joint control actions are generated and action instructions are output to drive the humanoid robot to complete multi-object rearrangement tasks corresponding to coherent behaviors such as grasping, transporting, placing, retreating, and turning.

[0119] Figure 9 This is a schematic diagram of the structure of a multi-object rearrangement model training device provided in an embodiment of this application. This device can be integrated into a model training device.

[0120] like Figure 9 As shown, the device may include: The processing module 801 is used to generate a first action instruction and a second action instruction for the first single-object rearrangement task in the first multi-object rearrangement task based on the first-view image of the simulation robot and the first natural language instruction of the first multi-object rearrangement task, using a student model and a pre-trained first teacher model respectively. The first action instruction is used to control the simulation robot to execute the first single-object rearrangement task in the first multi-object rearrangement task. The acquisition module 802 is used to acquire a second-view image of the simulated robot after the first single-object rearrangement task is completed in the first multi-object rearrangement task; The processing module 801 is also used to generate a third action instruction and a fourth action instruction for a non-first single object rearrangement task in the first multi-object rearrangement task based on the second perspective image and the first natural language instruction, using the student model and the pre-trained second teacher model respectively. The third action instruction is used to control the simulation robot to perform the non-first single object rearrangement task in the first multi-object rearrangement task. The processing module 801 is also used to supervise the training of the student model according to the first action instruction, the second action instruction, the third action instruction and the fourth action instruction to generate a multi-object rearrangement model.

[0121] In an optional implementation, the processing module 801 is further configured to: Based on the second natural language instruction of the single object rearrangement task and the viewpoint image corresponding to the single object rearrangement task, the preset rearrangement model is trained for the single object rearrangement task to obtain the first teacher model. Based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task, the preset rearrangement model is trained on the multi-object rearrangement task to obtain the second teacher model.

[0122] In an optional implementation, the processing module 801 is specifically used for: A preset rearrangement model is adopted, and a third action instruction is generated based on the second natural language instruction and the viewpoint image corresponding to the single object rearrangement task. The third action instruction is used to control the simulation robot to perform the single object rearrangement task. Obtain the first reward parameter when the simulated robot performs a single-object rearrangement task; Based on the first reward parameter, the preset rearrangement model is trained to obtain the first teacher model.

[0123] In an optional implementation, the acquisition module 802 is further configured to: Obtain the task completion status information after the simulated robot performs a single object rearrangement task. The task completion status information includes: the distance between the simulated robot's torso and the single object corresponding to the single object rearrangement task, the distance between the simulated robot's hand and the single object, and the distance between the single object and the rearrangement termination position. Based on the task completion status information, obtain the second reward parameters; Processing module 801 is specifically used for: Based on the first reward parameter and the second reward parameter, the preset rearrangement model is trained to obtain the first teacher model.

[0124] In an optional implementation, the processing module 801 is specifically used for: Based on the third natural language instructions and the third-person perspective image of the simulated robot, the first teacher model is used to generate the fourth action instructions; the fourth action instructions are used to control the simulated robot to perform the first single-object rearrangement task in the second multi-object rearrangement task. Acquire a fourth-view image of the simulated robot after performing the first single-object rearrangement task in the second multi-object rearrangement task; Based on the third natural language instruction and the fourth perspective image, a preset rearrangement model is used to generate a fifth action instruction. The fifth action instruction is used to control the simulation robot to perform the non-first single object rearrangement task in the second multi-object rearrangement task. Obtain the third reward parameter when the simulated robot performs a second multi-object rearrangement task that is not the first single-object rearrangement task; Based on the third reward parameter, the preset rearrangement model is trained to obtain the second teacher model.

[0125] In an optional implementation, the processing module 801 is specifically used for: Based on the first-view image, the first natural language command, and the first ontology historical data, the first action command and the second action command are generated using the student model and the first teacher model, respectively. The first ontology historical data is the ontological perception data of the simulated robot before the acquisition time corresponding to the first-view image. Based on the second-view image, the first natural language command, and the second ontology historical data, the third and fourth action commands are generated using the student model and the second teacher model, respectively. The second ontology historical data is the ontological perception data of the simulated robot before the acquisition time corresponding to the second-view image.

[0126] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0127] Figure 10 This is a schematic diagram of the structure of a multi-object rearrangement device based on a humanoid robot provided in an embodiment of this application. The device can be integrated into a humanoid robot.

[0128] like Figure 10 As shown, the device may include: The acquisition module 901 is used to acquire robot-view images of the humanoid robot and natural voice commands for the target multi-object rearrangement task. Processing module 902 is used to generate target action instructions for the humanoid robot based on the robot's perspective image and the natural speech instructions for the target multi-object rearrangement task, using a multi-object rearrangement model. The multi-object rearrangement model is a model trained using the above-mentioned multi-object rearrangement model training method. The control module 903 is used to control the humanoid robot to perform a target multi-object rearrangement task according to the target action instructions.

[0129] In an optional implementation, the processing module 902 is specifically used for: Based on the robot's perspective images, the natural speech commands for the target multi-object rearrangement task, and the target's historical data, a multi-object rearrangement model is used to generate target action commands. The target's historical data consists of the humanoid robot's proprioceptive data prior to the acquisition time corresponding to the robot's perspective images.

[0130] Figure 11 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 11 As shown, the device may include a processor 1001 and a memory 1002. The memory 1002 stores machine-readable instructions that can be executed by the processor 1001. When the electronic device is running, the processor 1001 executes the machine-readable instructions to perform the above-mentioned multi-object rearrangement model training method or the multi-object rearrangement method based on humanoid robots.

[0131] The electronic device can be either the aforementioned model training device or a humanoid robot.

[0132] This application also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the above-described method.

[0133] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.

[0134] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0136] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0137] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0138] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0139] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A training method for a multi-object rearrangement model, characterized in that, include: Based on the first-view image of the simulated robot and the first natural language instruction of the first multi-object rearrangement task, the first action instruction and the second action instruction of the first single-object rearrangement task in the first multi-object rearrangement task are generated by using the student model and the pre-trained first teacher model, respectively. The first action instruction is used to control the simulated robot to execute the first single-object rearrangement task in the first multi-object rearrangement task. Acquire a second-view image of the simulated robot after the first single-object rearrangement task in the first multi-object rearrangement task has been completed; Based on the second perspective image and the first natural language instruction, the student model and the pre-trained second teacher model are used respectively to generate the third action instruction and the fourth action instruction for the non-first single object rearrangement task in the first multi-object rearrangement task. The third action instruction is used to control the simulation robot to perform the non-first single object rearrangement task in the first multi-object rearrangement task. The student model is trained under supervision based on the first action instruction, the second action instruction, the third action instruction, and the fourth action instruction to generate a multi-object rearrangement model.

2. The method according to claim 1, characterized in that, The method further includes: Based on the second natural language instruction for the single object rearrangement task and the viewpoint image corresponding to the single object rearrangement task, the preset rearrangement model is trained for the single object rearrangement task to obtain the first teacher model. Based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task, the preset rearrangement model is trained on the multi-object rearrangement task to obtain the second teacher model.

3. The method according to claim 2, characterized in that, The step of training a preset rearrangement model for the single-object rearrangement task based on the second natural language instruction for the single-object rearrangement task and the viewpoint image corresponding to the single-object rearrangement task to obtain the first teacher model includes: Using the preset rearrangement model, a third action instruction is generated based on the second natural language instruction and the viewpoint image corresponding to the single object rearrangement task. The third action instruction is used to control the simulation robot to perform the single object rearrangement task. Obtain the first reward parameter when the simulated robot performs the single-object rearrangement task; Based on the first reward parameter, the preset rearrangement model is trained to obtain the first teacher model.

4. The method according to claim 3, characterized in that, The method further includes: The task completion status information of the simulated robot after performing the single object rearrangement task is obtained, wherein the task completion status information includes: the distance between the simulated robot's torso and the single object corresponding to the single object rearrangement task, the distance between the simulated robot's hand and the single object, and the distance between the single object and the rearrangement termination position. Based on the task completion status information, obtain the second reward parameter; The step of training the preset rearrangement model according to the first reward parameter to obtain the first teacher model includes: The preset rearrangement model is trained based on the first reward parameter and the second reward parameter to obtain the first teacher model.

5. The method according to claim 2, characterized in that, The step of training the preset rearrangement model for the multi-object rearrangement task based on the third natural language instruction of the second multi-object rearrangement task and the viewpoint image corresponding to the second multi-object rearrangement task to obtain the second teacher model includes: Based on the third natural language instruction and the third-view image of the simulated robot, the first teacher model is used to generate a fourth action instruction; the fourth action instruction is used to control the simulated robot to perform the first single-object rearrangement task in the second multi-object rearrangement task. Acquire a fourth-view image of the simulated robot after it has performed the first single-object rearrangement task in the second multi-object rearrangement task; Based on the third natural language instruction and the fourth perspective image, the fifth action instruction is generated using the preset rearrangement model. The fifth action instruction is used to control the simulation robot to perform a non-first single-object rearrangement task in the second multi-object rearrangement task. Obtain the third reward parameter when the simulated robot performs the second multi-object rearrangement task, which is not the first single-object rearrangement task; Based on the third reward parameter, the preset rearrangement model is trained to obtain the second teacher model.

6. The method according to claim 1, characterized in that, Based on the first-view image of the simulated robot and the first natural language instruction of the first multi-object rearrangement task, the first action instruction and the second action instruction of the first single-object rearrangement task in the first multi-object rearrangement task are generated using a student model and a pre-trained first teacher model, respectively, including: Based on the first perspective image, the first natural language instruction, and the first ontology history data, the first action instruction and the second action instruction are generated using the student model and the first teacher model, respectively. The first ontology history data is the ontological perception data of the simulated robot before the acquisition time corresponding to the first perspective image. The step involves generating third and fourth action instructions for non-first single-object rearrangement tasks in the first multi-object rearrangement task, based on the second-view image and the first natural language instruction, using the student model and the pre-trained second teacher model respectively, including: Based on the second perspective image, the first natural language instruction, and the second ontology history data, the third action instruction and the fourth action instruction are generated using the student model and the second teacher model, respectively. The second ontology history data is the ontological perception data of the simulated robot before the acquisition time corresponding to the second perspective image.

7. A method for rearranging multiple objects based on a humanoid robot, characterized in that, include: Acquire robot-perspective images of a humanoid robot and natural speech commands for a multi-object rearrangement task; Based on the robot's perspective image and the natural speech command of the target multi-object rearrangement task, a multi-object rearrangement model is used to generate the target action command of the humanoid robot. The multi-object rearrangement model is a model trained using the method described in any one of claims 1-6. According to the target action command, the humanoid robot is controlled to perform the target multi-object rearrangement task.

8. The method according to claim 7, characterized in that, The step of generating target action commands for the humanoid robot based on the robot's perspective image and the natural language speech commands for the target multi-object rearrangement task, using a multi-object rearrangement model, includes: Based on the robot's perspective image, the natural language command for the target multi-object rearrangement task, and the target's historical data, the target action command is generated using the multi-object rearrangement model. The target's historical data refers to the humanoid robot's proprioceptive data prior to the acquisition time corresponding to the robot's perspective image.

9. An electronic device, characterized in that, include: A processor and a memory, the memory storing machine-readable instructions executable by the processor, wherein when the electronic device is running, the processor executes the machine-readable instructions to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the method according to any one of claims 1 to 8.