Pseudo-label generation method and robot system for somatic robot self-learning

By acquiring multimedia data of the interaction between the embodied robot and the learning object, the system automatically decomposes the state multimedia data to generate components and state pseudo-labels, solving the problems of high cost and easy distortion of traditional label generation, and improving the recognition accuracy and interactive control stability of the embodied robot for the learning object.

CN122024240BActive Publication Date: 2026-06-23WOCAO TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WOCAO TECH (SHENZHEN) CO LTD
Filing Date
2026-04-10
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Traditional methods for generating tags for embodied robots are costly and prone to distortion, leading to task failures and impacting the success rate.

Method used

By acquiring interactive multimedia data between the embodied robot and the object to be learned, the state multimedia data is automatically decomposed to generate component and state pseudo-labels. The component and state pseudo-labels are combined to generate object pseudo-labels for the embodied robot's self-learning.

Benefits of technology

It significantly reduces data annotation costs, improves the recognition accuracy and state judgment reliability of embodied robots when learning objects, and enhances the stability and success rate of autonomous operation and interactive control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024240B_ABST
    Figure CN122024240B_ABST
Patent Text Reader

Abstract

The application relates to a pseudo-label generation method and a robot system for embodied robot self-learning. The method comprises the following steps: acquiring interactive multimedia data between an embodied robot and a to-be-learned object; determining state multimedia data from the interactive multimedia data; the state multimedia data reflects state changes of the to-be-learned object in the interaction process with the embodied robot; performing motion decomposition on the state multimedia data to obtain part multimedia data of key parts of the to-be-learned object, and generating part pseudo-labels about the to-be-learned object based on the part multimedia data; generating state pseudo-labels for reflecting the state changes of the to-be-learned object based on the state multimedia data; generating object pseudo-labels about the to-be-learned object based on the state pseudo-labels and the part pseudo-labels; and the object pseudo-labels are used for learning and training of the to-be-learned object by the embodied robot. The method can improve the quality of the pseudo-labels, and further improve the task success rate of the embodied robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied robot technology, and in particular to a pseudo-label generation method and robot system for self-learning of embodied robots. Background Technology

[0002] With the development of technologies in the field of artificial intelligence, embodied robots are being used more and more widely, and their functions are becoming increasingly powerful. For example, home service robots can perform tasks according to user instructions. When embodied robots perform tasks, they mainly rely on the labels of the interactive objects. In traditional technologies, label generation mainly takes two forms: one is to rely on manual pixel-by-pixel annotation; the other is to rely on model prediction or teacher network output.

[0003] However, in the two traditional approaches, the embodied robot faces diverse object forms and long-tailed categories, making manual pixel-by-pixel annotation costly and causing distortion in model predictions or teacher network outputs, which leads to task execution failures and affects the task success rate. Summary of the Invention

[0004] Therefore, it is necessary to provide a pseudo-label generation method and robot system for self-learning of embodied robots to address the above-mentioned technical problems, which can improve the quality of pseudo-labels and thus improve the success rate of interactive tasks of embodied robots.

[0005] In a first aspect, this application provides a method for generating pseudo-labels for self-learning in embodied robots, including:

[0006] Acquire multimedia data of the interaction between the embodied robot and the learning object;

[0007] State multimedia data is determined from interactive multimedia data; state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot.

[0008] Motion decomposition is performed on the state multimedia data to obtain the component multimedia data of the key components of the object to be learned, and component pseudo-labels about the object to be learned are generated based on the component multimedia data.

[0009] Based on state multimedia data, generate state pseudo-labels to reflect the state changes of the object to be learned.

[0010] Based on state pseudo-labels and component pseudo-labels, object pseudo-labels are generated for the object to be learned; these object pseudo-labels are used by the embodied robot to learn and train on the object to be learned.

[0011] Secondly, this application also provides a body-worn robot system, which includes:

[0012] The acquisition module is used to acquire multimedia data of the interaction between the embodied robot and the object to be learned;

[0013] A determination module is used to determine state multimedia data from the interactive multimedia data; the state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot;

[0014] The first generation module is used to perform motion decomposition on the state multimedia data to obtain component multimedia data of the key components of the object to be learned, and generate component pseudo-labels for the object to be learned based on the component multimedia data.

[0015] The second generation module is used to generate pseudo-state labels that reflect the state changes of the object to be learned based on the state multimedia data.

[0016] The third generation module is used to generate object pseudo-labels for the object to be learned based on the state pseudo-labels and the component pseudo-labels; the object pseudo-labels are used by the embodied robot to learn and train the object to be learned.

[0017] Thirdly, this application also provides a computer device, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above-described pseudo-tag generation method for self-learning of embodied robots.

[0018] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described pseudo-label generation method for self-learning of embodied robots.

[0019] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described pseudo-label generation method for self-learning of embodied robots.

[0020] The aforementioned pseudo-label generation method and robot system for embodied robot self-learning acquire interactive multimedia data between the embodied robot and the learning target, and determine state multimedia data from this data. Further, motion decomposition is performed on the state multimedia data to obtain component multimedia data of the key parts of the learning target, and component pseudo-labels are generated based on this data. Thus, state pseudo-labels reflecting state changes of the learning target are generated based on the state multimedia data. Finally, object pseudo-labels are generated based on the state and component pseudo-labels; these are used by the embodied robot for learning and training. In this process, by collecting interactive multimedia data between the embodied robot and the learning target, key component information and state change information are automatically decomposed, generating component, state, and object pseudo-labels without manual annotation, significantly reducing data annotation costs and manpower consumption. On the other hand, by fusing key component information and dynamic state information of an object to form a complete object pseudo-label, embodied robots can learn the composition structure, motion laws and state change characteristics of the object to be learned more comprehensively and accurately, thereby improving the recognition accuracy and state judgment reliability of different objects to be learned by embodied robots, and thus improving the stability and success rate of subsequent autonomous operation and interactive control. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is an application environment diagram of a pseudo-label generation method for self-learning of embodied robots in one embodiment;

[0023] Figure 2 This is a flowchart illustrating a pseudo-label generation method for self-learning of an embodied robot in one embodiment.

[0024] Figure 3 This is a flowchart illustrating the steps for determining multimedia data of a component in one embodiment;

[0025] Figure 4 This is a flowchart illustrating the steps for determining motion coordination relationships in one embodiment;

[0026] Figure 5 This is a flowchart illustrating a pseudo-label-based self-learning method for embodied robots in one embodiment.

[0027] Figure 6 This is a flowchart illustrating the steps for determining simulated multimedia data in one embodiment;

[0028] Figure 7 This is a flowchart illustrating the training steps for an embodied robot in one embodiment;

[0029] Figure 8 This is a structural block diagram of a pseudo-label generation device for self-learning of an embodied robot in one embodiment;

[0030] Figure 9 This is a structural block diagram of a pseudo-label-based embodied robot self-learning device in one embodiment;

[0031] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0033] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0034] The pseudo-label generation method for self-learning of embodied robots provided in this application embodiment can be applied to, for example... Figure 1The application scenario shown includes an embodied robot 102 and a target environment, with the embodied robot 102 situated within the target environment. The embodied robot possesses perception, movement, and interaction capabilities, and can interact with its environment in real time. It can capture information about its surroundings through sensory organs such as cameras, LiDAR, or tactile LiDAR installed on the embodied robot. The embodied robot can be, but is not limited to, cleaning robots, companion robots, and humanoid robots (in this embodiment, the embodied robot is a humanoid robot, and the target environment is a home indoor environment for illustration). The cleaning robot can be, but is not limited to, sweeping robots, mopping robots, and sweeping-and-mopping robots. The target environment can be any environment suitable for the application of the embodied robot, including, but is not limited to, home indoor environments, industrial manufacturing environments, medical environments, warehousing and logistics environments, agricultural production environments, public service environments, educational and research environments, or entertainment performance environments.

[0035] Currently, embodied robots face complex and varied application scenarios, serving diverse and long-tailed objects with dynamically changing placement positions over time. To enable embodied robots to perceive and understand target objects in their environment and support subsequent interactive tasks, traditional solutions primarily rely on manual annotation to obtain object labels or the generation of pseudo-labels through large-scale model predictions. However, when embodied robots interact with manipulable or variable-state objects, they require not only the object's shape features but also the characteristics of its manipulable components and its real-time state as decision-making criteria. Labels generated by traditional manual annotation and large-scale model predictions typically only represent the object's external shape and fail to provide effective information at the component and state levels, thus failing to meet the needs of sophisticated robot interaction.

[0036] Meanwhile, pixel-by-pixel manual annotation is not only costly, but also difficult to guarantee stable annotation accuracy. Relying solely on large model prediction is prone to problems such as misidentifying key components of interactive objects, such as handles and hinges, and unstable judgment of opening and closing states in unfamiliar environments, which directly reduces the success rate of task completion.

[0037] This application aims to address the aforementioned deficiencies in traditional technologies. Specifically, the core technical problems to be solved include: how to generate high-quality pseudo-labels for embodied robots to learn from, serving as data for subsequent interactions of the embodied robot, thereby improving the robot's object recognition accuracy, object state judgment stability, and promoting the robot's successful execution of interactive tasks.

[0038] In one exemplary embodiment, such as Figure 2 As shown, a pseudo-label generation method for self-learning in embodied robots is provided, which is then applied to... Figure 1 Taking the embodied robot in the example, the following steps are included:

[0039] S210: Acquire multimedia data of the interaction between the embodied robot and the object to be learned.

[0040] In this context, the learning target can be understood as the object that the embodied robot needs to learn in the target environment. By generating pseudo-labels for the learning target through multimedia data generated from the interaction between the robot and the learning target, the embodied robot can become familiar with or better understand the object's shape, operational state, and other characteristics. This allows the embodied robot to achieve task success and applicability when performing tasks related to the learning target. For example, the learning target can be a task object that the embodied robot is "unfamiliar" with or even "never seen before." The task object can be understood as the operational target that the embodied robot aims to achieve when performing an operational task.

[0041] Interactive multimedia data can be understood as multimedia data when the embodied robot interacts with the object to be learned, such as at least one of the following: images (RGB images, RGB-D depth maps), videos (continuous video frames), point cloud data, and audio.

[0042] The following describes the steps for acquiring interactive multimedia data:

[0043] In one optional implementation, the interactive multimedia data may include video data. When performing an operation task, the embodied robot can collect interactive video data between itself and the task object in real time through an RGB camera mounted on the embodied robot, and store the interactive video data in a preset interaction record database. If the task object is determined to be a learning object, the embodied robot can obtain the interactive video data between itself and the corresponding task object from the preset interaction record database as interactive multimedia data between the embodied robot and the learning object.

[0044] Optionally, interactive video data can also carry depth information. Accordingly, based on the RGB camera acquisition, color and depth information can be simultaneously acquired by an RGB-D depth camera to form interactive video data carrying depth information; or, depth estimation can be performed on RGB images using a monocular depth estimation model to generate interactive video data carrying depth information.

[0045] In one alternative implementation, the interactive multimedia data may further include interactive audio data. When performing interactive tasks, the embodied robot can collect interactive audio data in real time via a microphone and fuse the collected audio data with the aforementioned interactive video data to form interactive multimedia data between the robot and the task object.

[0046] For example, the interactive audio data may include friction sounds, collision sounds, etc. generated when interacting with the task object, which can be used as a reference for subsequent determination of pseudo-labels.

[0047] It is worth noting that, in order to reduce unnecessary computational overhead and thus improve the effectiveness and efficiency of pseudo-tag generation, in this embodiment, not all task objects are used as learning objects. Instead, only task objects that meet preset requirements are considered as learning objects. The process of determining learning objects is described below:

[0048] In one optional implementation, task multimedia data during the interaction between the embodied robot and the task object can be acquired; based on the task multimedia data, object feature information of the task object can be determined; based on the object feature information, the familiarity level of the embodied robot with respect to the task object can be determined; if the familiarity level is less than or equal to a preset threshold, the task object is determined as the learning object, and correspondingly, the task multimedia data is determined as the interactive multimedia data between the embodied robot and the learning object.

[0049] Among them, task multimedia data can be understood as one or more of the following: video, audio, RGB or RGB-D depth images, point cloud data, etc., collected by the embodied robot based on its own sensors during the interaction with the task object when performing interactive tasks.

[0050] Here, object feature information can be understood as information extracted from task multimedia data that can characterize the features of the task object. For example, it may include at least one of appearance feature information, structural feature information, state change feature information, and key component distribution feature information.

[0051] For example, during the interaction between the embodied robot and the task object, one or more of the following devices can be used: image acquisition device, audio acquisition device, depth sensing device, and point cloud acquisition device. These devices can collect video, audio, RGB-D images, and point cloud data during the interaction, serving as multimedia data for the task. Feature extraction is performed on this multimedia data, and the extracted results are used as the object feature information of the task object. This object feature information is then matched against existing object features in a preset object memory (including feature information of historical interaction objects, the number of interactions with each historical interaction object, the relationships between historical interaction objects, the number of learning attempts for each historical interaction object, and the recognition accuracy rate). Based on at least one of the feature matching degree, the number of historical learning attempts, and the historical recognition accuracy rate, the familiarity level of the embodied robot with the task object is determined. Furthermore, based on the relationship between the familiarity level and a preset threshold, it is determined whether the task object is a learning object. For example, if the familiarity level exceeds the preset threshold, the corresponding task object is determined to be a learning object.

[0052] It should be noted that the method for feature extraction of task multimedia data can be a common feature extraction method, such as inputting the task multimedia data into a feature extraction network to obtain the feature extraction results, which will not be elaborated here.

[0053] Optionally, there are many ways to determine the embodied robot's familiarity with the task object based on at least one of feature matching degree, historical learning count, and historical recognition accuracy. For example, the similarity score between the object feature information of the task object and the existing object feature information in a preset object memory can be calculated, and the resulting feature matching degree score can be used as the embodied robot's familiarity with the task object. Another example is to predetermine the correspondence between historical learning counts and reference familiarity, query the historical learning counts of the task object in the preset object memory, and use the reference familiarity corresponding to those historical learning counts as the embodied robot's familiarity with the task object. Yet another example is to predetermine the correspondence between historical recognition accuracy and reference familiarity, query the historical recognition accuracy corresponding to the task in the preset object memory, and use the reference familiarity corresponding to those historical recognition accuracy as the embodied robot's familiarity with the task object. Furthermore, the familiarity levels determined based on each dimension can be weighted and fused to obtain the final familiarity level. In this process, the weight coefficients corresponding to different dimensions can be determined based on human experience, and this application does not impose any limitations on this.

[0054] In some embodiments, the object feature information includes object type information and object mask information. Accordingly, if the object type information of the task object is the target object type, the object mask information of the task object is matched with each candidate mask information included in the memory of the embodied robot to obtain a mask matching result. If the mask matching result indicates that there is no candidate mask information in the memory that matches the object mask information, it is determined that the familiarity of the embodied robot with respect to the task object is less than or equal to a preset degree threshold.

[0055] The object type information is used to characterize the type of the task object. The target object type includes one or more of the following: operable object type and variable-state object type. Operable object type can be understood as an object category that possesses structural attributes that allow the embodied robot to perform physical operations and can interact with the embodied robot's end effector. Variable-state object type can be understood as an object category that can observably, quantify, and repeatedly switch between multiple stable states. Examples include objects with built-in handles, pulls, latches, and flap edges that can be located, grasped, and force-applied by the embodied robot; or objects containing discrete states such as closed, half-open, fully open, retracted, extended, and flap-closed / open, such as cabinet doors, drawers, refrigerator doors, washing machine doors, and storage cabinet flaps.

[0056] The object mask information is a mask used to characterize the outline, region, and pixel-level range of the task object. The candidate mask information in the memory can be the mask information from historically generated pseudo-labels. This is used to determine whether a pseudo-label for the task object has already been generated in a historical time period. If a pseudo-label for the task object has already been generated in a historical time period, it is not necessary to generate another pseudo-label; otherwise, it is necessary to generate a pseudo-label. The candidate mask information consists of one or more object mask data already stored in the preset memory.

[0057] S220, determine the state multimedia data from the interactive multimedia data.

[0058] Among them, state multimedia data is used to reflect the state changes of the learning object during the interaction with the embodied robot. The state multimedia data includes the complete process of the learning object changing from one state to another during the interaction with the embodied robot.

[0059] In one optional implementation, frame difference calculation can be performed on each adjacent video frame in the interactive multimedia data. For the current video frame: if the frame difference between the current video frame and its previous frame is greater than a preset state change threshold, then the previous frame of the current video frame is marked as the state change start frame; if the frame difference between the current video frame and its next frame is less than a preset stability threshold, then the current video frame is marked as the state change end frame; the continuous video frame segments between the state change start frame and the state change end frame are determined as state multimedia data reflecting the state change of the object to be learned.

[0060] In another alternative implementation, optical flow field calculations can be performed on consecutive RGB-D video frames in the interactive multimedia data to obtain pixel-level motion vectors for each frame. Based on the object mask information of the object to be learned, the background area is removed and only the pixel motion vectors of the object area to be learned are retained. When the amplitude of the pixel motion vector in this area is greater than a preset motion amplitude threshold and the motion direction remains stable and consistent, the corresponding RGB-D image sequence data within this continuous time period is extracted and used as state multimedia data.

[0061] In another alternative implementation, the contact time between the embodied robot and the learning object can be determined based on the real-time feedback information from the end effector of the embodied robot, and this time can be used as the start time of the interactive action; the separation time between the embodied robot and the learning object can be determined, and this time can be used as the end time of the interactive action; the multimedia data in the interactive multimedia data from the start time of the interactive action to the end time of the interactive action can be used as the state multimedia data.

[0062] It's important to note that the purpose of identifying state multimedia data from interactive multimedia data is to extract data showing genuine state changes in the object to be learned. For example, interactive multimedia data might include the entire process of an embodied robot's end effector approaching the object, making contact with it, manipulating it, and then leaving it. State multimedia data, on the other hand, only includes the process of the embodied robot manipulating the object. This process preserves the core segments of state changes in the object and eliminates redundant data with no movement or change, reducing subsequent data processing and thus improving the efficiency of pseudo-label generation.

[0063] S230, perform motion decomposition on the state multimedia data to obtain component multimedia data of the key components of the object to be learned, and generate component pseudo-labels for the object to be learned based on the component multimedia data.

[0064] Motion decomposition refers to the process of breaking down the overall motion of an object into multiple independent or interrelated component motions. There are many methods for performing motion decomposition on state multimedia data, such as at least one of optical flow, feature tracking, and frame difference methods.

[0065] In this context, key components can be understood as constituent parts of the object to be learned. When the object to be learned is a movable object, its corresponding key components refer to the functional parts that constitute the movable structure of the movable object, including at least moving parts and stationary parts. Component multimedia data refers to the multimedia data corresponding to the key components, such as images including the key components.

[0066] In one alternative implementation, motion decomposition of state multimedia data can be performed based on optical flow. For example, for consecutive video frames of state multimedia data, a pixel-by-pixel motion vector can be calculated using an optical flow algorithm, retaining the optical flow information within the region to be learned and eliminating background interference optical flow. Through direction or amplitude analysis of the optical flow vectors, the motion trajectory of key components in the region to be learned can be located, and the multimedia data corresponding to each key component region can be used as the corresponding component multimedia data.

[0067] In another alternative implementation, motion decomposition of the state multimedia data can be performed based on feature tracking. For example, key feature points of the object to be learned can be extracted from the state multimedia data, and these feature points can be continuously tracked in consecutive video frames to record their motion trajectories. Based on the motion trajectory of the feature point, the component type corresponding to that feature point is determined (e.g., if the motion trajectory indicates that the feature point is stationary, it is determined to be a stationary component; if the motion trajectory indicates that the feature point moves over time, it is determined to be a moving component), and the multimedia data corresponding to that component region is used as the multimedia data of the corresponding component. Optionally, the key feature points of the object to be learned in the multimedia data can be extracted using a common Scale-Invariant Feature Transform (SIFT) algorithm to extract unique features such as corners and edges, which will not be elaborated upon here.

[0068] In one optional implementation, the component multimedia data corresponding to each key component includes the image data of the corresponding key component. Accordingly, the image data of the corresponding key component can be masked to obtain the mask label of the corresponding key component, and the mask label of each key component can be used as the component mask pseudo label of the object to be learned.

[0069] It should be noted that there are many other ways to generate pseudo-labels for components of the object to be learned based on component multimedia data, which will be described in the following embodiments.

[0070] S240 generates pseudo-state labels based on state multimedia data to reflect changes in the state of the object to be learned.

[0071] In one optional implementation, the state multimedia data includes M frames of multimedia data, where M is a positive integer; correspondingly, the motion type and motion constraint data of the object to be learned can be obtained; for the i-th frame of multimedia data in the M frames, based on the motion type, motion constraint data, and state data of the object to be learned in the i-th frame of multimedia data, the motion parameters of the object to be learned in the i-th frame of multimedia data are determined; i is a positive integer less than or equal to M; based on the motion parameters and motion constraint data corresponding to the M frames of multimedia data, state pseudo-labels reflecting the state changes of the object to be learned are generated.

[0072] Understandably, since state multimedia data is used to reflect the state changes of the object to be learned during its interaction with the embodied robot, any frame of state multimedia data includes the state of the object to be learned at a specific moment during the transition from one state to another. For example, state multimedia data includes the states of a cabinet door at different moments as it transitions from being completely closed to being completely opened by the embodied robot. Similarly, state multimedia data includes the states of a drawer at different moments as it transitions from being completely closed to being completely pulled out by the embodied robot (i.e., the travel distance of the drawer in different frames when it is pulled by the embodied robot).

[0073] For example, for the i-th frame of multimedia data, the state data of the object to be learned in the current frame is extracted. Combined with the motion type and motion constraint data of the object to be learned, the state data of the current frame is substituted into the motion constraint rules of the object to be learned for calculation. The state data of the object to be learned is converted into quantized data, which serves as the motion parameters of the object to be learned in the i-th frame of multimedia data. For example, if the motion type of the object to be learned is rotation, the motion parameter represents the current opening / closing angle; if the motion type of the object to be learned is translation, the motion parameter represents the current pull-out stroke. The state data includes at least one of the following: the position of the object to be learned, its contour boundary, and the offset of the moving part relative to the stationary part.

[0074] Optionally, the motion parameters corresponding to each frame of multimedia data can be used as pseudo-labels for the state of the object to be learned.

[0075] Optionally, the motion parameters corresponding to each frame in the M frames can be arranged in frame order to form a continuous state change sequence of the object to be learned from the initial state to the final state. Then, according to a preset state interval division rule, the continuous motion parameters are mapped to discrete state identifiers corresponding to the motion constraint data (each motion parameter corresponds to a discrete state identifier). For example, the opening angle can be divided into three intervals: closed, half-open, and fully open; the translational stroke can be divided into three intervals: closed, half-pull, and fully pulled. Furthermore, the discrete state identifier of the motion parameters corresponding to each frame of multimedia data is used as the pseudo-label of the state corresponding to that frame.

[0076] In some embodiments, motion constraint data includes motion constraint direction and motion constraint range. Correspondingly, the movable direction and movable range of the object to be learned can be determined based on the motion constraint direction and the motion constraint range. Based on the motion parameters corresponding to the M-frame multimedia data, and the movable direction and movable range of the object to be learned, the motion constraint model of the object to be learned is determined. Based on the motion constraint model of the object to be learned, a pseudo-state label is generated to reflect the state changes of the object to be learned.

[0077] Among them, the motion constraint direction is used to characterize the reasonable direction of the actual executable movement of the object to be learned, and the motion constraint direction corresponds to the movable direction. The motion constraint range is used to characterize the allowed range of movement of the object to be learned in the movable direction, and the movable range corresponds to the motion constraint range. For example, if the motion constraint direction and motion constraint range of the object to be learned indicate that the object to be learned can open 90 degrees to the left, then the movable direction of the object to be learned is to the left, and the movable range is 90 degrees; as another example, if the motion constraint direction and motion constraint range of the object to be learned indicate that the object to be learned can be pulled forward 50 centimeters, then the movable direction of the object to be learned is forward, and the motion constraint range is 50 centimeters.

[0078] The motion constraint model can be understood as a model used to characterize the motion mode, movable range, and state change law of the object to be learned, reflecting the fixed motion rules of the object under the physical structure. For example, it includes at least one of the following: the direction of motion allowed, the maximum amplitude of motion, and the corresponding quantized state at each moment.

[0079] In some embodiments, the pseudo-state labels corresponding to the learning object under different motion states can be determined according to the motion constraint model of the learning object, and the obtained pseudo-state labels can be used as the set of pseudo-state labels of the learning object.

[0080] It should be noted that the motion type and motion constraint data of the learning object are obtained by inverse joint calculation of the learning object based on the state multimedia data and the motion interaction data of the embodied robot. The specific process will be further described in the following embodiments.

[0081] S250 generates object pseudo-labels for the object to be learned based on state pseudo-labels and component pseudo-labels.

[0082] The object pseudo-labels are used by the embodied robot to learn and train on the learning object. Optionally, an embodied robot can learn and train its own model based on pseudo-labels generated from its own interactive multimedia data; it can also perform joint learning and knowledge sharing based on pseudo-labels generated by other embodied robots, thereby improving learning efficiency and generalization ability. The embodied robot that generates pseudo-labels based on the interactive multimedia data (or the embodied robot in the interactive multimedia data) and the embodied robot that uses object pseudo-labels for learning and training can be the same embodied robot or different embodied robots. For example, if the embodied robot in the interactive multimedia data is embodied robot A, after generating object pseudo-labels for the learning object from the interactive multimedia data, these object pseudo-labels can be used to learn and train embodied robot A, or embodied robot B or other robots.

[0083] In one optional implementation, for any object to be learned, all its corresponding state pseudo-labels and the component pseudo-labels corresponding to key components can be used as the object pseudo-label of the object to be learned.

[0084] In another alternative implementation, component pseudo-labels and state pseudo-labels within the same frame can be aligned in time and position. For example, region information for each component such as door panels, drawers, handles, and hinges can be extracted from the component pseudo-labels, and state change information such as opening / closing angles, push / pull strokes, and on / off states can be extracted from the state pseudo-labels. These component structural and state information are then mapped one-to-one and integrated to form a complete label that simultaneously contains the object's structural composition and real-time state. This yields pseudo-labels for the learning object in different frames, which are then used as the object's pseudo-labels. Alternatively, the component pseudo-labels, state pseudo-labels, and object pseudo-labels of the learning object can be used. During subsequent training, the embodied robot can use the component pseudo-labels to understand the component functions and positional relationships of each component of the learning object, and further determine the state of the learning object in each frame using the state pseudo-labels. Based on the component pseudo-labels and state labels, the robot's motion joint parameters are determined to achieve accurate and adaptive execution of operational tasks related to the learning object.

[0085] In the aforementioned pseudo-label generation method for embodied robot self-learning, interactive multimedia data between the embodied robot and the learning target is acquired, and state multimedia data is determined from this data. Further, motion decomposition is performed on the state multimedia data to obtain component multimedia data of the key parts of the learning target. Based on this component multimedia data, component pseudo-labels for the learning target are generated, thereby generating state pseudo-labels reflecting state changes of the learning target based on the state multimedia data. Finally, based on the state and component pseudo-labels, object pseudo-labels for the learning target are generated; these object pseudo-labels are used by the embodied robot for learning and training. In this process, on the one hand, by collecting interactive multimedia data between the embodied robot and the learning target, key component information and state change information are automatically decomposed, generating component, state, and object pseudo-labels without manual annotation, significantly reducing data annotation costs and manpower consumption. On the other hand, by fusing key component information and dynamic state information of an object to form a complete object pseudo-label, embodied robots can learn the composition structure, motion laws and state change characteristics of the object to be learned more comprehensively and accurately, thereby improving the recognition accuracy and state judgment reliability of different objects to be learned by embodied robots, and thus improving the stability and success rate of subsequent autonomous operation and interactive control.

[0086] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the key components include moving components and stationary components. Accordingly, the steps of performing motion decomposition on the state multimedia data to obtain the component multimedia data of the key components of the object to be learned are described.

[0087] See Figure 3 The component multimedia data determination steps shown include:

[0088] S310 performs motion decomposition on the state multimedia data to obtain multimedia data of the moving region and multimedia data of the stationary region.

[0089] The "following area" refers to the area corresponding to the moving parts of the learning object, while the "stationary area" refers to the area corresponding to the stationary parts of the learning object. For example, taking doors and drawers as learning objects, during the process of the embodied robot opening and closing a cabinet door, the door panel will rotate and change position with the opening and closing action; this area is the following area. The door frame and cabinet body, which are connected to the cabinet and remain in a fixed position, are the stationary areas. Similarly, during the process of the embodied robot pushing and pulling a drawer, the drawer front and inner drawer will move back and forth with the pushing and pulling action; this area is the following area. The drawer slides and cabinet frame, which are fixed to the furniture and do not change position, are the stationary areas.

[0090] For example, in a series of consecutive frames of state multimedia data, pixel changes at the same position between different frames can be compared to identify areas where the position has moved significantly and areas where the position remains essentially unchanged. Then, areas whose position changes with the movement of the object to be learned are designated as moving regions, and the corresponding image information is extracted to obtain multimedia data for the moving regions. Conversely, areas whose position remains essentially fixed throughout the movement are designated as stationary regions, and the corresponding image information is extracted to obtain multimedia data for the stationary regions.

[0091] There are various ways to perform motion decomposition on state multimedia data, and the different decomposition methods have been described in detail in the foregoing embodiments. Therefore, the process of determining the moving region multimedia data and the stationary region multimedia data based on motion decomposition described above is only an exemplary implementation method provided by this application and does not constitute a limitation on the specific implementation method of motion decomposition.

[0092] Accordingly, the above-mentioned generation of pseudo-labels for the learning object based on component multimedia data includes: performing masking processing on key multimedia data of the moving region to obtain a moving component mask, and determining the moving component mask as the mask label of the moving component; performing masking processing on key multimedia data of the stationary region to obtain a stationary component mask, and determining the stationary component mask as the mask label of the stationary component; and generating pseudo-labels for the learning object based on the mask labels of the moving component and the stationary component.

[0093] For example, within the defined moving area, movable key component areas such as door panels and drawer panels can be located. These areas are then masked using pixel-level annotation or contour extraction to obtain a moving component mask containing only the moving component areas, which is then used as the mask label for the moving component. Similarly, within the stationary area, fixed key component areas such as cabinets, door frames, and furniture frames are located. These areas are masked in the same way to obtain a stationary component mask containing only the stationary component areas, which is then used as the mask label for the stationary component. Furthermore, the moving component mask labels and stationary component mask labels corresponding to the same learning object are spatially aligned and merged to form pseudo-labels that completely distinguish between moving and stationary components.

[0094] In some embodiments, to accurately locate and finely segment the components of the learning object, after determining the mask labels corresponding to each component, semantic category names can be assigned to the mask labels corresponding to each component, such as handle, door panel, cabinet, drawer panel, etc. For example, a text-based Grounded-Segment-Anything (SAM) model can be used, which uses text guidance to detect and segment key components at the pixel level, thereby automatically assigning corresponding semantic categories to the mask of each component.

[0095] In some embodiments, to further improve the accuracy and boundary regularity of the mask pseudo-labels, the boundaries of the mask labels for each key component can be refined based on a preset segmenter. The preset segmenter may include models with high-precision segmentation capabilities, such as the SAM model, which optimizes the edges and corrects the contours of the mask region, making the generated mask pseudo-labels more closely match the actual boundaries of the key components and improving label accuracy.

[0096] S320 generates component multimedia data for key components based on multimedia data from the moving region, the stationary region, and the status multimedia data.

[0097] In one alternative implementation, the multimedia data of the servo region can be directly used as the multimedia data of the moving part; and the multimedia data of the stationary region can be used as the multimedia data of the stationary part.

[0098] In another alternative implementation, the key components also include operable components. Accordingly, multimedia data of the end effector region of the embodied robot can be determined from the state multimedia data; and multimedia data of the operable components can be determined based on the intersection between the multimedia data of the end effector region and the multimedia data of the follower region.

[0099] In this context, operable parts can be understood as components on the object to be learned that can be directly touched, force-applied, and driven to move by the embodied robot, such as cabinet door handles, drawer pulls, and switch buttons. The end effector of the embodied robot can be understood as the execution structure used by the embodied robot to perform operations such as grasping and pushing, such as grippers and robotic arms. The corresponding end effector area is the region in the image corresponding to the end effector of the embodied robot.

[0100] For example, based on the appearance features of the robot end effector, the multimedia data corresponding to the end effector region can be identified and extracted from the state multimedia data first. Then, the multimedia data of the end effector region and the multimedia data of the follower region are spatially compared. The intersection area where the two overlap is taken as the region where the operable part is located, and the multimedia data of this region is taken as the multimedia data of the operable part.

[0101] In another alternative embodiment, the key component further includes a motion-coordinating component. Accordingly, the motion-coordinating relationship between the following region and the stationary region can be obtained, and the motion-coordinating region of the state multimedia data can be analyzed based on the motion-coordinating relationship to obtain the multimedia data of the motion-coordinating component.

[0102] In this context, kinematic components can be understood as parts on the object being learned that connect and transmit power between moving and stationary parts. Examples include hinges connecting a door panel to a cabinet, or drawer slides connecting a drawer to a cabinet. These kinematic components act as coordination and constraints during the movement of the object being learned.

[0103] In this context, motion coordination can be understood as the connection method of relative motion between the moving region and the stationary region, such as rotation and translation. The corresponding motion coordination region is the area in the image corresponding to the motion coordination component.

[0104] For example, the motion coordination relationship (e.g., rotation, translation, etc.) between the following region and the stationary region can be determined based on the relative motion trajectory of the following region and the stationary region in a continuous frame. Then, based on the position of the following region, the position of the stationary region and the motion coordination relationship, the motion coordination region can be determined. The motion coordination region is used as the region where the motion coordination component is located, and the multimedia data of the region is used as the multimedia data of the motion coordination component.

[0105] Optionally, the motion coordination region can be determined by: separately determining the boundary coordinates and spatial positions of the moving region and the stationary region; and then, based on the rotational or translational coordination relationship between them, finding the boundary region where the moving region and the stationary region connect and are close to each other, and using this region as the motion coordination region. For example, if the motion coordination relationship is determined to be a rotational relationship, the connecting region where the moving region rotates around the stationary region is taken as the motion coordination region; if the motion coordination relationship is determined to be a translational relationship, the contact region where the moving region slides relative to the stationary region is taken as the motion coordination region.

[0106] Furthermore, the multimedia data of the moving area, the multimedia data of the stationary area, the multimedia data of the operable parts, and the multimedia data of the moving parts can be identified as the component multimedia data of the key components.

[0107] Accordingly, the above-mentioned generation of pseudo-labels for components based on component multimedia data includes: performing masking processing on the multimedia data of operable components to obtain an operable component mask, and determining the operable component mask as the mask label of the operable component; performing masking processing on the multimedia data of kinematic components to obtain a kinematic component mask, and determining the kinematic component mask as the mask label of the kinematic component; and determining one or more of the mask labels of kinematic components, static components, operable components, and kinematic components as pseudo-labels for components related to the object to be learned.

[0108] In the above embodiments, by accurately distinguishing and extracting multimedia data of the following region, stationary region, operable parts, and motion-cooperating parts, the key parts of the learning object can be completely characterized, as well as the motion relationships between the key parts and the functions of each part (such as stationary parts having a support function, operable parts having operability, etc.). This makes the generated part pseudo-labels more comprehensive and accurate, and closer to the actual physical structure and motion mechanism of the learning object. This provides comprehensive and effective data support for the subsequent structural learning, motion understanding, and autonomous interaction of the embodied robot with the target object.

[0109] In some embodiments, the component multimedia data includes pixel position data; the component pseudo-label also includes position labels for each key component and relative positional relationship labels between different key components; correspondingly, for any key component, the position information of the key component can be determined based on the pixel position data of the key component, and the position label of the key component can be determined based on the position information; based on the position information of the key component and the position information of other key components, a relative positional relationship label between the key component and other key components can be generated.

[0110] In one optional implementation, the pixel position data corresponding to each pixel of the key component can be used as the pixel set of the corresponding key component, and the centroid of the pixel set, i.e., the centroid of the key component, can be determined as the position reference point of the key component. Based on the relative positional relationship between the position reference point corresponding to the key component and the preset initial position, the position label of the key component is determined.

[0111] In this context, the centroid of a pixel set can be understood as the geometric center of all pixels contained in the corresponding component mask within the image, equivalent to the position of the center pixel of the component region on the image. For example, a component mask corresponds to a 3×3 pixel set with pixel coordinates (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), (3,3). Summing the x-coordinates and y-coordinates of all pixels in this set yields 18. Dividing these by the total number of pixels (9) gives the centroid coordinates of the pixel set as (2,2).

[0112] For example, taking the household embodied robot grasping a cabinet door handle as an example, the position data of all pixels in the handle area can be extracted first to form the pixel set of the handle. The centroid of the pixel set is calculated as the position reference point of the handle. The position reference point is compared with the preset standard grasping center position. Based on the relative offset direction and offset amount of the two, a position label of center, left, right, top or bottom is generated for the handle.

[0113] Furthermore, based on information such as the relative distance, relative orientation, or relative angle between the centroids of each key component, the relative positional relationship labels between the key components are determined.

[0114] For example, taking a cabinet door as an example, first calculate the centroids of the handle and the hinge as reference points for their positions. Then, determine the orientation relationship based on the coordinate difference between the two reference points and generate relative position relationship labels such as "the handle is located on the opposite side of the hinge" or "the handle and the hinge are at the same horizontal height of the door panel".

[0115] In the above embodiments, component location tags are added to the component pseudo-tags, enriching the information in the component pseudo-tags. This allows the embodied robot to not only accurately locate the positions of key components after learning the object tags of the learning object, but also to clearly understand the positional and structural relationships between different key components, thereby improving self-learning efficiency and the success rate of subsequent interactive tasks.

[0116] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of obtaining the motion coordination relationship between the following region and the stationary region is described.

[0117] See Figure 4 The steps for determining the motion coordination relationship shown include:

[0118] S410, determine the motion interaction data of the embodied robot from the interactive multimedia data.

[0119] Among them, motion interaction data reflects the joint states of the embodied robot during its interaction with the object to be learned. For example, when the embodied robot performs opening, closing, pushing, and pulling actions, the joint angles, poses, and motion trajectories of the end effector are recorded.

[0120] In one optional implementation, the time period corresponding to the interactive multimedia data can be acquired, and the joint state data stream of the embodied robot during the corresponding time period can be used as the motion interaction data of the embodied robot. For example, the joint state data stream for the corresponding time period can be acquired from the historical control records of the embodied robot controller.

[0121] In another alternative implementation, the position of the end effector of the android in the interactive multimedia data can be identified, and the state changes of the end effector in consecutive frames can be determined as the motion interaction data of the android. The state changes of the end effector include at least one of position change, posture change, motion trajectory, motion speed, and contact state.

[0122] S420 performs joint inverse calculation on the object to be learned based on the following region, the stationary region, and the joint state to obtain the motion type and motion constraint data of the object to be learned.

[0123] The motion constraint data includes one or more of the motion constraint direction and motion constraint range.

[0124] Joint inverse calculation refers to deducing the motion structure and constraint relationships of an object from its observed motion state. In this embodiment, the embodied robot does not need to know the motion mode of objects such as cabinet doors and drawers in advance. It can deduce the pivot position, slide rail direction, rotation range, or sliding stroke simply by interacting with the object to be learned and observing its motion process. That is, the motion type and motion constraint data of the object to be learned can be determined based on the motion state of the robot and the object to be learned at different times during the interaction.

[0125] In one alternative implementation, a joint inverse model can be pre-trained, and different video frames can be input into the joint inverse model to obtain the motion type and motion constraint data of the object to be learned. The video frames are labeled with the following regions, stationary regions, and joint states.

[0126] In another alternative implementation, the correspondence between joint states and motion trajectories can be determined based on the motion trajectories of the following and stationary regions in consecutive frames, combined with the joint states of the embodied robot in consecutive frames. Based on this correspondence, the changes of the learning object as the embodied robot operates can be determined, thereby determining the motion type and motion constraint data of the learning object.

[0127] For example, if the follower region moves forward at a constant speed relative to the stationary region in consecutive frames, and the robot arm joints exhibit a smooth horizontal pushing motion, then the learning object can be determined to be of the translational motion type, and its motion constraint direction can be determined to be horizontal forward, and the motion constraint range can be determined to be the maximum displacement range of this interaction.

[0128] Accordingly, in the aforementioned process of determining the motion coordination region, when the motion type characterization of the learning object is rotational, the region corresponding to the motion constraint direction in the motion constraint data can be used as the motion coordination region. Furthermore, from the motion coordination region, the rotation axis of the learning object is determined; the multimedia data of the region where the rotation axis is located is then determined as the multimedia data of the motion coordination component.

[0129] For example, the position of the central axis of rotation in the motion coordination area can be used as the rotation axis of the object to be learned, and the multimedia data corresponding to the local image area where the rotation axis is located can be used as the multimedia data of the motion coordination component.

[0130] S430 determines the motion type and motion constraint data of the object to be learned as the motion coordination relationship between the following region and the stationary region.

[0131] In the above embodiments, by acquiring the joint states when the embodied robot interacts with the object to be learned, and combining the follower region and the stationary region to perform joint inverse calculation, the motion type and motion constraint data of the object to be learned are determined, and the motion coordination relationship between the follower region and the stationary region is determined. This can provide reliable constraints for the subsequent operation of the embodied robot, and improve the accuracy of operation and environmental adaptability.

[0132] It is worth noting that in this embodiment, by performing joint inverse kinematics on the learning object based on the following region, stationary region, and joint state, the motion type and motion constraint data of the learning object can be obtained, thereby accurately representing the motion structure of the learning object, rather than just describing the robot's own motion. For example, for objects such as cabinet doors and drawers, joint inverse kinematics can clarify their rotation axis, rotation range, translation direction, and maximum stroke, etc., bringing multiple beneficial effects: First, it can generate accurate state pseudo-labels for the learning object based on the motion type and motion constraint data, quantifying the motion of the learning object into indicators such as opening and closing angles and pull-out strokes, and discretizing them into state intervals such as closed, half-open, and fully open, achieving a refined description of different states of the learning object; second, it can refine the component pseudo-labels of the learning object, distinguishing motion cooperation areas such as the hinge side and handle side based on the motion type and motion constraint data, enabling the embodied robot not only to identify moving and stationary regions, but also to better... The system provides four key functions: First, it understands the functional types of different key components in the learning object. Second, it supports cross-view and cross-state propagation of pseudo-labels. Based on motion constraint data, it constructs a motion constraint model of the learning object, enabling the 3D projection and propagation of pseudo-labels from multiple perspectives. This simulates the pseudo-labels of the learning object in different opening and closing states, thus expanding the self-learning samples of the embodied robot. Third, it provides crucial support for the subsequent operation planning and task execution of the embodied robot. The motion type and motion constraint data of the learning object can provide accurate motion constraint basis for the embodied robot, guiding it to complete the operation task along a constraint trajectory that conforms to the motion law of the learning object, thereby improving the rationality and success rate of the embodied robot's interactive actions.

[0133] Based on the aforementioned pseudo-label generation method for self-learning of embodied robots, such as Figure 5 As shown, this application also provides a pseudo-label-based self-learning method for embodied robots, which can be applied to computers or embodied robots. This pseudo-label-based self-learning method for embodied robots can be applied to... Figure 1 The following steps are used as an example of the embodied robot shown:

[0134] S510 generates pseudo-labels for the state and components of the learning object based on the interactive multimedia data between the embodied robot and the learning object.

[0135] The process of generating pseudo-labels for the state and components of the object to be learned has been described in the above embodiments and will not be repeated here.

[0136] S520 generates cross-perspective pseudo-labels for the object to be learned based on state pseudo-labels, component pseudo-labels, and interactive multimedia data.

[0137] Cross-view pseudo-labels can be understood as pseudo-labels of the object to be learned from perspectives other than the perspective of interactive multimedia data acquisition. This means that the same cabinet will have different pseudo-labels when viewed from different perspectives. For example, the shape (mask label) may differ, or the relative positions of different key components may vary.

[0138] In one optional implementation, the interactive multimedia data is collected by the embodied robot from a first perspective. Correspondingly, based on component pseudo-labels and the interactive multimedia data, a 3D spatial reprojection process can be performed on the learning object to obtain a 3D semantic map. Based on the 3D semantic map, object multimedia data of the learning object is obtained from a second perspective. The second perspective is different from the first perspective. Based on the object multimedia data, component mask pseudo-labels of the learning object in the second perspective are determined. Based on the component mask pseudo-labels and state pseudo-labels, cross-perspective pseudo-labels for the learning object are determined. Optionally, when there is only one second perspective, the component mask pseudo-labels and state pseudo-labels are determined as cross-perspective pseudo-labels for the learning object. Optionally, when there are multiple second perspectives, based on the component mask pseudo-labels and state pseudo-labels of each second perspective, a corresponding perspective pseudo-label for that second perspective is generated, and the perspective pseudo-labels corresponding to the multiple second perspectives are determined as cross-perspective pseudo-labels for the learning object.

[0139] For example, interactive multimedia data is collected from a first-person perspective using an embodied robot to determine the pose information of the object to be learned from the first-person perspective. This pose information may include component masks of key parts of the object to be learned from the first-person perspective. Based on the pose information of the object to be learned from the first-person perspective, a 3D spatial reprojection process is performed on the object to be learned to obtain a 3D semantic map of the object to be learned. For example, the component masks of key parts of the object to be learned from the first-person perspective are projected into 3D to obtain a 3D semantic map of the object to be learned. It can be understood that the 3D semantic map includes a stereoscopic image of the object to be learned. Component pseudo-labels of the object to be learned from the first-person perspective are added to the 3D semantic map. At this time, each pixel of the object to be learned in the 3D semantic map corresponds to a component pseudo-label. Alternatively, the 3D semantic map records the semantic category of each position point in 3D space, such as: cabinet door, handle, wall, etc. Then, from one or more second perspectives different from the first perspective, the object multimedia data of the object to be learned is obtained from the 3D semantic map, and the corresponding component mask pseudo-labels are determined accordingly for each second perspective.

[0140] The number of second-viewpoints can be one or more. Second-viewpoint pseudo-labels are generated by combining the state pseudo-labels of the object to be learned under the first-viewpoint. For example, the state pseudo-labels of the object to be learned under the first-viewpoint can be used as the second-viewpoint pseudo-labels. When there is only one second-viewpoint, the component mask pseudo-label and the state pseudo-label under that viewpoint are directly used together as the cross-viewpoint pseudo-labels. When there are multiple second-viewpoints, a viewpoint pseudo-label containing both component mask pseudo-labels and state pseudo-labels is generated for each second-viewpoint, and these multiple viewpoint pseudo-labels are then merged to obtain the final cross-viewpoint pseudo-labels for the object to be learned.

[0141] In this process, a 3D semantic map is constructed by reprojecting interactive multimedia data from a first-person perspective. This allows for the generation of corresponding component mask pseudo-labels and state pseudo-labels from multiple different second-person perspectives, thereby obtaining cross-perspective pseudo-labels that cover multiple perspectives. This effectively expands the perspective diversity and data integrity of the pseudo-labels, improves the robustness of the pseudo-labels, and provides more comprehensive supervision information for subsequent model learning.

[0142] S530 performs motion simulation on the learning object based on interactive multimedia data, obtains the simulation multimedia data of the learning object, and generates enhanced pseudo-labels about the learning object based on the simulation multimedia data.

[0143] Simulated multimedia data can be understood as virtual multimedia data generated through motion simulation based on interactive multimedia data and the motion type and constraint data of the object to be learned. Its data type is consistent with that of real interactive multimedia data. Simulated multimedia data is used to reproduce the diverse motion processes of the object to be learned under the operation of an embodied robot, providing expanded samples and verification basis for generating enhanced pseudo-labels.

[0144] Among them, enhanced pseudo-labels can be understood as pseudo-labels generated based on simulated multimedia data, used to supplement and improve other pseudo-labels of the object to be learned in different motion states. Enhanced pseudo-labels are used to assist the self-learning of embodied robots.

[0145] For example, based on existing interactive multimedia data and the motion type and constraint data of the object to be learned, the interaction process between the embodied robot and the object to be learned can be reproduced in a simulation environment. This simulates the motion state of the object to be learned under different postures and actions, obtaining corresponding simulation images, simulation trajectories, and other simulation multimedia data. Based on the motion state, component positions, and morphological changes of the object to be learned in this simulation multimedia data, enhanced pseudo-labels containing richer motion details and perspective information can be generated. For example, pseudo-labels can be generated for the object to be learned when partial occlusion exists.

[0146] S540 verifies state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, and trains the embodied robot based on the verified pseudo-labels.

[0147] In one optional implementation, the pseudo-state labels can be validated for reasonableness to determine whether they conform to the state change rules of the corresponding learning object. If the reasonableness validation passes, they can be used as training data for the embodied robot. For example, for a cabinet door type learning object, the state change rule is a continuous and gradual change in opening and closing angles. If a pseudo-state label indicates that the cabinet door jumps directly from closed to fully open without any intermediate transition angle, the label is determined to be inconsistent with the state change rules, fails the reasonableness validation, and is discarded. Otherwise, the validation passes and the label is used for training.

[0148] The verified pseudo-labels can be added to the embodied robot's memory bank, which includes one or more instance entries and the graph relationships between each instance entry. For example, one instance entry can correspond to one learning object, and the identifier of the corresponding instance entry can be determined based on the object identifier of the learning object. For example, the identifier of an instance entry could be a cabinet, a door panel, etc. Instance sub-entries can also be created for instance entries, where the relationship between instance entries and sub-entries is parent-child. For example, if a cabinet contains multiple drawers, an instance entry can be created as "cabinet," and multiple instance sub-entries as "multiple drawers," with one instance sub-entry corresponding to one drawer. This allows subsequent interactions to automatically complete the pseudo-labels for its components and states with minimal effort. Once structural priors have been accumulated on the same cabinet or similar instances, only minimal interactions are needed to complete the pseudo-labels for new similar drawers.

[0149] In another alternative implementation, the pseudo-labels of components can be validated for reasonableness to determine whether they conform to the component design principles of the corresponding learning object. If the reasonableness validation passes, the pseudo-labels can be used as training data for the embodied robot. For example, for cabinet doors as learning objects, the component design principle is that rotating shaft components are usually located on the side of the door panel and are continuously and rigidly connected to the door panel. If a pseudo-label marks the rotating shaft in the center area of ​​the door panel, it is determined that the marking does not conform to the conventional component design principles, the reasonableness validation fails, and it is discarded. Otherwise, the validation passes and it can be used for training.

[0150] In another optional implementation, cross-view consistency verification can be performed on the pseudo-labels to determine whether the pseudo-labels corresponding to the same spatial point are consistent under different views. If the consistency verification passes, the pseudo-labels can be used as training data for the embodied robot. For example, for a cabinet door handle, if it is labeled as a handle component in the first view, the pixels at the same spatial location should still be labeled as handle components in the second view. If different component types are labeled under different views, the cross-view consistency verification is deemed to have failed and the pseudo-label is discarded; otherwise, the verification passes and the pseudo-label is used for training.

[0151] In another optional implementation, the enhanced pseudo-labels can be verified for realism through a comparison. The simulated enhanced pseudo-labels are compared with pseudo-labels obtained in real interaction scenarios in terms of temporal, spatial, and motion states to determine whether the enhanced pseudo-labels conform to the actual motion patterns of the object to be learned. If the verification is successful, the enhanced pseudo-labels are used as training data for the embodied robot. For example, in a real scenario, a cabinet door can only be opened by rotating around the hinges and cannot be moved horizontally. If the simulated enhanced pseudo-labels indicate that the cabinet door is directly pushed open horizontally, which is clearly inconsistent with reality, the verification is deemed unsuccessful and the label is removed. Otherwise, the verification is successful and the label is used for model training.

[0152] In the aforementioned pseudo-label-based self-learning method for embodied robots, pseudo-labels for the state and components of the learning object are generated based on the interactive multimedia data between the embodied robot and the learning object. Furthermore, cross-view pseudo-labels for the learning object are generated based on these pseudo-labels, component pseudo-labels, and the interactive multimedia data. Further, motion simulation is performed on the learning object based on the interactive multimedia data to obtain simulated multimedia data. Enhanced pseudo-labels for the learning object are generated based on this simulation data. The state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels are then validated. The embodied robot is trained based on the validated pseudo-labels. In this process, on the one hand, the state and component pseudo-labels are automatically generated from the real interaction data between the embodied robot and the learning object, eliminating the need for extensive manual annotation and significantly reducing annotation costs and manpower consumption. On the other hand, generating cross-view pseudo-labels and simulated enhanced pseudo-labels effectively expands the perspective diversity and sample richness of the pseudo-labels, improving their coverage and robustness, which is beneficial for the self-learning of the embodied robot. On the other hand, by verifying multiple types of pseudo-labels, abnormal and erroneous labels can be eliminated, ensuring the reliability and accuracy of training data. This enables the embodied robot to learn the composition, motion patterns, and state change characteristics of the object to be learned more comprehensively and accurately, improving the recognition accuracy and state judgment reliability of the embodied robot for different movable objects, thereby enhancing the stability and success rate of subsequent autonomous operation and interactive control.

[0153] In some embodiments, if a target task for the object to be learned is received, the motion constraint model of the object to be learned is obtained; the motion constraint model is learned by the embodied robot based on verified pseudo-labels; the target state that the object to be learned needs to achieve is determined based on the target task, and the execution joint parameters of the embodied robot are determined based on the target state, the object to be learned, and the real-time state of the motion constraint model; the embodied robot is controlled to execute the joint parameters until the target task is completed.

[0154] The motion constraint model of the object to be learned can be understood as the aforementioned motion constraint model, which has been described in the above embodiments and will not be repeated here. It should be added that the motion constraint model of the object to be learned can also characterize the correspondence between different motion states of the embodied robot and the states of the object to be learned.

[0155] For example, the target state that the learning object needs to achieve can be determined according to the objective task, such as a cabinet door fully open or a drawer fully closed. Based on the motion constraint model of the learning object, the target motion state that the embodied robot should possess when controlling the object to the target state can be planned, such as the position and pose of the end effector of the embodied robot. At the same time, combined with the real-time state of the learning object, the current motion state of the embodied robot when it comes into contact with the operable parts of the object can be determined. Based on the difference between the current motion state and the target motion state, the corresponding joint parameters are solved and output, controlling the embodied robot to smoothly adjust from the current motion state to the target motion state, thereby completing the objective task.

[0156] In the above embodiments, the motion constraint model of the learning object is obtained by learning based on the verified pseudo-labels. When the target task is received, the robot can automatically plan and determine the joint parameters of the robot based on the real-time state of the object, the target state and the motion constraint model, which can improve the interaction efficiency and operation accuracy of the embodied robot.

[0157] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the part about performing motion simulation on the learning object based on interactive multimedia data to obtain the simulation multimedia data of the learning object is described in detail.

[0158] See Figure 6 The simulated multimedia data determination steps shown include:

[0159] S610, based on interactive multimedia data, performs joint inverse calculation on the object to be learned to obtain the motion type and corresponding motion constraint data of the object to be learned.

[0160] The specific implementation method of performing joint inverse calculation on the object to be learned to obtain the motion type and motion constraint data of the object to be learned (the above-mentioned corresponding motion constraint data indicates that the motion constraint data corresponds to the object to be learned) has been introduced in the above embodiments and will not be repeated here.

[0161] S620, based on the motion type and corresponding motion constraint data of the object to be learned, performs multi-state simulation transformation on the object to be learned to obtain multimedia data of the object under different state simulation transformations.

[0162] Multi-state simulation transformation refers to simulating or transforming the current position, posture, or state of the learning object in a virtual simulation scenario to obtain simulated state data of the learning object in another posture or state. Correspondingly, transformed multimedia data refers to simulated state data generated after multi-state simulation transformation of the learning object, and its data type is consistent with the multimedia data collected through real interaction.

[0163] In one optional implementation, sampling can be performed within a reasonable range under the constraints of the motion type and corresponding motion constraint data of the object to be learned, generating multiple intermediate states between the initial state and the limit state of the object to be learned, and using the simulated visual images, pose information and other data corresponding to these intermediate states as the transformation multimedia data of the object to be learned under different state simulation transformations.

[0164] S630, the transformed multimedia data of the object to be learned under different state simulation transformations is determined as the simulation multimedia data of the object to be learned.

[0165] In the above embodiments, based on the motion type and corresponding motion constraint data of the object to be learned, multi-state simulation transformations are performed on the object to be learned, thereby determining the simulation multimedia data of the object to be learned. This can generate diverse and reliable simulation data without the need for a large amount of real interaction, providing a data foundation for the subsequent self-learning of the embodied robot.

[0166] In some embodiments, the simulation multimedia data includes transformed multimedia data under N state simulation transformations; N is a positive integer; correspondingly, for state simulation transformation j among the N state simulation transformations, based on the transformed multimedia data under state simulation transformation j, the motion constraint data of the object to be learned under state simulation transformation j can be determined; j is a positive integer less than or equal to N; based on the motion constraint data of the object to be learned under state simulation transformation j and the motion type of the object to be learned, the motion parameters of the object to be learned under state simulation transformation j are determined; based on the motion parameters of the object to be learned under state simulation transformation j and component pseudo-labels, simulated pseudo-labels of the object to be learned under state simulation transformation j are generated; the simulated pseudo-labels corresponding to the object to be learned under each of the N state simulation transformations are determined as enhanced pseudo-labels for the object to be learned. Here, the simulated pseudo-labels can be understood as pseudo-labels of the object to be learned in a virtual state, generated through state simulation transformation derivation based on existing pseudo-labels.

[0167] For example, the simulation multimedia data includes transformed multimedia data under N state simulation transformations. Taking a drawer with a translational joint motion type as an example, N can be set to 5 states, corresponding to the drawer fully closed, pulled out one-quarter, pulled out one-half, pulled out three-quarters, and fully pulled out. For state simulation transformation j, such as the one-half-pull-out state, the transformation multimedia data under this state determines the current drawer's translational displacement, guide rail limit, and other motion constraint data. The motion constraint data under each state simulation transformation can include the motion constraint range and motion constraint direction. The motion constraint range may be different under each state simulation transformation, such as the drawer fully closed, pulled out one-quarter, pulled out one-half, pulled out three-quarters, and fully pulled out. Combining the drawer's translational motion type, the corresponding translational direction, current travel, and other motion constraint data for this state are obtained. Then, based on these motion constraint data and existing pseudo-labels for components such as the drawer panel, guide rail, and handle, a simulated pseudo-label corresponding to the drawer in the one-half-pull-out state is generated. After generating simulated pseudo-labels for all five states in the same manner, these simulated pseudo-labels under different states are collectively determined as the enhanced pseudo-labels for that drawer.

[0168] For example, based on the motion parameters of the object to be learned under state simulation transformation j, simulated state pseudo-labels for the object can be determined; sample component masks are obtained from the embodied robot's memory bank, and the component pseudo-labels are occluded using the sample component masks to obtain occluded component pseudo-labels; the simulated state pseudo-labels and the occluded component pseudo-labels are determined as the simulated pseudo-labels of the object to be learned under state simulation transformation j. Here, the simulated state pseudo-labels can be understood as the state pseudo-labels of the object to be learned in the virtual state, derived from the pseudo-labels of the object to be learned in the real state, combined with the motion laws of the object to be learned, after performing a state simulation transformation.

[0169] Alternatively, the pseudo-labels of the components of the object to be learned can be determined as the pseudo-labels of the simulated components of the object under state simulation transformation j, and the pseudo-labels of the simulated state and components under state simulation transformation j can be determined as the pseudo-labels of the object under state simulation transformation j. For example, taking the object to be learned as a drawer and state simulation transformation j as half-pull-out, we can first generate a pseudo-label of the simulated state of the drawer in a half-open state based on the motion constraint data such as translation stroke and displacement coordinates in this state. Then, we can retrieve the corresponding sample component masks such as cabinet, drawer panel, and handle from the embodied robot's memory library, and use the masks to occlude the original component pseudo-labels to obtain occluded component pseudo-labels that only retain the currently visible area. Finally, we combine the simulated state pseudo-labels and the occluded component pseudo-labels together as the simulated pseudo-labels of the drawer in the half-pull-out state.

[0170] In this process, by generating corresponding motion constraint data and simulated state pseudo-labels under different state simulation transformations, the labeled data in multiple poses and states can be rapidly expanded, significantly improving the coverage and richness of the labels. Thus, the embodied robot can accurately identify and operate the learning object under different states and varying degrees of occlusion, improving the accuracy and success rate of task completion.

[0171] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of verifying the above-mentioned state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, and training the embodied robot based on the verified pseudo-labels is described in detail.

[0172] See Figure 7 The illustrated embodied robot training steps include:

[0173] S710 verifies the status pseudo-label, component pseudo-label, cross-view pseudo-label, and enhanced pseudo-label based on the preset verification logic, and obtains the pseudo-label verification results.

[0174] In one optional implementation, the preset verification logic includes motion constraint verification; correspondingly, if the state pseudo-label, component pseudo-label, cross-view pseudo-label, and enhanced pseudo-label all conform to the motion type and motion constraint data of the object to be learned, a pseudo-label verification result indicating that the verification has passed can be generated; if there is a pseudo-label among the state pseudo-label, component pseudo-label, cross-view pseudo-label, and enhanced pseudo-label that does not conform to the motion type or corresponding motion constraint data of the object to be learned, a pseudo-label verification result indicating that the verification has failed can be generated.

[0175] For example, taking a drawer as the learning object, its motion type is translation, and the motion constraint data is linear translation along the guide rail, without rotation, with a stroke within the range of 0-50cm (i.e., the motion constraint range). If the generated state pseudo-label indicates that the drawer pulls out with a stroke of 30cm, the component pseudo-label indicates that the handle and panel positions conform to the translation structure, the cross-view pseudo-label maintains linear motion characteristics at different angles, and the enhanced pseudo-label does not exceed the translation constraint in multiple states, then it is determined that all types of pseudo-labels conform to the motion type and motion constraint data, and the pseudo-label verification result is passed. If any pseudo-label shows that the drawer rotates or shifts, the pull-out stroke reaches 60cm and exceeds the motion constraint range, or the component position does not conform to the translation structure, then it is determined that at least one pseudo-label does not conform to the motion constraint, and the verification result is failed.

[0176] In another optional implementation, the preset verification logic includes multi-view consistency verification. Accordingly, pseudo-labels with positional features can be selected from state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels as selected pseudo-labels. Pseudo-labels under the first view are determined from the selected pseudo-labels to obtain first-view pseudo-labels. The first-view pseudo-labels are then 3D projected to obtain the 3D features of the object to be learned under the first view. Pseudo-labels under the second view are determined from the selected pseudo-labels to obtain second-view pseudo-labels. The second-view pseudo-labels are then 3D projected to obtain the 3D features of the object to be learned under the second view. If the 3D features under the first view are consistent with the 3D features under the second view, a pseudo-label verification result indicating that the verification has passed is generated.

[0177] The location features include at least one of spatial information, mask information, or location information.

[0178] For example, pseudo-labels containing location information, boundary mask information, or relative spatial location information can be selected as pseudo-labels with location features to ensure that the pseudo-labels contain spatial positioning information that can be used for 3D verification. In specific implementation, pseudo-labels carrying pixel coordinates, spatial orientation, or component masks can be selected from state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels as filter pseudo-labels. Then, first-view pseudo-labels corresponding to the first viewpoint are divided from the filter pseudo-labels and projected onto 3D space to obtain 3D features under the first viewpoint. Pseudo-labels under the first viewpoint refer to pseudo-labels generated from interactive multimedia data collected under the first viewpoint. Second-view pseudo-labels corresponding to the second viewpoint are divided from the filter pseudo-labels. Pseudo-labels generated from object multimedia data collected under the second viewpoint are obtained using the same projection method to obtain 3D features under the second viewpoint. If the 3D features under the first viewpoint and the 3D features under the second viewpoint are consistent in spatial location, size ratio, and component distribution, the pseudo-label verification is determined to be consistent, and a verification result that passes the verification is generated; if there is a significant deviation, the verification is determined to be unsuccessful, and abnormal pseudo-labels are removed. For example, if the 3D feature in the first view is door panel 1 and the 3D feature in the second view is door panel 2, and if there is no obvious deviation between door panel 1 and door panel 2, then the pseudo-label verification is determined to be consistent, and a pseudo-label verification result that has passed the verification is generated.

[0179] For example, taking a drawer as the learning object, with a translational motion, and motion constraint data representing a linear translation along a guide rail, the pseudo-labels with positional features obtained from the embodied robot in the first view (front) are projected into 3D space to obtain the 3D pose features of the drawer in that view. Then, the corresponding pseudo-labels with positional features from the second view (side-front) are projected into 3D space to obtain the 3D features in the second view. If the drawer's translational direction, opening / closing stroke, and component spatial position obtained from the two view projections are consistent, the pseudo-label verification is considered successful; if there are significant conflicts in the 3D features, such as one view showing linear translation while the other shows offset or rotation, the verification fails.

[0180] In another optional implementation, the preset verification logic includes temporal consistency verification; accordingly, the state pseudo-labels, component pseudo-labels and enhancement pseudo-labels of the object to be learned can be obtained sequentially in time order in multiple consecutive interaction frames, and the state changes corresponding to the pseudo-labels between adjacent frames can be judged according to the motion type and motion constraint data of the object to be learned.

[0181] For example, consider a drawer as the object to be learned, a translational motion, and a normal timing sequence of slow pull-out. If the pseudo-labels corresponding to consecutive frames are closed, one-quarter pull-out, one-half pull-out, and three-quarter pull-out respectively, with smooth transitions, continuous displacement, and no jumps, then the timing is considered consistent, and the verification passes. If a label in a frame jumps directly from closed to fully pulled out, or exhibits an unreasonable change such as pulling out and then inexplicably retracting, then the timing continuity is violated, and the verification fails.

[0182] S720 generates quality scores corresponding to status pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels based on the pseudo-label verification results.

[0183] For example, pseudo-labels of status, components, cross-viewpoints, and enhanced pseudo-labels can be scored based on a preset scoring logic and the pseudo-label verification results. The preset scoring logic can be determined based on human experience, and this application does not impose any limitations on it.

[0184] For example, for a label in any dimension among state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhancement pseudo-labels, a label can be given a full score if all corresponding pseudo-label verification results pass; if any corresponding verification result fails, a score of zero is assigned. Alternatively, different weights can be set for pseudo-label verification of different dimensions, and a weighted calculation can be performed based on the pass rate of each verification item to obtain the quality score of the corresponding label. For example, motion constraint consistency verification can be given a high weight, while temporal consistency verification and view consistency verification can be given normal weights. Finally, the comprehensive quality score of the label is obtained by weighted summing of each verification score and its corresponding weight. The weights corresponding to different verification dimensions can be determined based on human experience, and this application does not impose any restrictions on this. Of course, confidence levels can also be generated for state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhancement pseudo-labels based on the pseudo-label verification results, and pseudo-labels that pass verification can be determined as qualified pseudo-labels based on the confidence levels.

[0185] S730 determines the verified pseudo-labels based on quality scores, and these pseudo-labels are then considered qualified.

[0186] For example, a preset scoring threshold can be established. If the quality score of a pseudo-label exceeds the preset threshold, it is considered a qualified pseudo-label and used for subsequent embodied robot training; otherwise, it is discarded. The preset scoring threshold can be set based on specific circumstances, determined by human experience, or determined through extensive experimentation; this application does not impose any limitations on this. All pseudo-labels can correspond to a single preset scoring threshold, or each pseudo-label can have its own preset scoring threshold. For example, a state pseudo-label corresponds to a first preset scoring threshold, a component pseudo-label to a second preset scoring threshold, a cross-view pseudo-label to a third preset scoring threshold, and an enhancement pseudo-label to a fourth scoring threshold. When the quality score of a state pseudo-label exceeds (i.e., is greater than) the first preset scoring threshold, the state pseudo-label is determined to be a qualified pseudo-label; when the quality score of a state pseudo-label does not exceed (i.e., is less than or equal to) the first preset scoring threshold, the state pseudo-label is determined to be an unqualified pseudo-label. The first preset scoring threshold can be specifically set based on the object type and state characteristics of the object to be learned. The preset scoring threshold corresponding to each pseudo-label can be specifically set based on the object type of the object to be learned and the label characteristics of the corresponding pseudo-label.

[0187] The S740 responds to the self-learning instructions of the embodied robot and performs self-learning based on qualified pseudo-labels.

[0188] The self-learning instruction is generated when the self-learning conditions are met.

[0189] For example, the self-learning condition could be that the current time exceeds a preset self-learning cycle. For instance, when the system detects that the time since the last self-learning was completed has reached the preset cycle duration, it determines that the self-learning condition is met, automatically generates a self-learning instruction, and triggers the embodied robot to perform a new round of self-learning based on the currently accumulated qualified pseudo-labels.

[0190] For example, the self-learning condition could be that the number of currently qualified pseudo-tags exceeds a preset threshold (preset value). For instance, if the preset threshold is 50, when the total number of qualified pseudo-tags among the status pseudo-tags, component pseudo-tags, cross-view pseudo-tags, and enhancement pseudo-tags counted by the system reaches 51, it is determined that the self-learning condition is met, and the self-learning instructions for the embodied robot are automatically generated.

[0191] For example, the self-learning condition could be: the embodied robot enters a new environment. For instance, when the embodied robot detects through its environmental perception module that the spatial layout of the current scene does not match the spatial layout of the historical scene, and determines that it has entered a new environment that it has not learned before, the self-learning condition is met, and a self-learning instruction is automatically generated.

[0192] In the above embodiments, by verifying and scoring pseudo-labels, filtering qualified data, and triggering self-learning according to multiple conditions, the reliability of learning data is effectively guaranteed, enabling the embodied robot to achieve automated, adaptive, and high-quality continuous intelligent upgrades.

[0193] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0194] Based on the same inventive concept, embodiments of this application also provide an embodied robot system for implementing the aforementioned pseudo-tag generation method for embodied robot self-learning, and an embodied robot system for implementing the aforementioned pseudo-tag-based embodied robot self-learning method. The solutions provided by the above two systems are similar to the solutions described in their corresponding methods; therefore, the specific limitations in one or more embodied robot system embodiments provided below can be found in the limitations of their corresponding methods described above, and will not be repeated here.

[0195] Based on the same inventive concept, this application also provides a pseudo-tag generation apparatus for embodied robot self-learning, which implements the pseudo-tag generation method for embodied robot self-learning described above. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations of one or more pseudo-tag generation apparatus embodiments for embodied robot self-learning provided below can be found in the limitations of the pseudo-tag generation method for embodied robot self-learning described above, and will not be repeated here.

[0196] In one exemplary embodiment, such as Figure 8 As shown, a pseudo-label generation device for self-learning of embodied robots is provided, comprising: an acquisition module 810, a determination module 820, a first generation module 830, a second generation module 840, and a third generation module 850, wherein:

[0197] The acquisition module 810 is used to acquire multimedia data of the interaction between the embodied robot and the object to be learned;

[0198] The determination module 820 is used to determine the state multimedia data from the interactive multimedia data; the state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot.

[0199] The first generation module 830 is used to perform motion decomposition on the state multimedia data to obtain the component multimedia data of the key components of the object to be learned, and to generate component pseudo-labels for the object to be learned based on the component multimedia data.

[0200] The second generation module 840 is used to generate pseudo-labels of the state based on the state multimedia data to reflect the state changes of the object to be learned.

[0201] The third generation module 850 is used to generate object pseudo-labels for the object to be learned based on state pseudo-labels and component pseudo-labels; the object pseudo-labels are used by the embodied robot to learn and train on the object to be learned.

[0202] In one exemplary embodiment, such as Figure 9 As shown, a pseudo-label-based embodied robot self-learning device is provided, comprising: a fourth generation module 910, a fifth generation module 920, a sixth generation module 930, and a training module 940. Wherein:

[0203] The fourth generation module 910 is used to generate pseudo-labels of the state and pseudo-labels of the learning object based on the interactive multimedia data between the embodied robot and the learning object.

[0204] The fifth generation module 920 is used to generate cross-perspective pseudo-labels for the object to be learned based on state pseudo-labels, component pseudo-labels, and interactive multimedia data.

[0205] The sixth generation module 930 is used to perform motion simulation on the learning object based on interactive multimedia data, obtain the simulation multimedia data of the learning object, and generate enhanced pseudo-labels about the learning object based on the simulation multimedia data.

[0206] Training module 940 is used to verify state pseudo-labels, part pseudo-labels, cross-view pseudo-labels, and augmented pseudo-labels, and to train the embodied robot based on the verified pseudo-labels.

[0207] The modules in the aforementioned pseudo-tag generation device for self-learning of embodied robots can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0208] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a pseudo-tag generation method for self-learning of embodied robots. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0209] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0210] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0211] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0212] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0213] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0214] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0215] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0216] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating pseudo-labels for self-learning in embodied robots, characterized in that, The method includes: Acquire multimedia data of the interaction between the embodied robot and the learning object; State multimedia data is determined from the interactive multimedia data; the state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot; the state multimedia data includes M frames of multimedia data, where M is a positive integer; Motion decomposition is performed on the state multimedia data to obtain component multimedia data of the key components of the object to be learned, and component pseudo-labels for the object to be learned are generated based on the component multimedia data. Based on the motion type and motion constraint data of the object to be learned, and the state data of the object to be learned in the i-th frame of multimedia data in the M frames of multimedia data, the motion parameters of the object to be learned in the i-th frame of multimedia data are determined; i is a positive integer less than or equal to M; the motion type and motion constraint data of the object to be learned are obtained by joint inverse calculation of the object to be learned based on the state multimedia data and the motion interaction data of the embodied robot. Based on the motion parameters corresponding to the M frames of multimedia data and the motion constraint data, a pseudo-label of state is generated to reflect the state changes of the object to be learned. Based on the state pseudo-label and the component pseudo-label, an object pseudo-label is generated for the object to be learned; the object pseudo-label is used by the embodied robot to learn and train the object to be learned.

2. The method according to claim 1, characterized in that, The key components include moving parts and stationary parts; Accordingly, the motion decomposition of the state multimedia data to obtain component multimedia data of the key components of the object to be learned includes: The state multimedia data is subjected to motion decomposition to obtain multimedia data of the following region and multimedia data of the stationary region; wherein, the following region is the region corresponding to the moving part of the object to be learned; the stationary region is the region corresponding to the stationary part of the object to be learned. Based on the multimedia data of the moving region, the multimedia data of the stationary region, and the state multimedia data, component multimedia data of the key component is generated.

3. The method according to claim 2, characterized in that, The key components also include operable components and motion-cooperating components; Accordingly, generating component multimedia data for the key component based on the multimedia data of the moving region, the multimedia data of the stationary region, and the state multimedia data includes: From the state multimedia data, determine the multimedia data of the end effector region of the embodied robot; The multimedia data of the operable component is determined based on the intersection between the multimedia data of the end effector region and the multimedia data of the follower region. The motion coordination relationship between the following region and the stationary region is obtained, and the motion coordination region of the state multimedia data is analyzed based on the motion coordination relationship to obtain the multimedia data of the motion coordination component; The multimedia data of the following area, the multimedia data of the stationary area, the multimedia data of the operable component, and the multimedia data of the motion-cooperating component are identified as the component multimedia data of the key component.

4. The method according to claim 2, characterized in that, The step of generating pseudo-labels for components of the object to be learned based on the component multimedia data includes: The multimedia data of the following area is masked to obtain the motion component mask of the motion component, and the motion component mask is determined as the mask label of the motion component; The multimedia data of the static area is masked to obtain the static component mask, and the static component mask is then labeled with the static component mask label. Based on the mask labels of the moving parts and the mask labels of the stationary parts, pseudo-labels for the parts to be learned are generated.

5. The method according to claim 3, characterized in that, The step of generating pseudo-labels for components of the object to be learned based on the component multimedia data includes: The multimedia data of the operable component is masked to obtain an operable component mask, and the operable component mask is determined as the mask label of the operable component. The multimedia data of the motion-fitting component is masked to obtain a motion-fitting component mask, and the motion-fitting component mask is determined as the mask label of the motion-fitting component; One or more of the mask labels of the moving parts, the stationary parts, the operable parts, and the kinematic parts are identified as pseudo-labels of the parts related to the object to be learned.

6. The method according to claim 5, characterized in that, The component multimedia data includes pixel position data; the component pseudo-tag also includes position tags for each key component, as well as relative positional relationship tags between different key components; correspondingly, the method further includes: For any critical component, based on the pixel position data of the critical component, the position information of the critical component is determined, and based on the position information, the position label of the critical component is determined; Based on the location information of the key component and the location information of other key components, a relative positional relationship label between the key component and other key components is generated.

7. The method according to claim 3, characterized in that, The step of obtaining the motion coordination relationship between the following region and the stationary region includes: The motion interaction data of the embodied robot is determined from the interactive multimedia data; the motion interaction data reflects the joint state of the embodied robot during the interaction with the learning object; Based on the following region, the stationary region, and the motion interaction data, joint inverse calculation is performed on the object to be learned to obtain the motion type and motion constraint data of the object to be learned; the motion constraint data includes one or more of the motion constraint direction and motion constraint range. The motion type and motion constraint data of the object to be learned are determined as the motion coordination relationship between the following region and the stationary region.

8. The method according to claim 7, characterized in that, The analysis of the motion coordination region of the state multimedia data based on the motion coordination relationship to obtain the multimedia data of the motion coordination component includes: When the motion type characterizes the learning object as a rotation type, the region corresponding to the motion constraint direction in the motion constraint data is taken as the motion coordination region. The rotation axis of the object to be learned is determined from the motion coordination region; The multimedia data of the region where the rotation axis is located is determined as the multimedia data of the motion-fitting component.

9. The method according to claim 1, characterized in that, The motion constraint data includes the motion constraint direction and the motion constraint range; The step of generating state pseudo-labels reflecting the state changes of the object to be learned, based on the motion parameters corresponding to the M frames of multimedia data and the motion constraint data, includes: Based on the motion constraint direction and the motion constraint range, the movable direction and movable range of the object to be learned are determined. Based on the motion parameters corresponding to the M frames of multimedia data, as well as the movable direction and movable range of the object to be learned, the motion constraint model of the object to be learned is determined. Based on the motion constraint model of the object to be learned, a pseudo-label of state is generated to reflect the state changes of the object to be learned.

10. The method according to claim 1, characterized in that, The acquisition of multimedia interaction data between the embodied robot and the learning object includes: Acquire multimedia data of the task during the interaction between the embodied robot and the task object; Based on the multimedia data of the task, determine the object feature information of the task object; Based on the object feature information, determine the degree of familiarity of the embodied robot with respect to the task object; If the familiarity level is less than or equal to a preset threshold, the task object is identified as the learning object, and the task multimedia data is identified as interactive multimedia data between the embodied robot and the learning object.

11. The method according to claim 10, characterized in that, The object feature information includes object type information and object mask information; The step of determining the familiarity of the embodied robot with respect to the task object based on the object feature information includes: If the object type information of the task object is the target object type, then the object mask information is matched with the candidate mask information included in the memory bank of the embodied robot to obtain the mask matching result; the target object type includes one or more of the operable object type and the variable state object type. If the mask matching result indicates that there is no candidate mask information in the memory that matches the object mask information, then it is determined that the familiarity of the embodied robot with respect to the task object is less than or equal to a preset threshold.

12. A hymenoidae robot system, characterized in that, The embodied robotic system includes: The first acquisition module is used to acquire multimedia data of the interaction between the embodied robot and the object to be learned; The first determining module is used to determine state multimedia data from the interactive multimedia data; the state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot; the state multimedia data includes M frames of multimedia data, where M is a positive integer; The first generation module is used to perform motion decomposition on the state multimedia data to obtain component multimedia data of the key components of the object to be learned, and generate component pseudo-labels for the object to be learned based on the component multimedia data. The second generation module is used to determine the motion parameters of the learning object in the i-th frame of multimedia data based on the motion type and motion constraint data of the learning object, and the state data of the learning object in the i-th frame of multimedia data in the M frames of multimedia data; i is a positive integer less than or equal to M; the motion type and motion constraint data of the learning object are obtained by joint inverse calculation of the learning object based on the state multimedia data and the motion interaction data of the embodied robot; and generate state pseudo-labels to reflect the state changes of the learning object based on the motion parameters corresponding to the M frames of multimedia data and the motion constraint data. The third generation module is used to generate object pseudo-labels for the object to be learned based on the state pseudo-labels and the component pseudo-labels; the object pseudo-labels are used by the embodied robot to learn and train the object to be learned.

Citation Information

Patent Citations

  • Human-object interaction detection method based on AutoHOINet

    CN117373111A

  • Weak supervision RT-DETR target detection method based on pseudo label improvement

    CN121661428A