Pseudo-label based embodied robot self-learning method and embodied robot system
By generating and verifying pseudo-labels for embodied robot self-learning, the problems of high cost and label distortion in traditional self-learning methods are solved, achieving high efficiency, accuracy and stability of embodied robot self-learning and improving the success rate of task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WOCAO TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-23
AI Technical Summary
Traditional self-learning methods for embodied robots rely on manual labeling, which is costly, inefficient, and prone to label distortion, leading to task execution failures and making it difficult to meet diverse needs.
Based on the interactive multimedia data between the embodied robot and the learning object, state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels are generated. After verification, the embodied robot is trained, reducing manual annotation, expanding the perspective diversity and sample richness of pseudo-labels, eliminating abnormal labels, and improving the reliability of training data.
Reduce labeling costs, improve the coverage and robustness of pseudo-labels, enhance the recognition accuracy and state judgment reliability of embodied robots, and improve the stability and success rate of autonomous operation and interactive control.
Smart Images

Figure CN122021714B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embodied robot technology, and in particular to an embodied robot self-learning method and embodied robot system based on pseudo-labels. Background Technology
[0002] With the development of technologies in the field of artificial intelligence, embodied robots are being used more and more widely, and their functions are becoming increasingly powerful. For example, home service robots can perform tasks according to user instructions. In order to better complete their tasks, embodied robots need to perform self-learning regularly, and their self-learning relies on a large amount of labeled data.
[0003] Traditional methods for embodied robots to learn themselves have significant shortcomings: manual labeling is costly and inefficient, and it is difficult to meet diverse needs; the labels predicted by the model are prone to distortion, and directly using them for embodied robot self-learning will lead to insufficient understanding of the task object, resulting in task execution failure and affecting the task success rate. Summary of the Invention
[0004] Therefore, it is necessary to provide a pseudo-label-based self-learning method and system for embodied robots to address the aforementioned technical problems. This method can improve the richness and accuracy of pseudo-labels, thereby increasing the success rate of interactive tasks for embodied robots.
[0005] Firstly, this application provides a pseudo-label-based self-learning method for embodied robots, including:
[0006] Based on the interactive multimedia data between the embodied robot and the learning object, pseudo-labels for the state and components of the learning object are generated.
[0007] Based on state pseudo-labels, component pseudo-labels, and interactive multimedia data, generate cross-perspective pseudo-labels for the object to be learned;
[0008] Based on interactive multimedia data, motion simulation is performed on the learning object to obtain the simulation multimedia data of the learning object, and enhanced pseudo-labels about the learning object are generated based on the simulation multimedia data.
[0009] The pseudo-labels for state, parts, cross-view, and augmented are validated, and the embodied robot is trained based on the validated pseudo-labels.
[0010] Secondly, this application also provides a pseudo-label-based embodied robot self-learning system, including:
[0011] The first generation module is used to generate pseudo-labels of the state and pseudo-labels of the learning object based on the interactive multimedia data between the embodied robot and the learning object.
[0012] The second generation module is used to generate cross-perspective pseudo-labels for the object to be learned based on state pseudo-labels, component pseudo-labels, and interactive multimedia data.
[0013] The third generation module is used to perform motion simulation on the learning object based on interactive multimedia data, obtain the simulation multimedia data of the learning object, and generate enhanced pseudo-labels about the learning object based on the simulation multimedia data.
[0014] The training module is used to verify state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and augmented pseudo-labels, and to train the embodied robot based on the verified pseudo-labels.
[0015] Thirdly, this application also provides a computer device, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps in the above-described pseudo-tag generation method for self-learning of embodied robots.
[0016] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described pseudo-label generation method for self-learning of embodied robots.
[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described pseudo-label generation method for self-learning of embodied robots.
[0018] The aforementioned pseudo-label-based embodied robot self-learning method and system generate pseudo-labels for the state and components of the learning object based on interactive multimedia data between the embodied robot and the learning object. Furthermore, based on these pseudo-labels, cross-view pseudo-labels are generated for the learning object. Further, motion simulation is performed on the learning object based on the interactive multimedia data to obtain simulated multimedia data. Enhanced pseudo-labels for the learning object are generated based on this simulation data. The state, component, cross-view, and enhanced pseudo-labels are then validated, and the embodied robot is trained using the validated pseudo-labels. In this process, on the one hand, the automatic generation of state and component pseudo-labels from the real interaction data between the embodied robot and the learning object eliminates the need for extensive manual annotation, significantly reducing annotation costs and manpower consumption. On the other hand, generating cross-view and simulated enhanced pseudo-labels effectively expands the perspective diversity and sample richness of the pseudo-labels, improving their coverage and robustness, which is beneficial for the self-learning of the embodied robot. On the other hand, by verifying multiple types of pseudo-labels, abnormal and erroneous labels can be eliminated, ensuring the reliability and accuracy of training data. This enables the embodied robot to learn the composition, motion patterns, and state change characteristics of the object to be learned more comprehensively and accurately, improving the recognition accuracy and state judgment reliability of the embodied robot for different movable objects, thereby enhancing the stability and success rate of subsequent autonomous operation and interactive control. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is an application environment diagram of a pseudo-label-based embodied robot self-learning method in one embodiment;
[0021] Figure 2 This is a flowchart illustrating a pseudo-label-based self-learning method for embodied robots in one embodiment.
[0022] Figure 3 This is a flowchart illustrating the steps for determining simulated multimedia data in one embodiment;
[0023] Figure 4 This is a flowchart illustrating the training steps for an embodied robot in one embodiment;
[0024] Figure 5This is a flowchart illustrating a pseudo-label generation method for self-learning of an embodied robot in one embodiment.
[0025] Figure 6 This is a flowchart illustrating the steps for determining multimedia data of a component in one embodiment;
[0026] Figure 7 This is a flowchart illustrating the steps for determining motion coordination relationships in one embodiment;
[0027] Figure 8 This is a structural block diagram of a pseudo-label-based embodied robot self-learning device in one embodiment;
[0028] Figure 9 This is a structural block diagram of a pseudo-label generation device for self-learning of an embodied robot in one embodiment;
[0029] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0031] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0032] The pseudo-tag-based self-learning method for embodied robots provided in this application can be applied to, for example... Figure 1The application scenario shown includes an embodied robot 102 and a target environment, with the embodied robot 102 situated within the target environment. The embodied robot possesses perception, movement, and interaction capabilities, and can interact with its environment in real time. It can capture information about its surroundings through sensory organs such as cameras, LiDAR, or tactile LiDAR installed on the embodied robot. The embodied robot can be, but is not limited to, cleaning robots, companion robots, and humanoid robots (in this embodiment, the embodied robot is a humanoid robot, and the target environment is a home indoor environment for illustration). The cleaning robot can be, but is not limited to, sweeping robots, mopping robots, and sweeping-and-mopping robots. The target environment can be any environment suitable for the application of the embodied robot, including, but is not limited to, home indoor environments, industrial manufacturing environments, medical environments, warehousing and logistics environments, agricultural production environments, public service environments, educational and research environments, or entertainment performance environments.
[0033] Currently, embodied robots face complex and varied application scenarios, serving diverse and long-tailed service objects, whose placement dynamically changes over time. To enable embodied robots to perceive and understand target objects in the environment and achieve reliable autonomous interaction and task execution, existing self-learning methods for embodied robots mostly rely on manually labeled data or ordinary pseudo-labels for model training.
[0034] However, when embodied robots interact with manipulable objects and objects with variable states, they not only need basic object recognition capabilities, but also need to learn component relationships, motion constraints, and state change patterns based on reliable labels. Traditional self-learning methods use relatively simple label information, only reflecting the object's external features, which is insufficient to support a embodied robot's deep understanding of such objects. This leads to problems such as incorrect localization of key components, insufficient understanding of motion patterns, and unstable state judgments during autonomous learning, preventing the robot from achieving efficient and stable autonomous learning and refined interaction.
[0035] This application aims to address the aforementioned shortcomings in traditional technologies. Specifically, the core technical problems to be solved include: how to achieve efficient self-learning of embodied robots based on high-quality pseudo-labels, fully learn the characteristics of the learning object, improve the self-learning effect of embodied robots, and promote the successful execution of interactive tasks by robots.
[0036] In one exemplary embodiment, such as Figure 2 As shown, a pseudo-label-based self-learning method for embodied robots is provided, which is then applied to... Figure 1 Taking the embodied robot 102 as an example, the following steps are included:
[0037] S210 generates pseudo-labels for the state and components of the learning object based on the interactive multimedia data between the embodied robot and the learning object.
[0038] The process of generating pseudo-labels for the state and components of the object to be learned is described in detail in subsequent embodiments.
[0039] S220 generates cross-perspective pseudo-labels for the object to be learned based on state pseudo-labels, component pseudo-labels, and interactive multimedia data.
[0040] Cross-view pseudo-labels can be understood as pseudo-labels of the object to be learned from perspectives other than the perspective of interactive multimedia data acquisition. This means that the same cabinet will have different pseudo-labels when viewed from different perspectives. For example, the shape (mask label) may differ, or the relative positions of different key components may vary.
[0041] In one optional implementation, the interactive multimedia data is collected by the embodied robot from a first perspective. Correspondingly, based on component pseudo-labels and the interactive multimedia data, a 3D spatial reprojection process can be performed on the learning object to obtain a 3D semantic map. Based on the 3D semantic map, object multimedia data of the learning object is obtained from a second perspective. The second perspective is different from the first perspective. Based on the object multimedia data, component mask pseudo-labels of the learning object in the second perspective are determined. Based on the component mask pseudo-labels and state pseudo-labels, cross-perspective pseudo-labels for the learning object are determined. Optionally, when there is only one second perspective, the component mask pseudo-labels and state pseudo-labels are determined as cross-perspective pseudo-labels for the learning object. Optionally, when there are multiple second perspectives, based on the component mask pseudo-labels and state pseudo-labels of each second perspective, a corresponding perspective pseudo-label for that second perspective is generated, and the perspective pseudo-labels corresponding to the multiple second perspectives are determined as cross-perspective pseudo-labels for the learning object.
[0042] For example, interactive multimedia data is collected from a first-person perspective using an embodied robot to determine the pose information of the object to be learned from the first-person perspective. This pose information may include component masks of key parts of the object to be learned from the first-person perspective. Based on the pose information of the object to be learned from the first-person perspective, a 3D spatial reprojection process is performed on the object to be learned to obtain a 3D semantic map of the object to be learned. For example, the component masks of key parts of the object to be learned from the first-person perspective are projected into 3D to obtain a 3D semantic map of the object to be learned. It can be understood that the 3D semantic map includes a stereoscopic image of the object to be learned. Component pseudo-labels of the object to be learned from the first-person perspective are added to the 3D semantic map. At this time, each pixel of the object to be learned in the 3D semantic map corresponds to a component pseudo-label. Alternatively, the 3D semantic map records the semantic category of each position point in 3D space, such as: cabinet door, handle, wall, etc. Then, from one or more second perspectives different from the first perspective, the object multimedia data of the object to be learned is obtained from the 3D semantic map, and the corresponding component mask pseudo-labels are determined accordingly for each second perspective.
[0043] The number of second-viewpoints can be one or more. Second-viewpoint pseudo-labels are generated by combining the state pseudo-labels of the object to be learned under the first-viewpoint. For example, the state pseudo-labels of the object to be learned under the first-viewpoint can be used as the second-viewpoint pseudo-labels. When there is only one second-viewpoint, the component mask pseudo-label and the state pseudo-label under that viewpoint are directly used together as the cross-viewpoint pseudo-labels. When there are multiple second-viewpoints, a viewpoint pseudo-label containing both component mask pseudo-labels and state pseudo-labels is generated for each second-viewpoint, and these multiple viewpoint pseudo-labels are then merged to obtain the final cross-viewpoint pseudo-labels for the object to be learned.
[0044] In this process, a 3D semantic map is constructed by reprojecting interactive multimedia data from a first-person perspective. This allows for the generation of corresponding component mask pseudo-labels and state pseudo-labels from multiple different second-person perspectives, thereby obtaining cross-perspective pseudo-labels that cover multiple perspectives. This effectively expands the perspective diversity and data integrity of the labels, improves the robustness of the pseudo-labels, and provides more comprehensive supervision information for subsequent model learning.
[0045] S230, based on interactive multimedia data, performs motion simulation on the object to be learned, obtains the simulation multimedia data of the object to be learned, and generates enhanced pseudo-labels about the object to be learned based on the simulation multimedia data.
[0046] Simulated multimedia data can be understood as virtual multimedia data generated through motion simulation based on interactive multimedia data and the motion type and constraint data of the object to be learned. Its data type is consistent with that of real interactive multimedia data. Simulated multimedia data is used to reproduce the diverse motion processes of the object to be learned under the operation of an embodied robot, providing expanded samples and verification basis for generating enhanced pseudo-labels.
[0047] Among them, enhanced pseudo-labels can be understood as pseudo-labels generated based on simulated multimedia data, used to supplement and improve other pseudo-labels of the object to be learned in different motion states. Enhanced pseudo-labels are used to assist the self-learning of embodied robots.
[0048] For example, based on existing interactive multimedia data and the motion type and constraint data of the object to be learned, the interaction process between the embodied robot and the object to be learned can be reproduced in a simulation environment. This simulates the motion state of the object to be learned under different postures and actions, obtaining corresponding simulation images, simulation trajectories, and other simulation multimedia data. Based on the motion state, component positions, and morphological changes of the object to be learned in this simulation multimedia data, enhanced pseudo-labels containing richer motion details and perspective information can be generated. For example, pseudo-labels can be generated for the object to be learned when partial occlusion exists.
[0049] S240 verifies the state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, and trains the embodied robot based on the verified pseudo-labels.
[0050] In one optional implementation, the pseudo-state labels can be validated for reasonableness to determine whether they conform to the state change rules of the corresponding learning object. If the reasonableness validation passes, they can be used as training data for the embodied robot. For example, for a cabinet door type learning object, the state change rule is a continuous and gradual change in opening and closing angles. If a pseudo-state label indicates that the cabinet door jumps directly from closed to fully open without any intermediate transition angle, the label is determined to be inconsistent with the state change rules, fails the reasonableness validation, and is discarded. Otherwise, the validation passes and the label is used for training.
[0051] The verified pseudo-labels can be added to the embodied robot's memory bank, which includes one or more instance entries and the graph relationships between each instance entry. For example, one instance entry can correspond to one learning object, and the identifier of the corresponding instance entry can be determined based on the object identifier of the learning object. For example, the identifier of an instance entry could be a cabinet, a door panel, etc. Instance sub-entries can also be created for instance entries, where the relationship between instance entries and sub-entries is parent-child. For example, if a cabinet contains multiple drawers, an instance entry can be created as "cabinet," and multiple instance sub-entries as "multiple drawers," with one instance sub-entry corresponding to one drawer. This allows subsequent interactions to automatically complete the pseudo-labels for its components and states with minimal effort. Once structural priors have been accumulated on the same cabinet or similar instances, only minimal interactions are needed to complete the pseudo-labels for new similar drawers.
[0052] In another alternative implementation, the pseudo-labels of components can be validated for reasonableness to determine whether they conform to the component design principles of the corresponding learning object. If the reasonableness validation passes, the pseudo-labels can be used as training data for the embodied robot. For example, for cabinet doors as learning objects, the component design principle is that rotating shaft components are usually located on the side of the door panel and are continuously and rigidly connected to the door panel. If a pseudo-label marks the rotating shaft in the center area of the door panel, it is determined that the marking does not conform to the conventional component design principles, the reasonableness validation fails, and it is discarded. Otherwise, the validation passes and it can be used for training.
[0053] In another optional implementation, cross-view consistency verification can be performed on the pseudo-labels to determine whether the pseudo-labels corresponding to the same spatial point are consistent under different views. If the consistency verification passes, the pseudo-labels can be used as training data for the embodied robot. For example, for a cabinet door handle, if it is labeled as a handle component in the first view, the pixels at the same spatial location should still be labeled as handle components in the second view. If different component types are labeled under different views, the cross-view consistency verification is deemed to have failed and the pseudo-label is discarded; otherwise, the verification passes and the pseudo-label is used for training.
[0054] In another optional implementation, the enhanced pseudo-labels can be verified for realism through a comparison. The simulated enhanced pseudo-labels are compared with pseudo-labels obtained in real interaction scenarios in terms of temporal, spatial, and motion states to determine whether the enhanced pseudo-labels conform to the actual motion patterns of the object to be learned. If the verification is successful, the enhanced pseudo-labels are used as training data for the embodied robot. For example, in a real scenario, a cabinet door can only be opened by rotating around the hinges and cannot be moved horizontally. If the simulated enhanced pseudo-labels indicate that the cabinet door is directly pushed open horizontally, which is clearly inconsistent with reality, the verification is deemed unsuccessful and the label is removed. Otherwise, the verification is successful and the label is used for model training.
[0055] In the aforementioned pseudo-label-based self-learning method for embodied robots, pseudo-labels for the state and components of the learning object are generated based on the interactive multimedia data between the embodied robot and the learning object. Furthermore, cross-view pseudo-labels for the learning object are generated based on these pseudo-labels, component pseudo-labels, and the interactive multimedia data. Further, motion simulation is performed on the learning object based on the interactive multimedia data to obtain simulated multimedia data. Enhanced pseudo-labels for the learning object are generated based on this simulation data. The state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels are then validated. The embodied robot is trained based on the validated pseudo-labels. In this process, on the one hand, the state and component pseudo-labels are automatically generated from the real interaction data between the embodied robot and the learning object, eliminating the need for extensive manual annotation and significantly reducing annotation costs and manpower consumption. On the other hand, generating cross-view pseudo-labels and simulated enhanced pseudo-labels effectively expands the perspective diversity and sample richness of the labels, improving the coverage and robustness of the pseudo-labels, which is beneficial for the self-learning of the embodied robot. On the other hand, by verifying multiple types of pseudo-labels, abnormal and erroneous labels can be eliminated, ensuring the reliability and accuracy of training data. This enables the embodied robot to learn the composition, motion patterns, and state change characteristics of the object to be learned more comprehensively and accurately, improving the recognition accuracy and state judgment reliability of the embodied robot for different movable objects, thereby enhancing the stability and success rate of subsequent autonomous operation and interactive control.
[0056] In some embodiments, if a target task for the object to be learned is received, the motion constraint model of the object to be learned is obtained; the motion constraint model is learned by the embodied robot based on verified pseudo-labels; the target state that the object to be learned needs to achieve is determined based on the target task, and the execution joint parameters of the embodied robot are determined based on the target state, the object to be learned, and the real-time state of the motion constraint model; the embodied robot is controlled to execute the joint parameters until the target task is completed.
[0057] The explanation and determination process of the motion constraint model of the object to be learned can be found in the relevant descriptions in the following embodiments.
[0058] It should be added that the motion constraint model of the object to be learned can also characterize the correspondence between different motion states of the embodied robot and the state of the object to be learned.
[0059] For example, the target state that the learning object needs to achieve can be determined according to the objective task, such as a cabinet door fully open or a drawer fully closed. Based on the motion constraint model of the learning object, the target motion state that the embodied robot should possess when controlling the object to the target state can be planned, such as the position and pose of the end effector of the embodied robot. At the same time, combined with the real-time state of the learning object, the current motion state of the embodied robot when it comes into contact with the operable parts of the object can be determined. Based on the difference between the current motion state and the target motion state, the corresponding joint parameters are solved and output, controlling the embodied robot to smoothly adjust from the current motion state to the target motion state, thereby completing the objective task.
[0060] In the above embodiments, the motion constraint model of the learning object is obtained by learning based on the verified pseudo-labels. When the target task is received, the robot can automatically plan and determine the joint parameters of the robot based on the real-time state of the object, the target state and the motion constraint model, which can improve the interaction efficiency and operation accuracy of the embodied robot.
[0061] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the part about performing motion simulation on the learning object based on interactive multimedia data to obtain the simulation multimedia data of the learning object is described in detail.
[0062] See Figure 3 The simulated multimedia data determination steps shown include:
[0063] S310, based on interactive multimedia data, performs joint inverse calculation on the object to be learned to obtain the motion type and corresponding motion constraint data of the object to be learned.
[0064] The specific implementation of performing joint inverse calculation on the object to be learned to obtain the motion type and motion constraint data of the object to be learned (the above-mentioned corresponding motion constraint data indicates that the motion constraint data corresponds to the object to be learned) is described in the relevant description in the following embodiments.
[0065] S320, based on the motion type and corresponding motion constraint data of the object to be learned, performs multi-state simulation transformation on the object to be learned to obtain multimedia data of the object under different state simulation transformations.
[0066] Multi-state simulation transformation refers to simulating or transforming the current position, posture, or state of the learning object in a virtual simulation scenario to obtain simulated state data of the learning object in another posture or state. Correspondingly, transformed multimedia data refers to simulated state data generated after multi-state simulation transformation of the learning object, and its data type is consistent with the multimedia data collected through real interaction.
[0067] In one optional implementation, sampling can be performed within a reasonable range under the constraints of the motion type and corresponding motion constraint data of the object to be learned, generating multiple intermediate states between the initial state and the limit state of the object to be learned, and using the simulated visual images, pose information and other data corresponding to these intermediate states as the transformation multimedia data of the object to be learned under different state simulation transformations.
[0068] S330, the transformed multimedia data of the object to be learned under different state simulation transformations is determined as the simulation multimedia data of the object to be learned.
[0069] In the above embodiments, based on the motion type and corresponding motion constraint data of the object to be learned, multi-state simulation transformations are performed on the object to be learned, thereby determining the simulation multimedia data of the object to be learned. This can generate diverse and reliable simulation data without the need for a large amount of real interaction, providing a data foundation for the subsequent self-learning of the embodied robot.
[0070] In some embodiments, the simulation multimedia data includes transformed multimedia data under N state simulation transformations; N is a positive integer; correspondingly, for state simulation transformation j among the N state simulation transformations, based on the transformed multimedia data under state simulation transformation j, the motion constraint data of the object to be learned under state simulation transformation j can be determined; j is a positive integer less than or equal to N; based on the motion constraint data of the object to be learned under state simulation transformation j and the motion type of the object to be learned, the motion parameters of the object to be learned under state simulation transformation j are determined; based on the motion parameters of the object to be learned under state simulation transformation j and component pseudo-labels, simulated pseudo-labels of the object to be learned under state simulation transformation j are generated; the simulated pseudo-labels corresponding to the object to be learned under each of the N state simulation transformations are determined as enhanced pseudo-labels for the object to be learned. Here, the simulated pseudo-labels can be understood as pseudo-labels of the object to be learned in a virtual state, generated through state simulation transformation derivation based on existing pseudo-labels.
[0071] For example, the simulation multimedia data includes transformed multimedia data under N state simulation transformations. Taking a drawer with a translational joint motion type as an example, N can be set to 5 states, corresponding to the drawer fully closed, pulled out one-quarter, pulled out one-half, pulled out three-quarters, and fully pulled out. For state simulation transformation j, such as the one-half-pull-out state, the transformation multimedia data under this state determines the current drawer's translational displacement, guide rail limit, and other motion constraint data. The motion constraint data under each state simulation transformation can include the motion constraint range and motion constraint direction. The motion constraint range may be different under each state simulation transformation, such as the drawer fully closed, pulled out one-quarter, pulled out one-half, pulled out three-quarters, and fully pulled out. Combining the drawer's translational motion type, the corresponding translational direction, current travel, and other motion constraint data for this state are obtained. Then, based on these motion constraint data and existing pseudo-labels for components such as the drawer panel, guide rail, and handle, a simulated pseudo-label corresponding to the drawer in the one-half-pull-out state is generated. After generating simulated pseudo-labels for all five states in the same manner, these simulated pseudo-labels under different states are collectively determined as the enhanced pseudo-labels for that drawer.
[0072] For example, based on the motion parameters of the object to be learned under state simulation transformation j, simulated state pseudo-labels for the object can be determined; sample component masks are obtained from the embodied robot's memory bank, and the component pseudo-labels are occluded using the sample component masks to obtain occluded component pseudo-labels; the simulated state pseudo-labels and the occluded component pseudo-labels are determined as the simulated pseudo-labels of the object to be learned under state simulation transformation j. Here, the simulated state pseudo-labels can be understood as the state pseudo-labels of the object to be learned in the virtual state, derived from the pseudo-labels of the object to be learned in the real state, combined with the motion laws of the object to be learned, after performing a state simulation transformation.
[0073] Alternatively, the pseudo-labels of the components of the object to be learned can be determined as the pseudo-labels of the simulated components of the object under state simulation transformation j, and the pseudo-labels of the simulated state and components under state simulation transformation j can be determined as the pseudo-labels of the object under state simulation transformation j. For example, taking the object to be learned as a drawer and state simulation transformation j as half-pull-out, we can first generate a pseudo-label of the simulated state of the drawer in a half-open state based on the motion constraint data such as translation stroke and displacement coordinates in this state. Then, we can retrieve the corresponding sample component masks such as cabinet, drawer panel, and handle from the embodied robot's memory library, and use the masks to occlude the original component pseudo-labels to obtain occluded component pseudo-labels that only retain the currently visible area. Finally, we combine the simulated state pseudo-labels and the occluded component pseudo-labels together as the simulated pseudo-labels of the drawer in the half-pull-out state.
[0074] In this process, by generating corresponding motion constraint data and simulated state pseudo-labels under different state simulation transformations, the labeled data in multiple poses and states can be rapidly expanded, significantly improving the coverage and richness of the labels. Thus, the embodied robot can accurately identify and operate the learning object under different states and varying degrees of occlusion, improving the accuracy and success rate of task completion.
[0075] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of verifying the above-mentioned state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, and training the embodied robot based on the verified pseudo-labels is described in detail.
[0076] See Figure 4 The illustrated embodied robot training steps include:
[0077] S410, based on the preset verification logic, verifies the status pseudo-label, component pseudo-label, cross-view pseudo-label and enhanced pseudo-label respectively, and obtains the pseudo-label verification results.
[0078] In one optional implementation, the preset verification logic includes motion constraint verification; correspondingly, if the state pseudo-label, component pseudo-label, cross-view pseudo-label, and enhanced pseudo-label all conform to the motion type and corresponding motion constraint data of the object to be learned, a pseudo-label verification result indicating that the verification has passed can be generated; if there is a pseudo-label among the state pseudo-label, component pseudo-label, cross-view pseudo-label, and enhanced pseudo-label that does not conform to the motion type or corresponding motion constraint data of the object to be learned, a pseudo-label verification result indicating that the verification has failed can be generated.
[0079] For example, taking a drawer as the learning object, its motion type is translation, and the motion constraint data is linear translation along the guide rail, without rotation, with a stroke within the range of 0-50cm (i.e., the motion constraint range). If the generated state pseudo-label indicates that the drawer pull-out stroke is 30cm, the component pseudo-label indicates that the handle and panel positions conform to the translation structure, the cross-view pseudo-label maintains linear motion characteristics at different angles, and the enhanced pseudo-label does not exceed the translation constraint in multiple states, then it is determined that all types of pseudo-labels conform to the motion type and motion constraint data, and the pseudo-label verification result is passed. If one of the pseudo-labels shows that the drawer rotates or shifts, the pull-out stroke reaches 60cm exceeding the motion constraint range, or the component position does not conform to the translation structure, then it is determined that at least one pseudo-label does not conform to the motion constraint, and the verification result is failed.
[0080] In another optional implementation, the preset verification logic includes multi-view consistency verification. Accordingly, pseudo-labels with positional features can be selected from state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels as selected pseudo-labels. Pseudo-labels under the first view are determined from the selected pseudo-labels to obtain first-view pseudo-labels. The first-view pseudo-labels are then 3D projected to obtain the 3D features of the object to be learned under the first view. Pseudo-labels under the second view are determined from the selected pseudo-labels to obtain second-view pseudo-labels. The second-view pseudo-labels are then 3D projected to obtain the 3D features of the object to be learned under the second view. If the 3D features under the first view are consistent with the 3D features under the second view, a pseudo-label verification result indicating that the verification has passed is generated.
[0081] The location features include at least one of spatial information, mask information, or location information.
[0082] For example, pseudo-labels containing location information, boundary mask information, or relative spatial location information can be selected as pseudo-labels with location features to ensure that the pseudo-labels contain spatial positioning information that can be used for 3D verification. In specific implementation, pseudo-labels carrying pixel coordinates, spatial orientation, or component masks can be selected from state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels as filter pseudo-labels. Then, first-view pseudo-labels corresponding to the first viewpoint are divided from the filter pseudo-labels and projected onto 3D space to obtain 3D features under the first viewpoint. Pseudo-labels under the first viewpoint refer to pseudo-labels generated from interactive multimedia data collected under the first viewpoint. Second-view pseudo-labels corresponding to the second viewpoint are divided from the filter pseudo-labels. Pseudo-labels generated from object multimedia data collected under the second viewpoint are obtained using the same projection method to obtain 3D features under the second viewpoint. If the 3D features under the first viewpoint and the 3D features under the second viewpoint are consistent in spatial location, size ratio, and component distribution, the pseudo-label verification is determined to be consistent, and a verification result that passes the verification is generated; if there is a significant deviation, the verification is determined to be unsuccessful, and abnormal pseudo-labels are removed. For example, if the 3D feature in the first view is door panel 1 and the 3D feature in the second view is door panel 2, and if there is no obvious deviation between door panel 1 and door panel 2, then the pseudo-label verification is determined to be consistent, and a pseudo-label verification result that has passed the verification is generated.
[0083] For example, taking a drawer as the learning object, with a translational motion, and motion constraint data representing a linear translation along a guide rail, the pseudo-labels with positional features obtained from the embodied robot in the first view (front) are projected into 3D space to obtain the 3D pose features of the drawer in that view. Then, the corresponding pseudo-labels with positional features from the second view (side-front) are projected into 3D space to obtain the 3D features in the second view. If the drawer's translational direction, opening / closing stroke, and component spatial position obtained from the two view projections are consistent, the pseudo-label verification is considered successful; if there are significant conflicts in the 3D features, such as one view showing linear translation while the other shows offset or rotation, the verification fails.
[0084] In another optional implementation, the preset verification logic includes temporal consistency verification; accordingly, the state pseudo-labels, component pseudo-labels and enhancement pseudo-labels of the object to be learned can be obtained sequentially in time order in multiple consecutive interaction frames, and the state changes corresponding to the pseudo-labels between adjacent frames can be judged according to the motion type and motion constraint data of the object to be learned.
[0085] For example, consider a drawer as the object to be learned, a translational motion, and a normal timing sequence of slow pull-out. If the pseudo-labels corresponding to consecutive frames are closed, one-quarter pull-out, one-half pull-out, and three-quarter pull-out respectively, with smooth transitions, continuous displacement, and no jumps, then the timing is considered consistent, and the verification passes. If a label in a frame jumps directly from closed to fully pulled out, or exhibits an unreasonable change such as pulling out and then inexplicably retracting, then the timing continuity is violated, and the verification fails.
[0086] S420: Based on the pseudo-label verification results, generate quality scores corresponding to status pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, respectively.
[0087] For example, pseudo-labels of status, components, cross-viewpoints, and enhanced pseudo-labels can be scored based on a preset scoring logic and the pseudo-label verification results. The preset scoring logic can be determined based on human experience, and this application does not impose any limitations on it.
[0088] For example, for a label in any dimension among state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhancement pseudo-labels, a label can be given a full score if all corresponding pseudo-label verification results pass; if any corresponding verification result fails, a score of zero is assigned. Alternatively, different weights can be set for pseudo-label verification of different dimensions, and a weighted calculation can be performed based on the pass rate of each verification item to obtain the quality score of the corresponding label. For example, motion constraint consistency verification can be given a high weight, while temporal consistency verification and view consistency verification can be given normal weights. Finally, the comprehensive quality score of the label is obtained by weighted summing of each verification score and its corresponding weight. The weights corresponding to different verification dimensions can be determined based on human experience, and this application does not impose any restrictions on this. Of course, confidence levels can also be generated for state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhancement pseudo-labels based on the pseudo-label verification results, and pseudo-labels that pass verification can be determined as qualified pseudo-labels based on the confidence levels.
[0089] S430 determines the verified pseudo-labels based on quality scores, and these pseudo-labels are considered qualified.
[0090] For example, a preset scoring threshold can be established. If the quality score of a pseudo-label exceeds the preset threshold, it is considered a qualified pseudo-label and used for subsequent embodied robot training; otherwise, it is discarded. The preset scoring threshold can be set based on specific circumstances, determined by human experience, or determined through extensive experimentation; this application does not impose any limitations on this. All pseudo-labels can correspond to a single preset scoring threshold, or each pseudo-label can have its own preset scoring threshold. For example, a state pseudo-label corresponds to a first preset scoring threshold, a component pseudo-label to a second preset scoring threshold, a cross-view pseudo-label to a third preset scoring threshold, and an enhancement pseudo-label to a fourth scoring threshold. When the quality score of a state pseudo-label exceeds (i.e., is greater than) the first preset scoring threshold, the state pseudo-label is determined to be a qualified pseudo-label; when the quality score of a state pseudo-label does not exceed (i.e., is less than or equal to) the first preset scoring threshold, the state pseudo-label is determined to be an unqualified pseudo-label. The first preset scoring threshold can be specifically set based on the object type and state characteristics of the object to be learned. The preset scoring threshold corresponding to each pseudo-label can be specifically set based on the object type of the object to be learned and the label characteristics of the corresponding pseudo-label.
[0091] S440, in response to the self-learning instructions of the embodied robot, performs self-learning of the embodied robot based on qualified pseudo-labels.
[0092] The self-learning instruction is generated when the self-learning conditions are met.
[0093] For example, the self-learning condition could be that the current time exceeds a preset self-learning cycle. For instance, when the system detects that the time since the last self-learning was completed has reached the preset cycle duration, it determines that the self-learning condition is met, automatically generates a self-learning instruction, and triggers the embodied robot to perform a new round of self-learning based on the currently accumulated qualified pseudo-labels.
[0094] For example, the self-learning condition could be that the number of currently qualified pseudo-tags exceeds a preset threshold (preset value). For instance, if the preset threshold is 50, when the total number of qualified pseudo-tags among the status pseudo-tags, component pseudo-tags, cross-view pseudo-tags, and enhancement pseudo-tags counted by the system reaches 51, it is determined that the self-learning condition is met, and the self-learning instructions for the embodied robot are automatically generated.
[0095] For example, the self-learning condition could be: the embodied robot enters a new environment. For instance, when the embodied robot detects through its environmental perception module that the spatial layout of the current scene does not match the spatial layout of the historical scene, and determines that it has entered a new environment that it has not learned before, the self-learning condition is met, and a self-learning instruction is automatically generated.
[0096] In the above embodiments, by verifying and scoring pseudo-labels, filtering qualified data, and triggering self-learning according to multiple conditions, the reliability of learning data is effectively guaranteed, enabling the embodied robot to achieve automated, adaptive, and high-quality continuous intelligent upgrades.
[0097] To make the aforementioned self-learning method for embodied robots based on pseudo-labels more practical, such as... Figure 5 As shown, this application also provides a pseudo-label generation method for self-learning of embodied robots. This pseudo-label generation method for self-learning of embodied robots is applied to... Figure 1 The following steps are used as an example of the embodied robot shown:
[0098] S510 acquires multimedia data of the interaction between the embodied robot and the object to be learned.
[0099] In this context, the learning target can be understood as the object that the embodied robot needs to learn in the target environment. By generating pseudo-labels for the learning target through multimedia data generated from the interaction between the robot and the learning target, the embodied robot can become familiar with or better understand the object's shape, operational state, and other characteristics. This allows the embodied robot to achieve task success and applicability when performing tasks related to the learning target. For example, the learning target can be a task object that the embodied robot is "unfamiliar" with or even "never seen before." The task object can be understood as the operational target that the embodied robot aims to achieve when performing an operational task.
[0100] Interactive multimedia data can be understood as multimedia data when the embodied robot interacts with the object to be learned, such as at least one of the following: images (RGB images, RGB-D depth maps), videos (continuous video frames), point cloud data, and audio.
[0101] The following describes the steps for acquiring interactive multimedia data:
[0102] In one optional implementation, the interactive multimedia data may include video data. When performing an operation task, the embodied robot can collect interactive video data between itself and the task object in real time through an RGB camera mounted on the embodied robot, and store the interactive video data in a preset interaction record database. If the task object is determined to be a learning object, the embodied robot can obtain the interactive video data between itself and the corresponding task object from the preset interaction record database as interactive multimedia data between the embodied robot and the learning object.
[0103] Optionally, interactive video data can also carry depth information. Accordingly, based on the RGB camera acquisition, color and depth information can be simultaneously acquired by an RGB-D depth camera to form interactive video data carrying depth information; or, depth estimation can be performed on RGB images using a monocular depth estimation model to generate interactive video data carrying depth information.
[0104] In one alternative implementation, the interactive multimedia data may further include interactive audio data. When performing interactive tasks, the embodied robot can collect interactive audio data in real time via a microphone and fuse the collected audio data with the aforementioned interactive video data to form interactive multimedia data between the robot and the task object.
[0105] For example, the interactive audio data may include friction sounds, collision sounds, etc. generated when interacting with the task object, which can be used as a reference for subsequent determination of pseudo-labels.
[0106] It is worth noting that, in order to reduce unnecessary computational overhead and thus improve the effectiveness and efficiency of pseudo-tag generation, in this embodiment, not all task objects are used as learning objects. Instead, only task objects that meet preset requirements are considered as learning objects. The process of determining learning objects is described below:
[0107] In one optional implementation, task multimedia data during the interaction between the embodied robot and the task object can be acquired; based on the task multimedia data, object feature information of the task object can be determined; based on the object feature information, the familiarity level of the embodied robot with respect to the task object can be determined; if the familiarity level is less than or equal to a preset threshold, the task object is determined as the learning object, and correspondingly, the task multimedia data is determined as the interactive multimedia data between the embodied robot and the learning object.
[0108] Among them, task multimedia data can be understood as one or more of the following: video, audio, RGB or RGB-D depth images, point cloud data, etc., collected by the embodied robot based on its own sensors during the interaction with the task object when performing interactive tasks.
[0109] Here, object feature information can be understood as information extracted from task multimedia data that can characterize the features of the task object. For example, it may include at least one of appearance feature information, structural feature information, state change feature information, and key component distribution feature information.
[0110] For example, during the interaction between the embodied robot and the task object, one or more of the following devices can be used: image acquisition device, audio acquisition device, depth sensing device, and point cloud acquisition device. These devices can collect video, audio, RGB-D images, and point cloud data during the interaction, serving as multimedia data for the task. Feature extraction is performed on this multimedia data, and the extracted results are used as the object feature information of the task object. This object feature information is then matched against existing object features in a preset object memory (including feature information of historical interaction objects, the number of interactions with each historical interaction object, the relationships between historical interaction objects, the number of learning attempts for each historical interaction object, and the recognition accuracy rate). Based on at least one of the feature matching degree, the number of historical learning attempts, and the historical recognition accuracy rate, the familiarity level of the embodied robot with the task object is determined. Furthermore, based on the relationship between the familiarity level and a preset threshold, it is determined whether the task object is a learning object. For example, if the familiarity level exceeds the preset threshold, the corresponding task object is determined to be a learning object.
[0111] It should be noted that the method for feature extraction of task multimedia data can be a common feature extraction method, such as inputting the task multimedia data into a feature extraction network to obtain the feature extraction results, which will not be elaborated here.
[0112] Optionally, there are many ways to determine the embodied robot's familiarity with the task object based on at least one of feature matching degree, historical learning count, and historical recognition accuracy. For example, the similarity score between the object feature information of the task object and the existing object feature information in a preset object memory can be calculated, and the resulting feature matching degree score can be used as the embodied robot's familiarity with the task object. Another example is to predetermine the correspondence between historical learning counts and reference familiarity, query the historical learning counts of the task object in the preset object memory, and use the reference familiarity corresponding to those historical learning counts as the embodied robot's familiarity with the task object. Yet another example is to predetermine the correspondence between historical recognition accuracy and reference familiarity, query the historical recognition accuracy corresponding to the task in the preset object memory, and use the reference familiarity corresponding to those historical recognition accuracy as the embodied robot's familiarity with the task object. Furthermore, the familiarity levels determined based on each dimension can be weighted and fused to obtain the final familiarity level. In this process, the weight coefficients corresponding to different dimensions can be determined based on human experience, and this application does not impose any limitations on this.
[0113] In some embodiments, the object feature information includes object type information and object mask information. Accordingly, if the object type information of the task object is the target object type, the object mask information of the task object is matched with each candidate mask information included in the memory of the embodied robot to obtain a mask matching result. If the mask matching result indicates that there is no candidate mask information in the memory that matches the object mask information, it is determined that the familiarity of the embodied robot with respect to the task object is less than or equal to a preset degree threshold.
[0114] The object type information is used to characterize the type of the task object. The target object type includes one or more of the following: operable object type and variable-state object type. Operable object type can be understood as an object category that possesses structural attributes that allow the embodied robot to perform physical operations and can interact with the embodied robot's end effector. Variable-state object type can be understood as an object category that can observably, quantify, and repeatedly switch between multiple stable states. Examples include objects with built-in handles, pulls, latches, and flap edges that can be located, grasped, and force-applied by the embodied robot; or objects containing discrete states such as closed, half-open, fully open, retracted, extended, and flap-closed / open, such as cabinet doors, drawers, refrigerator doors, washing machine doors, and storage cabinet flaps.
[0115] The object mask information is a mask used to characterize the outline, region, and pixel-level range of the task object. The candidate mask information in the memory can be the mask information from historically generated pseudo-labels. This is used to determine whether a pseudo-label for the task object has already been generated in a historical time period. If a pseudo-label for the task object has already been generated in a historical time period, it is not necessary to generate another pseudo-label; otherwise, it is necessary to generate a pseudo-label. The candidate mask information consists of one or more object mask data already stored in the preset memory.
[0116] S520 determines the state multimedia data from the interactive multimedia data.
[0117] Among them, state multimedia data is used to reflect the state changes of the learning object during the interaction with the embodied robot. The state multimedia data includes the complete process of the learning object changing from one state to another during the interaction with the embodied robot.
[0118] In one optional implementation, frame difference calculation can be performed on each adjacent video frame in the interactive multimedia data. For the current video frame: if the frame difference between the current video frame and its previous frame is greater than a preset state change threshold, then the previous frame of the current video frame is marked as the state change start frame; if the frame difference between the current video frame and its next frame is less than a preset stability threshold, then the current video frame is marked as the state change end frame; the continuous video frame segments between the state change start frame and the state change end frame are determined as state multimedia data reflecting the state change of the object to be learned.
[0119] In another alternative implementation, optical flow field calculations can be performed on consecutive RGB-D video frames in the interactive multimedia data to obtain pixel-level motion vectors for each frame. Based on the object mask information of the object to be learned, the background area is removed and only the pixel motion vectors of the object area to be learned are retained. When the amplitude of the pixel motion vector in this area is greater than a preset motion amplitude threshold and the motion direction remains stable and consistent, the corresponding RGB-D image sequence data within this continuous time period is extracted and used as state multimedia data.
[0120] In another alternative implementation, the contact time between the embodied robot and the learning object can be determined based on the real-time feedback information from the end effector of the embodied robot, and this time can be used as the start time of the interactive action; the separation time between the embodied robot and the learning object can be determined, and this time can be used as the end time of the interactive action; the multimedia data in the interactive multimedia data from the start time of the interactive action to the end time of the interactive action can be used as the state multimedia data.
[0121] It's important to note that the purpose of identifying state multimedia data from interactive multimedia data is to extract data showing genuine state changes in the object to be learned. For example, interactive multimedia data might include the entire process of an embodied robot's end effector approaching the object, making contact with it, manipulating it, and then leaving it. State multimedia data, on the other hand, only includes the process of the embodied robot manipulating the object. This process preserves the core segments of state changes in the object and eliminates redundant data with no movement or change, reducing subsequent data processing and thus improving the efficiency of pseudo-label generation.
[0122] S530 performs motion decomposition on the state multimedia data to obtain component multimedia data of the key components of the object to be learned, and generates component pseudo-labels for the object to be learned based on the component multimedia data.
[0123] Motion decomposition refers to the process of breaking down the overall motion of an object into multiple independent or interrelated component motions. There are many methods for performing motion decomposition on state multimedia data, such as at least one of optical flow, feature tracking, and frame difference methods.
[0124] In this context, key components can be understood as constituent parts of the object to be learned. When the object to be learned is a movable object, its corresponding key components refer to the functional parts that constitute the movable structure of the movable object, including at least moving parts and stationary parts. Component multimedia data refers to the multimedia data corresponding to the key components, such as images including the key components.
[0125] In one alternative implementation, motion decomposition of state multimedia data can be performed based on optical flow. For example, for consecutive video frames of state multimedia data, a pixel-by-pixel motion vector can be calculated using an optical flow algorithm, retaining the optical flow information within the region to be learned and eliminating background interference optical flow. Through direction or amplitude analysis of the optical flow vectors, the motion trajectory of key components in the region to be learned can be located, and the multimedia data corresponding to each key component region can be used as the corresponding component multimedia data.
[0126] In another alternative implementation, motion decomposition of the state multimedia data can be performed based on feature tracking. For example, key feature points of the object to be learned can be extracted from the state multimedia data, and these feature points can be continuously tracked in consecutive video frames to record their motion trajectories. Based on the motion trajectory of the feature point, the component type corresponding to that feature point is determined (e.g., if the motion trajectory indicates that the feature point is stationary, it is determined to be a stationary component; if the motion trajectory indicates that the feature point moves over time, it is determined to be a moving component), and the multimedia data corresponding to that component region is used as the multimedia data of the corresponding component. Optionally, the key feature points of the object to be learned in the multimedia data can be extracted using a common Scale-Invariant Feature Transform (SIFT) algorithm to extract unique features such as corners and edges, which will not be elaborated upon here.
[0127] In one optional implementation, the component multimedia data corresponding to each key component includes the image data of the corresponding key component. Accordingly, the image data of the corresponding key component can be masked to obtain the mask label of the corresponding key component, and the mask label of each key component can be used as the component mask pseudo label of the object to be learned.
[0128] It should be noted that there are many other ways to generate pseudo-labels for components of the object to be learned based on component multimedia data, which will be described in the following embodiments.
[0129] S540 generates pseudo-state labels based on state multimedia data to reflect changes in the state of the object to be learned.
[0130] In one optional implementation, the state multimedia data includes M frames of multimedia data, where M is a positive integer; correspondingly, the motion type and motion constraint data of the object to be learned can be obtained; for the i-th frame of multimedia data in the M frames, based on the motion type, motion constraint data, and state data of the object to be learned in the i-th frame of multimedia data, the motion parameters of the object to be learned in the i-th frame of multimedia data are determined; i is a positive integer less than or equal to M; based on the motion parameters and motion constraint data corresponding to the M frames of multimedia data, state pseudo-labels reflecting the state changes of the object to be learned are generated.
[0131] Understandably, since state multimedia data is used to reflect the state changes of the object to be learned during its interaction with the embodied robot, any frame of state multimedia data includes the state of the object to be learned at a specific moment during the transition from one state to another. For example, state multimedia data includes the states of a cabinet door at different moments as it transitions from being completely closed to being completely opened by the embodied robot. Similarly, state multimedia data includes the states of a drawer at different moments as it transitions from being completely closed to being completely pulled out by the embodied robot (i.e., the travel distance of the drawer in different frames when it is pulled by the embodied robot).
[0132] For example, for the i-th frame of multimedia data, the state data of the object to be learned in the current frame is extracted. Combined with the motion type and motion constraint data of the object to be learned, the state data of the current frame is substituted into the motion constraint rules of the object to be learned for calculation. The state data of the object to be learned is converted into quantized data, which serves as the motion parameters of the object to be learned in the i-th frame of multimedia data. For example, if the motion type of the object to be learned is rotation, the motion parameter represents the current opening / closing angle; if the motion type of the object to be learned is translation, the motion parameter represents the current pull-out stroke. The state data includes at least one of the following: the position of the object to be learned, its contour boundary, and the offset of the moving part relative to the stationary part.
[0133] Optionally, the motion parameters corresponding to each frame of multimedia data can be used as pseudo-labels for the state of the object to be learned.
[0134] Optionally, the motion parameters corresponding to each frame in the M frames can be arranged in frame order to form a continuous state change sequence of the object to be learned from the initial state to the final state. Then, according to a preset state interval division rule, the continuous motion parameters are mapped to discrete state identifiers corresponding to the motion constraint data (each motion parameter corresponds to a discrete state identifier). For example, the opening angle can be divided into three intervals: closed, half-open, and fully open; the translational stroke can be divided into three intervals: closed, half-pull, and fully pulled. Furthermore, the discrete state identifier of the motion parameters corresponding to each frame of multimedia data is used as the pseudo-label of the state corresponding to that frame.
[0135] In some embodiments, motion constraint data includes motion constraint direction and motion constraint range. Correspondingly, the movable direction and movable range of the object to be learned can be determined based on the motion constraint direction and the motion constraint range. Based on the motion parameters corresponding to the M-frame multimedia data, and the movable direction and movable range of the object to be learned, the motion constraint model of the object to be learned is determined. Based on the motion constraint model of the object to be learned, a pseudo-state label is generated to reflect the state changes of the object to be learned.
[0136] Among them, the motion constraint direction is used to characterize the reasonable direction of the actual executable movement of the object to be learned, and the motion constraint direction corresponds to the movable direction. The motion constraint range is used to characterize the allowed range of movement of the object to be learned in the movable direction, and the movable range corresponds to the motion constraint range. For example, if the motion constraint direction and motion constraint range of the object to be learned indicate that the object to be learned can open 90 degrees to the left, then the movable direction of the object to be learned is to the left, and the movable range is 90 degrees; as another example, if the motion constraint direction and motion constraint range of the object to be learned indicate that the object to be learned can be pulled forward 50 centimeters, then the movable direction of the object to be learned is forward, and the motion constraint range is 50 centimeters.
[0137] The motion constraint model can be understood as a model used to characterize the motion mode, movable range, and state change law of the object to be learned, reflecting the fixed motion rules of the object under the physical structure. For example, it includes at least one of the following: the direction of motion allowed, the maximum amplitude of motion, and the corresponding quantized state at each moment.
[0138] In some embodiments, the pseudo-state labels corresponding to the learning object under different motion states can be determined according to the motion constraint model of the learning object, and the obtained pseudo-state labels can be used as the set of pseudo-state labels of the learning object.
[0139] It should be noted that the motion type and motion constraint data of the learning object are obtained by inverse joint calculation of the learning object based on the state multimedia data and the motion interaction data of the embodied robot. The specific process will be further described in the following embodiments.
[0140] S550 generates object pseudo-labels for the object to be learned based on state pseudo-labels and component pseudo-labels.
[0141] The object pseudo-labels are used by the embodied robot to learn and train on the learning object. Optionally, an embodied robot can learn and train its own model based on pseudo-labels generated from its own interactive multimedia data; it can also perform joint learning and knowledge sharing based on pseudo-labels generated by other embodied robots, thereby improving learning efficiency and generalization ability. The embodied robot that generates pseudo-labels based on interactive multimedia data (or the embodied robot in the interactive multimedia data) and the embodied robot that uses object pseudo-labels for learning and training can be the same embodied robot or different embodied robots. For example, if the embodied robot in the interactive multimedia data is embodied robot A, after generating object pseudo-labels for the learning object from the interactive multimedia data, these object pseudo-labels can be used to learn and train embodied robot A, or embodied robot B or other robots.
[0142] In one optional implementation, for any object to be learned, all its corresponding state pseudo-labels and the component pseudo-labels corresponding to key components can be used as the object pseudo-label of the object to be learned.
[0143] In another alternative implementation, component pseudo-labels and state pseudo-labels within the same frame can be aligned in time and position. For example, region information for each component such as door panels, drawers, handles, and hinges can be extracted from the component pseudo-labels, and state change information such as opening / closing angles, push / pull strokes, and on / off states can be extracted from the state pseudo-labels. These component structural and state information are then mapped one-to-one and integrated to form a complete label that simultaneously contains the object's structural composition and real-time state. This yields pseudo-labels for the learning object in different frames, which are then used as the object's pseudo-labels. Alternatively, the component pseudo-labels, state pseudo-labels, and object pseudo-labels of the learning object can be used. During subsequent training, the embodied robot can use the component pseudo-labels to understand the component functions and positional relationships of each component of the learning object, and further determine the state of the learning object in each frame using the state pseudo-labels. Based on the component pseudo-labels and state labels, the robot's motion joint parameters are determined to achieve accurate and adaptive execution of operational tasks related to the learning object.
[0144] In the aforementioned pseudo-label generation method for embodied robot self-learning, interactive multimedia data between the embodied robot and the learning target is acquired, and state multimedia data is determined from this data. Further, motion decomposition is performed on the state multimedia data to obtain component multimedia data of the key parts of the learning target. Based on this component multimedia data, component pseudo-labels for the learning target are generated, thereby generating state pseudo-labels reflecting state changes of the learning target based on the state multimedia data. Finally, based on the state and component pseudo-labels, object pseudo-labels for the learning target are generated; these object pseudo-labels are used by the embodied robot for learning and training. In this process, on the one hand, by collecting interactive multimedia data between the embodied robot and the learning target, key component information and state change information are automatically decomposed, generating component, state, and object pseudo-labels without manual annotation, significantly reducing data annotation costs and manpower consumption. On the other hand, by fusing key component information and dynamic state information of an object to form a complete object pseudo-label, embodied robots can learn the composition structure, motion laws and state change characteristics of the object to be learned more comprehensively and accurately, thereby improving the recognition accuracy and state judgment reliability of different objects to be learned by embodied robots, and thus improving the stability and success rate of subsequent autonomous operation and interactive control.
[0145] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the key components include moving components and stationary components. Accordingly, the steps of performing motion decomposition on the state multimedia data to obtain the component multimedia data of the key components of the object to be learned are described.
[0146] See Figure 6 The component multimedia data determination steps shown include:
[0147] S610 performs motion decomposition on the state multimedia data to obtain multimedia data of the moving region and multimedia data of the stationary region.
[0148] The "following area" refers to the area corresponding to the moving parts of the learning object, while the "stationary area" refers to the area corresponding to the stationary parts of the learning object. For example, taking doors and drawers as learning objects, during the process of the embodied robot opening and closing a cabinet door, the door panel will rotate and change position with the opening and closing action; this area is the following area. The door frame and cabinet body, which are connected to the cabinet and remain in a fixed position, are the stationary areas. Similarly, during the process of the embodied robot pushing and pulling a drawer, the drawer front and inner drawer will move back and forth with the pushing and pulling action; this area is the following area. The drawer slides and cabinet frame, which are fixed to the furniture and do not change position, are the stationary areas.
[0149] For example, in a series of consecutive frames of state multimedia data, pixel changes at the same position between different frames can be compared to identify areas where the position has moved significantly and areas where the position remains essentially unchanged. Then, areas whose position changes with the movement of the object to be learned are designated as moving regions, and the corresponding image information is extracted to obtain multimedia data for the moving regions. Conversely, areas whose position remains essentially fixed throughout the movement are designated as stationary regions, and the corresponding image information is extracted to obtain multimedia data for the stationary regions.
[0150] There are various ways to perform motion decomposition on state multimedia data, and the different decomposition methods have been described in detail in the foregoing embodiments. Therefore, the process of determining the moving region multimedia data and the stationary region multimedia data based on motion decomposition described above is only an exemplary implementation method provided by this application and does not constitute a limitation on the specific implementation method of motion decomposition.
[0151] Accordingly, the above-mentioned generation of pseudo-labels for the learning object based on component multimedia data includes: performing masking processing on key multimedia data of the moving region to obtain a moving component mask, and determining the moving component mask as the mask label of the moving component; performing masking processing on key multimedia data of the stationary region to obtain a stationary component mask, and determining the stationary component mask as the mask label of the stationary component; and generating pseudo-labels for the learning object based on the mask labels of the moving component and the stationary component.
[0152] For example, within the defined moving area, movable key component areas such as door panels and drawer panels can be located. These areas are then masked using pixel-level annotation or contour extraction to obtain a moving component mask containing only the moving component areas, which is then used as the mask label for the moving component. Similarly, within the stationary area, fixed key component areas such as cabinets, door frames, and furniture frames are located. These areas are masked in the same way to obtain a stationary component mask containing only the stationary component areas, which is then used as the mask label for the stationary component. Furthermore, the moving component mask labels and stationary component mask labels corresponding to the same learning object are spatially aligned and merged to form pseudo-labels that completely distinguish between moving and stationary components.
[0153] In some embodiments, to accurately locate and finely segment the components of the learning object, after determining the mask labels corresponding to each component, semantic category names can be assigned to the mask labels corresponding to each component, such as handle, door panel, cabinet, drawer panel, etc. For example, a text-based Grounded-Segment-Anything (SAM) model can be used, which uses text guidance to detect and segment key components at the pixel level, thereby automatically assigning corresponding semantic categories to the mask of each component.
[0154] In some embodiments, to further improve the accuracy and boundary regularity of the mask pseudo-labels, the boundaries of the mask labels for each key component can be refined based on a preset segmenter. The preset segmenter may include models with high-precision segmentation capabilities, such as the SAM model, which optimizes the edges and corrects the contours of the mask region, making the generated mask pseudo-labels more closely match the actual boundaries of the key components and improving label accuracy.
[0155] S620 generates component multimedia data for key components based on multimedia data from the moving region, multimedia data from the stationary region, and state multimedia data.
[0156] In one alternative implementation, the multimedia data of the servo region can be directly used as the multimedia data of the moving part; and the multimedia data of the stationary region can be used as the multimedia data of the stationary part.
[0157] In another alternative implementation, the key components also include operable components. Accordingly, multimedia data of the end effector region of the embodied robot can be determined from the state multimedia data; and multimedia data of the operable components can be determined based on the intersection between the multimedia data of the end effector region and the multimedia data of the follower region.
[0158] In this context, operable parts can be understood as components on the object to be learned that can be directly touched, force-applied, and driven to move by the embodied robot, such as cabinet door handles, drawer pulls, and switch buttons. The end effector of the embodied robot can be understood as the execution structure used by the embodied robot to perform operations such as grasping and pushing, such as grippers and robotic arms. The corresponding end effector area is the region in the image corresponding to the end effector of the embodied robot.
[0159] For example, based on the appearance features of the robot end effector, the multimedia data corresponding to the end effector region can be identified and extracted from the state multimedia data first. Then, the multimedia data of the end effector region and the multimedia data of the follower region are spatially compared. The intersection area where the two overlap is taken as the region where the operable part is located, and the multimedia data of this region is taken as the multimedia data of the operable part.
[0160] In another alternative embodiment, the key component further includes a motion-coordinating component. Accordingly, the motion-coordinating relationship between the following region and the stationary region can be obtained, and the motion-coordinating region of the state multimedia data can be analyzed based on the motion-coordinating relationship to obtain the multimedia data of the motion-coordinating component.
[0161] In this context, kinematic components can be understood as parts on the object being learned that connect and transmit power between moving and stationary parts. Examples include hinges connecting a door panel to a cabinet, or drawer slides connecting a drawer to a cabinet. These kinematic components act as coordination and constraints during the movement of the object being learned.
[0162] In this context, motion coordination can be understood as the connection method of relative motion between the moving region and the stationary region, such as rotation and translation. The corresponding motion coordination region is the area in the image corresponding to the motion coordination component.
[0163] For example, the motion coordination relationship (e.g., rotation, translation, etc.) between the following region and the stationary region can be determined based on the relative motion trajectory of the following region and the stationary region in a continuous frame. Then, based on the position of the following region, the position of the stationary region and the motion coordination relationship, the motion coordination region can be determined. The motion coordination region is used as the region where the motion coordination component is located, and the multimedia data of the region is used as the multimedia data of the motion coordination component.
[0164] Optionally, the motion coordination region can be determined by: separately determining the boundary coordinates and spatial positions of the moving region and the stationary region; and then, based on the rotational or translational coordination relationship between them, finding the boundary region where the moving region and the stationary region connect and are close to each other, and using this region as the motion coordination region. For example, if the motion coordination relationship is determined to be a rotational relationship, the connecting region where the moving region rotates around the stationary region is taken as the motion coordination region; if the motion coordination relationship is determined to be a translational relationship, the contact region where the moving region slides relative to the stationary region is taken as the motion coordination region.
[0165] Furthermore, the multimedia data of the moving area, the multimedia data of the stationary area, the multimedia data of the operable parts, and the multimedia data of the moving parts can be identified as the component multimedia data of the key components.
[0166] Accordingly, the above-mentioned generation of pseudo-labels for components based on component multimedia data includes: performing masking processing on the multimedia data of operable components to obtain an operable component mask, and determining the operable component mask as the mask label of the operable component; performing masking processing on the multimedia data of kinematic components to obtain a kinematic component mask, and determining the kinematic component mask as the mask label of the kinematic component; and determining one or more of the mask labels of kinematic components, static components, operable components, and kinematic components as pseudo-labels for components related to the object to be learned.
[0167] In the above embodiments, by accurately distinguishing and extracting multimedia data of the following region, stationary region, operable parts, and motion-cooperating parts, the key parts of the learning object can be completely characterized, as well as the motion relationships between the key parts and the functions of each part (such as stationary parts having a support function, operable parts having operability, etc.). This makes the generated part pseudo-labels more comprehensive and accurate, and closer to the actual physical structure and motion mechanism of the learning object. This provides comprehensive and effective data support for the subsequent structural learning, motion understanding, and autonomous interaction of the embodied robot with the target object.
[0168] In some embodiments, the component multimedia data includes pixel position data; the component pseudo-label also includes position labels for each key component and relative positional relationship labels between different key components; correspondingly, for any key component, the position information of the key component can be determined based on the pixel position data of the key component, and the position label of the key component can be determined based on the position information; based on the position information of the key component and the position information of other key components, a relative positional relationship label between the key component and other key components can be generated.
[0169] In one optional implementation, the pixel position data corresponding to each pixel of the key component can be used as the pixel set of the corresponding key component, and the centroid of the pixel set, i.e., the centroid of the key component, can be determined as the position reference point of the key component. Based on the relative positional relationship between the position reference point corresponding to the key component and the preset initial position, the position label of the key component is determined.
[0170] In this context, the centroid of a pixel set can be understood as the geometric center of all pixels contained in the corresponding component mask within the image, equivalent to the position of the center pixel of the component region on the image. For example, a component mask corresponds to a 3×3 pixel set with pixel coordinates (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), (3,3). Summing the x-coordinates and y-coordinates of all pixels in this set yields 18. Dividing these by the total number of pixels (9) gives the centroid coordinates of the pixel set as (2,2).
[0171] For example, taking the household embodied robot grasping a cabinet door handle as an example, the position data of all pixels in the handle area can be extracted first to form the pixel set of the handle. The centroid of the pixel set is calculated as the position reference point of the handle. The position reference point is compared with the preset standard grasping center position. Based on the relative offset direction and offset amount of the two, a position label of center, left, right, top or bottom is generated for the handle.
[0172] Furthermore, based on information such as the relative distance, relative orientation, or relative angle between the centroids of each key component, the relative positional relationship labels between the key components are determined.
[0173] For example, taking a cabinet door as an example, first calculate the centroids of the handle and the hinge as reference points for their positions. Then, determine the orientation relationship based on the coordinate difference between the two reference points and generate relative position relationship labels such as "the handle is located on the opposite side of the hinge" or "the handle and the hinge are at the same horizontal height of the door panel".
[0174] In the above embodiments, component location tags are added to the component pseudo-tags, enriching the information in the component pseudo-tags. This allows the embodied robot to not only accurately locate the positions of key components after learning the object tags of the learning object, but also to clearly understand the positional and structural relationships between different key components, thereby improving self-learning efficiency and the success rate of subsequent interactive tasks.
[0175] Based on the technical solutions of the above embodiments, this application also provides an optional embodiment. In this optional embodiment, the process of obtaining the motion coordination relationship between the following region and the stationary region is described.
[0176] See Figure 7 The steps for determining the motion coordination relationship shown include:
[0177] S710 determines the motion interaction data of the embodied robot from the interactive multimedia data.
[0178] Among them, motion interaction data reflects the joint states of the embodied robot during its interaction with the object to be learned. For example, when the embodied robot performs opening, closing, pushing, and pulling actions, the joint angles, poses, and motion trajectories of the end effector are recorded.
[0179] In one optional implementation, the time period corresponding to the interactive multimedia data can be acquired, and the joint state data stream of the embodied robot during the corresponding time period can be used as the motion interaction data of the embodied robot. For example, the joint state data stream for the corresponding time period can be acquired from the historical control records of the embodied robot controller.
[0180] In another alternative implementation, the position of the end effector of the embodied robot in the interactive multimedia data can be identified, and the state changes of the end effector in consecutive frames can be determined as the motion interaction data of the embodied robot. The state changes of the end effector include at least one of position change, posture change, motion trajectory, motion speed, and contact state.
[0181] S720 performs inverse joint calculation on the object to be learned based on the following region, the stationary region, and the joint state to obtain the motion type and motion constraint data of the object to be learned.
[0182] The motion constraint data includes one or more of the motion constraint direction and motion constraint range.
[0183] Joint inverse calculation refers to deducing the motion structure and constraint relationships of an object from its observed motion state. In this embodiment, the embodied robot does not need to know the motion mode of objects such as cabinet doors and drawers in advance. It can deduce the pivot position, slide rail direction, rotation range, or sliding stroke simply by interacting with the object to be learned and observing its motion process. That is, the motion type and motion constraint data of the object to be learned can be determined based on the motion state of the robot and the object to be learned at different times during the interaction.
[0184] In one alternative implementation, a joint inverse model can be pre-trained, and different video frames can be input into the joint inverse model to obtain the motion type and motion constraint data of the object to be learned. The video frames are labeled with the following regions, stationary regions, and joint states.
[0185] In another alternative implementation, the correspondence between joint states and motion trajectories can be determined based on the motion trajectories of the following and stationary regions in consecutive frames, combined with the joint states of the embodied robot in consecutive frames. Based on this correspondence, the changes of the learning object as the embodied robot operates can be determined, thereby determining the motion type and motion constraint data of the learning object.
[0186] For example, if the follower region moves forward at a constant speed relative to the stationary region in consecutive frames, and the robot arm joints exhibit a smooth horizontal pushing motion, then the learning object can be determined to be of the translational motion type, and its motion constraint direction can be determined to be horizontal forward, and the motion constraint range can be determined to be the maximum displacement range of this interaction.
[0187] Accordingly, in the aforementioned process of determining the motion coordination region, when the motion type characterization of the learning object is rotational, the region corresponding to the motion constraint direction in the motion constraint data can be used as the motion coordination region. Furthermore, from the motion coordination region, the rotation axis of the learning object is determined; the multimedia data of the region where the rotation axis is located is then determined as the multimedia data of the motion coordination component.
[0188] For example, the position of the central axis of rotation in the motion coordination area can be used as the rotation axis of the object to be learned, and the multimedia data corresponding to the local image area where the rotation axis is located can be used as the multimedia data of the motion coordination component.
[0189] S730 determines the motion type and motion constraint data of the object to be learned as the motion coordination relationship between the follower region and the stationary region.
[0190] In the above embodiments, by acquiring the joint states when the embodied robot interacts with the object to be learned, and combining the follower region and the stationary region to perform joint inverse calculation, the motion type and motion constraint data of the object to be learned are determined, and the motion coordination relationship between the follower region and the stationary region is determined. This can provide reliable constraints for the subsequent operation of the embodied robot, and improve the accuracy of operation and environmental adaptability.
[0191] It is worth noting that in this embodiment, by performing joint inverse kinematics on the learning object based on the following region, stationary region, and joint state, the motion type and motion constraint data of the learning object can be obtained, thereby accurately representing the motion structure of the learning object, rather than just describing the robot's own motion. For example, for objects such as cabinet doors and drawers, joint inverse kinematics can clarify their rotation axis, rotation range, translation direction, and maximum stroke, etc., bringing multiple beneficial effects: First, it can generate accurate state pseudo-labels for the learning object based on the motion type and motion constraint data, quantifying the motion of the learning object into indicators such as opening and closing angles and pull-out strokes, and discretizing them into state intervals such as closed, half-open, and fully open, achieving a refined description of different states of the learning object; second, it can refine the component pseudo-labels of the learning object, distinguishing motion cooperation areas such as the hinge side and handle side based on the motion type and motion constraint data, enabling the embodied robot not only to identify moving and stationary regions, but also to better... The system provides four key functions: First, it understands the functional types of different key components in the learning object. Second, it supports cross-view and cross-state propagation of pseudo-labels. Based on motion constraint data, it constructs a motion constraint model of the learning object, enabling the 3D projection and propagation of pseudo-labels from multiple perspectives. This simulates the pseudo-labels of the learning object in different opening and closing states, thus expanding the self-learning samples of the embodied robot. Third, it provides crucial support for the subsequent operation planning and task execution of the embodied robot. The motion type and motion constraint data of the learning object can provide accurate motion constraint basis for the embodied robot, guiding it to complete the operation task along a constraint trajectory that conforms to the motion law of the learning object, thereby improving the rationality and success rate of the embodied robot's interactive actions.
[0192] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0193] Based on the same inventive concept, embodiments of this application also provide an embodied robot system for implementing the aforementioned pseudo-tag-based embodied robot self-learning method, and an embodied robot system for implementing the aforementioned pseudo-tag generation method for embodied robot self-learning. The solutions provided by the above two systems are similar to the solutions described in their corresponding methods; therefore, the specific limitations in one or more embodied robot system embodiments provided below can be found in the limitations of their corresponding methods described above, and will not be repeated here.
[0194] Based on the same inventive concept, this application also provides a pseudo-tag-based embodied robot self-learning device for implementing the pseudo-tag-based embodied robot self-learning method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations of one or more pseudo-tag-based embodied robot self-learning device embodiments provided below can be found in the limitations of the pseudo-tag-based embodied robot self-learning method described above, and will not be repeated here.
[0195] In one exemplary embodiment, such as Figure 8 As shown, a pseudo-label-based embodied robot self-learning device is provided, comprising: a first generation module 810, a second generation module 820, a third generation module 830, and a training module 840. Wherein:
[0196] The first generation module 810 is used to generate pseudo-labels of the state and pseudo-labels of the learning object based on the interactive multimedia data between the embodied robot and the learning object.
[0197] The second generation module 820 is used to generate cross-perspective pseudo-labels for the object to be learned based on state pseudo-labels, component pseudo-labels and interactive multimedia data.
[0198] The third generation module 830 is used to perform motion simulation on the learning object based on interactive multimedia data, obtain the simulation multimedia data of the learning object, and generate enhanced pseudo-labels about the learning object based on the simulation multimedia data.
[0199] Training module 840 is used to verify state pseudo-labels, part pseudo-labels, cross-view pseudo-labels, and augmented pseudo-labels, and to train the embodied robot based on the verified pseudo-labels.
[0200] In one exemplary embodiment, such as Figure 9 As shown, a pseudo-label generation device for self-learning of embodied robots is provided, comprising: an acquisition module 910, a determination module 920, a fourth generation module 930, a fifth generation module 940, and a sixth generation module 950, wherein:
[0201] The acquisition module 910 is used to acquire multimedia data of the interaction between the embodied robot and the object to be learned;
[0202] The determination module 920 is used to determine the state multimedia data from the interactive multimedia data; the state multimedia data reflects the state changes of the learning object during the interaction with the embodied robot.
[0203] The fourth generation module 930 is used to perform motion decomposition on the state multimedia data to obtain the component multimedia data of the key components of the object to be learned, and to generate component pseudo-labels for the object to be learned based on the component multimedia data.
[0204] The fifth generation module 940 is used to generate pseudo-state labels that reflect changes in the state of the object to be learned based on state multimedia data.
[0205] The sixth generation module 950 is used to generate object pseudo-labels for the object to be learned based on state pseudo-labels and component pseudo-labels; the object pseudo-labels are used by the embodied robot to learn and train on the object to be learned.
[0206] The modules in the aforementioned pseudo-tag-based embodied robot self-learning device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0207] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements a pseudo-tag-based embodied robot self-learning method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0208] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0209] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0210] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0211] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0212] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0213] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0214] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0215] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A self-learning method for embodied robots based on pseudo-labels, characterized in that, The method includes: Based on the interactive multimedia data between the embodied robot and the learning object, pseudo-labels for the state and components of the learning object are generated; the interactive multimedia data is collected by the embodied robot from a first-person perspective. Based on the component pseudo-labels and the interactive multimedia data, the object to be learned is subjected to 3D spatial reprojection processing to obtain a 3D semantic map; based on the 3D semantic map, the object multimedia data of the object to be learned is obtained from a second perspective; the second perspective is different from the first perspective; based on the object multimedia data, component mask pseudo-labels of the object to be learned in the second perspective are determined; based on the component mask pseudo-labels and the state pseudo-labels, cross-perspective pseudo-labels for the object to be learned are determined. Based on the interactive multimedia data, motion simulation is performed on the object to be learned to obtain the simulated multimedia data of the object to be learned, and enhanced pseudo-labels about the object to be learned are generated based on the simulated multimedia data. The state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhancement pseudo-label are verified, and the embodied robot is trained based on the verified pseudo-labels.
2. The method according to claim 1, characterized in that, The step of determining cross-view pseudo-labels for the object to be learned based on the component mask pseudo-labels and the state pseudo-labels includes: When the number of second perspectives is one, the component mask pseudo-label and the state pseudo-label are determined as cross-perspective pseudo-labels for the object to be learned; When there are multiple second perspectives, based on the component mask pseudo-labels and the state pseudo-labels of each second perspective, a perspective pseudo-label corresponding to the second perspective is generated, and the perspective pseudo-labels corresponding to the multiple second perspectives are determined as cross-perspective pseudo-labels about the object to be learned.
3. The method according to claim 1, characterized in that, The step of performing motion simulation on the object to be learned based on the interactive multimedia data to obtain the simulation multimedia data of the object to be learned includes: Based on the interactive multimedia data, joint inverse calculation is performed on the object to be learned to obtain the motion type and corresponding motion constraint data of the object to be learned. Based on the motion type and corresponding motion constraint data of the object to be learned, multi-state simulation transformation is performed on the object to be learned to obtain the transformed multimedia data of the object to be learned under different state simulation transformations. The multimedia data of the object to be learned under different state simulation transformations are determined as the simulation multimedia data of the object to be learned.
4. The method according to claim 3, characterized in that, The simulated multimedia data includes transformed multimedia data under N state simulation transformations; N is a positive integer. The generation of enhanced pseudo-labels for the object to be learned based on the simulated multimedia data includes: For state simulation transformation j among the N state simulation transformations, based on the transformed multimedia data under state simulation transformation j, determine the motion constraint data of the object to be learned under state simulation transformation j; j is a positive integer less than or equal to N; Based on the motion constraint data of the object to be learned under the state simulation transformation j and the motion type of the object to be learned, the motion parameters of the object to be learned under the state simulation transformation j are determined. Based on the motion parameters of the object to be learned under the state simulation transformation j and the component pseudo-labels, generate simulated pseudo-labels of the object to be learned under the state simulation transformation j; The simulated pseudo-labels corresponding to the object to be learned under the N simulated state transformations are determined as enhanced pseudo-labels for the object to be learned.
5. The method according to claim 4, characterized in that, The step of generating simulated pseudo-labels for the object to be learned under the state simulation transformation j based on the motion parameters of the object to be learned under the state simulation transformation j and the pseudo-labels of the components includes: Based on the motion parameters of the object to be learned under the state simulation transformation j, determine the simulation state pseudo-label of the object to be learned; Sample component masks are obtained from the memory bank of the embodied robot, and the sample component masks are used to mask the component pseudo-labels to obtain masked component pseudo-labels; The simulated state pseudo-label and the occlusion component pseudo-label are determined as the simulated pseudo-labels of the object to be learned under the simulated state transformation j.
6. The method according to claim 1, characterized in that, The process of verifying the state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhanced pseudo-labels, and training the embodied robot based on the verified pseudo-labels, includes: Based on the preset verification logic, the state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhanced pseudo-label are verified respectively to obtain the pseudo-label verification results. Based on the pseudo-label verification results, quality scores are generated for the state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhancement pseudo-label, respectively. Based on the quality score, pseudo-labels that pass verification are identified as qualified pseudo-labels. In response to the self-learning instruction of the embodied robot, the embodied robot performs self-learning based on the qualified pseudo-label; wherein the self-learning instruction is generated when the self-learning conditions are detected.
7. The method according to claim 6, characterized in that, The preset verification logic includes motion constraint verification; Based on preset verification logic, the state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhanced pseudo-label are verified respectively to obtain pseudo-label verification results, including at least one of the following: When the state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhancement pseudo-label all conform to the motion type and corresponding motion constraint data of the object to be learned, a pseudo-label verification result is generated to indicate that the verification has passed. If any of the state pseudo-labels, component pseudo-labels, cross-view pseudo-labels, and enhancement pseudo-labels do not conform to the motion type or corresponding motion constraint data of the object to be learned, a pseudo-label verification result is generated to indicate that the verification has failed.
8. The method according to claim 6, characterized in that, The preset verification logic includes multi-perspective consistency verification; Based on preset verification logic, the state pseudo-label, the component pseudo-label, the cross-view pseudo-label, and the enhanced pseudo-label are verified respectively to obtain pseudo-label verification results, including: From the state pseudo-labels, the component pseudo-labels, the cross-view pseudo-labels, and the enhanced pseudo-labels, pseudo-labels with location features are selected as the filtered pseudo-labels; The pseudo-labels under the first viewpoint are determined from the filtered pseudo-labels to obtain the first viewpoint pseudo-labels. The first viewpoint pseudo-labels are then projected in three dimensions to obtain the three-dimensional features of the object to be learned under the first viewpoint. From the filtered pseudo-labels, pseudo-labels under the second perspective are determined to obtain the second perspective pseudo-labels. The second perspective pseudo-labels are then projected in three dimensions to obtain the three-dimensional features of the object to be learned under the second perspective. If the 3D features in the first viewpoint are consistent with the 3D features in the second viewpoint, a pseudo-label verification result is generated to indicate that the verification has passed.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: If a target task is received regarding the object to be learned, then the motion constraint model of the object to be learned is obtained; the motion constraint model is learned by the embodied robot based on the verified pseudo-labels; Based on the target task, the target state that the learning object needs to achieve is determined, and based on the target state, the current state of the learning object, and the motion constraint model, the execution joint parameters to be executed by the embodied robot are determined. Control the embodied robot to execute the joint parameters until the target task is completed.
10. A hymenoidae robot system, characterized in that, The embodied robot is used to implement the steps of the pseudo-label-based embodied robot self-learning method as described in any one of claims 1-9; the embodied robot system includes: The first generation module is used to generate pseudo-labels for the state and pseudo-labels for the learning object based on the interactive multimedia data between the embodied robot and the learning object; the interactive multimedia data is collected by the embodied robot from a first-person perspective; The second generation module is used to perform three-dimensional spatial reprojection processing on the object to be learned based on the component pseudo-labels and the interactive multimedia data to obtain a three-dimensional semantic map; based on the three-dimensional semantic map, acquire object multimedia data of the object to be learned from a second perspective; the second perspective is different from the first perspective; based on the object multimedia data, determine the component mask pseudo-labels of the object to be learned in the second perspective; and based on the component mask pseudo-labels and the state pseudo-labels, determine cross-perspective pseudo-labels for the object to be learned. The third generation module is used to perform motion simulation on the object to be learned based on the interactive multimedia data, obtain the simulation multimedia data of the object to be learned, and generate enhanced pseudo-labels for the object to be learned based on the simulation multimedia data. The training module is used to verify the state pseudo-labels, the component pseudo-labels, the cross-view pseudo-labels, and the enhancement pseudo-labels, and to train the embodied robot based on the verified pseudo-labels.
Citation Information
Patent Citations
Bird-eye view semantic segmentation label generation method based on multi-frame semantic point cloud splicing
CN114445593A
Cross-view-angle image geographic positioning method based on dynamic threshold value pseudo label self-training learning
CN120950725A