Method, device, equipment and medium for managing state of robot
By using machine learning models and skeleton image coding technology, the problem of difficult to manage and predict the state of high-degree-of-freedom robots in the prior art is solved, and more accurate and effective state management and prediction are achieved.
Patent Information
- Application Number
- CN202510324594.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively manage and predict the state of robots with high degrees of freedom and complex structures, especially in the case of insufficient training data and inaccurate action coding.
Using a machine learning model, the state image of the robot at the initial time point and the action image of the future time point (represented as a skeleton image) is determined by receiving the state image of the robot at the initial time point. The model includes a skeleton encoder, a video encoder, a diffusion model and a video decoder, which can generate more accurate action representations and state predictions.
It realizes more accurate and effective robot state management and prediction, can handle robot actions with different structures and high degrees of freedom, and improves the generalization ability and prediction accuracy of the model.
Smart Images

Figure CN120163994A_ABST
Abstract
Description
Technical Field
[0001] Implementations of the present disclosure generally relate to robot management, and particularly to methods, apparatuses, devices, and computer-readable storage media for managing the state of a robot using a machine learning model. Background Art
[0002] In recent years, robotics technology has developed rapidly and has been widely used in multiple technical fields. For example, on a factory production line, robotic arms can be used to perform various actions such as processing, grasping, sorting, and packaging. Further, machine learning technology has also been widely applied in multiple application scenarios, and in different application scenarios, robots can be controlled to perform different actions in order to complete expected tasks respectively. At this time, it is desirable to combine robotics technology and machine learning technology to represent actions in a more accurate manner, and then manage the operation of the robot in a more accurate and efficient manner. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for managing the state of a robot is provided. In this method, a first image describing the first state of the robot at a first time point is received. An action image is received, the action image specifying an action to be performed by the robot at a second time point after the first time point, and the action is represented by a skeleton image of the robot. Using a machine learning model, based on the first image and the action image, a second image of the second state of the robot at the second time point is determined.
[0004] In a second aspect of the present disclosure, an apparatus for managing the state of a robot is provided. The apparatus includes: a state image receiving module configured to receive a first image describing the first state of the robot at a first time point; an action image receiving module configured to receive an action image, the action image specifying an action to be performed by the robot at a second time point after the first time point, and the action is represented by a skeleton image of the robot; and a determining module configured to use a machine learning model to determine, based on the first image and the action image, a second image of the second state of the robot at the second time point.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processor; and at least one memory, the at least one memory being coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to the first aspect of the present disclosure when executed by the at least one processor.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, the computer program causing a processor to implement the method according to the first aspect of the present disclosure when executed by the processor.
[0007] In a fifth aspect of the present disclosure, there is provided a computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this section is not intended to limit the key features or important features of the implementation manners of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In the following, with reference to the accompanying drawings and the following detailed description, the above and other features, advantages and aspects of the implementation manners of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A block diagram showing an application environment of a robot according to an implementation manner of the present disclosure;
[0011] Figure 2 A block diagram showing a state management of a robot according to some implementation manners of the present disclosure;
[0012] Figure 3A A block diagram showing an architecture of a machine learning model according to some implementation manners of the present disclosure;
[0013] Figure 3B A block diagram showing an architecture of a machine learning model according to some implementation manners of the present disclosure;
[0014] Figure 4 A block diagram showing a video encoding process and a decoding process according to some implementation manners of the present disclosure;
[0015] Figure 5 A block diagram showing an architecture of a machine learning model according to some implementation manners of the present disclosure;
[0016] Figure 6 A block diagram showing a determination of a reference sample based on human data according to some implementation manners of the present disclosure;
[0017] Figure 7 A block diagram showing an algorithm for determining a hand according to some implementation manners of the present disclosure;
[0018] Figure 8 A block diagram showing a determination of a reference sample based on robot data according to some implementation manners of the present disclosure;
[0019] Figure 9 A block diagram showing a prediction process according to some implementation manners of the present disclosure;
[0020] Figure 10 A block diagram showing a process for training a machine learning model according to some implementations of the present disclosure;
[0021] Figure 11 A block diagram showing a process for generating a video of a robot according to some implementations of the present disclosure;
[0022] Figure 12 A block diagram showing a process for using a machine learning model according to some implementations of the present disclosure;
[0023] Figure 13 A flowchart showing a method for managing the state of a robot according to some implementations of the present disclosure;
[0024] Figure 14 A block diagram showing an apparatus for managing the state of a robot according to some implementations of the present disclosure; and
[0025] Figure 15 A block diagram showing a device capable of implementing multiple implementations of the present disclosure. Detailed Implementations
[0026] Implementations of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some implementations of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. On the contrary, these implementations are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0027] In the description of the implementations of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". There may also be other explicit and implicit definitions below. As used herein, the term "model" may represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.
[0028] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0029] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.
[0030] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0031] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0032] It is understandable that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0033] As used herein, the term "in response to" means the state in which a corresponding event occurs or a condition is satisfied. It will be understood that the execution timing of subsequent actions performed in response to the event or condition and the time when the event occurs or the condition is established are not necessarily strongly correlated. For example, in some cases, the subsequent action can be immediately executed when the event occurs or the condition is established; while in other cases, the subsequent action can be executed after a period of time after the event occurs or the condition is established.
[0034] Example environment
[0035] In recent years, robotics technology has developed rapidly and has been widely used in multiple technical fields. The interaction state of a robot can be predicted through visual signals and a video format output can be provided, whereby a generative algorithm can be used to predict the future state changes of the robot. However, there are currently only a small amount of image and / or video data of the robot, which results in a serious shortage of training data. Further, the cost of collecting visual data exclusive to the robot is relatively high, and a real robot platform and multi-view sensor devices need to be deployed. In addition, the scene where the robot is located is closely coupled with the tasks performed by the robot, and it is difficult to obtain training data for performing multiple different tasks.
[0036] See Figure 1Describe the application environment according to some implementations of the present disclosure, the Figure 1 shows a block diagram 100 of the application environment of a robot according to one implementation of the present disclosure. As Figure 1 shown, an image 110 describing the current state of the robot can be provided. Further, an action 120 for controlling the robot can be obtained. Here, the action 120 represents one or more actions for controlling the robot so as to control the robot to execute the above one or more actions within a future time period. At this time, the pre-trained machine learning model 130 can output a video 140, and the video 140 can include one or more frames so as to represent the state of the robot within a future time period. Here, there can be a one-to-one relationship between the action and the video frame. Alternatively and / or additionally, based on interpolation techniques and / or other techniques, one or more intermediate actions between two actions can be determined, thereby generating more video frames. At this time, there can be a many-to-many relationship, a one-to-many relationship, etc. between the action and the video frame.
[0037] During the process of training the machine learning model 130, a large number of reference samples (that is, training samples) need to be collected. However, existing open-source data sets are mainly based on simulation environments, and there are domain differences between simulation data and real data. Further, robots involve different structures. For example, robot A can include 3 robotic arms and 3 fingers, and robot B can include 2 robotic arms and 5 fingers, etc. This results in different degrees of freedom of actions from robots with different hardware structures, making it difficult to utilize them uniformly. Further, the dimensions of the action spaces of different robot structures (number of degrees of freedom and / or type of end effector) do not match, and there is a lack of implicit encoding for representing the dynamic parameters (joint torque limits, kinematic constraints) of different robots in a unified manner.
[0038] Currently, frameworks for generating robot videos have been proposed. These frameworks use action sequences as conditional inputs to the video generation model to generate video frames depicting the results of the robot executing actions. However, the existing frameworks lack a unified but precise action representation and cannot effectively model high-degree-of-freedom and heterogeneous actions across different domains and applications, thus hindering the effective training of a unified cross-domain model.
[0039] Although a large amount of video data of human interactions with daily objects and the environment has been collected currently, due to the huge differences between the human hand and the robot structure, this part of human data cannot be directly utilized. The actions of robots have a high degree of freedom and high complexity, but current generation algorithms often only consider the prediction of low degrees of freedom of the execution subject and cannot generate control signals with a high degree of freedom for controlling robots with complex structures.
[0040] To accurately represent the action 120, various encoding methods have been proposed. From the perspective of action encoding, the following encoding methods exist: (1) Direct action encoding, that is, directly inputting the action of the end effector into the machine learning model. However, direct action encoding can only represent the translation and / or rotation of a single joint (such as the end effector), but it is difficult to represent the control of a multi-joint robotic arm, and it is difficult for the machine learning model to learn action-related knowledge. (2) Simple low-dimensional action encoding, that is, instead of considering complex actions, representing the actions of the robot as simple movement instructions such as up, down, left, right, etc., or inputting the above movement instructions into a small neural network to elevate the original low-dimensional instructions into high-dimensional vectors. However, simple low-dimensional action encoding can only represent very simple actions and is not practical for actual robot operations. (3) Using a two-dimensional mask as action encoding. This method can segment the observed subject when the robotic arm or human hand makes a certain action, and use the foreground mask as action encoding. Although the segmented visual encoding can represent relatively complex actions, this encoding cannot handle the problem of occlusion. In other words, when self-occlusion and / or occlusion by an object occur in various parts of the executing subject (human or robotic arm), the accuracy of the encoding will be severely reduced.
[0041] From the perspective of pre-training, the following encoding methods exist: 1) Without pre-training, only performing training through the true value actual data of the robot. Any of the above action encoding methods can be selected. At this time, only one type of robot structure is used, but cross-structure capabilities are not supported. However, the above technical solution requires a large amount of robot data of different structures to train the machine learning model in order to obtain the desired generalization. 2) Pre-training with robot data of multiple structures. The action spaces of multiple robots can be combined into a high-dimensional vector through a blank filling method. Specifically, assuming there are N robot structures, the high-dimensional vector can include N dimensions, and each dimension corresponds to the action encoding of a robot of one structure. For structure 1, only its corresponding dimension has valid values, while the corresponding dimensions of the other N - 1 structures are invalid values (for example, empty, filled with 0, or other predetermined symbols). Although the blank filling method using multiple robots can combine data of multiple structures, it cannot handle structures that have not been seen before, and when the number of structures increases, the high-dimensional vector will become larger and larger and increasingly difficult to expand.
[0042] At this time, it is expected to combine robotics and machine learning technologies to represent actions in a more accurate way, and then predict the operations of the robot in a more accurate and effective way.
[0043] Overview of robot state prediction
[0044] To at least partially address the deficiencies in the prior art, according to one implementation of the present disclosure, a method for managing the state of a robot is proposed. Generally speaking, the present disclosure proposes a new action representation format that can support the use of data of a large number of humans and data of robots with different hardware structures, thereby obtaining a machine learning model with higher performance. Further, a prediction model implemented using a machine learning model is proposed. The prediction model includes two inputs: an image of the observed previous frame and an action control signal for the subsequent N frames; the output is a video of N frames at a future time.
[0045] See Figure 2 Describe the overview according to one implementation of the present disclosure, Figure 2 FIG. 200 is a block diagram showing a method for managing the state of a robot according to some implementations of the present disclosure. As Figure 2 shown, a method for managing the state of a robot is proposed. Specifically, a first image 210 describing the first state (i.e., the current state) of the robot 212 at a first time point (i.e., the current time point) can be received. An action image 220 can be received. The action image 220 can specify an action to be performed by the robot 212 at a second time point (e.g., one or more future time points) after the first time point, and the action is represented by the skeleton image 222 of the robot 212. Further, the first image 210 and the action image 220 can be input to the machine learning model 230. At this time, the machine learning model 230 can determine a second image 240 of the second state of the robot 212 at the second time point based on the first image 210 and the action image 220. It should be understood that the second time point here can include one or more future time points. When there are multiple future time points, the action image can include multiple action images (i.e., a sequence of action images). At this time, the machine learning model 230 can output a video of the robot 212 performing multiple actions.
[0046] Using the implementation of the present disclosure, the skeleton image can be used as an accurate action encoding, and the actions of robots with different structures can be represented in a unified manner in the scenario of high-degree-of-freedom actions. In this way, the action encoding of the skeleton image can be used as a control condition for the machine learning model, and thus a more accurate prediction model can be provided for the interaction of robots with high degrees of freedom.
[0047] It should be understood that the actions herein can have multiple sources. For example, a pre-trained policy model can be utilized to provide an action image or a sequence of action images, and the policy model can be invoked in real time to respectively obtain action images corresponding to different time points. Alternatively and / or additionally, an action sequence can be provided based on rules. In the context of the present disclosure, the specific manner of obtaining the action sequence is not limited, but the action sequence can be specified based on the policy model, action rules, and / or in other ways. Herein, an action image can represent one or more actions to be performed at future time points. The prediction model can receive different action images and generate different outputs.
[0048] According to some implementations of the present disclosure, the proposed prediction model can be directly applied. Direct application refers to the tasks that can be achieved by the prediction model trained using the method of the present disclosure. The most important application is data augmentation. Specifically, during the data augmentation process, the source of the actions can be either a trained policy model or manually set policy rules. Alternatively and / or additionally, the proposed prediction model can be used to perform downstream tasks. Downstream tasks refer to the tasks that need to combine the prediction model and the policy model to be completed. For example, it can be evaluated which policy model is better among multiple policy models, a certain policy model can be used to perform inference on the prediction model to achieve online planning in a simulation environment, data can be continuously generated and the generated data can be used to train the policy model, and so on.
[0049] According to some implementations of the present disclosure, an exact visual action encoding can be used as a unified control signal for an interactive generation model driven by complex actions while maintaining cross-domain generality. The visual action encoding is generated by "rendering" the 3D structure of the agent state caused by the action into the image space, which can have different forms, such as a rough mask, a color rendering, a depth map, and a 2D skeleton. They can effectively represent the actions of high-degree-of-freedom entities (such as human hands, robot grippers, and dexterous hands) with high precision.
[0050] Using some implementations of the present disclosure, visual action encoding can serve as a more accurate and easier-to-learn control signal for actions in a video model. Additionally, visual action encoding can train a unified model across multi-domain data (including human-object interaction (HOI) and robot model usage), which promotes interaction-driven dynamic cross-domain knowledge transfer. Specifically, in scenarios involving complex high-degree-of-freedom actions, it is proposed that using precise visual action encoding (especially skeletons) can serve as a unified action representation for an action-driven generative model. An extensible strategy is proposed to recover skeleton-based visual action encoding for rich interaction datasets including HOI and robot manipulation. Using some implementations of the present disclosure, visual action encoding is easier to learn, has high precision and generality, thus supporting joint training based on heterogeneous data and promoting knowledge conversion.
[0051] Detailed process of robot state prediction
[0052] A summary of some implementations according to the present disclosure has been described. Below, more details regarding managing a robot will be provided. As Figure 2 shown, the skeleton image 222 is a type of action encoding represented by the skeleton of an action execution subject (e.g., including a human or a robot). In other words, the positions of the bones and joints of the skeleton under the current action can be projected onto the two-dimensional space of the image as a form of action encoding. In the context of the present disclosure, this action encoding method can be referred to as skeleton action encoding.
[0053] According to some implementations of the present disclosure, when representing a human action with a skeleton image, the skeleton image can include at least one line segment corresponding to at least one bone of the human. For example, for a human hand, the skeleton image can include line segments corresponding to the bones of the arm (e.g., the upper arm, the lower arm), the palm, and each finger. Each line segment can be connected by key points to represent the joints in the human skeleton. Again, for a robot, the skeleton image can include at least one line segment corresponding to at least one robotic arm. Here, the robotic arm can include the bones in one or more arms of the robot, as well as the end effector of the robot (e.g., a robotic hand, a fixture, a tool, etc.). Each line segment can be connected by key points to represent the joints in the robot skeleton.
[0054] Using some implementations of the present disclosure, the actions of a human arm and multiple robots with various different structures can be represented in a unified manner, that is, the complex actions of the arm are normalized into a two-dimensional image space represented by a skeleton. In this way, data from different execution subjects can be used to train a machine learning model, thereby improving the accuracy of the machine learning model.
[0055] See Figure 3ADescribe the basic architecture of a machine learning model, which Figure 3A shows a block diagram 300A of the architecture of a machine learning model according to some implementations of the present disclosure. As Figure 3A shown, the machine learning model 230 may include: a skeleton encoder 334 for converting an action image 220 into action features in an implicit space; a video encoder 330 for converting a first image 210 into image features in an implicit space; a diffusion model 350 for generating features of a second image based on the action features and the image features; and a video decoder 332 for converting the features of the second image into a second image 240.
[0056] According to some implementations of the present disclosure, in the process of determining the second image, a machine learning model may be used to determine second image features corresponding to the second image based on the first image and the action image; and the video decoder in the machine learning model may be used to determine the second image based on the second image features. Specifically, the skeleton encoder 334, the video encoder 330, and the diffusion model 330 in the machine learning model 230 may be used to determine the second image features. Subsequently, the video decoder 332 may be used to determine the second image 240. Using some implementations of the present disclosure, since the skeleton encoder 334 can obtain action features represented in a unified format regarding the action, the action features may be used as a constraint condition for the diffusion model 350 to determine future images at subsequent time points from the first image 210.
[0057] It should be understood that the skeleton encoder 334 here may process one or more action images (i.e., process one action image or an action video including multiple action images), and the video encoder 330 and the video decoder 332 may process a video including one or more image frames. Figure 4 shows a block diagram 400 of a video encoding process and a decoding process according to some implementations of the present disclosure. As Figure 4 shown, the video encoder 330 may process a video 410 (including one or more image frames) to generate video features 420, and the video decoder 332 may decode the video features 420 and output a reconstructed video 430. The video encoder 330 and the video decoder 332 may be trained using a reference video. For example, the video encoder 330 and the video decoder 332 may be updated in a direction that minimizes the difference between the video 410 and the reconstructed video 430.
[0058] According to some implementations of the present disclosure, the machine learning model 230 may have different structures, return Figure 3B Describe more details. Figure 3B shows a block diagram 300B of the architecture of a machine learning model according to some implementations of the present disclosure. As Figure 3BAs shown, the machine learning model 230 may include multiple skeleton encoders (e.g., the first skeleton encoder 310 and the second skeleton encoder 320). Here, the first skeleton encoder 310 may determine a first action feature based on the action image 220, and the first action feature may be input to the control network 340 to output the conditional feature of the diffusion model 350. Alternatively and / or additionally, the diffusion model 350 may include Low-Rank Adaptation (LoRA) 352. The main part of the diffusion model 350 may be frozen, and only the parameters of LoRA 352 may be fine-tuned during the training process.
[0059] According to some implementations of the present disclosure, the second image feature may be determined based on Figure 3B the machine learning model 230 shown, and then the second image 240 may be determined. Specifically, during the process of determining the second image feature, the initial feature of the diffusion model in the machine learning model may be determined based on the first image 210; the first action feature corresponding to the action image may be determined using the first skeleton encoder in the machine learning model; and the second image feature may be determined using the diffusion model based on the initial feature and the first action feature.
[0060] Referring to Figure 3B , the video encoder 330 may be used to determine the initial feature 360 of the diffusion model in the machine learning model based on the first image 210. The first action feature 362 corresponding to the action image 220 may be determined using the first skeleton encoder 310 in the machine learning model. The second image feature may be determined using the diffusion model 350 based on the initial feature 360 and the first action feature 362.
[0061] Alternatively and / or additionally, the machine learning model 230 may further include a control network 340. The first action feature 362 from the first skeleton encoder 310 may be directly input to the diffusion model 350, or the first action feature 362 may be input to the control network 340 to output the feature 366 used as the constraint condition of the diffusion model 350. Here, the control network 340 may include multiple network layers, and each network layer may output the constraint conditions for different denoising stages of the diffusion model 350.
[0062] According to some implementations of the present disclosure, the conditional feature of the diffusion model may be determined based on the first action feature using the control network model in the machine learning model; and the second image feature may be determined using the diffusion model based on the initial feature and the conditional feature. Continuing to refer to Figure 3B, the control network 340 can determine the conditional feature 366 of the diffusion model based on the first action feature. Further, the diffusion model 350 can determine the second image feature based on the initial feature 360 and the conditional feature 366. With some implementations of the present disclosure, the diffusion model 350 can be controlled in a more accurate manner, so that the image generated by the machine learning model 230 is more matched to the action image 220.
[0063] With some implementations of the present disclosure, the video encoder in the machine learning model can be used to determine the first image feature of the first image. The first image feature from the video encoder 330 can be directly used as the initial feature; alternatively and / or additionally, noise data can be added to the first image feature to determine the initial feature. With some implementations of the present disclosure, a perturbation factor can be added to the initial feature of the diffusion model, so as to use the denoising process of the diffusion model to generate the second image 240 that matches the first image 210 and the action image 220.
[0064] With some implementations of the present disclosure, the second skeleton encoder 320 can provide information for fine-tuning LoRA 352. At this time, in the process of determining the initial feature, the second skeleton encoder 320 in the machine learning model can be used to determine the second action feature corresponding to the action image; and the initial feature can be updated using the second action feature. At this time, the initial feature of the diffusion model can include knowledge about the action image, so that the initial feature as the starting point of the reverse denoising process can consider not only the relevant information of the first image, but also the relevant information of the action image, so that the generated second image feature is more matched to the action specified in the action image 220.
[0065] With some implementations of the present disclosure, the machine learning model 230 can have a more complex architecture, see Figure 5 for more details. The Figure 5 shows a block diagram 500 of the architecture of a machine learning model according to some implementations of the present disclosure. As Figure 5 shown, the control network can include multiple network layers, and the diffusion model 350 can include multiple network layers. The legend 530 represents the part that can be updated during the training process, and the legend 532 represents the frozen part.
[0066] As Figure 5As shown, the first skeleton encoder 310 can generate features 520 based on the action image 220, and can directly input the features 520 to the control network 340. Alternatively and / or additionally, the features 520 can be updated using the features 522, and the updated features are input to the control network 340. In other words, in the process of determining the first action features, the first action features can be further updated based on the initial features of the diffusion model. Using some implementations of the present disclosure, by considering the relevant information of the first image 210 in the constraints, the second image features generated by the diffusion model can be made to better match the first image 210.
[0067] According to some implementations of the present disclosure, the features 522 are determined based on the first image 210 and the noise 510. Specifically, the video encoder 330 can be used to determine the first image features of the first image 210, and the noise 510 can be added to the first image features to generate the features 522. Further, the features 520 can be updated using the features 522. According to some implementations of the present disclosure, the symbol “+” can represent a summation operation, and the symbol “C” can represent the concatenation of features after alignment by dimension. According to some implementations of the present disclosure, the second skeleton encoder 320 can be used to determine the features 524 of the action image 220. The features 526 can be generated based on the features 522 and 524, and then the features 526 are input to the diffusion network 350 as the initial features.
[0068] According to some implementations of the present disclosure, a video variational autoencoder (VAE) can be provided to implement the various encoder models described above. The prediction model can be implemented based on the diffusion model. At this time, the skeleton action encoding can be used as the control condition and the currently captured image can be used as the initial image of the diffusion model. Alternatively and / or additionally, controllable generation techniques (e.g., control networks) and LoRA fine-tuning techniques can be used. In the context of the present disclosure, there is no strong binding relationship between the training method and the model structure, but rather the prediction model can be implemented based on different model architectures, and the prediction model can be trained using the methods of the present disclosure.
[0069] First, visual action encoding is described. The input data of the machine learning model of the present disclosure includes: the initial observation s0 and the specified complex action sequence a 0:t-1 , and the output data includes: the output video s representing the reasonable scene dynamics and interaction results 1:t ∈S. The problem can be described as follows:
[0070] s 1:t ~P(s 1:t |s0,a 0:t-1 ) Formula 1 In the above formula, s0 represents the initial observation, that is, the image that is the start frame of the output video; a0:t-1 represents an action sequence; s 1:t represents the output video; and P represents the conditional probability distribution. The above formula needs to meet the following requirements: (1) adapt to various configurations of complex action spaces while ensuring cross-task compatibility; and (2) maintain the convertibility of visual dynamics under precise action control, so that the model can learn knowledge from an expandable dataset spanning various domains. To address these challenges, the core problem is to map the action sequence a 0:t-1 to the visual action encoding v 1:t :
[0071]
[0072] In the above formula, T represents the length of the action trajectory, H and W represent the image height and width respectively, C represents the number of channels determined by a specific visual representation, and R indicates the rendering operation according to known camera parameters. Grid-based rendering (e.g., color image, depth map) and 2D skeletons can be considered as visual action encodings. Considering the related challenges of fine-grained grids for data recovery from natural scenes, 2D skeletons can be selected as an example of our visual encoding.
[0073] According to some implementations of the present disclosure, the training and / or inference processes can be performed for the machine learning model described above. It should be understood that the data flow in the training process and the inference process is similar, so it will not be elaborated. According to some implementations of the present disclosure, reference samples (also referred to as training samples) can be collected to train the machine learning model. Specifically, the machine learning model is determined based on: obtaining reference samples, where the reference samples include: a first reference image of the first reference state of the subject performing the reference action at the first reference time point, a reference action image, and a second reference image, the reference action image specifying the reference action performed by the subject at the second reference time point after the first reference time point, the reference action being represented by the reference skeleton image of the subject, and the second reference image representing the second reference state of the subject at the second reference time point.
[0074] According to some implementations of the present disclosure, given an image observation as the initial frame, by using a series of precise action encodings from a human hand or a robot gripper, the generation of a video depicting the interaction result can be precisely controlled. According to some implementations of the present disclosure, there can be two execution agents: a human hand and a robot gripper. Although there are kinematic differences between the two agents, skeletons can be used to uniformly encode the actions of the two agents as the visual encoding of the model.
[0075] To train the model, two types of datasets can be processed and annotated: (1) human hand skeletons extracted from HOI videos in natural scenes via motion capture; and (2) robot gripper skeletons synthesized in robot operation scenes through joint state rendering. Using a large amount of sample data (including, a pair of skeleton - video), a pre - trained video generation model can be finely tuned to achieve visual action encoding control.
[0076] Specifically, reference samples can be determined based on various methods. For example, reference samples can be determined using human action images or robot action images. It should be understood that since there is a large amount of human data in the public dataset, in this way, the number of reference samples can be greatly increased, thus supporting the machine learning model to obtain relevant knowledge about more actions.
[0077] According to some implementations of the present disclosure, human data can be used to determine reference samples. For the human hand, first, the hand region is obtained through a hand detection algorithm, then hand pose detection is performed in the hand region where there is a hand to obtain the key points of each joint of the hand. Finally, through known camera intrinsics or estimated camera intrinsics, the positions of the joints are projected onto the camera view and drawn into a skeleton diagram.
[0078] See Figure 6 for more details, the Figure 6 shows a block diagram 600 for determining reference samples based on human hand data according to some implementations of the present disclosure. As Figure 6 shown, video data 610 including a human hand can be collected, and a hand detection algorithm 622 can be used to detect the hand region 620 in each video frame of the video data 610. Further, a hand key - point estimation algorithm 632 can be used to determine a skeleton 630 including the human hand skeleton. It should be understood that the skeleton 630 is merely illustrative and includes multiple skeletons corresponding to different video frames respectively. Skeleton images corresponding to each video frame can be generated based on the skeleton 630. For example, one skeleton image can include the skeleton 632, one skeleton image can include the skeleton 634, and another skeleton image can include the skeleton 636, and so on.
[0079] In HOI videos in natural scenes, rich HOI scene videos can be obtained, which are ideal for training visual dynamic models. HOI videos may suffer from severe self-occlusions and / or occlusions caused by interactions, resulting in unreliable and / or incomplete estimation of 2D hand poses. Alternatively and / or additionally, a 3D hand mesh recovery method can be employed, which can match 2D hand bounding boxes and handedness frame by frame before reconstructing the mesh from local crops. However, these methods often fail due to missed detections, false positives, incorrect handedness, and temporal jitter.
[0080] To address this problem, a multi-stage pipeline for hand mesh trajectory extraction is proposed. See Figure 7 for more details on the algorithm for tracking and associating hands, which Figure 7 shows a block diagram 700 of an algorithm 710 for determining hands according to some implementations of the present disclosure. As Figure 7 shown, initialization operations can be performed in lines 1 to 3. In lines 5 to 8, all potential hands in each frame can be detected and sorted. In lines 9 to 15, temporally consistent hand trajectory segments can be constructed to address detection errors and handedness inconsistencies. In lines 16 to 20, mesh refinement can be performed, i.e., re-estimating missing meshes in the trajectory segments. In lines 21 to 23, trajectory smoothing can be performed, i.e., applying a filter algorithm to eliminate temporal jitter in the hand trajectories.
[0081] According to some implementations of the present disclosure, robot data can be utilized to determine reference images. Specifically, for a robot with a given hardware structure, Unified Robot Description Format (URDF) data can be obtained, the position of each joint can be obtained through the URDF data and the current action, and finally the position of the joint can be projected onto the camera view through the internal and external parameters of the camera, thereby obtaining a skeleton graph. Alternatively and / or additionally, text descriptions of each video segment can be obtained by annotation to increase the information content of the image data.
[0082] See Figure 8 for more details, which Figure 8 shows a block diagram 800 for determining reference samples based on robot data according to some implementations of the present disclosure. As Figure 8As shown, video data 810 including a robotic arm can be collected, and the URDF 812 data of the robot can be used to determine the joint positions 820 in three-dimensional space. Further, the projection 830 of the joint positions (i.e., projecting the robotic arm from three-dimensional space to two-dimensional space) can be determined, and then the skeleton 840 of the robotic arm can be determined. It should be understood that the skeleton 830 is merely illustrative and includes multiple skeletons corresponding to different video frames respectively. Skeleton images corresponding to each video frame can be generated based on the skeleton 830. For example, one skeleton image can include the skeleton 832, one skeleton image can include the skeleton 834, and another skeleton image can include the skeleton 836, and so on.
[0083] Regarding robotic manipulation scenarios, in addition to HOI videos, datasets related to robotic manipulation scenarios have emerged, which provide data for learning interactions and scene dynamics. The robot state log supports direct 3D skeleton trajectory construction in egocentric coordinates, thus simplifying 2D skeleton projection at known camera poses. Even in the absence of state data, the robot skeleton can be recovered via key-point estimation similar to human models, feature matching, or distinguishable rendering algorithms. To facilitate efficient data processing, a dataset with pre-calibrated camera poses can be utilized, which can include different robotic manipulators and environmental settings. For each robot configuration, specific kinematic key points and their connections can be defined, and 2D skeletons can be generated through trajectory playback in the simulator. To address substantial variations in camera calibration quality across robotic scenarios and mitigate camera parameter drift during operation, a vision-based pipeline is provided to offer filtering and correction functions:
[0084] (1) Scene filtering: Matching can be performed between the robot mesh rendering and the real observation, and scenes with significant coordinate differences in the match can be filtered out. (2) Holographic correction: For scenes with camera drift, based on image matching and applying per-frame homography warping, the initial 2D skeleton rendering can be adjusted to ensure precise alignment with real-world observations.
[0085] According to some implementations of the present disclosure, a visual dynamics model with precise control is proposed. Specifically, the machine learning model of the present disclosure can be constructed based on existing video generation models. For example, the existing video generation model can be a text-to-video generation model pre-trained on large-scale (text, video) pairs, and further fine-tuned into a (text, image)->video model using (text, initial frame, video) triples. The model architecture mainly includes: a pre-trained text encoder, a video VAE, and a diffusion model with full attention for spatio-temporal video tagging and text tagging processing. The above existing video generation model can be used as a pre-trained base model to implement a visualization dynamic model for processing skeleton data.
[0086] To integrate visual action encoding, the control signal can be encoded. Specifically, to achieve control in the form of a skeleton and / or a mesh, the skeleton and / or the mesh can be rendered as an RGB image sequence where C = 3. Then these sequences are fed into a 3D convolutional trajectory encoder to obtain the latent state For depth control, v 1:t (where C = 3) can be directly fed into an encoder with the same architecture.
[0087] In a large-scale dataset, direct supervised fine-tuning of a pre-trained video generation model may lead to overfitting or loss of generalized knowledge. Therefore, a ControlNet can be used to inject visual action encoding. In particular, multiple (e.g., 14) trainable copies of the front part of a pre-trained diffusion model with zero-initialized linear layers can be created, and visual action encoding s is injected into these blocks 1:t / 4 . In addition, s 1:t / 4 is injected into the main diffusion model to implement a two-branch conditional mechanism by merging video and action tags and fine-tuning the diffusion model backbone using LoRA.
[0088] During training, the loss value around the hand / gripper area can be amplified to preferentially learn interactions and the dynamics caused by them. To mitigate the dominance of self-motion in the robot video performing a long task over the interaction dynamics, more data can be collected around the time stamps when the gripper state changes.
[0089] According to some implementations of the present disclosure, a machine learning model can be used to determine the predicted features of a second reference image based on a first reference image and a reference action image; and update the machine learning model based on the features of the second reference image and the predicted features of the second reference image. Using some implementations of the present disclosure, the machine learning model can be updated in a direction that minimizes the difference, thereby improving the accuracy of the machine learning model. See Figure 9Description for more information on determining predictions, the Figure 9 shows a block diagram 900 of a prediction process according to some implementations of the present disclosure. As Figure 9 shown, a first reference image 920 corresponding to a human hand and / or a robot, noise 922, and a reference action image 924 can be input to a machine learning model 230. The diffusion model 350 in the machine learning model 230 can be utilized, and a denoising process 910 can be used to determine video features 930. Further, the machine learning model 230 can be updated in a direction that minimizes the difference between the video features 930 and the features of a second reference image.
[0090] In the context of the present disclosure, different parts in the machine learning model can be updated in multiple stages. In the pre-training stage, only human hand data can be used. Through the above data processing method, a skeleton map of the human hand can be obtained, and training can be performed with the annotated video data. Only the control network and the LoRA part in the prediction model in the model structure diagram and the independent encoders before these two branches are trained. Other parts can be frozen, that is, the weights determined using ordinary training data are loaded. The video variational autoencoder also remains frozen and only predetermined weights are loaded. In the joint training stage, human data and robot data of various different hardware structures can be used to obtain skeleton action encodings in the data preprocessing manner described above. At this time, similar to the pre-training stage, some network layers in the prediction model can be trained, and other parts can be frozen.
[0091] Specifically, in the first stage (e.g., the pre-training stage), the machine learning model can be initially trained using data from humans. Specifically, the machine learning model is updated in the first stage, the subject includes a human subject, and updating the machine learning model includes updating at least any one of the following: a control network model, a fine-tuning model in the diffusion model, a first skeleton encoder, and a second skeleton encoder. See Figure 10 Description for more details, the Figure 10 shows a block diagram 1000 of a process for training a machine learning model according to some implementations of the present disclosure.
[0092] As Figure 10As shown, in the pre-training stage, hand data 1012 (e.g., including the first reference image, the second reference image, and the reference action image described above) can be obtained. The skeleton data preprocessing module 1020 can be used to generate corresponding hand skeleton features 1030, and then pre-training 1040 can be performed. Alternatively and / or additionally, additional data of the hand data 1012 (e.g., video-text pair data 1010, etc.) can be further obtained, and the additional data can be used to perform pre-training 1040. By using some implementations of the present disclosure, hand data that is easier to collect can be pre-obtained to perform the pre-training process. Since a robot arm usually simulates the actions of a human arm, in this way, the machine learning model can master the relevant knowledge of hand actions, thereby facilitating the generalization of this knowledge to the robot field.
[0093] According to some implementations of the present disclosure, in the second stage (e.g., the joint training stage), data from the robot can be used to jointly train the machine learning model. Specifically, the machine learning model is updated in the second stage after the first stage, and the subjects include at least any one of a human subject and a robot, and updating the machine learning model includes updating at least any one of the following: the control network model, the fine-tuning model in the diffusion model, the first skeleton encoder, and the second skeleton encoder. Continuing to refer to Figure 10 , in the joint training stage, training data from different robots can be obtained. For example, robot data 1014, 1016, and 1018 from robots with different structures can be obtained respectively. Here, the robot data can include the first reference image, the second reference image, and the reference action image. The above robot data can be processed by the skeleton data preprocessing module 1020, and then corresponding robot skeleton features 1032, 1034, and 1036 can be generated. Further, the first reference image, the features of the second reference image, and the robot skeleton features can be used to perform joint training 1060.
[0094] In this process, only the control network 340 in the machine learning model, LoRA 352 in the diffusion model, and the first skeleton encoder 310 and the second skeleton encoder 320 can be updated. By using some implementations of the present disclosure, more accurate ground-truth robot data can be used to train the machine learning model to improve the accuracy of the machine learning model in determining the robot state.
[0095] Alternatively and / or additionally, hand data can be used to perform joint training 1060. For example, the proportion of hand data in the entire training data set can be specified in advance (e.g., 10% or other proportions). In this way, the training data set in the joint sequence can include hand data and robot data, thereby further improving the accuracy of the machine learning model.
[0096] The training process of a machine learning model has been described. Further, the trained machine learning model can be used to predict the state of a robot. In the inference stage, the current image frame at the current moment and the action sequence corresponding to the next N frames can be input into the trained machine learning model, and the machine learning model can predict the robot interaction effect of the next N frames by gradually denoising through a diffusion model.
[0097] See Figure 11 For more details, the Figure 11 shows a block diagram 1100 of a process for generating a video of a robot according to some implementations of the present disclosure. As Figure 11 shown, a first image 1110 describing the first state of the robot at a first time point can be received. An action image 1130 can be received, and the action image 1130 specifies an action to be performed by the robot at a second time point after the first time point, and the action is represented by the skeleton image of the robot. Further, the machine learning model 230 can be used to determine a second image 1120 of the second state of the robot at the second time point based on the first image 1110 and the action image 1130. At this time, since the machine learning model 230 already includes rich knowledge about human hand actions and robot arm actions, in this way, the second image 1120 can be determined in a more accurate manner.
[0098] It should be understood that although Figure 11 only schematically shows that the action image includes the action image at a future second time point, alternatively and / or additionally, the action image can include multiple action images, and the multiple action images specify multiple actions to be performed by the robot at multiple second time points after the first time point. In this way, by specifying multiple actions within a future time period (including multiple second time points), a video sequence of the robot within that future time period can be generated. Using some implementations of the present disclosure, the state of the robot within a future time period can be predicted.
[0099] According to some implementations of the present disclosure, the first image can include multiple first images, for example, including the current robot video within a current time period. The robot video and the action image (or action video) can be input into the machine learning model to obtain the image (or video) of the relevant robot at a future time point (or time period). In this way, the action trend of the robot can be determined from the current robot video, thereby improving the accuracy of prediction.
[0100] According to some implementations of the present disclosure, the machine learning model can be directly applied. Alternatively and / or additionally, more downstream tasks can be performed based on the machine learning model. Figure 12 shows a block diagram 1200 of a process for using a machine learning model according to some implementations of the present disclosure. AsFigure 12 As shown, the current observation 120 (e.g., an image or video) and the action image 1220 of the robot can be used as inputs, and the machine learning model 230 can be used to predict changes in the robot interaction and scene dynamics at the next moment (or time period) (e.g., output video 1230). Using some implementations of the present disclosure, in the direct application 1240 mode, video results of different evolutions 1242 can be generated for subsequent training of the robot policy model. This process can amplify a large amount of data. For example, given the current input video frame, by specifying multiple different actions, multiple different prediction results of robot interactions can be generated.
[0101] According to some implementations of the present disclosure, in the indirect application 1250 mode, multiple downstream tasks can be performed. For example, in the policy evaluation 1252 scenario, the action sequences 1 and 2 can be obtained from the existing robot policy models 1 and 2 respectively, and the proposed machine learning model can be used to determine the scene changes corresponding to the action sequences 1 and 2 respectively, and different policy models can be evaluated without relying on real robot devices. In the policy learning 1254 scenario, the machine learning model can output different future scene changes according to different action sequences, and in this way, the world model can be repeatedly called to train the policy model. In the online planning 1256 scenario, actions can be output through any trained robot policy model, and the proposed machine learning model can be used to output future scene changes, and online planning of the robot policy can be realized.
[0102] The proposed machine learning model can be tested in multiple datasets. Specifically, the experiments can prove two core issues: (1) Visual action encoding is superior to existing control signals, such as text control signals or agent-centered raw action / state signals; (2) Visual action encoding can improve the generality of cross-agent configuration and joint training on different datasets. In addition, ablation studies prove the effectiveness of the proposed model design and present the results of different visual action encodings.
[0103] In summary, a visual action encoding is proposed to be used as a general action representation for action-to-video generation, which effectively represents complex high-degree-of-freedom actions and at the same time retains the cross-domain dynamic transformation ability of the video scene model. A robust pipeline for constructing visual action encoding from heterogeneous data sources for training is proposed, and lightweight fine-tuning is used to inject visual action encoding into a pre-trained video generation model. Experimental data proves that the proposed technical solution has made improvements in both interaction fidelity and domain adaptability, and the proposed machine learning model has higher performance.
[0104] Different from existing technical solutions that only use robot data as training data, in the context of the present disclosure, human data and robot data can be used, thereby improving the generalization of the prediction model. Using some implementations of the present disclosure, by converting the motion data of a human hand and the motion data of a robot into the form of a unified skeleton image, the two-dimensional skeleton image can be used as a motion encoding method. A general and unified motion encoding that is independent of the structures of humans and robots and independent of position control and force control during robot data acquisition is achieved. Further, joint training of human hand interaction data and robot interaction data is achieved.
[0105] According to some implementations of the present disclosure, it is proposed to use a skeleton as an accurate motion encoding to serve as a unified motion representation of a prediction model in a high-degree-of-freedom motion scenario. The proposed prediction model has better generality, and the motions of various execution entities such as a simple single-joint robotic arm, a multi-joint robotic arm, a dual-arm robot, a robotic dexterous hand, and a human hand can all be represented by a unified skeleton. Using some implementations of the present disclosure, it is easier to obtain the training data of the prediction model. Specifically, the current joint point positions can be calculated in real time through the model file of the robot structure, and then the original motion of the robot can be obtained. Human hand data can be obtained through algorithms such as motion capture and human hand key point estimation, and so on. In this way, a large amount of publicly available interactive videos and skeleton motion encodings can be used to greatly expand the training data available for training the prediction model.
[0106] Using the skeleton as a unified motion encoding provides the advantage of multi-source data training compared to single-data training. The skeleton motion encoding has better generality than existing methods and is easier to obtain and learn. Further, using the skeleton motion encoding as a control condition allows the prediction model to generate corresponding scene changes according to different motions and through video generation, thereby providing higher realism and stronger motion rationality.
[0107] Example process
[0108] Figure 13 A flowchart of a method 700 for managing the state of a robot according to some implementations of the present disclosure is shown. At block 1310, a first image describing a first state of the robot at a first time point is received; at block 1320, an action image is received, the action image specifying an action to be performed by the robot at a second time point after the first time point, the action being represented by a skeleton image of the robot; and at block 1330, a machine learning model is used to determine, based on the first image and the action image, a second image of a second state of the robot at the second time point.
[0109] According to some implementations of the present disclosure, determining the second image includes: using a machine learning model to determine second image features corresponding to the second image based on the first image and the action image; and using a video decoder in the machine learning model to determine the second image based on the second image features.
[0110] According to some implementations of the present disclosure, determining the second image features includes: determining initial features of a diffusion model in the machine learning model based on the first image; using a first skeleton encoder in the machine learning model to determine first action features corresponding to the action image; and using the diffusion model to determine the second image features based on the initial features and the first action features.
[0111] According to some implementations of the present disclosure, determining the second image features based on the initial features and the first action features includes: using a control network model in the machine learning model to determine conditional features of the diffusion model based on the first action features; and using the diffusion model to determine the second image features based on the initial features and the conditional features.
[0112] According to some implementations of the present disclosure, determining the first action features further includes: updating the first action features based on the initial features of the diffusion model.
[0113] According to some implementations of the present disclosure, determining the initial features includes: using a video encoder in the machine learning model to determine first image features of the first image; and adding noise information to the first image features to determine the initial features.
[0114] According to some implementations of the present disclosure, determining the initial features further includes: using a second skeleton encoder in the machine learning model to determine second action features corresponding to the action image; and updating the initial features using the second action features.
[0115] According to some implementations of the present disclosure, the machine learning model is determined based on: obtaining reference samples, the reference samples including: a first reference image of a first reference state of a subject performing a reference action at a first reference time point, a reference action image, and a second reference image, the reference action image specifying a reference action performed by the subject at a second reference time point after the first reference time point, the reference action being represented by a reference skeleton image of the subject, and the second reference image representing a second reference state of the subject at the second reference time point; using the machine learning model to determine predicted features of the second reference image based on the first reference image and the reference action image; and updating the machine learning model based on the features of the second reference image and the predicted features of the second reference image.
[0116] According to some implementations of the present disclosure, a machine learning model is updated in a first stage, the subject includes a human subject, and updating the machine learning model includes updating at least any one of the following: a control network model, a fine-tuning model in a diffusion model, a first skeleton encoder, and a second skeleton encoder.
[0117] According to some implementations of the present disclosure, the machine learning model is updated in a second stage after the first stage, the subject includes at least any one of a human subject and a robot, and updating the machine learning model includes updating at least any one of the following: a control network model, a fine-tuning model in a diffusion model, a first skeleton encoder, and a second skeleton encoder.
[0118] According to some implementations of the present disclosure, the skeleton image includes at least one line segment corresponding to at least one robotic arm of the robot.
[0119] According to some implementations of the present disclosure, the action image includes a plurality of action images, and the plurality of action images specify a plurality of actions to be performed by the robot at a plurality of second time points after the first time point.
[0120] Example device and equipment
[0121] Figure 14 A block diagram of an apparatus 1400 for managing the state of a robot according to some implementations of the present disclosure is shown. The apparatus 1400 includes: a state image receiving module 1410 configured to receive a first image describing a first state of the robot at a first time point; an action image receiving module 1420 configured to receive an action image, the action image specifying an action to be performed by the robot at a second time point after the first time point, the action being represented by a skeleton image of the robot; and a determining module 1430 configured to use a machine learning model to determine, based on the first image and the action image, a second image of a second state of the robot at the second time point.
[0122] According to some implementations of the present disclosure, the determining module 1430 is further configured to: use the machine learning model to determine a second image feature corresponding to the second image based on the first image and the action image; and use a video decoder in the machine learning model to determine the second image based on the second image feature.
[0123] According to some implementations of the present disclosure, the determining module 1430 is further configured to: determine an initial feature of a diffusion model in the machine learning model based on the first image; use a first skeleton encoder in the machine learning model to determine a first action feature corresponding to the action image; and use the diffusion model to determine a second image feature based on the initial feature and the first action feature.
[0124] According to some implementations of the present disclosure, the determination module 1430 is further configured to: determine conditional features of the diffusion model based on the first action features by using a control network model in the machine learning model; and determine second image features based on the initial features and the conditional features by using the diffusion model.
[0125] According to some implementations of the present disclosure, the determination module 1430 is further configured to: update the first action features based on the initial features of the diffusion model.
[0126] According to some implementations of the present disclosure, the determination module 1430 is further configured to: determine first image features of the first image by using a video encoder in the machine learning model; and add noise information to the first image features to determine the initial features.
[0127] According to some implementations of the present disclosure, the determination module 1430 is further configured to: determine second action features corresponding to the action image by using a second skeleton encoder in the machine learning model; and update the initial features by using the second action features.
[0128] According to some implementations of the present disclosure, the machine learning model is determined based on: obtaining reference samples, the reference samples including: a first reference image of a first reference state of a subject performing a reference action at a first reference time point, a reference action image, and a second reference image, the reference action image specifying a reference action performed by the subject at a second reference time point after the first reference time point, the reference action being represented by a reference skeleton image of the subject, and the second reference image representing a second reference state of the subject at the second reference time point; using the machine learning model to determine prediction features of the second reference image based on the first reference image and the reference action image; and updating the machine learning model based on the features of the second reference image and the prediction features of the second reference image.
[0129] According to some implementations of the present disclosure, the machine learning model is updated in a first stage, the subject includes a human subject, and updating the machine learning model includes updating at least any one of the following: a control network model, a fine-tuning model in the diffusion model, a first skeleton encoder, and a second skeleton encoder.
[0130] According to some implementations of the present disclosure, the machine learning model is updated in a second stage after the first stage, the subject includes at least any one of a human subject and a robot, and updating the machine learning model includes updating at least any one of the following: a control network model, a fine-tuning model in the diffusion model, a first skeleton encoder, and a second skeleton encoder.
[0131] According to some implementations of the present disclosure, the skeleton image includes at least one line segment corresponding to at least one robotic arm of the robot.
[0132] According to some implementations of the present disclosure, the motion images include a plurality of motion images, and the plurality of motion images specify a plurality of actions to be performed by the robot at a plurality of second time points after the first time point.
[0133] Figure 15 A block diagram of a device 1500 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that Figure 15 the illustrated computing device 1500 is merely exemplary and should not constitute any limitation on the functions and scope of the implementations described herein. Figure 15 The illustrated computing device 1500 can be used to implement the methods described above.
[0134] As Figure 15 shown, the computing device 1500 is in the form of a general-purpose computing device. The components of the computing device 1500 may include, but are not limited to, one or more processors 1510, a memory 1520, a storage device 1530, one or more communication units 1540, one or more input devices 1550, and one or more output devices 1560. The processor 1510 can be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 1520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 1500.
[0135] The computing device 1500 generally includes multiple computer storage media. Such media can be any available media accessible to the computing device 1500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1520 can be a volatile memory (such as registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1530 can be a removable or non-removable medium and can include machine-readable media, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device 1500.
[0136] The computing device 1500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 15As shown, a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data medium interfaces. The memory 1520 can include a computer program product 1525 having one or more program modules configured to perform the various methods or actions of the various implementations of the present disclosure.
[0137] The communication unit 1540 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 1500 can be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, the computing device 1500 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0138] The input device 1550 can be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 1560 can be one or more output devices such as a display, speaker, printer, etc. The computing device 1500 can also communicate with one or more external devices (not shown) as needed via the communication unit 1540, such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the computing device 1500, or communicate with any device that enables the computing device 1500 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0139] According to an implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, and the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0140] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0141] These computer-readable program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0142] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0143] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0144] The foregoing has described various implementations of the present disclosure. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The selection of the terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.
Claims
1. A method for managing the state of a robot, comprising: receiving a first image describing a first state of the robot at a first point in time; receiving an action image, the action image specifying an action to be performed by the robot at a second time point after the first time point, the action represented by a skeleton image of the robot; as well as A second image of the robot in a second state at the second time point is determined based on the first image and the action image using a machine learning model.
2. The method of claim 1 , wherein determining the second image comprises: Determine, using the machine learning model, a second image feature corresponding to the second image based on the first image and the action image; as well as The second image is determined based on the second image features using a video decoder in the machine learning model.
3. The method of claim 2, wherein determining the second image feature comprises: Based on the first image, determining initial features of a diffusion model in the machine learning model; Determining, using a first skeleton encoder in the machine learning model, a first action feature corresponding to the action image; and The second image feature is determined based on the initial feature and the first motion feature using the diffusion model.
4. The method according to claim 3, wherein determining the second image feature based on the initial feature and the first action feature comprises: Determining the conditional features of the diffusion model based on the first action features using a control network model in the machine learning model; as well as The second image feature is determined by using the diffusion model based on the initial feature and the conditional feature.
5. The method according to claim 3, wherein determining the first action feature further comprises: The first motion feature is updated based on the initial feature of the diffusion model.
6. The method of claim 3, wherein determining the initial features comprises: Determining, using a video encoder in the machine learning model, a first image feature of the first image; as well as Noise information is added to the first image feature to determine the initial feature.
7. The method according to claim 4, wherein determining the initial features further comprises: Determining, using a second skeleton encoder in the machine learning model, a second action feature corresponding to the action image; as well as The initial feature is updated using the second action feature.
8. The method of claim 7, wherein the machine learning model is determined based on: Obtain a reference sample, wherein the reference sample includes: a first reference image of a subject performing a reference action at a first reference state at a first reference time point, a reference action image, and a second reference image, the reference action image specifying a reference action performed by the subject at a second reference time point after the first reference time point, the reference action represented by a reference skeleton image of the subject, and the second reference image representing a second reference state of the subject at the second reference time point; Determining, using the machine learning model, predicted features of the second reference image based on the first reference image and the reference action image; as well as The machine learning model is updated based on the features of the second reference image and the predicted features of the second reference image.
9. A method according to claim 8, wherein the machine learning model is updated in a first stage, the subject includes a human subject, and updating the machine learning model includes updating at least any one of the following: the control network model, the fine-tuning model in the diffusion model, the first skeleton encoder, and the second skeleton encoder.
10. A method according to claim 9, wherein the machine learning model is updated in a second stage after the first stage, the subject includes at least any one of a human subject and a robot, and updating the machine learning model includes updating at least any one of the following: the control network model, the fine-tuning model in the diffusion model, the first skeleton encoder, and the second skeleton encoder.
11. The method of claim 1, wherein the skeleton image includes at least one line segment corresponding to at least one robot arm of the robot. 12 . The method according to claim 1 , wherein the action image includes a plurality of action images, and the plurality of action images specify a plurality of actions to be performed by the robot at a plurality of second time points after the first time point.
13. A device for managing the state of a robot, comprising: a state image receiving module, configured to receive a first image describing a first state of the robot at a first time point; an action image receiving module configured to receive an action image, the action image specifying an action performed by the robot at a second time point after the first time point, the action being represented by a skeleton image of the robot; as well as A determination module is configured to determine a second image of the second state of the robot at the second time point based on the first image and the action image by using a machine learning model.
14. An electronic device comprising: at least one processor; as well as At least one memory, the at least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processor.
15. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 12.
16. A computer instruction product comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 12.
Citation Information
Cited By
Control method of mechanical arm, robot, storage medium and computer program product
CN121798625A
Unified reinforcement learning control method and system for modular multi-form foot-type robot
CN122275014A