Robot control method, device, robot, electronic device and storage medium

By using scene flow as the dynamic representation of the world model, combining color depth images and robot motion information for prediction, the problem of inaccurate robot manipulation in the prior art is solved, and the robot's manipulation accuracy and response speed are improved in complex scenarios.

CN120002673BActive Publication Date: 2025-08-05BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Patent Information

Application Number
CN202510477069.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-05
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the prior art, world models cannot accurately predict the object motion trajectory and the interaction effect between the robot and the environment in rapidly changing or complex interactive scenarios, resulting in inaccurate robot manipulation strategies and increasing the risk of collision.

Method used

The scene flow is used as the dynamic representation of the world model. By obtaining the color depth image and robot execution action information at the current moment, the dynamic prediction module and diffusion module are used to process the prediction scene flow and color depth image, and accurately predict the scene characteristics and robot action characteristics.

Benefits of technology

It improves the accuracy and response speed of robot manipulation strategies, and enhances the operation reliability of robots in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120002673B_ABST
    Figure CN120002673B_ABST
Patent Text Reader

Abstract

The present application discloses a robot control method, device, robot, electronic device and storage medium, which belongs to the field of artificial intelligence technology. The robot control method includes obtaining a first color depth image of the target scene at the current moment and the execution action information of the robot; inputting the first color depth image and the execution action information into the dynamic prediction module of the target world model, and obtaining the predicted scene flow of the target scene at the next moment output by the dynamic prediction module; inputting the predicted scene flow and the first color depth image into the diffusion module of the target world model, and obtaining the second color depth image of the target scene at the next moment output by the diffusion module. This method can effectively predict the motion trajectory of an object and the law of change of the environmental state, and control the robot according to the prediction results, thereby improving the accuracy and response speed of the robot control strategy and enhancing the reliability of the robot in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a robot control method, device, robot, electronic device and storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, the practical application scenarios of robots are expanding. World models, a technology that can simulate and predict the dynamic changes in a robot's environment, are also becoming increasingly widely used. However, in scenarios involving rapid changes or complex interactions, world models often cannot accurately predict the motion trajectory of objects or the interaction between the robot and the environment due to the lack of a clear dynamic representation.

[0003] In robot control tasks, this prediction bias will accumulate over time, affecting the robot's operational reliability, easily leading to robot misoperation, and increasing the robot's collision risk. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems existing in the prior art. To this end, this application proposes a robot control method, device, robot, electronic device, and storage medium that can effectively predict the motion trajectory of an object and the changing patterns of the environmental state, thereby improving the accuracy and response speed of the robot control strategy.

[0005] In a first aspect, the present application provides a robot manipulation method, the method comprising:

[0006] Acquire a first color depth image of a target scene at a current moment and execution action information of a robot, wherein the robot is in the target scene;

[0007] Inputting the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtaining a predicted scene flow of the target scene at the next moment output by the dynamics prediction module;

[0008] Inputting the predicted scene stream and the first color depth image into the diffusion module of the target world model, obtaining a second color depth image of the target scene at the next moment output by the diffusion module, wherein the second color depth image is used to determine the manipulation strategy of the robot at the next moment;

[0009] The target world model is trained based on sample color depth images and sample scene flows.

[0010] According to the robot control method of the present application, based on the color depth image of the target scene at the current moment and the robot's execution action information, the predicted scene flow at the next moment can be obtained; based on the current color depth image and the predicted scene flow, the color depth image at the next moment can be obtained to determine the robot's control strategy. This method uses the scene flow as the dynamic representation of the world model, and can combine the scene characteristics and the robot's action characteristics to accurately predict the scene changes and object movements, and control the robot according to the prediction results, thereby improving the accuracy and response speed of the robot's control strategy and enhancing the robot's operation reliability in complex scenes.

[0011] According to one embodiment of the present application, the sample scene stream is obtained by the following steps:

[0012] Performing depth estimation on the sample color image to obtain a sample scene depth image;

[0013] A scene flow estimation is performed based on the sample color image and the sample scene depth image to obtain a sample scene flow.

[0014] According to one embodiment of the present application, performing depth estimation on the sample color image to obtain a sample scene depth image includes:

[0015] Perform relative depth estimation based on video data corresponding to the sample color image to obtain a first sample depth image;

[0016] Performing absolute depth estimation based on the frame data corresponding to the sample color image to obtain a second sample depth image;

[0017] The first sample depth image and the second sample depth image are aligned to obtain the sample scene depth image.

[0018] According to one embodiment of the present application, the performing scene flow estimation based on the sample color image and the sample scene depth image to obtain the sample scene flow includes:

[0019] Inputting the sample color image and the sample scene depth image into a scene flow estimation model to obtain an initial sample scene flow output by the scene flow estimation model;

[0020] performing optical flow estimation on the sample color image to determine a non-moving object mask;

[0021] Based on the non-moving object mask, mask processing is performed on the initial sample scene stream to obtain the sample scene stream.

[0022] According to one embodiment of the present application, inputting the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtaining a predicted scene flow of the target scene at the next moment output by the dynamics prediction module, includes:

[0023] Converting the first color depth image into first point cloud data of the target scene at a current moment;

[0024] Inputting the first point cloud data and the execution action information into the dynamics prediction module to obtain a first visual feature and a first action feature;

[0025] The predicted scene flow output by the dynamics prediction module is obtained based on the first visual feature and the first motion feature.

[0026] According to one embodiment of the present application, obtaining the predicted scene flow output by the dynamics prediction module based on the first visual feature and the first motion feature includes:

[0027] performing cross-attention processing on the first visual feature and the first motion feature;

[0028] Based on the cross-attention processing result, the predicted scene flow is obtained.

[0029] According to one embodiment of the present application, inputting the predicted scene stream and the first color depth image into the diffusion module of the target world model, and obtaining a second color depth image of the target scene at the next moment output by the diffusion module, includes:

[0030] performing splicing processing on the predicted scene stream and the first color depth image to obtain a first fused scene feature;

[0031] The first fused scene feature is input into the diffusion module to obtain the second color depth image output by the diffusion module.

[0032] According to one embodiment of the present application, the diffusion module includes an image decoder and a depth estimation network, and the step of inputting the predicted scene stream and the first color depth image into the diffusion module of the target world model to obtain a second color depth image of the target scene at the next moment output by the diffusion module includes:

[0033] Inputting the predicted scene stream and the first color depth image into the diffusion module to obtain a predicted color image output by the image decoder and a predicted depth image output by the depth estimation network;

[0034] The second color depth image output by the diffusion module is obtained based on the predicted color image and the predicted depth image.

[0035] In a second aspect, the present application provides a robot manipulation device, comprising:

[0036] an acquisition module, configured to acquire a first color depth image of a target scene at a current moment and information about an execution action of the robot, wherein the robot is in the target scene;

[0037] a first processing module, configured to input the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtain a predicted scene flow of the target scene at a next moment output by the dynamics prediction module;

[0038] a second processing module, configured to input the predicted scene stream and the first color depth image into a diffusion module of the target world model, obtain a second color depth image of the target scene at a next moment output by the diffusion module, and use the second color depth image to determine a manipulation strategy of the robot at a next moment;

[0039] The target world model is trained based on sample color images and sample scene flows.

[0040] According to the robot control device of the present application, based on the color depth image of the target scene at the current moment and the robot's execution action information, the predicted scene flow at the next moment can be obtained; based on the current color depth image and the predicted scene flow, the color depth image at the next moment can be obtained to determine the robot's control strategy. The device uses the scene flow as the dynamic representation of the world model, and can combine the scene characteristics and the robot's action characteristics to accurately predict the scene changes and object movements, and control the robot according to the prediction results, thereby improving the accuracy and response speed of the robot's control strategy and enhancing the robot's operation reliability in complex scenes.

[0041] In a third aspect, the present application provides a robot comprising the robot manipulation device as described in the second aspect.

[0042] According to the robot of the present application, based on the color depth image of the target scene at the current moment and the robot's execution action information, the predicted scene flow at the next moment can be obtained; based on the current color depth image and the predicted scene flow, the color depth image at the next moment can be obtained to determine the robot's control strategy. The robot uses the scene flow as the dynamic representation of the world model, and can combine the scene characteristics and the robot's action characteristics to accurately predict the scene changes and object movements, and control the robot according to the prediction results, thereby improving the accuracy and response speed of the robot's control strategy and enhancing the robot's operation reliability in complex scenes.

[0043] In a fourth aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the robot manipulation method as described in the first aspect above is implemented.

[0044] In a fifth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the robot manipulation method as described in the first aspect above.

[0045] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0047] Figure 1 This is one of the flow charts of the robot control method provided in the embodiment of the present application;

[0048] Figure 2 This is the second flow chart of the robot control method provided in the embodiment of the present application;

[0049] Figure 3 This is the third flow chart of the robot control method provided in the embodiment of the present application;

[0050] Figure 4 Schematic diagram of the structure of the robot control device provided in the embodiment of the present application;

[0051] Figure 5 is a schematic structural diagram of a robot provided in an embodiment of the present application;

[0052] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application.

[0053] Reference numerals:

[0054] Robot control device 400, acquisition module 410, first processing module 420, second processing module 430,

[0055] Robot 500,

[0056] Electronic device 600 , processor 610 , memory 620 . DETAILED DESCRIPTION

[0057] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0058] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0059] With the development of artificial intelligence technology, the application scenarios of robot 500 are increasing. The world model is used to simulate and predict changes in the environment in which robot 500 is located. However, it lacks a clear dynamic representation in rapidly changing or complex interactive scenarios, making it difficult to accurately predict object motion or interaction effects, limiting its practical application value.

[0060] An embodiment of the present application provides a robot control method that uses scene flow as the dynamic representation of the world model, which can effectively predict the motion trajectory of objects and the changing laws of environmental states, improve the accuracy and response speed of the robot 500 control strategy, and enhance the reliability of the robot 500 in complex scenarios.

[0061] like Figure 1 As shown, the robot manipulation method includes: step 110, step 120 and step 130.

[0062] Step 110 : Acquire a first color depth image of a target scene at the current moment and information about the execution action of the robot 500 , where the robot 500 is in the target scene.

[0063] The robot 500 is a mechanical device that can be controlled by a computer program and is used to perform various tasks, such as manufacturing, assembling, handling, cleaning, etc.

[0064] The target scene may be the specific environment in which the robot 500 performs a task. The target scene may include a variety of objects, such as objects, vehicles, humans, or other robots 500.

[0065] A color depth image is an image data containing color information and depth information.

[0066] In actual implementation, a sensor (such as an RGB-D camera) can simultaneously obtain a color image of the target scene, that is, color information, and the distance from each pixel in the image to the sensor, that is, depth information.

[0067] The execution action information of the robot 500 may refer to the specific actions and related data performed by the robot 500 in the process of completing a task.

[0068] The execution action information may include but is not limited to the action type, action parameters, and sensor data of the robot 500.

[0069] For example, the action type can be,moving, grabbing, operating, etc.;

[0070] Action parameters can be position, force, speed, etc.

[0071] Sensor data can be joint angle, motor torque, motor speed, etc.

[0072] In actual execution, the execution action information of the robot 500 can be obtained by setting specific sensors or obtaining feedback from the control system.

[0073] In this step, the first color depth image at the current moment can be obtained to represent the color information and depth information in the target scene at the current moment, and the execution action information at the current moment can be obtained to represent the motion state of the robot 500 at the current moment.

[0074] Step 120: Input the first color depth image and the execution action information into the dynamics prediction module of the target world model to obtain a predicted scene flow of the target scene at the next moment output by the dynamics prediction module.

[0075] The scene flow contains the motion information of objects in the target scene in three dimensions.

[0076] For example, the coordinates of an object at the current moment are (x1, y1, z1), and the coordinates at the next moment are (x2, y2, z2). The scene flow can be expressed as the displacement of the object between the two moments, that is, (x1-x2, y1-y2, z1-z2).

[0077] The target world model is a world model of the robot 500 trained based on specific sample data.

[0078] The dynamics prediction module is a functional module in the target world model, which can predict the scene flow of the target scene at the next moment, that is, predict the scene flow.

[0079] It can be understood that the predicted scene flow reflects the possible trend of object movement in the target scene at the next moment.

[0080] In this step, by inputting the first color depth image and the execution action information into the dynamics prediction module, the predicted scene flow of the target scene at the next moment can be obtained.

[0081] Step 130: Input the predicted scene flow and the first color depth image into the diffusion module of the target world model, and obtain the second color depth image of the target scene at the next moment output by the diffusion module. The second color depth image is used to determine the manipulation strategy of the robot 500 at the next moment.

[0082] The target world model is trained based on sample color depth images and sample scene flows.

[0083] It can be understood that the sample color depth image and the sample scene stream are sample data for training the target world model, and the sample color depth image can be collected by an RGB-D camera.

[0084] By splitting the sample color depth image, a sample color image and a sample depth image can be obtained.

[0085] The diffusion module is a functional module in the target world model, which can output the predicted scene flow and the first color depth image as the color depth image of the target scene at the next moment, that is, the second color depth image, through the diffusion model.

[0086] It is understandable that the second color depth image reflects the possible trend of color information and depth information in the target scene at the next moment, and can be used to determine the manipulation strategy of the robot 500 at the next moment.

[0087] Based on the second color depth image of the target scene at the next moment, the robot 500 can make decisions and judgments in advance according to the trends and changes reflected in the second color depth image to determine the control strategy at the next moment.

[0088] In actual implementation, the manipulation strategy may include but is not limited to path planning, motion control, and action instructions of the robot 500.

[0089] In this step, by inputting the predicted scene flow and the first color depth image into the diffusion module, a second color depth image can be obtained, and the second color depth image can improve the accuracy and response speed of the robot 500's control strategy at the next moment.

[0090] In related technologies, the world model can be used to control the robot 500. The world model is based on the observation data and the actions of the robot 500, and directly predicts the observation scene at the next moment through a generative model. However, this type of technology often focuses on whether the generated visual effects are realistic, and ignores the dynamic modeling of the scene, which makes the prediction results of the world model prone to deviations, affecting the accuracy and response speed of the robot 500's control strategy, easily leading to the robot 500's misoperation, and increasing the collision risk of the robot 500.

[0091] In an embodiment of the present application, the world model used in the robot manipulation method can be predicted based on the scene flow as a dynamic representation. First, a color depth image of the target scene at the current moment and information about the robot 500's execution of the action are obtained and input into the dynamic prediction module to obtain a predicted scene flow of the target scene at the next moment. The predicted scene flow and color depth image are then input into the diffusion module to obtain a color depth image of the target scene at the next moment. This method uses the scene flow as the dynamic representation of the world model. It can combine the scene characteristics and the robot 500's own motion characteristics to accurately predict scene changes and object motion, and manipulate the robot 500 based on the prediction results, thereby improving the accuracy and response speed of the robot 500's manipulation strategy and enhancing the robot 500's operational reliability in complex scenarios.

[0092] According to the robot manipulation method provided in the embodiments of the present application, a predicted scene flow for the next moment can be obtained based on the color depth image of the target scene at the current moment and information about the actions performed by robot 500. Based on the current color depth image and the predicted scene flow, a color depth image for the next moment can be obtained to determine the manipulation strategy of robot 500. This method uses the scene flow as a dynamic representation of the world model, combining scene characteristics with the characteristics of the robot's 500 actions to accurately predict scene changes and object motion. The robot 500 is then manipulated based on the predicted results, improving the accuracy and responsiveness of the robot's manipulation strategy and enhancing its operational reliability in complex scenarios.

[0093] In some embodiments, as Figure 2 As shown, the sample scene flow is obtained through the following steps:

[0094] Perform depth estimation on the sample color image to obtain a sample scene depth image;

[0095] The scene flow is estimated based on the sample color image and the sample scene depth image to obtain the sample scene flow.

[0096] The sample color image may be composed of at least two video frames containing color information of the same scene.

[0097] In actual implementation, an existing data set can be used as a sample color image, or a camera or other device can be used to shoot and scan, and the results of the shooting and scanning can be used as the sample color image.

[0098] It should be noted that scene flow data cannot be directly obtained through devices such as cameras or sensors, but scene flow can be obtained through scene flow estimation methods.

[0099] Depth estimation of the sample color image and scene flow estimation based on the sample color image and the sample scene depth image can be achieved through a specific model.

[0100] In this embodiment, depth estimation is performed on the sample color image, and the depth information of the pixel point, that is, the distance of the pixel point from the observation point, can be estimated from the sample color image to obtain a sample scene depth image. Based on the sample color image and the sample scene depth image, scene flow estimation can be performed. The color information and depth information are combined to estimate the high-precision sample scene flow, which can be used to train the target world model.

[0101] In actual implementation, by calculating the mean square error between the predicted scene flow and the sample scene flow, the similarity between the two predictions can be evaluated. The lower the mean square error, the closer the predicted scene flow is to the sample scene flow.

[0102] It is understandable that the training goal of the target world model is the error between the predicted scene flow and the sample scene flow (for example, the mean square error). By continuously adjusting the relevant parameters of the dynamics prediction module, the error can be made lower and lower, thereby improving the prediction accuracy of the dynamics prediction module.

[0103] In some embodiments, performing depth estimation on a sample color image to obtain a sample scene depth image includes:

[0104] Perform relative depth estimation based on video data corresponding to the sample color image to obtain a first sample depth image;

[0105] Performing absolute depth estimation based on frame data corresponding to the sample color image to obtain a second sample depth image;

[0106] The first sample depth image and the second sample depth image are aligned to obtain a sample scene depth image.

[0107] The video data may be two or more video frame data in the sample color image, and the video data can reflect motion information in the sample color image.

[0108] Relative depth may refer to the relative distance relationship between different objects in a scene. After performing relative depth estimation on video data, a first sample depth image may be obtained.

[0109] In the first sample depth image, the value of a pixel represents the depth order of the pixel relative to other pixels (ie, which object is closer to the observation point and which object is farther away from the observation point).

[0110] Relative depth estimation based on video data corresponding to sample color images can better estimate the relative depth relationship between objects in the scene.

[0111] In this embodiment, relative depth estimation can be performed on the video data using a Video Depth Anything model.

[0112] The frame data is a single video frame data in the sample color image, and the frame data can reflect the precise scale information in the sample color image.

[0113] Absolute depth may refer to a specific depth distance of an object in a scene. After performing absolute depth estimation on the frame data, a second sample depth image may be obtained.

[0114] In the second sample depth image, the value of a pixel represents a specific depth distance (such as 1 meter, 0.5 meter, etc.) of the pixel from the observation point.

[0115] Absolute depth estimation based on the frame data corresponding to the sample color image can better estimate the specific depth distance of objects in the scene.

[0116] In this embodiment, absolute depth estimation can be performed on the frame data using the Depth Anything v2 model.

[0117] It should be noted that relative depth estimation outputs depth order, which indicates the relative distance relationship between objects and has no actual physical meaning. Absolute depth estimation outputs specific physical distance and has actual scale information. Both depth estimation methods have errors caused by estimation.

[0118] In actual implementation, the image output by the relative depth estimation (ie, the first sample depth image) and the image output by the absolute depth estimation (ie, the second sample depth image) may be aligned based on the scale parameter and the offset parameter.

[0119] Among them, the scale parameter is used to map the range of relative depths (for example, object A is closer and object B is farther away) to the range of absolute depths (for example, object A is 10 meters closer to the observation point than object B).

[0120] The offset parameter is used to correct the error between relative depth and absolute depth.

[0121] In this embodiment, by aligning the first sample depth image with the second sample depth image, a sample scene depth image can be obtained, thereby improving the accuracy of depth estimation.

[0122] In some embodiments, as Figure 2 As shown, scene flow estimation is performed based on the sample color image and the sample scene depth image to obtain the sample scene flow, including:

[0123] Inputting the sample color image and the sample scene depth image into the scene flow estimation model to obtain an initial sample scene flow output by the scene flow estimation model;

[0124] Perform optical flow estimation on the sample color image to determine the non-moving object mask;

[0125] Based on the non-moving object mask, the initial sample scene flow is masked to obtain the sample scene flow.

[0126] The scene flow estimation model is used to output an initial sample scene flow based on a sample color image and a sample scene depth image.

[0127] In actual implementation, scene flow estimation can be performed on the sample color image and the sample scene depth image through the RAFT-3D model.

[0128] It should be noted that the initial sample scene flow output by the scene flow estimation model may contain unnecessary scene flow information such as noise.

[0129] In this embodiment, mask processing may be used to process unnecessary information in the initial sample scene stream.

[0130] First, optical flow estimation is performed on the sample color image to determine the non-moving object mask.

[0131] Among them, optical flow can reflect the motion information of objects in the scene and can be expressed as the motion vector of pixel points on the image plane.

[0132] In actual implementation, the RAFT model can be used to analyze the pixel changes between different frames of the sample color image, estimate the movement direction and speed of the object, and realize optical flow estimation based on the sample color image.

[0133] A mask is a data structure that can be used to select or restrict the area of an image where an operation is to be performed.

[0134] In actual implementation, the mask can be a binary image or matrix, in which each element (such as 0 or 1) indicates whether the corresponding part of the image is selected for operation.

[0135] In this embodiment, the sample color image may be divided into a portion where the optical flow is 0 (ie, a non-moving object) and a portion where the optical flow is not 0 (ie, a moving object).

[0136] The part with an optical flow of 0 can be used for background segmentation to help identify the static background part; the part with an optical flow of not 0 can be used for target detection and tracking to help identify and analyze moving objects. Through optical flow recognition, the mask corresponding to the area where the non-moving object is located in the sample color image (i.e., the non-moving object mask) can be determined.

[0137] Then, based on the non-moving object mask, the initial sample scene stream is masked to obtain the sample scene stream.

[0138] The mask processing performed on the initial sample scene stream may be to remove noise in the region of the initial sample scene stream by using a non-moving object mask.

[0139] In this embodiment, mask processing is performed on the initial sample scene stream to obtain a sample scene stream, which can further refine the initial sample scene stream and remove noise in non-moving object areas in the initial sample scene stream, making the sample scene stream more accurate.

[0140] In some embodiments, the first color depth image and the executed action information are input into a dynamics prediction module of a target world model, and a predicted scene stream of the target scene at the next moment is obtained from the dynamics prediction module, including:

[0141] Converting the first color depth image into first point cloud data of the target scene at the current moment;

[0142] Inputting the first point cloud data and the execution action information into a dynamics prediction module to obtain a first visual feature and a first action feature;

[0143] A predicted scene flow output by a dynamics prediction module is obtained based on the first visual feature and the first motion feature.

[0144] Among them, point cloud data can be a data set composed of a large number of points in three-dimensional space, which contains the position information of each point in the three-dimensional coordinate system, namely (X, Y, Z). In addition to position information, it can also contain information such as color, intensity or direction.

[0145] In actual implementation, point cloud data can be stored in the form of lists or matrices.

[0146] In this embodiment, the two-dimensional pixels in the first color depth image can be mapped to the three-dimensional space by observing the intrinsic parameter matrix of the camera used and combining it with the depth information in the first color depth image to obtain the coordinate point of each two-dimensional pixel in the three-dimensional space.

[0147] It can be understood that the intrinsic parameter matrix is a parameter matrix that can include main parameters such as focal length, principal point coordinates and tilt factor. It is used to describe the internal geometric characteristics of the camera and can reflect the mapping relationship between the two-dimensional image pixel coordinate system and the three-dimensional coordinate system.

[0148] By observing the intrinsic parameter matrix of the camera used and combining it with the depth information, the coordinate point corresponding to each two-dimensional pixel in the first color depth image in the three-dimensional space can be obtained, that is, the first point cloud data.

[0149] In this embodiment, the dynamics prediction module can obtain a first visual feature and a first motion feature based on the first point cloud data and the executed action information, and obtain a predicted scene flow based on the first visual feature and the first motion feature.

[0150] In actual implementation, the dynamics prediction module can achieve feature fusion and scene flow prediction through a conditional U-Net network structure.

[0151] First, feature extraction is performed on the first point cloud data to extract key features such as the object's geometric shape (such as edges, surface curvature), positional relationship (such as distance between objects) and motion information (such as speed, direction), namely the first visual features.

[0152] Then, feature extraction is performed on the execution action information to extract key features such as the actuator posture (position + attitude), joint angle or motion speed / acceleration of the robot 500, namely the first action feature.

[0153] Finally, scene flow prediction is performed based on the first visual feature and the first action feature to obtain the predicted scene flow.

[0154] In this embodiment, the scene flow prediction is performed by the dynamics prediction module in combination with the execution action information of the robot 500, which can improve the accuracy and reliability of the predicted scene flow.

[0155] In some embodiments, obtaining a predicted scene flow output by a dynamics prediction module based on the first visual feature and the first motion feature includes:

[0156] Perform cross-attention processing on the first visual feature and the first action feature;

[0157] Based on the cross-attention processing results, the predicted scene flow is obtained.

[0158] In this embodiment, the first visual feature and the first action feature may be processed using a cross-attention mechanism.

[0159] It can be understood that the cross-attention mechanism is an attention mechanism for processing multi-sequence data. It can adjust the attention weight of visual features through action features to guide the dynamics prediction module to focus on the visual features.

[0160] For example, if the robot 500 is accelerating, the cross-attention mechanism will increase the attention weight of the dynamics prediction module on the object in front; if the robot 500 is turning, the cross-attention mechanism will guide the dynamics prediction module to focus on the movement changes of the objects on the side.

[0161] In this embodiment, the predicted scene flow obtained based on the cross-attention processing results can dynamically combine the robot 500 execution action information and visual features, and can more accurately predict the movement of objects in the scene.

[0162] In some embodiments, as Figure 3 As shown, the predicted scene flow and the first color depth image are input into the diffusion module of the target world model to obtain the second color depth image of the target scene at the next moment output by the diffusion module, including:

[0163] splicing the predicted scene stream and the first color depth image to obtain a first fused scene feature;

[0164] The first fused scene feature is input into the diffusion module to obtain a second color depth image output by the diffusion module.

[0165] In actual implementation, Figure 3 As shown in Figure 2, the diffusion module can use the pre-trained stable diffusion model as the basic model and perform end-to-end training in conjunction with the dynamics prediction module.

[0166] The overall loss function can be composed of the diffusion loss of the diffusion module and the scene flow prediction loss of the dynamics prediction module, and the two are balanced by the weight coefficient.

[0167] It should be noted that the latent space of the stable diffusion model has a specific size, and the size of the latent space can be determined by the model architecture.

[0168] Here, the latent space refers to the intermediate representation space used by the stable diffusion model during the generation process.

[0169] In the process of generating a stable diffusion model, the input data can be first mapped to the latent space to obtain the latent representation of the input data, and then the data can be gradually restored from the latent space according to preset conditions and methods, and finally the image can be output.

[0170] In practice, the input data can be downsampled to keep the dimensions consistent.

[0171] In this embodiment, the predicted scene stream and the first color depth image are spliced to obtain a first fused scene feature.

[0172] The stitching process is to stitch the predicted scene stream and the first color depth image in the channel dimension.

[0173] For example, the dimension of the first color depth image is H×W×4 (i.e., height H, width W, 4 channels), the dimension of the predicted scene flow is H×W×3 (i.e., height H, width W, 3 channels), and the dimension of the first fused scene feature obtained after splicing is H×W×(4+3).

[0174] Based on the first fused scene features, the diffusion module can output a second color depth image of the target scene at the next moment.

[0175] In some embodiments, as Figure 3 As shown, the diffusion module includes an image decoder and a depth estimation network, which inputs the predicted scene stream and the first color depth image into the diffusion module of the target world model, and obtains the second color depth image of the target scene at the next moment output by the diffusion module, including:

[0176] Input the predicted scene stream and the first color depth image into the diffusion module to obtain a predicted color image output by the image decoder and a predicted depth image output by the depth estimation network;

[0177] Based on the predicted color image and the predicted depth image, a second color depth image output by the diffusion module is obtained.

[0178] The image decoder is configured to output a predicted color image based on the predicted scene stream and the first color depth image;

[0179] The depth estimation network is configured to output a predicted depth image based on the predicted scene flow and the first color depth image.

[0180] In actual implementation, the diffusion model can gradually predict the potential representation of the image at the next moment through a multi-step denoising process, and then restore the potential representation to a high-resolution color image through the image decoder, that is, predict the color image.

[0181] In addition, the diffusion model can also predict and output the depth image at the next moment through the depth estimation network.

[0182] In this embodiment, through the image decoder and the depth estimation network, the diffusion module can generate a predicted color image and a predicted depth image based on the predicted scene flow and the first color depth image, and combine the predicted color image and the predicted depth image to output a second color depth image of the target scene at the next moment.

[0183] The robot control method provided in the embodiment of the present application can be executed by the robot control device 400. In the embodiment of the present application, the robot control device 400 is taken as an example to illustrate the robot control method provided in the embodiment of the present application.

[0184] The embodiment of the present application also provides a robot control device 400.

[0185] like Figure 4 As shown, the device includes: an acquisition module 410 , a first processing module 420 and a second processing module 430 .

[0186] The acquisition module 410 is used to acquire a first color depth image of the target scene at the current moment and the execution action information of the robot 500, where the robot 500 is in the target scene;

[0187] The first processing module 420 is configured to input the first color depth image and the execution action information into the dynamics prediction module of the target world model, and obtain a predicted scene flow of the target scene at the next moment output by the dynamics prediction module;

[0188] The second processing module 430 is used to input the predicted scene flow and the first color depth image into the diffusion module of the target world model, and obtain the second color depth image of the target scene at the next moment output by the diffusion module. The second color depth image is used to determine the manipulation strategy of the robot 500 at the next moment.

[0189] The target world model is trained based on sample color depth images and sample scene flows.

[0190] According to the robot manipulation device 400 of the present application, based on the color depth image of the target scene at the current moment and the action information executed by the robot 500, the predicted scene flow at the next moment can be obtained; based on the current color depth image and the predicted scene flow, the color depth image at the next moment can be obtained to determine the manipulation strategy of the robot 500. The device uses the scene flow as the dynamic representation of the world model, and can combine the scene characteristics and the action characteristics of the robot 500 to accurately predict the scene changes and object movements, and manipulate the robot 500 according to the prediction results, thereby improving the accuracy and response speed of the robot 500 manipulation strategy and enhancing the operation reliability of the robot 500 in complex scenes.

[0191] In some embodiments, the sample scene stream is obtained by the following steps:

[0192] Perform depth estimation on the sample color image to obtain a sample scene depth image;

[0193] The scene flow is estimated based on the sample color image and the sample scene depth image to obtain the sample scene flow.

[0194] In some embodiments, performing depth estimation on a sample color image to obtain a sample scene depth image includes:

[0195] Perform relative depth estimation based on video data corresponding to the sample color image to obtain a first sample depth image;

[0196] Performing absolute depth estimation based on frame data corresponding to the sample color image to obtain a second sample depth image;

[0197] The first sample depth image and the second sample depth image are aligned to obtain a sample scene depth image.

[0198] In some embodiments, performing scene flow estimation based on the sample color image and the sample scene depth image to obtain the sample scene flow includes:

[0199] Inputting the sample color image and the sample scene depth image into the scene flow estimation model to obtain an initial sample scene flow output by the scene flow estimation model;

[0200] Perform optical flow estimation on the sample color image to determine the non-moving object mask;

[0201] Based on the non-moving object mask, the initial sample scene flow is masked to obtain the sample scene flow.

[0202] In some embodiments, the first processing module 420 is configured to input the first color depth image and the executed action information into a dynamics prediction module of a target world model, and obtain a predicted scene stream of the target scene at the next moment output by the dynamics prediction module, including:

[0203] Converting the first color depth image into first point cloud data of the target scene at the current moment;

[0204] Inputting the first point cloud data and the execution action information into a dynamics prediction module to obtain a first visual feature and a first action feature;

[0205] A predicted scene flow output by a dynamics prediction module is obtained based on the first visual feature and the first motion feature.

[0206] In some embodiments, obtaining a predicted scene flow output by a dynamics prediction module based on the first visual feature and the first motion feature includes:

[0207] Perform cross-attention processing on the first visual feature and the first action feature;

[0208] Based on the cross-attention processing results, the predicted scene flow is obtained.

[0209] In some embodiments, the second processing module 430 inputs the predicted scene stream and the first color depth image into the diffusion module of the target world model to obtain the second color depth image of the target scene at the next moment output by the diffusion module, including:

[0210] splicing the predicted scene stream and the first color depth image to obtain a first fused scene feature;

[0211] The first fused scene feature is input into the diffusion module to obtain a second color depth image output by the diffusion module.

[0212] In some embodiments, the diffusion module includes an image decoder and a depth estimation network. The second processing module 430 inputs the predicted scene stream and the first color depth image into the diffusion module of the target world model, and obtains a second color depth image of the target scene at the next moment output by the diffusion module, including:

[0213] Input the predicted scene stream and the first color depth image into the diffusion module to obtain a predicted color image output by the image decoder and a predicted depth image output by the depth estimation network;

[0214] Based on the predicted color image and the predicted depth image, a second color depth image output by the diffusion module is obtained.

[0215] The embodiment of the present application also provides a robot 500.

[0216] like Figure 5 As shown, the robot 500 includes the above-mentioned robot control device 400. The robot 500 can implement each process of the above-mentioned robot control method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0217] According to the robot 500 of the present application, based on the color depth image of the target scene at the current moment and the information on the action performed by the robot 500, the predicted scene flow at the next moment can be obtained; based on the current color depth image and the predicted scene flow, the color depth image at the next moment can be obtained to determine the control strategy of the robot 500. The robot 500 uses the scene flow as the dynamic representation of the world model, and can combine the scene characteristics and the action characteristics of the robot 500 to accurately predict the scene changes and object movements, and control the robot 500 according to the prediction results, thereby improving the accuracy and response speed of the robot 500 control strategy and enhancing the operation reliability of the robot 500 in complex scenes.

[0218] The embodiment of the present application also provides an electronic device 600 .

[0219] like Figure 6As shown, the electronic device 600 includes a memory 620, a processor 610, and a computer program stored in the memory 620 and executable on the processor 610. When the processor 610 executes the computer program, each process of the above-mentioned robot control method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0220] An embodiment of the present application also provides a non-transitory computer-readable storage medium.

[0221] The non-transitory computer-readable storage medium stores a computer program, which, when executed by the processor, implements the various processes of the above-mentioned robot control method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0222] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of this application.

[0224] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

[0225] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0226] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A robot control method, characterized in that: include: Acquire a first color depth image of a target scene at a current moment and execution action information of a robot, wherein the robot is in the target scene; Inputting the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtaining a predicted scene flow of the target scene at the next moment output by the dynamics prediction module; Inputting the predicted scene stream and the first color depth image into the diffusion module of the target world model, obtaining a second color depth image of the target scene at the next moment output by the diffusion module, wherein the second color depth image is used to determine the manipulation strategy of the robot at the next moment; Wherein, the target world model is trained based on sample color depth images and sample scene flows; Inputting the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtaining a predicted scene flow of the target scene at a next moment output by the dynamics prediction module, includes: Converting the first color depth image into first point cloud data of the target scene at a current moment; Inputting the first point cloud data and the execution action information into the dynamics prediction module to obtain a first visual feature and a first action feature; The predicted scene flow output by the dynamics prediction module is obtained based on the first visual feature and the first motion feature.

2. The robot control method according to claim 1, characterized in that: The sample scene stream is obtained by the following steps: Perform depth estimation on the sample color image to obtain a sample scene depth image; A scene flow estimation is performed based on the sample color image and the sample scene depth image to obtain a sample scene flow.

3. The robot control method according to claim 2, characterized in that: The performing depth estimation on the sample color image to obtain the sample scene depth image includes: Perform relative depth estimation based on video data corresponding to the sample color image to obtain a first sample depth image; Performing absolute depth estimation based on the frame data corresponding to the sample color image to obtain a second sample depth image; The first sample depth image and the second sample depth image are aligned to obtain the sample scene depth image.

4. The robot control method according to claim 2, characterized in that: The performing scene flow estimation based on the sample color image and the sample scene depth image to obtain the sample scene flow includes: Inputting the sample color image and the sample scene depth image into a scene flow estimation model to obtain an initial sample scene flow output by the scene flow estimation model; performing optical flow estimation on the sample color image to determine a non-moving object mask; Based on the non-moving object mask, mask processing is performed on the initial sample scene stream to obtain the sample scene stream.

5. The robot control method according to claim 1, characterized in that: The obtaining, based on the first visual feature and the first motion feature, the predicted scene flow output by the dynamics prediction module includes: performing cross-attention processing on the first visual feature and the first motion feature; Based on the cross-attention processing result, the predicted scene flow is obtained.

6. The robot control method according to any one of claims 1 to 4, characterized in that: Inputting the predicted scene stream and the first color depth image into a diffusion module of the target world model, and obtaining a second color depth image of the target scene at a next moment output by the diffusion module, comprises: performing splicing processing on the predicted scene stream and the first color depth image to obtain a first fused scene feature; The first fused scene feature is input into the diffusion module to obtain the second color depth image output by the diffusion module.

7. The robot control method according to any one of claims 1 to 4, characterized in that: The diffusion module includes an image decoder and a depth estimation network. The step of inputting the predicted scene stream and the first color depth image into the diffusion module of the target world model and obtaining a second color depth image of the target scene at a next moment output by the diffusion module includes: Inputting the predicted scene stream and the first color depth image into the diffusion module to obtain a predicted color image output by the image decoder and a predicted depth image output by the depth estimation network; The second color depth image output by the diffusion module is obtained based on the predicted color image and the predicted depth image.

8. A robot control device, characterized in that: include: an acquisition module, configured to acquire a first color depth image of a target scene at a current moment and information about an execution action of the robot, wherein the robot is in the target scene; a first processing module, configured to input the first color depth image and the execution action information into a dynamics prediction module of a target world model, and obtain a predicted scene flow of the target scene at a next moment output by the dynamics prediction module; a second processing module, configured to input the predicted scene stream and the first color depth image into a diffusion module of the target world model, obtain a second color depth image of the target scene at a next moment output by the diffusion module, and use the second color depth image to determine a manipulation strategy of the robot at a next moment; Wherein, the target world model is trained based on sample color images and sample scene flows; The first processing module is further configured to convert the first color depth image into first point cloud data of the target scene at a current moment; Inputting the first point cloud data and the execution action information into the dynamics prediction module to obtain a first visual feature and a first action feature; The predicted scene flow output by the dynamics prediction module is obtained based on the first visual feature and the first motion feature.

9. A robot, characterized in that: Comprising the robot manipulation device as claimed in claim 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the robot manipulation method according to any one of claims 1 to 7 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot manipulation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Scene flow estimation method and device and scene flow estimation model training method and device

    CN113160278A

  • Depth image acquisition method, device and equipment

    CN117173232A

  • Robot operation method based on world model and progressive reasoning

    CN118288274A

  • World model disturbance method, device and equipment and storage medium

    CN119129695A

Cited By

  • Robot training system and robot control system

    CN121683861A