Motion feature prediction method, task execution device, electronic device and medium
By combining color image and depth image features into the motion model, the problem of poor accuracy in motion feature prediction is solved, and the accuracy and success rate of robot task execution are improved.
Patent Information
- Application Number
- CN202510767801.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-16
AI Technical Summary
Existing motion feature prediction methods have the problem of poor accuracy in robot task execution, especially under different light intensities, which makes the robot prone to empty clamping, collision and other phenomena, affecting the success rate of task execution.
By obtaining the intermediate features of the color image features and combining them with the depth image features and noise motion features to input the motion model, the depth image is used to provide depth information to improve the accuracy of predicting motion features and enhance the model's ability to resist light interference.
The accuracy of motion feature prediction is improved, the probability of empty clamping and collision during task execution is reduced, and the accuracy and success rate of task execution are improved.
Smart Images

Figure CN120655940A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robot vision technology, and in particular to an action feature prediction method, task execution equipment, electronic equipment and medium. Background Art
[0002] With the continuous advancement of science and technology, people's lives and work are gradually moving towards intelligence. For example, in scenarios where a task execution device is used to perform a target task, the motion characteristics of the task execution device at a future moment can be predicted based on the scene image of the target task at the current moment. In this way, the task execution device can be controlled to perform the target task based on the predicted motion characteristics. However, the motion characteristics predicted by current motion feature prediction methods suffer from poor accuracy. Summary of the Invention
[0003] In view of this, embodiments of the present application provide an action feature prediction method, a task execution device, an electronic device, and a medium, which can improve the accuracy of predicting action features.
[0004] In the first aspect, an embodiment of the present application provides an action feature prediction method, comprising: obtaining intermediate features output by an image processing model, wherein the intermediate features are obtained by the image processing model processing input data, and the input data includes color image features, and the color image features are used to characterize the color image collected during the process of the task execution device performing the target task, and the intermediate features are used to characterize the information in the color image features; inputting the intermediate features, depth image features and noise action features into the action model to obtain predicted action features output by the action model, wherein the predicted action features are used to characterize the motion trajectory of the task execution device in the future time period, and the depth image features and the color image features correspond in time.
[0005] In the second aspect, an embodiment of the present application provides a model training method, including: obtaining sample intermediate features output by an image processing model, wherein the sample intermediate features are obtained by the image processing model processing sample input data, and the sample input data includes sample color image features, and the sample color image features are used to characterize the sample color images collected during the task execution device performing the target task, and the sample intermediate features are used to characterize the information in the sample color image features; inputting the sample intermediate features, sample depth image features, and sample noise motion features into a pre-trained motion model to obtain sample predicted motion features output by the pre-trained motion model, wherein the sample predicted motion features are used to characterize the motion trajectory of the task execution device in the target time period, and the sample depth image features and the sample color image features correspond in time; post-training the pre-trained motion model based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model, wherein the target motion features are used to characterize the actual motion trajectory of the task execution device in the target time period.
[0006] In the third aspect, an embodiment of the present application provides an action feature prediction device, including: an acquisition module for acquiring intermediate features output by an image processing model, wherein the intermediate features are obtained by the image processing model processing input data, and the input data includes color image features, and the color image features are used to characterize the color image collected during the task execution device performing the target task, and the intermediate features are used to characterize the information in the color image features; a prediction module for inputting the intermediate features, depth image features and noise action features into the action model to obtain predicted action features output by the action model, wherein the predicted action features are used to characterize the motion trajectory of the task execution device in the future time period, and the depth image features and the color image features correspond in time.
[0007] In a fourth aspect, an embodiment of the present application provides a model training device, comprising: an acquisition module for acquiring sample intermediate features output by an image processing model, wherein the sample intermediate features are obtained by the image processing model processing sample input data, and the sample input data includes sample color image features, and the sample color image features are used to characterize the sample color images collected during the task execution device performing the target task, and the sample intermediate features are used to characterize the information in the sample color image features; a prediction module for inputting the sample intermediate features, sample depth image features, and sample noise motion features into a pre-trained motion model to obtain sample predicted motion features output by the pre-trained motion model, wherein the sample predicted motion features are used to characterize the motion trajectory of the task execution device in the target time period, and the sample depth image features and the sample color image features correspond in time; a prediction module for post-training the pre-trained motion model based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model, wherein the target motion features are used to characterize the actual motion trajectory of the task execution device in the target time period.
[0008] In a fifth aspect, an embodiment of the present application provides a task execution device, including a control module, which is used to execute the action feature prediction method described in the first aspect above.
[0009] In a sixth aspect, an embodiment of the present application provides an electronic device comprising: a processor; and a memory for storing processor-executable instructions, wherein the processor is used to execute the motion feature prediction method described in the first aspect or the model training method described in the second aspect.
[0010] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is used to execute the action feature prediction method described in the first aspect or the model training method described in the second aspect.
[0011] In the eighth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor of a computer device, the computer device is able to execute the action feature prediction method described in the first aspect or the model training method described in the second aspect.
[0012] In the ninth aspect, an embodiment of the present application provides a chip comprising: a processor; and a memory for storing processor executable instructions, wherein the processor is used to execute the motion feature prediction method described in the first aspect or the model training method described in the second aspect.
[0013] The embodiments of the present application provide a method for predicting motion features, a task execution device, an electronic device, and a medium. By obtaining intermediate features based on color image features, and inputting the intermediate features, depth image features, and noise motion features into the motion model, the predicted motion features output by the motion model are obtained, thereby combining the scene content information contained in the color image features and the depth information contained in the depth image features to improve the accuracy of the predicted motion features. In the embodiments of the present application, the depth image features can provide depth information or distance information. In particular, under different light intensities, the motion model can obtain more accurate depth information based on the depth image features, thereby improving the model's ability to resist light interference and improving the accuracy of the predicted motion features. Furthermore, controlling the robot to perform the target task based on the predicted motion features can reduce the probability of empty clips, collisions, and the like, thereby improving the execution accuracy and success rate of the target task. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Shown is a schematic diagram of the system architecture of an action feature prediction system provided by an exemplary embodiment of the present application.
[0015] Figure 2 Shown is a flow chart of an action feature prediction method provided by an exemplary embodiment of the present application.
[0016] Figure 3A Shown is a schematic structural diagram of a depth encoder provided by an exemplary embodiment of the present application.
[0017] Figure 3B Shown is a schematic structural diagram of a depth encoder provided by another exemplary embodiment of the present application.
[0018] Figure 3C Shown is a schematic structural diagram of a depth encoder provided by another exemplary embodiment of the present application.
[0019] Figure 4 Shown is a flow chart of an action feature prediction method provided by another exemplary embodiment of the present application.
[0020] Figure 5 Shown is a structural diagram of a vision-language-action model provided by an exemplary embodiment of the present application.
[0021] Figure 6 Shown is a flow chart of a model training method provided by an exemplary embodiment of the present application.
[0022] Figure 7 Shown is a flowchart of a model training method provided by another exemplary embodiment of the present application.
[0023] Figure 8FIG2 is a schematic diagram of the structure of an action feature prediction device provided by an exemplary embodiment of the present application.
[0024] Figure 9 Shown is a structural schematic diagram of a model training device provided by an exemplary embodiment of the present application.
[0025] Figure 10 Shown is a block diagram of an electronic device for executing an action feature prediction method or a model training method provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] Application Overview
[0028] With the vigorous development of artificial intelligence technology, the integration of task execution equipment, such as robots, and artificial intelligence technology continues to deepen. Among them, image processing models (especially vision-language-action models) provide new solutions for robots to perform tasks in complex environments.
[0029] Taking the vision-language-action model as an example, by inputting image features and task instruction features into the vision-language-action model, the model outputs action features for controlling the robot's actions. The vision-language-action model is usually trained based on large-scale color images (such as RGB images). The trained vision-language-action model can predict action features at future moments based on task instruction features and existing image features. However, the action features currently obtained based on this vision-language-action model are prone to poor accuracy, which in turn leads to a low success rate in task execution. For example, during the execution of a task, the robot is prone to empty clips, collisions, sensitivity to light intensity, etc., and in some cases, the robot or other hardware equipment is prone to damage.
[0030] In response to the above technical problems, an embodiment of the present application provides a method for predicting motion features, which obtains intermediate features based on color image features, and inputs the intermediate features, depth image features, and noise motion features into the motion model to obtain predicted motion features output by the motion model, thereby combining the scene content information contained in the color image features and the depth information contained in the depth image features to improve the accuracy of the predicted motion features. In an embodiment of the present application, the depth image features can provide depth information or distance information. In particular, under different light intensities, the motion model can obtain more accurate depth information based on the depth image features, thereby improving the model's ability to resist light interference and improving the accuracy of the predicted motion features. Furthermore, controlling the robot to perform the target task based on the predicted motion features can reduce the probability of empty clips, collisions, and other phenomena, thereby improving the execution accuracy and success rate of the target task.
[0031] Exemplary Systems
[0032] Figure 1 FIG. 1 is a schematic diagram of the system architecture of an action feature prediction system provided by an exemplary embodiment of the present application. Figure 1 As shown, the motion feature prediction system 100 may include a task execution device 110. This task execution device 110 can be a robot, vehicle, drone, or other device, and is used to perform a target task. For example, a robot performs a series of actions to complete a target task such as sweeping the floor, cooking, walking the dog, watering the flowers, serving food, or playing Go. A robot on a production line performs a series of actions to perform target tasks such as welding, assembly, and transportation. A robot in a football game performs a series of actions to complete a target task such as intercepting a pass from an opposing robot. A drone completes its flight mission by adjusting its flight path to avoid collisions with obstacles.
[0033] In one example, an image processing model 111 and an action model 112 may be deployed on the task execution device 110 .
[0034] Furthermore, in one example, the task execution device 110 may be provided with an image acquisition device 113 for acquiring images around the task execution device 110 , such as color images.
[0035] In one example, the task execution device 110 can be used to implement the motion feature prediction method provided in an embodiment of the present application. Specifically, while the task execution device 110 is performing a target task, the task execution device 110 can capture a color image using the image acquisition device 113 and obtain color image features based on the color image. The task execution device 110 can input the color image features into the image processing model 111. The image processing model 111 can understand and extract features from the color image features and output intermediate features, which are used to represent the information in the color image features. Furthermore, the task execution device 110 can input the intermediate features, depth image features, and noise motion features into the motion model 112 to obtain predicted motion features output by the motion model 112. Here, the depth image features and the color image features correspond in time. The predicted motion features can be used to represent the motion trajectory of the task execution device 110 in a future time period. For example, when the task execution device 110 is a robot, the predicted motion features can be used to represent the motion trajectory of the robot's arm end, head end, and / or leg end.
[0036] In one example, the task execution device 110 may control its movements based on the predicted motion features to continue executing the target task. For example, if the task execution device 110 is a robot and the predicted motion features are used to represent the motion trajectory of the robot's end arm, the robot may control the movement of the robot's end arm based on the predicted motion features.
[0037] In other examples, at least one of the image processing model 111 and the motion model 112 can be deployed on a computer device, which can be a server or a terminal device, such as a mobile phone, laptop, or desktop computer. For example, the image processing model 111 can be deployed on a computer device, and the motion model 112 can be deployed on the task execution device 110. The computer device can input color image features into the image processing model 111, obtain intermediate features output by the image processing model 111, and send the intermediate features to the task execution device 110. The task execution device 110 can input the intermediate features, depth image features, and noise motion features into the motion model 112, and obtain predicted motion features output by the motion model 112. For another example, the motion model 112 can be deployed on a computer device, and the image processing model 111 can be deployed on the task execution device 110. The task execution device 110 can input color image features into the image processing model 111, obtain intermediate features output by the image processing model 111, and send the intermediate features to the computer device. The computer device may input the intermediate features, depth image features, and noise motion features into the motion model 112 to obtain the predicted motion features output by the motion model 112, and send the predicted motion features to the task execution device 110. For another example, if both the image processing model 111 and the motion model 112 are deployed on the computer device, the computer device may input the color image features into the image processing model 111 to obtain the intermediate features output by the image processing model 111, input the intermediate features, depth image features, and noise motion features into the motion model 112 to obtain the predicted motion features output by the motion model 112, and send the predicted motion features to the task execution device 110.
[0038] In one example, the task execution device 110 is a vehicle, and the task to be performed by the vehicle is to travel to a target location. The color image features may include features corresponding to the color image collected by the vehicle at the current location. The predicted action features may characterize the motion trajectory of the vehicle in the future time period, and the vehicle may control the vehicle to move from the current location to the target location based on the predicted action features. Alternatively, the task execution device 110 is a dual-arm robot, and the task to be performed by the dual-arm robot is to grab a target object. The color image features may include features corresponding to the color image collected by the dual-arm robot at the current location (or the current moment). The predicted action features may characterize the motion trajectory of the end of the dual-arm robot's arm in the future time period, and the dual-arm robot may control the movement of the dual-arm robot's arm based on the predicted action features to grab the target object.
[0039] Exemplarily, the image acquisition device 113 may include a monocular camera, a binocular camera, or other types of cameras.
[0040] It should be understood that the above application scenario examples are only provided to facilitate understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited thereto. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0041] Exemplary Methods
[0042] Figure 2 Shown is a flow chart of an action feature prediction method provided by an exemplary embodiment of the present application. Figure 2 The method can be Figure 1 The task execution device 110 (such as a robot) or the above-mentioned computer device is executed. For the convenience of description, the following description is given by taking the robot executing the method as an example. Figure 2 As shown, the action feature prediction method may include the following contents.
[0043] 210: Obtain the intermediate features output by the image processing model.
[0044] Specifically, the intermediate features are obtained by processing the input data by the image processing model. The input data includes color image features. The color image features are used to represent the color images collected during the task execution device performing the target task. The intermediate features are used to represent the information in the color image features.
[0045] In one example, an image processing model can be a model used to understand and extract features from images or image features. Images can be represented in the form of image features to facilitate processing by the image processing model. For example, image features can be the image itself, or they can be image features obtained after preprocessing the image. The preprocessing here can include image dimensionality conversion, normalization, or other processing.
[0046] In one example, a color image can record color and brightness in space. The color image can be a three-channel two-dimensional raster image, i.e., an RGB image. In one example, the color image can be acquired by a monocular camera or other type of camera.
[0047] In one example, a camera (such as a monocular camera or other types of cameras) may be installed on the robot. When the robot performs a target task, the robot may use the camera to collect a color image at the current moment.
[0048] In one example, the color image features corresponding to a color image can be input into an image processing model. The image processing model can then interpret and extract the information contained in the color image features to produce intermediate features. Intermediate features can essentially be understood as image features that represent the information contained in the color image features. The emphasis of the intermediate features on this information is determined by the capabilities of the image processing model.
[0049] 220: Input the intermediate features, deep image features, and noise motion features into the motion model to obtain the predicted motion features output by the motion model.
[0050] Specifically, the predicted motion features are used to characterize the motion trajectory of the task execution device in the future period, and the depth image features correspond to the color image features in time.
[0051] In one example, the motion model can also be called an action expert model, i.e., an Action Expert model. The motion model can predict the motion features of the robot in the future time period based on the image features at the current moment, such as the intermediate features. For example, the robot's execution of the target task requires the participation of the target part of the robot, and the target part may include at least one of the arms, head, legs, etc. The intermediate features at the current moment can represent the information in the color image or color image features collected at the current moment. Based on the intermediate features, the action model can obtain the predicted motion features of the target part in the future time period, which is equivalent to predicting the motion trajectory of the target part of the robot in the future time period.
[0052] In one example, the depth image features correspond to the color image features in time, and the depth image features can represent the depth image. For example, the depth image features can be the depth image itself, or image features obtained after the depth image is preprocessed. The preprocessing here can be image dimension conversion processing, normalization processing or other processing.
[0053] Similarly, a depth image can correspond to a color image. The depth image can be single-channel, and each pixel in the depth image can represent a distance or depth value. In one example, while a camera is capturing a color image at the current moment, a depth image at the current moment can be simultaneously acquired using a camera or other means. In one example, the color image corresponds to a target location on the task execution device, and the depth image corresponds to at least a portion of the target location.
[0054] In one example, the noise action feature can be noise such as Gaussian noise or other noise, or can be a feature of the noise after preprocessing. The preprocessing here can be dimensionality conversion, normalization, or other processing. The noise action feature is a motion feature to be predicted that is represented by noise. In one example, the intermediate features, depth image features, and noise action features are input into the motion model. The motion model can denoise the noise action features based on the intermediate features and depth image features to obtain a predicted motion feature.
[0055] In some cases, if the predicted motion features are derived solely based on color image features, the predicted motion features may be inaccurate. For example, color image features do not explicitly provide depth information, making it difficult for the motion model to accurately estimate depth information based on color image features. This is particularly true under varying lighting intensities, where the depth information estimated by the motion model is less accurate, indicating sensitivity to lighting intensity. This results in poorly accurate output of the predicted motion features.
[0056] This embodiment adds depth image features to the motion model input, providing the motion model with relatively accurate depth or distance information, thereby improving the accuracy of predicted motion features. Specifically, the intermediate features serve as prior information, representing information extracted from the input data of the image processing model, such as color image features. Based on this prior information and the depth information provided by the depth image features, the motion model can more accurately denoise noisy motion features, thereby improving the accuracy of predicted motion features.
[0057] The embodiment of the present application provides a method for predicting motion features, which obtains intermediate features based on color image features, and inputs the intermediate features, depth image features, and noise motion features into the motion model to obtain predicted motion features output by the motion model, thereby combining the scene content information contained in the color image features and the depth information contained in the depth image features to improve the accuracy of the predicted motion features. In the embodiment of the present application, the depth image features can provide depth information or distance information. In particular, under different light intensities, the motion model can obtain more accurate depth information based on the depth image features, thereby improving the model's ability to resist light interference and improving the accuracy of the predicted motion features. Furthermore, controlling the robot to perform the target task based on the predicted motion features can reduce the probability of empty clips, collisions, and other phenomena, thereby improving the execution accuracy and success rate of the target task.
[0058] According to an embodiment of the present application, the motion feature prediction method further includes: normalizing the depth image to obtain a grayscale image, wherein the depth image and the color image correspond in time; and obtaining depth image features based on the grayscale image.
[0059] In one example, a grayscale image may refer to an image consisting of only black and white colors, and their transitional colors. The grayscale value of each pixel in the grayscale image represents the brightness of that pixel. For example, a larger grayscale value may represent a higher brightness. The grayscale value range in the grayscale image can be set as needed. For example, the grayscale value range can be 0 to 255, or another suitable range.
[0060] In one example, each pixel in a depth image may represent a distance or a depth value. If the depth value range in the depth image is wide, it is difficult to highlight the difference between different depth values. For example, the depth value range in a depth image is 0 to 2000. If the pixel with a larger depth value in the depth image appears darker, and the pixel with a smaller depth value appears whiter, then the depth values from 1500 to 2000 may all appear very black visually, making it difficult to distinguish the difference between the two. Through normalization, the original depth values contained in the depth image can be uniformly converted into a grayscale value range that is easy to distinguish. If the range of the original depth value is large, then the grayscale value range can be a smaller range than the range of the original depth value. If the range of the original depth value is small, then the grayscale value range can be a larger range than the range of the original depth value.
[0061] In one example, feature extraction may be performed on the grayscale image to obtain depth image features.
[0062] In this embodiment, the depth image is normalized to obtain a grayscale image, and depth image features are obtained based on the grayscale image, so that the depth information in the depth image can be extracted as fully as possible, thereby improving the accuracy of predicting motion features.
[0063] According to one embodiment of the present application, the color image corresponds to the target part on the task execution device, the depth image corresponds to at least part of the target part, and the depth image feature is obtained based on the grayscale image, including: inputting the grayscale image into the encoder to obtain a first coding feature, and inputting the color image in the color image corresponding to the depth image into the encoder to obtain a second coding feature; splicing the first coding feature and the second coding feature to obtain a spliced image feature, and obtaining a depth image feature based on the spliced image feature.
[0064] In one example, an encoder can be used to extract depth information or depth features contained in a grayscale image to obtain depth image features. For example, the encoder can include a visual model, or a large visual model. The visual model can be a Vision Transformer model. In this case, the encoder can be called a Vision Transformer encoder, i.e., a ViT Encoder or a ViT encoder. More specifically, the visual model can be a Dino-v2 model, or a model of another structure or form.
[0065] Specifically, Figure 3A FIG. 1 is a schematic diagram of the structure of a depth encoder provided by an exemplary embodiment of the present application. Figure 3A The depth encoder shown may include a ViT encoder. Figure 3AAs shown, the grayscale image can be input into the ViT encoder to obtain the first encoding feature, and the first encoding feature can be input into the action model as a depth image feature. Alternatively, Figure 3A The depth encoder shown may include a ViT encoder and a multilayer perceptron (MLP). The first encoding feature can be input into the multilayer perceptron for dimension adjustment or dimension conversion, and the dimensionally converted encoding feature can be input into the action model as a deep image feature.
[0066] Furthermore, in one example, the depth image corresponds to at least a portion of the target area. Figure 3A As shown, a grayscale image can be input into a ViT encoder to obtain a first encoded feature, and a color image in the color image corresponding to the depth image (i.e., at least a portion of the color image in the color image) can be input into the ViT encoder to obtain a second encoded feature. The first and second encoded features are then concatenated to obtain a concatenated image feature. The concatenated image feature can be input into the action model as a depth image feature. Alternatively, the concatenated image feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion, and the converted image feature can be input into the action model as a depth image feature. Here, the color image can include an image of a target part of the robot captured by a camera. When the target part includes multiple parts, such as the head, left arm, and right arm, a grayscale image obtained based on the depth image corresponding to at least a portion of the part can be input into the ViT encoder. For example, a grayscale image obtained based on the depth image corresponding to the right arm can be input into the ViT encoder to obtain the first encoded feature. Similarly, the color image corresponding to the right arm can be input into the ViT encoder to obtain the second encoded feature. In other words, at least a portion of the color image in the color image can include a color image corresponding to at least a portion of the part.
[0067] In one example, the predicted motion features can be used to characterize the motion trajectory of a target part of a robot, such as the motion trajectory of an arm, head, and / or leg, in the future. Specifically, the predicted motion features may include the position and posture parameters of the target part at each moment in the future. Furthermore, when the target part includes an arm, the predicted motion features may include the position and posture parameters and the opening and closing state of the arm at each moment in the future.
[0068] In this embodiment, the grayscale image is input into the encoder to obtain a first coding feature, the color image in the color image corresponding to the depth image is input into the encoder in parallel to obtain a second coding feature, the first coding feature and the second coding feature are spliced to obtain a spliced image feature, and the depth image feature is obtained based on the spliced image feature. In this way, the action model can refer to or combine the scene content information contained in the color image in the process of extracting depth information based on the depth image feature, thereby improving the model's ability to resist light interference, and further improving the accuracy of predicting action features.
[0069] In one example, the aforementioned visual model can be a pre-trained large visual model, e.g., one pre-trained using a large number of training samples. The pre-trained large visual model has powerful feature extraction capabilities and can effectively capture key information in grayscale and color images. In one example, the training samples may include training samples other than those corresponding to the target task, or may include only training samples corresponding to other tasks, excluding the target task. This increases the number of training samples for the model and improves its performance.
[0070] According to an embodiment of the present application, the motion feature prediction method further includes: converting the depth image into point cloud data, wherein the depth image and the color image correspond in time; and obtaining depth image features based on the point cloud data.
[0071] In one example, a depth image can be converted into point cloud data using camera intrinsics, such as depth camera intrinsics. For example, the depth value corresponding to each pixel in the depth image can be converted into a point, thereby obtaining the corresponding point cloud data. Feature extraction of the point cloud data can obtain depth image features.
[0072] In this embodiment, by deriving depth image features from point cloud data, a new method for acquiring depth image features can be provided, which can improve the motion model's understanding of depth information. Furthermore, by converting depth images into point cloud data, compared to directly acquiring point cloud data from a camera, the cost of camera equipment can be reduced, and computational complexity can be lowered to a certain extent.
[0073] According to one embodiment of the present application, the color image corresponds to the target part on the task execution device, the depth image corresponds to at least part of the target part, and the depth image feature is obtained based on the point cloud data, including: downsampling the point cloud data, or fusing the point cloud data and the color image corresponding to the depth image in the color image and then downsampling to obtain the downsampled point cloud data; encoding the downsampled point cloud data to obtain a third encoding feature; and obtaining the depth image feature based on the third encoding feature.
[0074] In one example, after converting a depth image into point cloud data, the point cloud data can be downsampled to reduce the number of points in the point cloud data. Derivation of depth image features based on the downsampled point cloud data can increase computation speed and reduce computing resource usage.
[0075] In one example, the downsampling method may be a voxel sampling method. For example, a voxel may be collected at intervals of one or more voxels. In this way, the collected voxels are downsampled voxels, that is, downsampled point cloud data.
[0076] Figure 3B FIG. 1 is a schematic diagram of the structure of a depth encoder provided by another exemplary embodiment of the present application. Figure 3B As shown, the point cloud data is downsampled to obtain the downsampled point cloud data. The downsampled point cloud data is input into an encoder (such as a point cloud encoder), and the encoder extracts the depth information or depth features contained in the downsampled point cloud data to obtain a third encoding feature, which can be used as a depth image feature, that is, Figure 3B The depth encoder shown may include a point cloud encoder. Alternatively, the third encoded feature may be input into a multilayer perceptron (MLP) for dimension adjustment or dimension conversion, and the dimension-converted encoded feature may be input into the action model as a depth image feature, i.e. Figure 3B The depth encoder shown may include a point cloud encoder and a multi-layer perceptron. Of course, the depth encoder may further include a module related to downsampling or other suitable modules.
[0077] Figure 3B The point cloud encoder in can be Point Net, or a model of other structures or forms.
[0078] Optionally, in one example, Figure 3B As shown, the point cloud data and the color image corresponding to the depth image in the color image can be fused and then downsampled to obtain the downsampled point cloud data. Specifically, for the explanation of the color image corresponding to the depth image in the color image, please refer to the description of the relevant part above. In one example, the method of fusing the point cloud data and the color image corresponding to the depth image in the color image may include vector splicing or other suitable methods. The data obtained after fusion can be downsampled. For example, for the point cloud data in the data obtained after fusion, one voxel can be collected at intervals of one or more voxels. Correspondingly, for the color image in the data obtained after fusion, one pixel can be collected at intervals of one or more pixels. In this way, the same position or corresponding position in the point cloud data and the color image can be sampled.
[0079] In this example, the data obtained after downsampling can also be called downsampled point cloud data. The fusion of the color image and the point cloud data is equivalent to increasing the dimension of the point cloud data, that is, using the color image to enrich the point cloud data, which facilitates the subsequent accurate extraction of depth information. The downsampled point cloud data is input into the point cloud encoder, and the depth information or depth features contained in the downsampled point cloud data are extracted by the point cloud encoder to obtain a third encoded feature, which can be used as a depth image feature. Alternatively, the third encoded feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion, and the dimensionally converted encoded feature can be input into the action model as a depth image feature.
[0080] In this embodiment, downsampling can reduce the number of points in the point cloud data, thereby increasing computational speed, reducing computing resource usage, and ultimately improving the robot's responsiveness. Furthermore, in this embodiment, the motion model can reference or incorporate scene content information contained in the color image when extracting depth information based on depth image features, thereby improving the model's resistance to illumination interference and further enhancing the accuracy of predicted motion features.
[0081] According to one embodiment of the present application, the color image corresponds to the target part on the task execution device, the depth image corresponds to at least part of the target part, and the depth image features are obtained based on the point cloud data, including: downsampling the point cloud data to obtain downsampled point cloud data; encoding the downsampled point cloud data to obtain a fourth encoding feature; extracting features of the color image corresponding to the depth image in the color image to obtain at least part of the color image features; and fusing the fourth encoding features and at least part of the color image features through a cross-attention mechanism to obtain a depth image feature.
[0082] Specifically, the specific process of downsampling the point cloud data to obtain the downsampled point cloud data can be found in the description in the above embodiment. The explanation of the color image corresponding to the depth image in the color image can also be found in the description of the relevant part above. To avoid repetition, it will not be repeated here.
[0083] Figure 3C FIG. 1 is a schematic diagram of the structure of a depth encoder provided by another exemplary embodiment of the present application. Figure 3C As shown, the point cloud data is downsampled to obtain downsampled point cloud data. The downsampled point cloud data can be input into an encoder (such as a point cloud encoder) to obtain a fourth encoding feature. Figure 3C The point cloud encoder in can be Point Net, or a model of other structures or forms.
[0084] Furthermore, if Figure 3CAs shown, the color image corresponding to the depth image in the color image (ie, at least a portion of the color image in the color image) can be input into the ViT encoder for feature extraction to obtain at least a portion of the color image features.
[0085] In one example, if Figure 3C As shown, the fourth coding feature and at least part of the color image feature can be fused through a cross-attention mechanism to obtain a depth image feature. For example, the fusion process may include a unidirectional cross-attention calculation, or a bidirectional cross-attention calculation. In one example, the unidirectional cross-attention calculation may be to perform an attention calculation based on the query vector of the fourth coding feature and the value vector of at least part of the color image feature to obtain a depth image feature; or to perform an attention calculation based on the query vector of at least part of the color image feature and the value vector of the fourth coding feature to obtain a depth image feature. In one example, the bidirectional cross-attention calculation may be to perform an attention calculation based on the query vector of the fourth coding feature and the value vector of at least part of the color image feature to obtain an image feature, and to perform an attention calculation based on the query vector of at least part of the color image feature and the value vector of the fourth coding feature to obtain another image feature. The two image features are stitched or otherwise processed as appropriate to obtain a depth image feature. In this example, Figure 3C The depth encoder shown may include a point cloud encoder, a ViT encoder, and a module related to cross-attention calculation. Of course, the depth encoder may further include a module related to downsampling or other suitable modules.
[0086] In this embodiment, three-dimensional point cloud coding and two-dimensional color image coding can be combined to obtain depth image features, which can provide a new way to obtain depth image features, improve the model's ability to resist light interference, and further improve the accuracy of predicted motion features.
[0087] According to one embodiment of the present application, the intermediate features include a key vector and a value vector obtained based on the input data, wherein the intermediate features, the depth image features and the noise motion features are input into the action model to obtain the predicted action features output by the action model, including: using the action model to splice the depth image features and the noise motion features to obtain a query vector; performing cross-attention calculation based on the key vector, the value vector and the query vector to obtain the predicted action features.
[0088] Specifically, the intermediate features obtained by the image processing model based on the input data may include intermediate attention vectors cached by the image processing model, such as key vectors and value vectors, which can be represented as a KV Cache. The key vectors and value vectors can be input into the action model.
[0089] In one example, the deep image features and the noisy motion features can be concatenated, and the concatenated features can be used as a query vector. A cross-attention calculation is performed based on the query vector, the key vector, and the value vector to obtain the predicted motion features. Specifically, the concatenation of the deep image features and the noisy motion features can be performed before or after the deep image features and the noisy motion features are input into the motion model.
[0090] In one example, the motion model can be a Transformer-based model. Intermediate features, deep image features, and noise motion features can serve as input vectors for the motion model. A multi-head self-attention mechanism is used to learn the feature representations of the input vectors in parallel, and outputs predicted motion features.
[0091] In this embodiment, the key vector and value vector obtained by the image processing model based on the input data can serve as prior information. The action model performs attention calculations based on this prior information and the depth image features, inheriting the image processing model's ability to understand the input data. In this way, the action model can combine the prior information and the depth information provided by the depth image features, and after multiple denoising processes, output predicted action features, thereby improving the accuracy of the predicted action features.
[0092] According to one embodiment of the present application, the image processing model includes a visual language model, and the input data also includes task instruction features. The task instruction features are used to describe the target task. The intermediate features are obtained by the visual language model fusing the color image features and the task instruction features. The intermediate features are also used to characterize the information in the task instruction features.
[0093] Specifically, the visual-language model can be a pre-trained visual-language-large model, and the task instruction features can include text features or audio features. Taking text features as an example, the task instruction features can be used to describe the target task, such as grasping a target object. In one example, the pre-trained visual-language-large model is used to extract visual and language features. Through the forward propagation process, intermediate features for the model's attention module, namely the intermediate vector KV cache, are obtained.
[0094] In one example, task instruction features and color image features are input into a visual language model, which can understand and extract features from the task instruction features and color image features to output intermediate features, that is, the intermediate features can be used to represent information in the task instruction features and the color image features.
[0095] In one example, the fusion processing of the color image features and the task instruction features may include at least one of a convolution operation, attention calculation, splicing, or other processing methods.
[0096] In one example, color image features can be obtained by performing preliminary feature extraction based on the color image. For example, the color image can be input into a visual model (such as a Vision Transformer model) for preliminary feature extraction, and the obtained preliminary image features can be used as color image features. Alternatively, the preliminary image features can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion, and the dimensionality converted image features can be used as color image features.
[0097] In this embodiment, the vision-language-action model and the action model form a complete model, namely, a vision-language-action model.
[0098] In this embodiment, the intermediate features are obtained by combining the color image features and the task instruction features, which can improve the richness and accuracy of the intermediate features, thereby improving the accuracy of the predicted action features obtained subsequently.
[0099] Figure 4 Shown is a flow chart of an action feature prediction method provided by another exemplary embodiment of the present application. Figure 4 The embodiment is Figure 2 For the example of embodiment, in order to avoid repetition, the same points can be referred to the description in the above embodiment, which will not be repeated here. Figure 4 As shown, the motion feature prediction method may include the following contents.
[0100] 410: Inputting color images corresponding to multiple parts of the robot into the visual model, and the features output by the visual model are passed through a multi-layer perceptron to obtain color image features.
[0101] Specifically, the target parts that need to be observed during the robot's execution of the target task may include multiple parts, such as the robot's head, left arm, and right arm. The robot's camera can be used to capture color images corresponding to the multiple parts at the current time or at historical moments.
[0102] Performing preliminary feature extraction on these color images can obtain color image features. For example, see Figure 5 The color image can be input into the visual model, and the features output by the visual model can be passed through the multi-layer perceptron to obtain the color image features. It should be understood that the specific content of the visual model can be found in the relevant description of the above embodiment, and to avoid repetition, it will not be repeated here.
[0103] 420: Input the color image features and task instruction features into the visual language model to obtain intermediate features.
[0104] See also Figure 5 , color image features and task instruction features serve as input data of the visual language model, and the intermediate features output by the visual language model can be input into the action model.
[0105] It should be understood that the specific contents of the visual language model and the intermediate features can be found in the relevant descriptions in the above embodiments, and will not be described again here to avoid repetition.
[0106] 430: Obtain depth image features based on the depth image, input the intermediate features, the depth image features, and the noise motion features into the motion model, and obtain the predicted motion features output by the motion model.
[0107] It should be understood that the specific contents of the depth image, depth image features, noise motion features and motion model can be found in the relevant descriptions in the above embodiments. To avoid repetition, they will not be repeated here.
[0108] Specifically, the depth image may include depth images corresponding to at least some of the multiple parts. The acquisition process of the depth image features can refer to the description of the above-mentioned related embodiments, for example, based on Figure 3A 、 Figure 3B or Figure 3C The illustrated embodiment obtains depth image features.
[0109] See also Figure 5 , based on the depth image and the color image corresponding to at least some of the multiple parts (i.e., the color image corresponding to the depth image), the depth image features can be obtained through the encoding process of the depth encoder. The specific structure of the depth encoder can be found in Figure 3A 、 Figure 3B or Figure 3C .
[0110] An embodiment of the present application also provides a model training method, which corresponds to the above-mentioned action feature prediction method. To avoid repetition, the similarities can be referred to the description in the above-mentioned embodiment and will not be repeated here. Figure 6 The method can be Figure 1 The task execution device 110 (such as a robot) or the above-mentioned computer device is executed. For the convenience of description, the following description is given by taking the robot executing the method as an example. Figure 6 As shown, the model training method may include the following contents.
[0111] 610: Obtain sample intermediate features output by the image processing model.
[0112] Specifically, the sample intermediate features are obtained by processing the sample input data by the image processing model. The sample input data includes the sample color image features. The sample color image features are used to characterize the sample color images collected during the task execution device performing the target task. The sample intermediate features are used to characterize the information in the sample color image features.
[0113] In one example, the specific content of the image processing model can be found in the relevant description in the above embodiment, and will not be repeated here to avoid repetition.
[0114] In one example, sample color images are similar to color images. Color images may include images captured in real time during the actual use of the action model, i.e., during inference. Sample color images may also include images captured during the execution of a target task at a historical moment, such as images captured during the execution of the target task via teleoperation or other means. These captured images can serve as training samples. Specifically, sample color images can record color and brightness in space. Sample color images can be three-channel two-dimensional raster images, i.e., RGB images. In one example, sample color images can be captured using a monocular camera or other type of camera.
[0115] In one example, the sample intermediate feature is similar to the intermediate feature. The meaning and acquisition method of the sample intermediate feature can refer to the above description of the meaning and acquisition method of the intermediate feature. To avoid repetition, it will not be repeated here.
[0116] 620: Input the sample intermediate features, the sample depth image features, and the sample noise motion features into the pre-trained motion model to obtain the sample predicted motion features output by the pre-trained motion model.
[0117] Specifically, the sample predicted motion features are used to characterize the motion trajectory of the task execution device in the target period, and the sample depth image features correspond to the sample color image features in time.
[0118] In one example, the model training method provided in this embodiment is equivalent to fine-tuning a pre-trained motion model, i.e., performing post-training. The post-trained motion model can be used to execute the motion feature prediction method provided in any of the above embodiments. For the specific content of the motion model in this embodiment, reference can be made to the relevant description in the above embodiments.
[0119] In one example, the sample depth image feature is similar to the depth image feature, and the sample depth image feature can represent the sample depth image. For example, the sample depth image feature can be the sample depth image itself, or the image feature obtained after the sample depth image is preprocessed. The preprocessing here can be image dimension conversion processing, normalization processing or other processing.
[0120] Similarly, the sample depth image can also correspond to the sample color image. The sample depth image can be single-channel, and each pixel in the sample depth image can represent a distance or depth value. In one example, when the sample color image at the current moment is captured by a camera, the sample depth image at the current moment can be simultaneously obtained by the camera or other means. In one example, the target time period can be located after the current moment.
[0121] In one example, the sample noise motion feature can be noise, such as Gaussian noise or other noise, or can be a feature of the noise after preprocessing, where the preprocessing can include dimensionality conversion, normalization, or other processing. The sample noise motion feature can be a motion feature to be predicted characterized by noise. Alternatively, the sample noise motion feature can be obtained by adding noise to the target motion feature, where the target motion feature is used to represent the actual motion trajectory of the task execution device during the target time period.
[0122] In one example, target motion features, similar to predicted motion features, can be used to characterize the motion trajectory of a target end-point of a robot's target component, such as an arm, head, and / or leg, during a target period. Specifically, the target motion features may include the pose parameters of the target end-point at each moment in the target period. Furthermore, when the target component includes an arm, the target motion features may include the pose parameters and opening and closing states of the arm end at each moment in the target period.
[0123] In one example, the sample intermediate features, sample depth image features, and sample noise motion features are input into a pre-trained motion model. The pre-trained motion model can denoise the sample noise motion features based on the sample intermediate features and sample depth image features to obtain sample predicted motion features. The sample predicted motion features are used to represent the motion features predicted by the pre-trained motion model.
[0124] 630: Post-training the pre-trained motion model based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model.
[0125] Specifically, the target motion feature is used to characterize the actual motion trajectory of the task execution device during the target period.
[0126] In one example, during the post-training of a pre-trained motion model based on the difference between the sample predicted motion features and the target motion features, the image processing model may or may not be trained. For example, the image processing model may be a pre-trained model or an open-source model, and the parameters of the image processing model may not need to be adjusted in the model training method provided in this embodiment, or in other words, the parameters of the image processing model may be frozen.
[0127] In one example, the pre-trained motion model can be obtained by training using a large number of training samples. For example, the large number of training samples can include training samples other than the training samples corresponding to the target task, and may not include the training samples corresponding to the target task, but only include training samples corresponding to other tasks. In this way, the number of training samples of the model can be increased, and the performance of the model can be improved. Specifically, the training samples can include color image features, target motion features, and noise motion features for other tasks. The image processing model can output intermediate features based on the color image features. The intermediate features can be input into the untrained motion model. The noise motion features can also be input into the untrained motion model. In this way, the predicted motion features output by the untrained motion model can be obtained. The untrained motion model can be pre-trained based on the difference between the predicted motion features and the target motion features to obtain a pre-trained motion model.
[0128] Through pre-training, the action model can have certain parameters. Later, the action model can be post-trained by using the training data corresponding to the target task (sample input data, sample depth image features, sample noise action features and target action features), that is, executing steps 610 to 630. This can reduce the sample requirement for the target task and make the action model converge quickly.
[0129] The embodiment of the present application provides a model training method, which performs post-training on a pre-trained action model based on sample color image features and sample depth image features, so that the action model can infer the action features in combination with the scene content information contained in the color image features and the depth information contained in the depth image features, thereby improving the accuracy of the predicted action features output by the action model in the inference stage. Furthermore, the introduction of sample depth image features in the post-training process can shorten the training time of the model, reduce the occupancy of GPU resources, and reduce the demand for a large number of high-quality training samples for target tasks, thereby reducing the training cost. For example, through hundreds of depth image data sequences, within hundreds of card hours of training time, the depth understanding ability of the model can be significantly improved, the success rate of embodied operations is greatly improved, and the model's ability to resist light interference is greatly enhanced.
[0130] In one example, the image processing model may include a visual language model, and the sample input data may also include task instruction features. The task instruction features are used to describe the target task. The sample intermediate features are obtained by the visual language model by fusing the sample color image features and the task instruction features. The sample intermediate features are also used to represent the information in the task instruction features. For the specific content of the visual language model and the task instruction features, please refer to the relevant description in the above embodiments.
[0131] According to one embodiment of the present application, the model training method further includes: normalizing the sample depth image to obtain a sample grayscale image, wherein the sample depth image corresponds to the sample color image in time; and obtaining sample depth image features based on the sample grayscale image.
[0132] In one example, the specific process of obtaining the sample grayscale image by normalizing the sample depth image can refer to the specific process of obtaining the grayscale image by normalizing the depth image mentioned above, and will not be repeated here to avoid repetition.
[0133] In this embodiment, the sample depth image is normalized to obtain a sample grayscale image, and the sample depth image features are obtained based on the sample grayscale image, so that the depth information in the sample depth image can be extracted as fully as possible, thereby improving the accuracy of the predicted action features output by the action model in the subsequent reasoning stage.
[0134] According to one embodiment of the present application, the sample color image corresponds to the target part on the task execution device, the sample depth image corresponds to at least part of the target part, and the sample depth image feature is obtained based on the sample grayscale image, including: inputting the sample grayscale image into the encoder to obtain the sample first coding feature, and inputting the sample color image corresponding to the sample depth image in the sample color image into the encoder to obtain the sample second coding feature; splicing the sample first coding feature and the sample second coding feature to obtain the sample splicing image feature, and obtaining the sample depth image feature based on the sample splicing image feature.
[0135] In one example, the encoder may include a visual model, or a visual big model, and the visual model may be a Vision Transformer model. In this case, the encoder may be called a Vision Transformer encoder, i.e., a ViTEncoder or a ViT encoder. More specifically, the visual model may be a Dino-v2 model, or may be a model of other structures or forms. In one example, the visual model may be a pre-trained visual big model, such as one obtained by pre-training with a large number of training samples. The pre-trained visual big model has a powerful feature extraction capability and can effectively capture key information in sample grayscale images and sample color images. In one example, the training samples may be training samples other than the training samples corresponding to the target task, and may not include the training samples corresponding to the target task, but only include training samples corresponding to other tasks. In this way, the number of training samples of the model can be increased, and the performance of the model can be improved.
[0136] Specifically, the sample depth image corresponds to at least a portion of the target part. Figure 3AThe sample grayscale image can be input into the ViT encoder to obtain the sample first encoded feature, which can be used as the sample deep image feature to input into the pre-trained motion model. Alternatively, the sample first encoded feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimension conversion, and the dimensionally converted encoded feature can be used as the sample deep image feature to input into the pre-trained motion model.
[0137] Furthermore, in one example, Figure 3A As shown, a sample grayscale image can be input into a ViT encoder to obtain a sample first encoded feature, and a sample color image corresponding to a sample depth image in a sample color image (i.e., at least a portion of the sample color image in the sample color image) can be input into the ViT encoder to obtain a sample second encoded feature. The sample first encoded feature and the sample second encoded feature are then concatenated to obtain a sample concatenated image feature. The sample concatenated image feature can be input into a pre-trained motion model as a sample depth image feature. Alternatively, the sample concatenated image feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion. The dimensionality-converted image feature can be input into the pre-trained motion model as a sample depth image feature. Here, the sample color image can include an image of a target part of the robot captured by a camera. When the target part includes multiple parts, such as the head, left arm, and right arm, a sample grayscale image obtained based on the sample depth image corresponding to at least a portion of the part can be input into the ViT encoder. For example, a sample grayscale image obtained based on the sample depth image corresponding to the right arm can be input into the ViT encoder to obtain a sample first encoded feature. Similarly, the sample color image corresponding to the right arm can be input into the ViT encoder to obtain a sample second encoded feature. That is, at least part of the sample color images in the sample color image may include the sample color image corresponding to at least part of the portion.
[0138] In this embodiment, the sample grayscale image is input into the encoder to obtain the sample first coding feature, the sample color image corresponding to the sample depth image in the sample color image is input into the encoder in parallel to obtain the sample second coding feature, the sample first coding feature and the sample second coding feature are spliced to obtain the sample spliced image feature, and the sample depth image feature is obtained based on the sample spliced image feature. In this way, the trained motion model can refer to or combine the scene content information contained in the color image in the process of extracting depth information based on the depth image feature, thereby improving the model's ability to resist light interference, and further improving the accuracy of predicted motion features.
[0139] According to one embodiment of the present application, the training method of the model also includes: converting the sample depth image into sample point cloud data, wherein the sample depth image and the sample color image correspond in time; and obtaining the sample depth image features based on the sample point cloud data.
[0140] In one example, a sample depth image can be converted into sample point cloud data using camera intrinsic parameters, such as depth camera intrinsic parameters. For example, the depth value corresponding to each pixel in the sample depth image can be converted into a point, thereby obtaining the corresponding sample point cloud data. Feature extraction of the sample point cloud data can obtain sample depth image features.
[0141] In this embodiment, by obtaining sample depth image features based on sample point cloud data, a new method for acquiring sample depth image features can be provided, which can improve the motion model's understanding of depth information. Furthermore, by converting sample depth images into sample point cloud data, compared to directly acquiring sample point cloud data from a camera, the cost of camera equipment can be reduced, and computational complexity can be reduced to a certain extent.
[0142] According to one embodiment of the present application, the sample color image corresponds to the target part on the task execution device, the sample depth image corresponds to at least part of the target part, and the sample depth image feature is obtained based on the sample point cloud data, including: downsampling the sample point cloud data, or fusing the sample point cloud data and the sample color image corresponding to the sample depth image in the sample color image and then downsampling to obtain the downsampled sample point cloud data; encoding based on the downsampled sample point cloud data to obtain the sample third coding feature; and obtaining the sample depth image feature based on the sample third coding feature.
[0143] Specifically, the specific process of downsampling the sample point cloud data can refer to the specific process of downsampling the point cloud data described above, and will not be described again here to avoid repetition.
[0144] In one example, reference Figure 3B , downsample the sample point cloud data to obtain downsampled sample point cloud data. The downsampled sample point cloud data is input into a point cloud encoder, which extracts depth information or depth features contained in the downsampled sample point cloud data to obtain a sample third encoded feature. The sample third encoded feature can be used as a sample depth image feature. Alternatively, the sample third encoded feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion. The dimensionally converted encoded feature can be input into the motion model as a sample depth image feature.
[0145] Optionally, in one example, Figure 3BAs shown, the sample point cloud data and the sample color image corresponding to the sample depth image in the sample color image can be fused and then downsampled to obtain the downsampled sample point cloud data. In one example, the specific method for fusing the sample point cloud data and the sample color image corresponding to the sample depth image in the sample color image can refer to the specific method for fusing the point cloud data and the color image corresponding to the depth image in the color image. The specific process for downsampling the data obtained after fusion can also refer to the relevant description in the above embodiment. To avoid repetition, it is not repeated here.
[0146] In this example, the data obtained after downsampling can also be referred to as downsampled sample point cloud data. The downsampled sample point cloud data is input into a point cloud encoder, and the point cloud encoder extracts the depth information or depth features contained in the downsampled sample point cloud data to obtain a sample third encoded feature. The sample third encoded feature can be used as a sample depth image feature. Alternatively, the sample third encoded feature can be input into a multilayer perceptron (MLP) for dimensionality adjustment or dimensionality conversion. The dimensionally converted encoded feature can be input into the motion model as a sample depth image feature.
[0147] In this embodiment, downsampling can reduce the number of points in the sample point cloud data, thereby increasing computational speed, reducing computing resource usage, and improving training efficiency. Furthermore, in this embodiment, the motion model can reference or incorporate scene content information contained in the sample color image when extracting depth information based on the sample depth image features. This improves the motion model's ability to withstand illumination interference and further enhances the accuracy of the predicted motion features output by the motion model.
[0148] According to one embodiment of the present application, the sample color image corresponds to the target part on the task execution device, the sample depth image corresponds to at least part of the target part, and the sample depth image feature is obtained based on the sample point cloud data, including: downsampling the sample point cloud data to obtain the downsampled sample point cloud data; encoding the downsampled sample point cloud data to obtain the sample fourth encoding feature; extracting features of the sample color image corresponding to the sample depth image in the sample color image to obtain at least part of the sample color image feature; and fusing the sample fourth encoding feature and at least part of the sample color image feature through a cross-attention mechanism to obtain the sample depth image feature.
[0149] In one example, reference Figure 3C , downsample the sample point cloud data to obtain the downsampled sample point cloud data. The downsampled sample point cloud data can be input into the point cloud encoder to obtain the fourth encoded feature of the sample.
[0150] Further, refer to Figure 3C The sample color image corresponding to the sample depth image in the sample color image (that is, at least part of the sample color image in the sample color image) can be input into the ViT encoder for feature extraction to obtain at least part of the sample color image features.
[0151] In one example, reference Figure 3C The fourth encoded feature of the sample and at least part of the sample color image feature can be fused through a cross-attention mechanism to obtain the sample depth image feature. The specific process of fusing the fourth encoded feature of the sample and at least part of the sample color image feature to obtain the sample depth image feature can be referred to the specific process of fusing the fourth encoded feature and at least part of the color image feature to obtain the depth image feature. To avoid repetition, it is not repeated here.
[0152] In this embodiment, the three-dimensional sample point cloud encoding and the two-dimensional sample color image encoding can be combined to obtain the sample depth image features, which can provide a new way to obtain the sample depth image features, and can improve the model's ability to resist light interference, thereby further improving the accuracy of the predicted motion features output by the motion model.
[0153] In one example, the sample intermediate features may include a sample key vector and a sample value vector obtained based on the sample input data, wherein the sample intermediate features, the sample depth image features, and the sample noise motion features are input into a pre-trained motion model to obtain a sample predicted motion feature output by the pre-trained motion model, including: using the pre-trained motion model to splice the sample depth image features and the sample noise motion features to obtain a sample query vector; performing cross-attention calculation based on the sample key vector, the sample value vector, and the sample query vector to obtain the sample predicted motion feature. The specific process of performing cross-attention calculation based on the sample key vector, the sample value vector, and the sample query vector to obtain the sample predicted motion feature can refer to the above-mentioned specific process of obtaining the predicted motion feature based on the cross-attention calculation based on the key vector, the value vector, and the query vector. To avoid repetition, it will not be repeated here.
[0154] According to one embodiment of the present application, the model training method also includes: using multiple methods to obtain sample depth image features based on the sample depth image, wherein the sample depth image corresponds to the sample color image in time; evaluating the trained motion model obtained based on the sample depth image features obtained by each of the multiple methods to obtain an evaluation result.
[0155] Specifically, multiple pre-trained motion models can be prepared, and the pre-trained motion models can be post-trained using multiple methods, with the multiple methods corresponding to the multiple pre-trained motion models.
[0156] For example, there can be multiple image processing models (such as a visual language model) and multiple pre-trained motion models, with a one-to-one correspondence between the multiple image processing models and the multiple pre-trained motion models, that is, multiple sets of image processing models and pre-trained motion models. For a set of image processing models and pre-trained motion models, the intermediate features of the samples output by the image processing model can be input into the pre-trained motion model. Furthermore, for a set of image processing models and pre-trained motion models, a method can be used to post-train the pre-trained motion model.
[0157] Alternatively, there can be one image processing model and multiple motion models, and the sample intermediate features output by the image processing model can be input into multiple pre-trained motion models. Moreover, for each pre-trained motion model, a method can be used to post-train the pre-trained motion model.
[0158] Specifically, post-training a pre-trained motion model using a method may include: obtaining sample depth image features based on a sample depth image using the method (e.g., an encoding method), inputting the sample depth image features, the sample noise motion features, and the sample intermediate features output by the image processing model into the pre-trained motion model for post-training, thereby obtaining a trained motion model.
[0159] The training samples corresponding to the various methods (such as sample color images, sample depth images, sample noise motion features, and target motion features) may be the same, that is, they may be training samples for the same target task.
[0160] In one example, the trained motion models obtained by various methods can be evaluated manually or by algorithms to obtain evaluation results corresponding to each motion model, so as to facilitate the selection of appropriate motion models for the actual reasoning process based on the evaluation results.
[0161] Furthermore, the model training method also includes: determining a target motion model from multiple trained motion models based on the evaluation results of the trained motion model obtained by each of the multiple methods, wherein the multiple trained motion models correspond one-to-one to the multiple methods.
[0162] In one example, the evaluation result can be used to characterize the task completion of the action model. A higher task completion indicates a better evaluation result. The target action model can be the action model with the best evaluation result. This target action model can be used in the actual reasoning process, i.e., to actually execute the target task.
[0163] In this embodiment, multiple methods are used to obtain sample depth image features based on sample depth images, and the trained motion models obtained based on the sample depth image features obtained by each of the multiple methods are evaluated to obtain evaluation results. This facilitates the selection of target motion models with better performance based on the evaluation results, thereby improving the success rate of subsequent target tasks.
[0164] According to an embodiment of the present application, the sample color image corresponds to the target part of the task execution device, and the sample depth image corresponds to at least part of the target part. The multiple methods include at least two of the following methods:
[0165] In the first method, the sample depth image is normalized to obtain a sample grayscale image; and the sample depth image features are obtained based on the sample grayscale image.
[0166] Specifically, the specific content of the first method can refer to the relevant description in the above embodiment. In one example, the encoder structure corresponding to the first method can refer to Figure 3A The structure of the deep encoder is shown.
[0167] The second method converts the sample depth image into sample point cloud data; downsamples the sample point cloud data, or fuses the sample point cloud data and the sample color image corresponding to the sample depth image in the sample color image and then downsamples to obtain the downsampled sample point cloud data; encodes the downsampled sample point cloud data to obtain the sample third encoding feature; and obtains the sample depth image feature based on the sample third encoding feature.
[0168] Specifically, the specific content of the second method can refer to the relevant description in the above embodiment. In one example, the encoder structure corresponding to the second method can refer to Figure 3B The structure of the deep encoder is shown.
[0169] The third method converts the sample depth image into sample point cloud data; downsamples and encodes the sample point cloud data to obtain the sample fourth encoding feature; extracts features of the sample color image corresponding to the sample depth image in the sample color image to obtain at least part of the sample color image feature; and fuses the sample fourth encoding feature and at least part of the sample color image feature through a cross-attention mechanism to obtain the sample depth image feature.
[0170] Specifically, the specific content of the third method can refer to the relevant description in the above embodiment. In one example, the encoder structure corresponding to the third method can refer to Figure 3C The structure of the deep encoder is shown.
[0171] In an embodiment of the present application, the depth encoder is decoupled from the pre-trained visual language model and connected to the action model, which makes it convenient to flexibly process depth image data based on a variety of different depth encoder structures.
[0172] Figure 7 Shown is a flowchart of a model training method provided by another exemplary embodiment of the present application. Figure 7 The embodiment is Figure 6 For the example of embodiment, in order to avoid repetition, the same points can be referred to the description in the above embodiment, which will not be repeated here. Figure 7 As shown, the model training method may include the following contents.
[0173] 710: Input the sample color images corresponding to multiple parts of the robot into the visual model, and the features output by the visual model are passed through a multi-layer perceptron to obtain the sample color image features.
[0174] Specifically, the sample color image is subjected to preliminary feature extraction to obtain the sample color image features. The specific process of obtaining the sample color image features can be referred to above. Figure 4 The specific process of obtaining color image features in the embodiment.
[0175] Figure 7 The model structure corresponding to the model training method can refer to the above Figure 5 The structure shown.
[0176] 720: Input the sample color image features and task instruction features into the visual language model to obtain the sample intermediate features.
[0177] 730: Obtain a sample depth image feature based on the sample depth image using a method.
[0178] 740: Input the sample intermediate features, the sample noise motion features, and the sample depth image features into the pre-trained motion model to obtain the sample predicted motion features output by the pre-trained motion model.
[0179] 750: Post-training the pre-trained motion model based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model.
[0180] 760: Evaluate the motion models that are post-trained based on the sample depth image features obtained by multiple methods to obtain evaluation results.
[0181] 770: Determine a target motion model from the multiple trained motion models based on the evaluation results of the trained motion models obtained by each of the multiple methods.
[0182] Specifically, the various methods in this embodiment can refer to the relevant descriptions in the above embodiments, and to avoid repetition, they will not be described again here.
[0183] In one example, the multiple methods refer to multiple encoding methods, such as including three encoding methods, which may correspond to three encoder structures, for example, Figure 3A 、 Figure 3B as well as Figure 3C The deep encoder structure shown.
[0184] Each encoder structure can form an overall model structure with the visual language model and the action model, for example Figure 5 The model structure shown.
[0185] For each model structure, hundreds of data points with depth images can be used to post-train the vision-language-action model. For example, testing or evaluation is performed on three target tasks: operating a microwave oven, cleaning stains, and bagging goods in a supermarket. The evaluation results are shown in Table 1 below.
[0186] The numbers in Table 1 are the evaluation results, which represent the degree of task completion. The baseline model can be Figure 5 In the model structure shown, the model structure that removes the input of the depth image feature can input the intermediate features and the noise motion features into the motion model to obtain the predicted motion features output by the motion model.
[0187] Table 1 Evaluation results of the models based on three encoder structures
[0188]
[0189] The evaluation results in Table 1 show that post-training based on depth images can flexibly improve the depth understanding ability of the vision-language-action model, and can significantly improve the accuracy of the model's embodied operations and its ability to resist light interference. Specifically, the task completion of the model obtained by post-training based on depth images is higher than that of the model obtained by post-training without the participation of depth images. Moreover, the encoder structures with high task completion corresponding to different target tasks may be different. For example, the supermarket packaging task has the highest task completion of the model obtained by post-training the action model with the third or third encoder structure, while the microwave oven operation task has the highest task completion of the model obtained by post-training the action model with the second encoder structure.
[0190] The method provided in the embodiment of the present application can improve the operational accuracy and success rate of the robot in the execution of complex tasks by introducing depth images into the model post-training process. Moreover, the method provided in the embodiment of the present application is based on the pre-training model, introduces depth images in the post-training stage, and provides flexible and diverse depth map encoding methods (adaptable to a variety of target tasks). In this way, through hundreds of depth image data sequences, within hundreds of card hours of training time, the model's depth understanding ability can be significantly improved, the success rate of embodied operations is greatly improved, and the model's ability to resist light interference is greatly enhanced.
[0191] Exemplary devices
[0192] Figure 8 FIG. 1 is a schematic diagram of the structure of an action feature prediction device provided by an exemplary embodiment of the present application. Figure 8 As shown, the motion feature prediction device 800 includes: an acquisition module 810 and a prediction module 820.
[0193] Acquisition module 810 is used to obtain intermediate features output by the image processing model. Intermediate features are obtained by the image processing model by processing input data. The input data includes color image features, which are used to represent color images captured during the task execution device's execution of the target task. Intermediate features are used to represent the information contained in the color image features. Prediction module 820 is used to input the intermediate features, depth image features, and noise motion features into the action model to obtain predicted motion features output by the action model. The predicted motion features are used to represent the motion trajectory of the task execution device in the future. The depth image features correspond to the color image features in time.
[0194] The embodiment of the present application provides an action feature prediction device, which obtains intermediate features based on color image features, and inputs the intermediate features, depth image features and noise action features into the action model to obtain predicted action features output by the action model, thereby combining the scene content information contained in the color image features and the depth information contained in the depth image features to improve the accuracy of the predicted action features. In the embodiment of the present application, the depth image features can provide depth information or distance information. Especially under different light intensities, the action model can obtain more accurate depth information based on the depth image features, which can improve the model's ability to resist light interference and improve the accuracy of the predicted action features. Furthermore, controlling the robot to perform the target task based on the predicted action features can reduce the probability of empty clips, collisions and other phenomena, thereby improving the execution accuracy and success rate of the target task.
[0195] According to an embodiment of the present application, the acquisition module 810 is further configured to: normalize the depth image to obtain a grayscale image, wherein the depth image and the color image correspond in time; and obtain depth image features based on the grayscale image.
[0196] According to one embodiment of the present application, the color image corresponds to a target portion on the task execution device, and the depth image corresponds to at least a portion of the target portion. Acquisition module 810 is configured to: input the grayscale image into an encoder to obtain a first encoded feature, and input the color image corresponding to the depth image into an encoder to obtain a second encoded feature; concatenate the first encoded feature and the second encoded feature to obtain a concatenated image feature, and obtain a depth image feature based on the concatenated image feature.
[0197] According to an embodiment of the present application, the acquisition module 810 is further used to: convert the depth image into point cloud data, wherein the depth image and the color image correspond in time; and obtain depth image features based on the point cloud data.
[0198] According to one embodiment of the present application, the color image corresponds to a target portion on the task execution device, and the depth image corresponds to at least a portion of the target portion. Acquisition module 810 is configured to: downsample the point cloud data, or fuse the point cloud data with a color image in the color image that corresponds to the depth image and then downsample the resultant point cloud data; encode the downsampled point cloud data to obtain a third encoded feature; and obtain a depth image feature based on the third encoded feature.
[0199] According to one embodiment of the present application, the color image corresponds to a target portion on the task execution device, and the depth image corresponds to at least a portion of the target portion. Acquisition module 810 is configured to: downsample the point cloud data to obtain downsampled point cloud data; encode the downsampled point cloud data to obtain a fourth encoded feature; extract features from the color image corresponding to the depth image in the color image to obtain at least a portion of the color image features; and fuse the fourth encoded feature with at least a portion of the color image features using a cross-attention mechanism to obtain a depth image feature.
[0200] According to one embodiment of the present application, the intermediate features include a key vector and a value vector obtained based on the input data, wherein the prediction module 820 is used to: use the motion model to splice the depth image features and the noise motion features to obtain a query vector; perform cross-attention calculation based on the key vector, the value vector and the query vector to obtain the predicted motion features.
[0201] According to one embodiment of the present application, the image processing model includes a visual language model, and the input data also includes task instruction features. The task instruction features are used to describe the target task. The intermediate features are obtained by the visual language model fusing the color image features and the task instruction features. The intermediate features are also used to characterize the information in the task instruction features.
[0202] It should be understood that the operations and functions of the acquisition module 810 and the prediction module 820 in the above embodiment can refer to the above Figure 2 or Figure 4 To avoid repetition, the description of the motion feature prediction method provided in the embodiment will not be repeated here.
[0203] Figure 9 The figure shows a schematic diagram of the structure of a model training device provided by an exemplary embodiment of the present application. Figure 9 As shown, the model training device 900 includes: an acquisition module 910, a prediction module 920 and a training module 930.
[0204] Acquisition module 910 is used to acquire sample intermediate features output by the image processing model, wherein the sample intermediate features are obtained by the image processing model by processing sample input data, wherein the sample input data includes sample color image features, which are used to represent the sample color images collected during the task execution device's execution of the target task, and the sample intermediate features are used to represent the information in the sample color image features. Prediction module 920 is used to input the sample intermediate features, sample depth image features, and sample noise motion features into a pre-trained motion model to obtain sample predicted motion features output by the pre-trained motion model, wherein the sample predicted motion features are used to represent the motion trajectory of the task execution device during the target time period, and the sample depth image features and the sample color image features correspond in time. Prediction module 830 is used to post-train the pre-trained motion model based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model, wherein the target motion features are used to represent the actual motion trajectory of the task execution device during the target time period.
[0205] The embodiment of the present application provides a model training device, which performs post-training on a pre-trained action model based on sample color image features and sample depth image features, so that the action model can infer the action features in combination with the scene content information contained in the color image features and the depth information contained in the depth image features, thereby improving the accuracy of the predicted action features output by the action model in the inference stage. Furthermore, the introduction of sample depth image features in the post-training process can shorten the training time of the model, reduce the occupancy of GPU resources, and reduce the demand for a large number of high-quality training samples for target tasks, thereby reducing the training cost. For example, through hundreds of depth image data sequences, within hundreds of card hours of training time, the model's depth understanding ability can be significantly improved, the success rate of embodied operations is greatly improved, and the model's ability to resist light interference is greatly enhanced.
[0206] According to an embodiment of the present application, the acquisition module 910 is further used to: obtain a sample grayscale image by normalizing the sample depth image, wherein the sample depth image and the sample color image correspond in time; and obtain sample depth image features based on the sample grayscale image.
[0207] According to one embodiment of the present application, the sample color image corresponds to a target portion on the task execution device, and the sample depth image corresponds to at least a portion of the target portion. Acquisition module 910 is configured to: input the sample grayscale image into an encoder to obtain a sample first encoded feature, and input the sample color image corresponding to the sample depth image in the sample color image into an encoder to obtain a sample second encoded feature; concatenate the sample first encoded feature and the sample second encoded feature to obtain a sample concatenated image feature, and obtain a sample depth image feature based on the sample concatenated image feature.
[0208] According to an embodiment of the present application, the acquisition module 910 is further used to: convert the sample depth image into sample point cloud data, wherein the sample depth image and the sample color image correspond in time; and obtain sample depth image features based on the sample point cloud data.
[0209] According to one embodiment of the present application, the sample color image corresponds to a target portion on the task execution device, and the sample depth image corresponds to at least a portion of the target portion. Acquisition module 910 is configured to: downsample the sample point cloud data, or fuse the sample point cloud data with a sample color image corresponding to the sample depth image in the sample color image and then downsample the sample point cloud data to obtain downsampled sample point cloud data; encode the downsampled sample point cloud data to obtain a sample third encoded feature; and obtain a sample depth image feature based on the sample third encoded feature.
[0210] According to one embodiment of the present application, the sample color image corresponds to a target portion on the task execution device, and the sample depth image corresponds to at least a portion of the target portion. Acquisition module 910 is configured to: downsample the sample point cloud data to obtain downsampled sample point cloud data; encode the downsampled sample point cloud data to obtain a fourth encoded sample feature; extract features from the sample color image corresponding to the sample depth image in the sample color image to obtain at least a portion of the sample color image features; and fuse the fourth encoded sample feature and at least a portion of the sample color image features using a cross-attention mechanism to obtain a sample depth image feature.
[0211] According to one embodiment of the present application, acquisition module 910 is further configured to obtain sample depth image features based on the sample depth image using multiple methods, where the sample depth image and the sample color image correspond in time. Model training device 900 also includes evaluation module 940, configured to evaluate the trained motion model obtained based on the sample depth image features obtained by each of the multiple methods to obtain an evaluation result.
[0212] According to one embodiment of the present application, the evaluation module 940 is also used to: determine the target motion model from multiple trained motion models based on the evaluation results of the trained motion model obtained by each method in multiple methods, wherein the multiple trained motion models correspond one-to-one to the multiple methods.
[0213] According to one embodiment of the present application, the sample color image corresponds to a target part on the task execution device, and the sample depth image corresponds to at least part of the target part. The various methods include at least two of the following methods: a first method, normalizing the sample depth image to obtain a sample grayscale image; and obtaining sample depth image features based on the sample grayscale image. A second method, converting the sample depth image into sample point cloud data; downsampling the sample point cloud data, or fusing the sample point cloud data and the sample color image corresponding to the sample depth image in the sample color image and then downsampling to obtain downsampled sample point cloud data; encoding the downsampled sample point cloud data to obtain a sample third encoding feature; and obtaining a sample depth image feature based on the sample third encoding feature. A third method, converting the sample depth image into sample point cloud data; downsampling and encoding the sample point cloud data to obtain a sample fourth encoding feature; extracting features from the sample color image corresponding to the sample depth image in the sample color image to obtain at least part of the sample color image features; and fusing the sample fourth encoding feature and at least part of the sample color image features through a cross-attention mechanism to obtain a sample depth image feature.
[0214] It should be understood that the operations and functions of the acquisition module 910, prediction module 920, training module 930, and evaluation module 940 in the above embodiment can refer to the above Figure 6 or Figure 7 To avoid repetition, the description of the model training method provided in the embodiment will not be repeated here.
[0215] An embodiment of the present application further provides a task execution device, which includes a control module, and the control module is used to execute the action feature prediction method provided by any of the above embodiments.
[0216] In some embodiments, the task performance device may be a robot, such as a dual-arm robot, or may be a vehicle or other device that can be used to perform a specific task.
[0217] The operation and functions of the task execution device provided in the embodiment of the present application can refer to the above Figure 2 or Figure 4 To avoid repetition, the description of the motion feature prediction method provided in the embodiment will not be repeated here.
[0218] Figure 10 FIG2 is a block diagram of an electronic device 1000 for executing an action feature prediction method or a model training method according to an exemplary embodiment of the present application. The electronic device 1000 may be a server, a task execution device, a control device for a task execution device, a server interacting with a task execution device, or other device.
[0219] Reference Figure 10 The electronic device 1000 includes a processing component 1010, which further includes one or more processors, and a memory resource represented by a memory 1020 for storing instructions executable by the processing component 1010, such as an application. The application stored in the memory 1020 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1010 is configured to execute the instructions to perform the above-mentioned motion feature prediction method or model training method.
[0220] The electronic device 1000 may further include a power supply component configured to perform power management of the electronic device 1000, a wired or wireless network interface configured to connect the electronic device 1000 to a network, and an input / output (I / O) interface. The electronic device 1000 may be operated based on an operating system stored in the memory 1020, such as Windows Server 2003. TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or similar.
[0221] A non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the above-mentioned electronic device 1000, enables the above-mentioned electronic device 1000 to perform an action feature prediction method or a model training method.
[0222] A computer program product includes a computer program. When the computer program is executed by a processor of a computer device, the computer device is enabled to execute the action feature prediction method or model training method provided in any of the above embodiments.
[0223] All of the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, and will not be described in detail here.
[0224] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0225] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0226] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0227] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0228] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0229] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program check codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0230] It should be noted that, in the description of this application, the terms "first," "second," "third," etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0231] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0232] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A motion feature prediction method, characterized in that: include: Obtaining intermediate features output by the image processing model, wherein the intermediate features are obtained by the image processing model by processing input data, the input data including color image features, the color image features being used to characterize a color image acquired during the process of the task execution device performing a target task, and the intermediate features being used to characterize information in the color image features; The intermediate features, depth image features and noise motion features are input into the motion model to obtain the predicted motion features output by the motion model, wherein the predicted motion features are used to characterize the motion trajectory of the task execution device in the future time period, and the depth image features correspond to the color image features in time.
2. The motion feature prediction method according to claim 1, wherein: Also includes: Normalizing the depth image to obtain a grayscale image, wherein the depth image and the color image correspond in time; The depth image feature is obtained based on the grayscale image.
3. The motion feature prediction method according to claim 2, wherein: The color image corresponds to a target part on the task execution device, the depth image corresponds to at least a portion of the target part, and obtaining the depth image feature based on the grayscale image includes: Inputting the grayscale image into an encoder to obtain a first encoding feature, and inputting a color image in the color image corresponding to the depth image into the encoder to obtain a second encoding feature; The first coding feature and the second coding feature are spliced together to obtain a spliced image feature, and the depth image feature is obtained based on the spliced image feature.
4. The motion feature prediction method according to claim 1, wherein: Also includes: Converting a depth image into point cloud data, wherein the depth image and the color image correspond in time; The depth image features are obtained based on the point cloud data.
5. The motion feature prediction method according to claim 4, characterized in that: The color image corresponds to a target part on the task execution device, the depth image corresponds to at least a portion of the target part, and obtaining the depth image features based on the point cloud data includes: Downsampling the point cloud data, or fusing the point cloud data and a color image in the color image corresponding to the depth image and then downsampling the point cloud data to obtain downsampled point cloud data; Encoding the downsampled point cloud data to obtain a third encoding feature; The depth image feature is obtained based on the third encoding feature.
6. The motion feature prediction method according to claim 4, wherein: The color image corresponds to a target part on the task execution device, the depth image corresponds to at least a portion of the target part, and obtaining the depth image features based on the point cloud data includes: Downsampling the point cloud data to obtain downsampled point cloud data; Encoding the downsampled point cloud data to obtain a fourth encoding feature; Performing feature extraction on a color image in the color image corresponding to the depth image to obtain at least part of the color image features; The fourth encoding feature and the at least part of the color image feature are fused through a cross-attention mechanism to obtain the depth image feature.
7. The motion feature prediction method according to claim 1, wherein: The intermediate features include a key vector and a value vector obtained based on the input data, The step of inputting the intermediate features, the depth image features, and the noise motion features into the motion model to obtain the predicted motion features output by the motion model includes: Using the motion model, concatenate the depth image features and the noise motion features to obtain a query vector; A cross-attention calculation is performed according to the key vector, the value vector, and the query vector to obtain the predicted action feature.
8. The motion feature prediction method according to any one of claims 1 to 7, characterized in that: The image processing model includes a visual language model, and the input data also includes task instruction features, which are used to describe the target task. The intermediate features are obtained by the visual language model fusing the color image features and the task instruction features, and the intermediate features are also used to characterize the information in the task instruction features.
9. A model training method, characterized in that: include: Obtaining sample intermediate features output by the image processing model, wherein the sample intermediate features are obtained by the image processing model by processing sample input data, the sample input data including sample color image features, the sample color image features being used to characterize sample color images collected during the process of the task execution device performing the target task, and the sample intermediate features being used to characterize information in the sample color image features; Inputting the sample intermediate features, sample depth image features, and sample noise motion features into a pre-trained motion model to obtain sample predicted motion features output by the pre-trained motion model, wherein the sample predicted motion features are used to characterize the motion trajectory of the task execution device in the target time period, and the sample depth image features correspond to the sample color image features in time; The pre-trained motion model is post-trained based on the difference between the sample predicted motion features and the target motion features to obtain a trained motion model, wherein the target motion features are used to characterize the actual motion trajectory of the task execution device in the target time period.
10. The model training method according to claim 9, characterized in that: Also includes: Obtaining the sample depth image features based on the sample depth image using a plurality of methods, wherein the sample depth image and the sample color image correspond in time; The trained motion model obtained based on the sample depth image features obtained by each of the multiple methods is evaluated to obtain an evaluation result.
11. The model training method according to claim 10, characterized in that: Also includes: According to the evaluation results of the trained motion models obtained by each of the multiple methods, a target motion model is determined from the multiple trained motion models, wherein the multiple trained motion models correspond one-to-one to the multiple methods.
12. The model training method according to claim 10 or 11, characterized in that: The sample color image corresponds to the target part of the task execution device, the sample depth image corresponds to at least part of the target part, and the multiple methods include at least two of the following methods: The first method is to obtain a sample grayscale image by normalizing the sample depth image; and obtain the sample depth image feature based on the sample grayscale image. The second method is to convert the sample depth image into sample point cloud data; downsample the sample point cloud data, or fuse the sample point cloud data and a sample color image in the sample color image corresponding to the sample depth image and then downsample to obtain downsampled sample point cloud data; and encode the downsampled sample point cloud data to obtain a sample third encoding feature; Obtaining the sample depth image feature based on the sample third encoding feature, A third method is to convert the sample depth image into sample point cloud data; downsample and encode the sample point cloud data to obtain a fourth encoding feature of the sample; Feature extraction is performed on the sample color image corresponding to the sample depth image in the sample color image to obtain at least part of the sample color image feature; and the sample fourth encoding feature and the at least part of the sample color image feature are fused through a cross-attention mechanism to obtain the sample depth image feature.
13. A task execution device, characterized in that: It comprises a control module, which is used to execute the action feature prediction method according to any one of claims 1 to 8.
14. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor, Wherein, the processor is used to execute the action feature prediction method described in any one of claims 1 to 8 or the model training method described in any one of claims 9 to 12.
15. A computer-readable storage medium, characterized in that The storage medium stores a computer program, which is used to execute the action feature prediction method described in any one of claims 1 to 8 or the model training method described in any one of claims 9 to 12.
16. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by the processor of a computer device, the computer device is enabled to execute the action feature prediction method described in any one of claims 1 to 8 or the model training method described in any one of claims 9 to 12.
Citation Information
Cited By
Depth information processing method and device based on VLA large model
CN121962786A