Multi-modal flow matching-based motion prediction method and device for robot with body
Through the multimodal flow matching method, combined with image features and depth information, the problem of information loss in embodied robot motion prediction is solved, and higher prediction accuracy and environmental perception capabilities are achieved.
Patent Information
- Application Number
- CN202511259401.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-17
AI Technical Summary
Existing methods for predicting the motion of embodied robots are mainly based on single-frame image input, without considering video sequence information and depth information, resulting in poor motion prediction accuracy and an inability to accurately perceive the operating environment.
By acquiring the instruction text and the image feature set collected by the robot, feature stitching and refinement are performed, and feature fusion is carried out in combination with the depth image set. The robot arm action prediction is performed using a multimodal flow matching mechanism, including feature stitching of the image sequence feature set, depth image feature fusion, and determination of text visual modal feature information.
It improves the accuracy of motion prediction and environmental perception of embodied robots, enhances the model's ability to adapt motion prediction to different paths, and improves the model's generalization performance and execution accuracy.
Smart Images

Figure CN120791793A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent robots, in particular to a somatic robot action prediction method and device based on multi-modal flow matching. BACKGROUND
[0002] At present, the related technology proposes that the somatic robot mainly makes modal input based on a single frame image when performing action prediction, and adopts a fixed strategy to predict the robot action, but the above scheme is prone to information loss due to the failure to consider the information between video sequences, thereby causing poor action prediction accuracy of the model. In addition, the model does not add depth information when predicting, which also causes the model to fail to accurately perceive the environment being operated, thereby affecting the accuracy of the model in predicting the robot action. SUMMARY
[0003] Therefore, the purpose of the present application is to provide a somatic robot action prediction method and device based on multi-modal flow matching, which can significantly improve the accuracy of somatic robot action prediction.
[0004] In a first aspect, an embodiment of the present application provides a somatic robot action prediction method based on multi-modal flow matching, which comprises: acquiring an instruction text and an image feature set collected by a robot, and a depth image set corresponding to the image features at each time position; performing feature splicing processing and feature refining processing on the image feature set to obtain an image sequence feature set, and performing feature fusion processing on the depth image set based on the image sequence feature set to determine a target visual feature; fusing the text feature in the instruction text with the target visual feature to determine text visual modal feature information, and predicting the action of the robot mechanical arm based on the text visual modal feature information to determine a motion pose prediction feature, wherein the motion pose prediction feature is used to predict the motion pose of the mechanical arm at the next time.
[0005] In an embodiment, the step of performing feature splicing processing and feature refining processing on the image feature set to obtain the image sequence feature set comprises: obtaining an input feature set based on the image feature set, and performing feature splicing processing on the input feature set by using a three-view three-frame spatio-temporal cross-attention model to determine an environmental feedback feature; refining the environmental feedback feature by a positive feedback module to determine the image sequence feature set.
[0006] In an embodiment, the step of obtaining the input feature set based on the image feature set comprises: adding the image feature set collected by the camera on the head and the left and right arms of the robot at consecutive three time positions to the time coding feature corresponding to the time positions to obtain the input feature set.
[0007] In an implementation, the step of determining the target visual feature based on the image sequence feature set includes: performing semantic segmentation on the depth image set, the instruction text and the image feature set respectively to obtain semantic segmentation results, and determining the class with the most votes as the current mask class by voting on the semantic segmentation results; extracting the visual feature in the time channel range based on the current mask class, determining the category semantic feature, and performing feature fusion on the semantic feature and the image sequence feature set by the visual feature fusion module to determine the target visual feature.
[0008] In an implementation, the step of performing semantic segmentation on the depth image set, the instruction text and the image feature set respectively to obtain semantic segmentation results includes: rendering the depth image set and sending it to an unsupervised segmentation network to obtain a first semantic mask, and performing mask recognition on the image features and the instruction text corresponding to each depth image by an open vocabulary semantic segmentation network to obtain a second semantic mask, wherein the semantic segmentation results include the first semantic mask and the second semantic mask.
[0009] In an implementation, the step of determining the text visual modal feature information by fusing the text feature in the instruction text with the target visual feature includes: performing feature extraction and feature fusion on the instruction text by a single-modal text encoder and a multi-modal text encoder to obtain a text feature, and fusing the text feature with the target visual feature by a cross-attention model and a self-attention model to determine the fused text visual modal feature information.
[0010] In an implementation, the step of predicting the motion of the robot arm based on the text visual modal feature information to determine the motion pose prediction feature includes: encoding the motion feature of the robot arm according to the text visual modal feature information to obtain a motion feature set, wherein the motion feature set includes Gaussian noise encoding, arm motion frequency encoding, diffusion time step encoding, arm motion feature encoding and arm current pose encoding; determining the conditional encoding of the flow matching diffusion model according to the arm current pose encoding, and combining the remaining motion features to obtain input encoding features, so that the flow matching diffusion model determines the motion pose prediction feature according to the conditional encoding and the input encoding features.
[0011] In a second aspect, the embodiment of the present application also provides a somatic robot action prediction device based on multi-modal flow matching, which comprises: an information acquisition module, which acquires instruction text and an image feature set collected by a robot, and a depth image set corresponding to the image features at each position; a feature extraction module, which performs feature splicing processing and feature refining processing on the image feature set to obtain an image sequence feature set, and performs feature fusion processing on the depth image set based on the image sequence feature set to determine target visual features; and a pose prediction module, which fuses text features in the instruction text with the target visual features to determine text visual modal feature information, and predicts the action of a robot arm based on the text visual modal feature information to determine motion pose prediction features, wherein the motion pose prediction features are used to predict the motion pose of the robot arm at the next moment.
[0012] In a third aspect, the embodiment of the present application also provides a server, which comprises a processor and a memory, the memory stores computer executable instructions capable of being executed by the processor, and the processor executes the computer executable instructions to implement the method of any one of the first aspect.
[0013] In a fourth aspect, the embodiment of the present application also provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions, when called and executed by a processor, cause the processor to implement the method of any one of the first aspect.
[0014] The embodiment of the present application has the following beneficial effects: The somatic robot action prediction method and device based on multi-modal flow matching provided by the embodiment of the present application, after acquiring instruction text and an image feature set collected by a robot, and a depth image set corresponding to the image features at each position, performs feature splicing processing and feature refining processing on the image feature set to obtain an image sequence feature set, performs feature fusion processing on the depth image set based on the image sequence feature set to determine target visual features, and finally fuses text features in the instruction text with the target visual features to determine text visual modal feature information, and predicts the action of a robot arm based on the text visual modal feature information to determine motion pose prediction features. The embodiment of the present application can improve the ability of the model to perceive the environment by adding image depth information and image sequence information, provide environmental feedback information, improve the accuracy of the model in predicting the action of the robot, and improve the accuracy of the model in predicting the action under different paths through a flow matching mechanism network.
[0015] Other features and advantages of the present application will be set forth in the descriptions that follow, and in part will be apparent from the description, or can be learned by practice of the application. The purposes and other advantages of the present application will be realized and attained by the structures particularly pointed out in the description, claims and drawings.
[0016] In order to make the above objectives, features and advantages of the present application more apparent, the following will specifically describe preferred embodiments of the present application, and the accompanying drawings will be described in detail as follows. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0018] Figure 1 A flowchart of a body-possessed robot action prediction method based on multi-modal flow matching provided by an embodiment of the present application; Figure 2 A flowchart of a model training and reasoning method provided by an embodiment of the present application; Figure 3 A structural schematic diagram of a body-possessed robot action prediction device based on multi-modal flow matching provided by an embodiment of the present application; Figure 4 A structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the embodiments of the present application more apparent, the technical solutions of the present application will be described clearly and completely in the following with reference to the embodiments. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.
[0020] Currently, the related technology proposes that the embodied robot mainly makes modal input based on a single frame image when performing action prediction, and a fixed strategy is used to predict the robot action, but the above scheme is easy to cause information loss due to not considering the information between video sequences, and further makes the action prediction accuracy of the model poor, in addition, the model does not add depth information when predicting, which also causes the model to not accurately perceive the environment operated, thereby affecting the accuracy of the model in predicting the robot action, that is, the problems of the existing robot action prediction scheme mainly concentrate in that the model only considers single frame visual information, and part of the visual modal input does not consider depth information, based on this, the embodied robot action prediction method and device based on multi-modal flow matching provided by the embodiment of the present application can improve the ability of the model to perceive the environment by adding image depth information and image sequence information, at the same time, the environment feedback information is provided, the accuracy of the model in predicting the robot action is improved, in addition, the flow matching mechanism network is also used to improve the accuracy of the model in predicting the action under different paths.
[0021] Referring to Figure 1 A flowchart of an embodied robot action prediction method based on multi-modal flow matching is shown, and the method mainly includes the following steps S102 to S106: Step S102, acquiring instruction text and image feature set collected by the robot, and depth image set corresponding to image features at each time position, wherein the image feature set includes: RGB image obtained from the left arm camera of the robot , the head camera obtains the RGB image , the right arm camera obtains the RGB image , the depth image includes corresponding: left arm depth image , head depth image , right arm depth image In an embodiment, it is necessary to acquire the RGB image and the corresponding depth image at different times obtained by each position camera at t, t+1, t+2 time points on the time sequence axis respectively, wherein t, t+1, t+2 are separated by 30 frames.
[0022] In step S104, feature stitching processing and feature refining processing are performed on the image feature set to obtain an image sequence feature set, and based on the image sequence feature set, feature fusion processing is performed on the depth image set to determine the target visual feature. In an embodiment, the input feature set can be obtained based on the image feature set, and the spatiotemporal cross-attention model of three views and three frames is used to perform feature stitching processing on the input feature set to determine the environmental feedback feature. Then, the environmental feedback feature is refined by the positive feedback module to determine the image sequence feature set. In this embodiment, the image features collected by the camera on the head and the left and right arms of the robot at three consecutive time points (30 frames at one time point) can be added to the time encoding features at the corresponding time points to obtain the input feature set.
[0023] In actual applications, the RGB images obtained at different positions at different time points can be input into the BLIP2 image feature extractor ImageEncoder. The parameters in the feature extractor are frozen and do not participate in model training, and the left arm image features at time t , the left arm image features at time t+1 , the left arm image features at time t+2 , the head image features at time t , the head image features at time t+1 , the head image features at time t+2 , the right arm image features at time t , the right arm image features at time t+1 , and the right arm image features at time t+2 are obtained. The RGB features at different positions are input into the VSE Module module (i.e., the spatiotemporal cross-attention model of three views and three frames) as a unit (i.e., the input feature set), which includes the features at time t, time t+1, and time t+2.
[0024] First, the features at time t are added to the time encoding features at time t to obtain features , the image features at time t+1 are added to the time encoding features at time t+1 to obtain features , the image features at time t+2 are added to the time encoding features at time t+2 to obtain features , and , , and the features are merged, and then tiled to obtain features , while setting the learned parameters learned latent quries and the features Concate (used to concatenate two or more tensors together in a certain dimension, generate a larger tensor) to get the feature , the parameter as key, value and as a question input into the cross attention Cross-Attention module to get the feature , that is, the environmental feedback feature, which more integrates the visual information at different times, is conducive to the model to obtain the feedback information of the environment when performing actions, and is conducive to the model to predict the next step information.
[0025] Further, in order to make the feature better for learning and fitting, so as to be fused with the subsequent features, it is necessary to input these features into the FFW positive feedback module (level, refine the "time-multiple view" visual features fused by cross attention again, and use the residual structure to prevent information loss, provide clean, compact and high expressive visual features for the higher order fusion / decision network), which is composed of multiple MLP, activation function Relu and Norm normalization module in series, and residual mechanism is added to prevent model information forgetting. The specific operation of this module is that the environmental feedback feature is added to the feature to get the feature , which is input into the first FFW module to get the feature , which is added to the feature to get the feature , and the feature is input into the second FFW module to get the feature . Because the image sequence feature set of the robot's left arm, head and right arm is input into the module, the three features are respectively , , .
[0026] In step S106, the text feature in the instruction text is fused with the target visual feature to determine the text visual modal feature information, and the motion pose prediction feature of the robot manipulator is predicted based on the text visual modal feature information, wherein the motion pose prediction feature is used to predict the motion pose of the manipulator at the next time. In one embodiment, the motion feature set of the robot manipulator can be obtained by encoding the text visual modal feature information, wherein the motion feature set includes: Gaussian noise encoding, manipulator motion frequency encoding, diffusion time step encoding, manipulator action feature encoding and current pose encoding of the manipulator. Then, the condition code of the flow matching diffusion model is determined according to the current pose encoding of the manipulator, and the remaining motion features are combined to obtain the input encoding feature. The flow matching diffusion model determines the motion pose prediction feature according to the condition code and the input encoding feature.
[0027] In practical applications, the fused text visual modal feature information can be input into the ACUFM Module, in which Gaussian noise is encoded to obtain a feature , and the motion frequency of the robot arm is encoded to obtain a feature , the diffusion time step is encoded to obtain a feature , the motion feature of the robot arm is encoded to obtain a feature , and finally the current pose of the robot arm is encoded to obtain a feature .
[0028] In an embodiment, the step of combining the motion features to obtain the input encoded feature includes: first, inputting the motion frequency feature and the feature into the MLP module to obtain a feature , using to obtain a feature , using to obtain a feature , then adding and the feature to obtain a feature , i.e., the input encoded feature, and then encoding the current pose of the robot arm through the MLP to obtain as the condition code of the flow matching diffusion model.
[0029] Further, the previously obtained is Concate-merged with the feature through the MLP module to obtain a feature , and then the feature is obtained by passing through another MLP module, the previously obtained feature is Concate-merged with the feature through the MLP module, and then the feature is obtained by passing through the FMUNetBlock module. The FMUNetBlock module first inputs the obtained feature into the Reshape Module1 module and the Reshape Module2 module, respectively, to obtain a feature and a feature through Reshape operation, inputs the feature into two down-sampling modules DownSample Module, and obtains a feature , and then compare this feature with the previously obtained feature Perform cross-attention operation to obtain features , and then get the feature after upsampling once . Further, the feature is input into three upsampling modules UnSampBlock module to obtain the feature As k, v and the features obtained before As q, it is input into the Cross-Attention module to obtain features , and then input it into the self-attention mechanism to get the feature The above UnsampBlock consists of multiple deconvolution, normalization and activation functions, and is finally input into the Head layer to obtain the feature The head layer consists of multiple MLPs, activation functions, and normalization (i.e., motion pose prediction features). This feature is used to predict the next pose of the robot arm. Therefore, the loss function is calculated based on this value. Here, we choose MSE as the loss function.
[0030] The above-mentioned embodied robot motion prediction method based on multimodal stream matching provided by an embodiment of the present invention can improve the overall generalization performance of the model, so that the model can still have a good execution effect in scenarios with a small number of training samples. In addition, through unsupervised segmentation information, the model's spatial semantic understanding ability of objects that need to be operated is improved, and the accuracy of the model's prediction execution is improved. Through depth image and video sequence information, the model's environmental perception ability and environmental feedback information are improved, and the accuracy of the model's prediction execution is improved. Finally, through the stream matching mechanism, the accuracy of the model's motion prediction under different paths of adaptation is improved.
[0031] The embodiment of the present invention further provides an implementation method for determining target visual features, for details, see (1) to (3) below: (1) Semantic segmentation processing is performed on the depth image set, the instruction text and the image feature set respectively to obtain semantic segmentation results, and the semantic segmentation results are voted to determine the category with the most votes as the current mask category. In one embodiment, the depth image set can be rendered and sent to an unsupervised segmentation network to obtain a first semantic mask, and an open vocabulary semantic segmentation network is used to perform mask recognition processing on the image features and instruction text corresponding to each depth image to obtain a second semantic mask, wherein the semantic segmentation results include: a first semantic mask and a second semantic mask.
[0032] In practical applications, for the depth image information, first convert it into Colormaps to obtain the rendered depth map, which is recorded as , and input it into the SAM2 unsupervised segmentation network to obtain the first semantic mask , the network weight is completely frozen and does not participate in model training, and the RGB image and the instruction text at the corresponding time position are input into the OVSeg (Open-Vocabulary Semantic) network to obtain the second semantic mask , the OVSeg performs zero-shot semantic segmentation on the RGB image, and according to the object categories contained in the input text information, the category mask identification is completed, then the category of each SAM mask is voted according to the semantic segmentation result of the points in the current mask, and the category with the most points is selected as the category of the current mask (mask, which is a matrix or image with the same size as the original image, used to indicate or identify the pixels or regions of a specific region in the original image), and the feature is denoted as , that is, the visual information feature, which not only contains visual semantic category information, but also contains spatial position information between different objects in vision, providing more accurate spatial prior information for the model to accurately grasp.
[0033] Further, the features obtained by the mechanical arm at different positions are subjected to Concate merging operation and input into the 3DCNN convolution module, which can extract visual information at different time points, and the operator of the module can also obtain the change feature of the visual information in the time channel range, which is denoted as category semantic feature, and finally the left arm image sequence feature, the head image sequence feature, the right arm image sequence feature and the category semantic feature are input into the Vision FeatureFusion Module (that is, the visual feature fusion module).
[0034] (2) Based on the current mask category, the change feature of the visual information in the time channel range is extracted, the category semantic feature is determined, and the semantic feature and the image sequence feature set are subjected to feature fusion processing through the visual feature fusion module to determine the target visual feature. In an embodiment, for the visual feature fusion module, first, the image sequence features are respectively input into the respective CBAM channel space, attention mechanism modules, to obtain features , and feature , respectively, and then input them into the ConvBlock convolution module, which includes a combination of multiple convolution operations, activation functions, normalization operations and pooling operations, to finally obtain features , and feature , respectively. Then, the residual mechanism is added to prevent information forgetting, and the obtained features 、 、 and feature , feature and feature are obtained by performing a Concate merging operation on the features 、 、 , which respectively represent more semantically informative visual information of the left arm view, visual information of the head view, and visual information of the right arm view. Since the visual information of the head view has more comprehensive visual information, and the visual information of the left and right arms has more detailed visual information, Cross-Attention cross-attention operations are performed on the head visual information and the detailed visual information of the left and right arms, respectively, with the head visual information as q, and the left and right arm visual information as k and v. Finally, the feature is and feature , and a Concate merging operation is performed on the three features to obtain feature . Then, a Cross-Attention cross-attention operation is performed on the previously obtained category semantic feature and feature to obtain feature , which better integrates spatial visual information and semantic spatial information to obtain more rich visual information. Here, the category semantic feature is used as k and v, and the feature is used as q. Then, the obtained feature is input into a Self-Attention self-attention mechanism to obtain feature (i.e., the target visual feature).
[0035] (3) The instruction text is processed by a single-modal text encoder and a multi-modal text encoder to obtain text features, and the text features and the target visual features are fused by a Cross-Attention cross-attention model and a Self-Attention self-attention model to determine the fused text visual modal feature information. In an implementation, for text information, the text instruction is first input into a BLIP2 single-modal text encoder and a multi-modal text encoder to obtain features and feature . Then, a Concate merging operation is performed on the two features to obtain feature (i.e., the text feature). Then, the text is considered as a visual modal fusion, and the feature is input into an IT Fusion Module module to obtain feature . The feature considered as q, and the target visual feature obtained before considered as k and v are input into the cross attention mechanism to perform initial fusion of the text mode and the visual mode to obtain a feature The feature is then subjected to a Flatten display operation to obtain a feature The feature is then input into a GAP global average pooling operation to obtain a feature .
[0036] Further, the feature is continuously input into an MLP module to obtain a feature And a Sigmoid activation function is connected to obtain a weight feature Because it is considered that the weight of the fusion of the text and visual modes needs to pay attention to the weight of the final feature, therefore, the feature better balances the importance of the two modes, so the feature obtained before And the weight feature are multiplied to obtain a feature And the feature obtained before is added to obtain a feature And input into a Self-Attention self-attention module to obtain a feature And an MLP module to obtain a feature The feature f integrates the text and visual information features.
[0037] Referring to the flowchart of a model training and reasoning method shown in Figure 2 The embodiment of the present application also provides an implementation of model training and reasoning. In the process of model training, the data set is divided into a training set and a validation set. The training set is used during training, and the validation set is used to observe the model training convergence condition. The weight with the lowest loss function of the validation set is selected as the final weight of the model.
[0038] In actual application, MSE can be selected as the model training loss function. At the same time, during reasoning, the current mechanical arm action is set as Gaussian random noise as the input of the model. The prediction of the model uses the Gaussian random noise plus the position coding of the model prediction multiplied by the time diffusion step t as the final prediction result of the model. In addition, random mask and image augmentation are also performed during training, including color jitter, arbitrary cropping, etc. Thus, the model can be adapted to different lighting conditions, and the model can not only rely too much on single camera information, but also use the combined information of the three cameras to predict behavior. Finally, in the prediction execution stage, the mechanical arm will traverse and execute according to the motion sequence predicted by the model. After that, the model reasoning is continuously performed until the model completes the whole action.
[0039] In summary, the present invention can improve the accuracy and generalization of the model in new scenarios through the embodied algorithm diffusion model AFM Network; through the 3DSS spatial semantic prior feature extraction module, the algorithm is provided with image detail features and visual depth information, so that the robot arm can move accurately and improve the accuracy of motion execution; through the VSE Module module, the visual sequence information is more fully extracted, and the model's ability to perceive the environment is improved by adding image depth information and image sequence information, while providing environmental feedback information to improve the success rate of model prediction execution; and through the FMUNet flow matching mechanism network, the accuracy of the model in adapting to different paths for action prediction is improved.
[0040] Regarding the embodied robot motion prediction method based on multimodal flow matching provided in the aforementioned embodiment, an embodiment of the present invention provides an embodied robot motion prediction device based on multimodal flow matching, see Figure 3 The schematic diagram of the structure of an embodied robot motion prediction device based on multimodal flow matching is shown in FIG. The device includes the following parts: The information acquisition module 302 acquires the instruction text and the image feature set collected by the robot, as well as the depth image set corresponding to the image features at each time position; The feature extraction module 304 performs feature splicing and feature refinement on the image feature set to obtain an image sequence feature set, and performs feature fusion processing on the depth image set based on the image sequence feature set to determine the target visual features; The posture prediction module 306 determines the text visual modal feature information by fusing the text features in the instruction text with the target visual features, and predicts the movement of the robot arm based on the text visual modal feature information to determine the motion posture prediction feature, wherein the motion posture prediction feature is used to predict the motion posture of the robot arm at the next moment.
[0041] The above-mentioned embodied robot motion prediction device based on multimodal flow matching provided in the embodiments of the present application can significantly improve the accuracy of embodied robot motion prediction.
[0042] In one embodiment, when performing feature splicing and feature refinement processing on the image feature set to obtain the image sequence feature set, the above-mentioned feature extraction module 304 is also used to: obtain the input feature set based on the image feature set, and use the three-view three-frame spatiotemporal cross-attention model to perform feature splicing processing on the input feature set to determine the environmental feedback feature; and refine the environmental feedback feature through the positive feedback module to determine the image sequence feature set.
[0043] In an implementation, when the step of obtaining the input feature set based on the image feature set is performed, the feature extraction module 304 is further configured to: add the image feature set collected by the camera of the robot head and the left and right arms at the three consecutive time points to the time coding feature of the corresponding time point to obtain the input feature set.
[0044] In an implementation, when the step of performing feature fusion processing on the depth image set based on the image sequence feature set to determine the target visual feature is performed, the feature extraction module 304 is further configured to: perform semantic segmentation processing on the depth image set, the instruction text, and the image feature set respectively to obtain a semantic segmentation result, and determine the class with the most votes as the current mask class by performing voting processing on the semantic segmentation result; extract the visual feature in the time channel range based on the current mask class to determine the class semantic feature, and perform feature fusion processing on the semantic feature and the image sequence feature set by the visual feature fusion module to determine the target visual feature.
[0045] In an implementation, when the step of performing semantic segmentation processing on the depth image set, the instruction text, and the image feature set respectively to obtain a semantic segmentation result is performed, the feature extraction module 304 is further configured to: render the depth image set and send it to an unsupervised segmentation network to obtain a first semantic mask, and perform mask recognition processing on the image feature and the instruction text corresponding to each depth image by using an open vocabulary semantic segmentation network to obtain a second semantic mask, wherein the semantic segmentation result includes the first semantic mask and the second semantic mask.
[0046] In an implementation, when the step of determining the text visual modal feature information by fusing the text feature in the instruction text with the target visual feature is performed, the pose prediction module 306 is further configured to: perform feature extraction and feature fusion processing on the instruction text by a single-modal text encoder and a multi-modal text encoder to obtain a text feature, and fuse the text feature with the target visual feature by a cross-attention model and a self-attention model to determine the fused text visual modal feature information.
[0047] In an implementation, when the step of determining the motion pose prediction feature based on the text visual modal feature information and predicting the motion of the robot arm is performed, the pose prediction module 306 is further configured to: encode the motion feature of the robot arm according to the text visual modal feature information to obtain a motion feature set, wherein the motion feature set includes: Gaussian noise encoding, robot arm motion frequency encoding, diffusion time step encoding, robot arm motion feature encoding, and robot arm current pose encoding; determine the condition encoding of the flow matching diffusion model according to the robot arm current pose encoding, and combine the remaining motion features to obtain input encoding features, so that the flow matching diffusion model determines the motion pose prediction feature according to the condition encoding and the input encoding features.
[0048] The device provided by the embodiment of the present application has the same implementation principle and technical effects as the foregoing method embodiment. For brevity, the part not mentioned in the device embodiment can be referred to the corresponding content in the foregoing method embodiment.
[0049] The embodiment of the present application provides a server, specifically, the server includes a processor and a storage device; the storage device stores a computer program, and the computer program executes the method of any one of the foregoing embodiments when the processor runs.
[0050] Figure 4 A structural diagram of a server provided by the embodiment of the present application is provided, and the server 100 includes a processor 40, a memory 41, a bus 42, and a communication interface 43, the processor 40, the communication interface 43, and the memory 41 are connected through the bus 42; the processor 40 is configured to execute an executable module stored in the memory 41, for example, a computer program.
[0051] The memory 41 can include a high-speed random access memory (RAM) and can also include a non-volatile memory, for example, at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0052] The bus 42 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 4 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0053] The memory 41 is configured to store a program, and the processor 40 executes the program after receiving an execution instruction. The method performed by the device for defining a flow process according to any of the foregoing embodiments of the application can be applied to the processor 40 or implemented by the processor 40.
[0054] The processor 40 can be an integrated circuit chip with a signal processing capability. In the implementation process, each step of the foregoing method can be completed by an integrated logic circuit of hardware in the processor 40 or an instruction in the form of software. The foregoing processor 40 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), and the like; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. Each method, step and logic block diagram disclosed in the embodiments of the application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, or other mature storage medium in the art. The storage medium is located in the memory 41, and the processor 40 reads information in the memory 41 and combines hardware to complete the steps of the foregoing method.
[0055] The computer program product of the readable storage medium provided by the embodiments of the application includes a computer readable storage medium storing a program code, and the program code includes instructions for executing the method described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments, which will not be described here.
[0056] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0057] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for predicting embodied robot motion based on multimodal flow matching, characterized in that: The method comprises: Obtain the command text and the image feature set collected by the robot, as well as the depth image set corresponding to the image features at each moment; Performing feature splicing and feature refinement processing on the image feature set to obtain an image sequence feature set, and performing feature fusion processing on the depth image set based on the image sequence feature set to determine target visual features; By fusing the text features in the instruction text with the target visual features, the text visual modal feature information is determined, and based on the text visual modal feature information, the movement of the robot manipulator is predicted to determine the motion posture prediction feature, wherein the motion posture prediction feature is used to predict the motion posture of the manipulator at the next moment.
2. The method for embodied robot motion prediction based on multimodal flow matching according to claim 1, characterized in that: The step of performing feature splicing processing and feature refinement processing on the image feature set to obtain an image sequence feature set includes: Based on the image feature set, an input feature set is obtained, and a spatiotemporal cross-attention model of three views and three frames is used to perform feature splicing processing on the input feature set to determine environmental feedback features; The environmental feedback features are refined by a positive feedback module to determine the image sequence feature set.
3. The method for embodied robot motion prediction based on multimodal flow matching according to claim 2, characterized in that: The step of obtaining an input feature set based on the image feature set includes: The image feature set collected by the cameras of the robot head and left and right arms in three consecutive moments is added to the time coding features of the corresponding moments to obtain the input feature set.
4. The method for embodied robot motion prediction based on multimodal flow matching according to claim 1, characterized in that: The step of performing feature fusion processing on the depth image set based on the image sequence feature set to determine target visual features includes: Performing semantic segmentation processing on the depth image set, the instruction text, and the image feature set respectively to obtain semantic segmentation results, and performing voting processing on the semantic segmentation results to determine the category with the most votes as the current mask category; Based on the current mask category, the visual change features in the time channel range are extracted to determine the category semantic features, and the semantic features and the image sequence feature set are subjected to feature fusion processing through a visual feature fusion module to determine the target visual features.
5. The method for embodied robot motion prediction based on multimodal flow matching according to claim 4, characterized in that: The step of performing semantic segmentation processing on the depth image set, the instruction text, and the image feature set to obtain a semantic segmentation result includes: After rendering, the depth image set is sent to an unsupervised segmentation network to obtain a first semantic mask, and an open vocabulary semantic segmentation network is used to perform mask recognition processing on the image features corresponding to each depth image and the instruction text to obtain a second semantic mask, wherein the semantic segmentation result includes: the first semantic mask and the second semantic mask.
6. The method for embodied robot motion prediction based on multimodal flow matching according to claim 1, characterized in that: The step of determining text visual modality feature information by fusing the text features in the instruction text with the target visual features includes: The instruction text is subjected to feature extraction and feature fusion processing respectively by a unimodal text encoder and a multimodal text encoder to obtain text features, and the text features are fused with the target visual features through a cross-attention model and a self-attention model to determine the fused text visual modality feature information.
7. The method for embodied robot motion prediction based on multimodal flow matching according to claim 1, characterized in that: The steps of predicting the motion of the robot arm based on the text visual modal feature information and determining the motion posture prediction features include: According to the text visual modal feature information, the robot's manipulator motion features are encoded to obtain a motion feature set, wherein the motion feature set includes: Gaussian noise encoding, manipulator motion frequency encoding, diffusion time step encoding, manipulator action feature encoding, and manipulator current posture encoding; The conditional coding of the flow matching diffusion model is determined according to the current posture coding of the robotic arm, and the remaining motion features are combined to obtain the input coding features, so that the flow matching diffusion model is used to determine the motion posture prediction features according to the conditional coding and the input coding features.
8. A device for predicting embodied robot motion based on multimodal flow matching, characterized in that: The device comprises: An information acquisition module acquires the command text and the image feature set collected by the robot, as well as the depth image set corresponding to the image features at each moment; a feature extraction module that performs feature splicing and feature refinement processing on the image feature set to obtain an image sequence feature set, and performs feature fusion processing on the depth image set based on the image sequence feature set to determine target visual features; The posture prediction module determines text visual modal feature information by fusing the text features in the instruction text with the target visual features, and predicts the movement of the robot manipulator based on the text visual modal feature information to determine the motion posture prediction feature, wherein the motion posture prediction feature is used to predict the motion posture of the manipulator at the next moment.
9. A server, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Industrial part instance segmentation method based on voting mechanism
CN114627289A
Robot motion skill learning method fusing text instruction and motion information
CN117428780A
Mechanical arm control method, device and equipment based on large visual model and storage medium
CN118143940A
Cited By
Fast image restoration method based on stream matching
CN121563840A