Processing method and device for robot with body, electronic equipment and storage medium

By using a global state intermediate supervision network model for visual prediction and action prediction, the problem of low execution accuracy of embodied robots is solved, and their adaptability and execution accuracy in complex operation scenarios are improved.

CN120791805AActive Publication Date: 2025-10-17SHENZHEN SHIHE ROBOTIC TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511301180.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

In the prior art, embodied robots only use vision modules when performing tasks, resulting in low execution accuracy.

Method used

A global state intermediate supervision network model is adopted to obtain the current visual information, task instruction information and robotic arm state information of the embodied robot, and use the global state intermediate supervision network model to perform visual prediction and action prediction, and predict the visual information and action block sequence at the next moment until the task is completed.

Benefits of technology

The adaptability and execution accuracy of embodied robots in complex operation scenarios have been significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120791805A_ABST
    Figure CN120791805A_ABST
Patent Text Reader

Abstract

The invention provides a processing method and device for a robot with a body, electronic equipment and a storage medium. The processing method comprises the steps that current visual information, task instruction information and current mechanical arm state information of the robot with the body for a target scene are obtained; the current visual information, the task instruction information and the current mechanical arm state information are input into the global state intermediate supervision network model for visual prediction processing and action prediction processing, and the visual information of the robot with the body at the next moment and an action block sequence of the robot with the body are predicted; after the robot with the body executes the action block sequence, the visual information of the next moment and the state of the mechanical arm after executing the action block sequence are input into the global state intermediate supervision network model to continue visual prediction processing and action prediction processing; and the prediction is stopped until the robot completes the target task corresponding to the task instruction information. And the adaptability and the execution precision of the robot with the body in a complex operation scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent decision-making, in particular to a processing method and device of a body-equipped robot, an electronic device and a storage medium. BACKGROUND

[0002] With the explosive iteration of artificial intelligence technology, body-equipped intelligence, as a typical method and technology in the artificial intelligence era, has ushered in a huge development opportunity. Body-equipped intelligence refers to a method of combining an intelligent system with a physical entity, which enables the physical entity to perceive the environment, make decisions and perform corresponding actions, so the physical entity combined with body-equipped intelligence is also called an agent. "Body-equipped" not only refers to abstract algorithms and data, but also interacts with the world through physical form. The application scenarios of body-equipped intelligence are also extremely wide, and the agent can be integrated into intelligent manufacturing, service industry and other vertical fields, such as industrial inspection and housekeeping services, so that body-equipped intelligence leads the upgrading of new manufacturing, service and other industries. In the prior art, image data is collected in real time during the execution of the current task of the body-equipped robot, and the image data is input into a pre-trained improved YOLO model to obtain inference information features of a target object. In the prior art, only a vision module is used, which leads to the technical problem of low execution accuracy of the body-equipped intelligent robot. SUMMARY

[0003] Therefore, the purpose of the present application is to provide a processing method and device of a body-equipped robot, an electronic device and a storage medium, which can directly determine the visual information and action block sequence of the next moment by using a global state intermediate supervision network model, until the task is successfully completed, thereby significantly improving the adaptability and execution accuracy of the body-equipped robot in complex operation scenarios.

[0004] The embodiment of the present application provides a processing method of a body-equipped robot, which comprises the following steps: obtaining current visual information, task instruction information and current mechanical arm state information of the body-equipped robot for a target scene; inputting the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model for visual prediction processing and action prediction processing, and predicting next moment visual information of the body-equipped robot and an action block sequence of the body-equipped robot; when the body-equipped robot executes the action block sequence, inputting the next moment visual information and the mechanical arm state after executing the action block sequence into the global state intermediate supervision network model to continue visual prediction processing and action prediction processing, until the body-equipped robot completes the target task corresponding to the task instruction information.

[0005] In a possible implementation, the inputting the current visual information, task instruction information and the current robot arm state information into a pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, predicting the next time visual information of the embodied robot and the action block sequence of the robot, comprises: performing visual prediction processing on the current visual information and task instruction information based on the global state intermediate supervision network model, to predict the next time visual information; wherein the current visual information is composed of a plurality of sub visual images; performing action prediction processing on the current visual information, the task instruction information and the current robot arm state based on the global state intermediate supervision network model, to predict the action block sequence.

[0006] In a possible implementation, the performing visual prediction processing on the current visual information and task instruction information based on the global state intermediate supervision network model to predict the next time visual information, comprises: performing depth image feature extraction, bidirectional cross-attention fusion processing and gate cycle fusion processing on the sub visual images respectively acquired by the left robot arm, the right robot arm and the head of the embodied robot in the global state intermediate supervision network model, to output a first visual feature corresponding to each of the sub visual images; performing multi-modal segmentation processing and full attention mechanism processing on the task instruction information, a plurality of the sub visual images and the first visual feature corresponding to each of the sub visual images in the global state intermediate supervision network model, to output the next time visual information.

[0007] In a possible implementation, for the sub visual image of the left robot arm, the outputting the first visual feature corresponding to each of the sub visual images comprises: fusing the features extracted from the sub visual image by the feature extraction network and the visual feature extraction network based on the global state intermediate supervision network model, to obtain an RGB visual feature; performing feature extraction on the RGB visual feature by a residual depth visual feature extraction network layer composed of a plurality of convolutional layers to generate a plurality of bottom features of the RGB visual feature, and performing depth feature extraction on the plurality of bottom features by a depth feature extraction network to obtain a depth feature corresponding to each of the bottom features; performing bidirectional cross-attention processing on each of the bottom features and the corresponding depth feature to obtain a plurality of first fusion features; pyramid pooling is performed on each of the bottom features and the corresponding depth features to generate a plurality of sub-features of different receptive fields, and a plurality of second fusion features are obtained by performing gated recurrent fusion processing on a plurality of sub-features of each of the bottom features and the corresponding depth features; The first fusion features are up-sampled and spliced to generate a two-dimensional visual feature, and the second fusion features are sampled and spliced to generate a depth visual feature. The two-dimensional visual feature and the depth visual feature are dimensionally converted to generate a first visual feature corresponding to the sub-visual image.

[0008] In one possible implementation, the multi-modal segmentation processing and the full attention mechanism processing of the task instruction information, the plurality of sub-visual images and the first visual feature corresponding to each of the sub-visual images in the global state intermediate supervision network model output the next time visual information, including: Text feature extraction is performed on the task instruction information. For any sub-visual image, the task instruction information and the sub-visual image are subjected to multi-modal segmentation processing to determine a segmented image, and cross-attention processing is performed on the features in the segmented image and the first visual feature of the sub-visual image to obtain a second visual feature. The first visual feature, the second visual feature and the text feature are subjected to full attention mechanism processing based on a transformer network layer to output the next time visual information.

[0009] In one possible implementation, the action prediction processing of the current visual information, the task instruction information and the current robot state in the global state intermediate supervision network model predicts the action block sequence, including: Feature extraction is performed on the current robot state to determine a plurality of current robot motion features. The first visual feature, the second visual feature and the plurality of current robot motion features are subjected to causal attention mechanism processing based on a transformer network layer to predict the action block sequence.

[0010] In one possible implementation, the global state intermediate supervision network model is determined by the following steps: The sample visual information, sample task instruction information and sample robot state information of the embodied robot are input into an initial global state intermediate supervision network model for visual prediction processing and action prediction processing to predict the predicted next time visual information of the embodied robot and the predicted action block sequence of the embodied robot. Based on the cross entropy loss value between the actual next time visual information of the embodied robot, the actual action block sequence, the predicted next time visual information and the predicted action block sequence, the network parameters of the initial global state intermediate supervision network model are iteratively trained to determine the global state intermediate supervision network model.

[0011] Embodiments of the present application also provide a processing device of an embodied robot, and the processing device comprises: The acquisition module is configured to acquire current visual information of the embodied robot for a target scene, task instruction information and current mechanical arm state information. The first determination module is configured to input the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model for visual prediction processing and action prediction processing, and predict next time visual information of the embodied robot and an action block sequence of the embodied robot. The second determination module is configured to input next time visual information and a mechanical arm state after the action block sequence is executed into the global state intermediate supervision network model for continuous visual prediction processing and action prediction processing until the embodied robot completes a target task corresponding to the task instruction information.

[0012] Embodiments of the present application also provide an electronic device, which comprises a processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor and the memory communicate through the bus, and the machine readable instructions are executed by the processor to perform the steps of the processing method of the embodied robot as described above.

[0013] Embodiments of the present application also provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the steps of the processing method of the embodied robot as described above.

[0014] The embodiment of the present application provides a processing method and device of a body robot, electronic equipment and a storage medium, and the processing method comprises the following steps: acquiring current visual information, task instruction information and current mechanical arm state information of the body robot for a target scene; inputting the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, predicting next time visual information of the body robot and an action block sequence of the body robot; when the body robot executes the action block sequence, inputting the next time visual information and the mechanical arm state after executing the action block sequence into the global state intermediate supervision network model to continue the visual prediction processing and the action prediction processing until the body robot completes a target task corresponding to the task instruction information.

[0015] In order to make the above objectives, characteristics and advantages of the present application more apparent, clear and easy to understand, the following will specifically describe the preferred embodiments in combination with the accompanying drawings, and the detailed description is as follows. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0017] Figure 1 A flow chart of a processing method of a body robot provided by the embodiment of the present application; Figure 2 A structure schematic diagram of a processing device of a body robot provided by the embodiment of the present application; Figure 3 A structure schematic diagram of a processing device of a body robot provided by the embodiment of the present application; Figure 4 A structure schematic diagram of an electronic equipment provided by the embodiment of the present application. DETAILED DESCRIPTION

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings in the embodiments of the present application to make a clear and complete description of the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by a person skilled in the art without creative work belongs to the scope of protection of the present application.

[0019] Firstly, the application scenarios applicable to the present application are introduced. The present application can be applied to the technical field of intelligent decision-making.

[0020] It is found through research that, with the explosive iteration of artificial intelligence technology, embodied intelligence, as a typical method and technology in the artificial intelligence era, has ushered in a great development opportunity. Embodied intelligence refers to a method of combining an intelligent system with a physical entity, which enables the physical entity to perceive the environment, make decisions and perform corresponding actions, so the physical entity combined with embodied intelligence is also called an agent. "Embodiment" is not only abstract algorithms and data, but also interacts with the world through physical form. The application scenarios of embodied intelligence are also extremely wide, and the agent can be integrated into intelligent manufacturing, service industry and other vertical fields, such as industrial inspection, housekeeping service, etc., so that embodied intelligence leads the upgrading of new manufacturing, service and other industries. In the prior art, image data is collected in real time during the execution of the current task by the embodied robot, and the image data is input into a pre-trained improved YOLO model to obtain inference information features of the target object. In the prior art, only a vision module is used, which leads to the technical problem of low execution accuracy of the embodied intelligent robot.

[0021] Based on this, the embodiments of the present application provide a processing method and device of an embodied robot, electronic equipment and storage medium, which utilize a global state intermediate supervision network model to directly determine the visual information and action block sequence at the next moment until the task is successfully completed, thereby significantly improving the adaptability and execution accuracy of the embodied robot in complex operation scenarios.

[0022] Please refer to Figure 1 , Figure 1 The flowchart of the processing method of the embodied robot provided by the embodiments of the present application is shown in FIG. 1. As shown in FIG. 1, the processing method provided by the embodiments of the present application comprises the following steps. Figure 1 S101: acquiring current visual information, task instruction information and current mechanical arm state information of the embodied robot for a target scene.​

[0023] In this step, the current visual information of the embodied robot for the target scene, the task instruction information and the current robot state information are acquired.

[0024] Here, the current visual information includes a plurality of sub-visual images, which are collected by different robot arms of the embodied robot. The RGB images acquired by the left robot arm, the right robot arm and the head at the current time are sequentially denoted as , and , and their depths are sequentially denoted as , and .

[0025] S102: inputting the current visual information, the task instruction information and the current robot state information into a pre-trained global state intermediate supervision network model for visual prediction processing and action prediction processing, and predicting the next time visual information of the embodied robot and the action block sequence of the embodied robot.

[0026] In this step, the current visual information, the task instruction information and the current robot state information are input into the global state intermediate supervision network model for visual prediction processing and action prediction processing, and the next time visual information of the embodied robot and the action block sequence of the embodied robot are predicted. It should be noted that the intermediate supervision network model (Intermediate Supervision) refers to introducing an additional supervision signal (such as a loss function) in the middle layer of the neural network to guide the model to learn more robust feature expression, and avoid problems such as gradient disappearance and unstable training in the deep network. The global state intermediate supervision network model introduces the visual state at the next time (t+1) as a supervision signal on the basis of the intermediate supervision network model, that is, the model not only predicts the action, but also predicts the visual change after the action is performed.

[0027] In one possible implementation, the inputting the current visual information, the task instruction information and the current robot state information into the pre-trained global state intermediate supervision network model for visual prediction processing and action prediction processing, and predicting the next time visual information of the embodied robot and the action block sequence of the embodied robot, comprises: A: performing visual prediction processing on the current visual information and the task instruction information based on the global state intermediate supervision network model, and predicting the next time visual information; wherein the current visual information is composed of a plurality of sub-visual images.

[0028] Here, the current visual information and the task instruction information are subjected to visual prediction processing according to the global state intermediate supervision network model, and the next time visual information is predicted.

[0029] The next time visual information includes next time visual information of the left mechanical arm, next time visual information of the right mechanical arm, and next time visual information of the robot head.

[0030] In one possible implementation, the visual prediction processing of the current visual information and the task instruction information based on the global state intermediate supervision network model to predict the next time visual information includes: (1) The sub-visual images respectively acquired by the left mechanical arm, the right mechanical arm, and the head of the embodied robot are subjected to depth image feature extraction, bidirectional cross-attention fusion processing, and gate recurrent fusion processing in the global state intermediate supervision network model, and the first visual feature corresponding to each sub-visual image is output.

[0031] Here, the sub-visual images respectively acquired by the left mechanical arm, the right mechanical arm, and the head of the embodied robot are subjected to depth image feature extraction, bidirectional cross-attention fusion processing, and gate recurrent fusion processing in the global state intermediate supervision network model, and the first visual feature corresponding to each sub-visual image is output.

[0032] In one possible implementation, for the sub-visual image of the left mechanical arm, the output of the first visual feature corresponding to each sub-visual image includes: a: The features extracted from the sub-visual image by the feature extraction network and the visual feature extraction network based on the global state intermediate supervision network model are fused to obtain the RGB visual feature.

[0033] Here, the feature extraction network (DINOv2 feature extraction network) is used to extract features from the sub-visual image, and the visual feature extraction network (visual Transformer network layer) is used to extract context features from the sub-visual image. Then, the features extracted by the two feature extraction network layers are fused to obtain the RGB visual feature.

[0034] It should be noted that the visual Transformer network layer divides the image into fixed-size "image blocks" (patches), flattens these image blocks as input sequences, and then processes them using the Transformer architecture, modeling long-range dependencies through self-attention mechanisms. VIT introduces this mechanism into the visual field to model image features from a global perspective, thereby capturing global context information in the image.

[0035] b: performing feature extraction on the RGB visual feature based on a residual deep visual feature extraction network layer composed of multiple convolution layers to generate multiple bottom layer features of the RGB visual feature, and performing deep feature extraction on the multiple bottom layer features based on a deep feature extraction network to obtain a deep feature corresponding to each bottom layer feature.

[0036] Here, the residual deep visual feature extraction network layer includes a plurality of convolution modules in series, respectively convolution Block_x1, convolution Block_x2, convolution Block_x3, convolution Block_x4, and convolution Block_x5. The RGB visual feature is processed by the convolution Block_x2 to the convolution Block_x5 layer to generate four bottom layer feature maps of different scales, denoted as R1, R2, R3, and R4.

[0037] Each bottom layer feature map is connected to a channel attention module (CAM) to enhance key semantic channels. For the bottom layer feature map, a depth-specific extraction network symmetric to the RGB branch is used (the first layer convolution is replaced by a 1x1 convolution to adapt to single-channel input), and four depth feature maps of corresponding scales are output (denoted as D1, D2, D3, and D4).

[0038] c: performing bidirectional cross-attention processing on each bottom layer feature and corresponding deep feature to obtain multiple first fusion features.

[0039] Here, bidirectional cross-attention processing is performed on each bottom layer feature and corresponding deep feature to obtain multiple first fusion features.

[0040] For the bottom layer features (R1 and D1, R2 and D2, R3 and D3, R4 and D4), a bidirectional cross-attention mechanism is designed: RGB guides Depth enhancement. For R2 and D2, R2 is compressed in channel by 1x1 convolution, then matrix multiplication is performed with D2 to generate a spatial attention map, and D2 is weighted to obtain D2'; Depth guides RGB enhancement: D2 is compressed in channel by 1x1 convolution, then matrix multiplication is performed with R2 to generate a spatial attention map, and R2 is weighted to obtain R2'; feature interaction fusion: R2' and D2' are subjected to "element-level multiplication + channel concatenation", and are fused into a first fusion feature by 3x3 grouped convolution (group number = 2). The above steps are sequentially performed on each bottom layer feature and corresponding deep feature to determine the corresponding first fusion feature.

[0041] d: pyramid pooling is performed on each of the bottom features and the corresponding depth features to generate a plurality of sub-features of different receptive fields, and a gated recurrent fusion process is performed on the plurality of sub-features of each of the bottom features and the corresponding depth features to obtain a plurality of second fusion features.

[0042] Here, pyramid pooling is performed on each of the bottom features and the corresponding depth features to generate a plurality of sub-features of different receptive fields, and a gated recurrent fusion process is performed on the plurality of sub-features of each of the bottom features and the corresponding depth features to obtain a plurality of second fusion features.

[0043] Wherein, a multi-scale gated recurrent fusion unit is designed: scale decomposition: R2 and D2 are respectively decomposed into 4 sub-features of different receptive fields (1x1, 3x3, 5x5, 7x7) through pyramid pooling; Gated interaction: for each scale of sub-feature pair (R3_k, D3_k), a gating mechanism Gk=σ(WrR3_k +Wd D3_k) +b) is used to generate a gating coefficient, where σ is a sigmoid function, Wr,Wd are learnable parameters; Intra-scale fusion: F3_k=Gk• R3_k +(1-Gk)• D3_k); Inter-scale fusion: the F3_k of the 4 scales are spliced after being uniformly sized by upsampling, and F3 is obtained by 1x1 convolution dimension reduction. The gating coefficient Gk introduces the modal reliability estimation, and by online calculating the feature entropy of R3_k and D3_k in the current sample (the lower the entropy value, the more reliable the feature), the initial bias b is dynamically adjusted to make the reliable modal obtain a higher initial weight. For R3 and D3, R4 and D4, the same way is used to obtain second fusion features F4, F5, F6 respectively.

[0044] e: upsampling and splicing are performed on the plurality of first fusion features to generate a two-dimensional visual feature, and sampling and splicing are performed on the plurality of second fusion features to generate a depth visual feature; dimension conversion is performed on the two-dimensional visual feature and the depth visual feature to generate a first visual feature corresponding to the sub-visual image.

[0045] Here, upsampling and splicing are performed on the plurality of first fusion features to generate a two-dimensional visual feature, and sampling and splicing are performed on the plurality of second fusion features to generate a depth visual feature; dimension conversion is performed on the two-dimensional visual feature and the depth visual feature to generate a first visual feature corresponding to the sub-visual image.

[0046] wherein the feature F2 and the feature F3 are up-sampled and merged with the feature F1 to obtain a feature FTE, and the features F5 and F6 are up-sampled and merged with the feature F4 to obtain a feature FTM, and then the depth visual feature is obtained by feature fusion between the depth visual feature and the two-dimensional visual feature based on the cross attention mechanism. The two-dimensional visual feature and the depth visual feature are subjected to dimension conversion processing to generate the first visual feature corresponding to the sub visual image.

[0047] (2) In the global state intermediate supervision network model, the task instruction information, the plurality of sub visual images and the first visual feature corresponding to each sub visual image are subjected to multi-modal segmentation processing and full attention mechanism processing, and the next time visual information is output.

[0048] Here, in the global state intermediate supervision network model, the task instruction information, the plurality of sub visual images and the first visual feature corresponding to each sub visual image are subjected to multi-modal segmentation processing and full attention mechanism processing, and the next time visual information is output.

[0049] In a possible implementation, the multi-modal segmentation processing and the full attention mechanism processing of the task instruction information, the plurality of sub visual images and the first visual feature corresponding to each sub visual image in the global state intermediate supervision network model, and the output of the next time visual information, include: I: text feature extraction of the task instruction information.

[0050] II: for any sub visual image, the task instruction information and the sub visual image are subjected to multi-modal segmentation processing to determine the segmented image, and the features in the segmented image and the first visual feature of the sub visual image are subjected to cross attention processing to obtain a second visual feature.

[0051] Here, the task instruction of the robot operation needs to be input into the T5 network layer (Text-to-Text Transfer Transformer) of the model to obtain the text feature. For the left arm, the text instruction and the corresponding RGB two-dimensional image captured by the camera are input into the model to find the object to be executed in the task in the two-dimensional image and perform reasonable semantic segmentation. After that, the foreground of the segmented image is retained with the RGB information, the background information is covered with the RGB pixel (0, 0, 0), and the picture is sent to the convolution module for feature extraction to obtain the feature. The feature and the first visual feature of the left arm camera are subjected to cross attention mechanism to finally obtain the second visual feature.

[0052] III: performing full attention mechanism processing on the plurality of first visual features, the plurality of second visual features, and the text feature based on a transformer network layer, to output the next-time visual information.

[0053] Here, the next-time visual information is output by performing full attention mechanism processing on the plurality of first visual features, the plurality of second visual features, and the text feature based on a transformer network layer.

[0054] B: predicting the action block sequence based on action prediction processing on the current visual information, the task instruction information, and the current robot state in the global state intermediate supervision network model.

[0055] Here, the action block sequence is predicted by performing action prediction processing on the current visual information, the task instruction information, and the current robot state in the global state intermediate supervision network model.

[0056] The current robot state includes a left robot motion state and a right robot motion state.

[0057] In one possible implementation, the predicting the action block sequence based on action prediction processing on the current visual information, the task instruction information, and the current robot state in the global state intermediate supervision network model includes: extracting features from the current robot state to determine a plurality of current robot motion features, and performing causal attention mechanism processing on the plurality of first visual features, the plurality of second visual features, and the plurality of current robot motion features based on a transformer network layer, to predict the action block sequence.

[0058] Here, the action block sequence is predicted by performing causal attention mechanism processing on the plurality of first visual features, the plurality of second visual features, and the plurality of current robot motion features based on a transformer network layer.

[0059] In one possible implementation, the global state intermediate supervision network model is determined by: The sample visual information, sample task instruction information and sample mechanical arm state information of the embodied robot are input into an initial global state intermediate supervision network model for visual prediction processing and action prediction processing, and the predicted next time visual information of the embodied robot and the predicted action block sequence of the embodied robot are predicted; the network parameters of the initial global state intermediate supervision network model are iteratively trained based on the cross entropy loss value between the actual next time visual information, the actual action block sequence, the predicted next time visual information and the predicted action block sequence of the embodied robot, and the global state intermediate supervision network model is determined.

[0060] Here, in the training process, in order to cope with different robots with different mechanical arms, other tokens are introduced, which are obtained by combining the current mechanical arm pose through embedding and the current mechanical arm motion frequency through embedding and then passing through an MLP. At the same time, in order to introduce intermediate supervision information in the subsequent process, PAD tokens of corresponding size are added. Finally, in order to predict the final execution action, Tokenaction is added, and then the above obtained tokens are successively input into the transformer network to continuously predict action tokens. In the process of predicting visual information, PAD tokens are replaced by other position visual tokens. The purpose of this is to introduce feedback supervision information, so that the model can predict what changes will occur in the vision after the execution of the predicted action, which can make the model better understand the direction of action execution and improve the accuracy of model action prediction. The visual information generated by the left mechanical arm after executing the predicted action at time+1 needs to be predicted, and similarly, the visual information generated by the head and the right arm at time+1 also needs to be predicted. These information is represented by tokens, and then the visual tokens are decoded based on RQ-VAE and the real visual pictures are calculated by MSE loss. The design here not only makes the model sensitive to its own predicted results, but also enables different position visual information to learn from each other and improve the understanding of spatial position, thereby improving the accuracy of model prediction. In order to effectively perform the next time model visual prediction, the visual information at t+1 time can be based on the mutual attention of the visual token at t time, but the visual information at t time cannot see the visual information at t+1 time. Similarly, in order to perform action prediction, the causal attention mechanism is still used for action tokens. Finally, the dimensions of the final action token are discretized, and each continuous action dimension is mapped to 256 discrete bins. During training, the cross entropy loss of action prediction is minimized.

[0061] S103: When the embodied robot executes the action block sequence, the next time visual information and the mechanical arm state after executing the action block sequence are input into the global state intermediate supervision network model to continue visual prediction processing and action prediction processing until the embodied robot completes the target task corresponding to the task instruction information.

[0062] In this step, when the embodied robot executes the action block sequence, the next time visual information and the mechanical arm state after executing the action block sequence are input into the global state intermediate supervision network model to continue visual prediction processing and action prediction processing until the embodied robot completes the target task corresponding to the task instruction information.

[0063] It should be noted that the global state intermediate supervision network model includes a visual processing network layer and an action processing network layer. The action processing network layer is composed of a transformer network layer. The visual processing network layer includes a feature extraction network, a visual feature extraction network, a residual deep visual feature extraction network layer composed of multiple convolution layers, a channel attention module, a multi-scale gated recurrent fusion unit, a T5 network layer (Text-to-Text Transfer Transformer), a convolution module, and a transformer network layer. The feature extraction network and the visual feature extraction network are two parallel extraction networks. The residual deep visual feature extraction network layer composed of multiple convolution layers, the channel attention module, and the multi-scale gated recurrent fusion unit connected in sequence are arranged after the feature extraction network and the visual feature extraction network. The T5 network layer (Text-to-Text Transfer Transformer) in the visual processing network layer is connected with the convolution module, and the convolution module is connected with the transformer network layer and the multi-scale gated recurrent fusion unit, respectively.

[0064] The embodiment of the present application provides a processing method of a body robot, and the processing method comprises the following steps: acquiring current visual information, task instruction information and current mechanical arm state information of the body robot for a target scene; inputting the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, and predicting next time visual information of the body robot and an action block sequence of the body robot; after the body robot performs the action block sequence, inputting the next time visual information and a mechanical arm state after performing the action block sequence into the global state intermediate supervision network model to continue the visual prediction processing and the action prediction processing until the body robot completes a target task corresponding to the task instruction information.

[0065] Please refer to Figure 2 、 Figure 3 , Figure 2 Figure 1 is a structural schematic diagram of a processing device of a body robot provided by the embodiment of the present application. Figure 3 Figure 2 is another structural schematic diagram of the processing device of the body robot provided by the embodiment of the present application. Figure 2 As shown in Figure 1, the processing device 200 of the body robot comprises: An acquisition module 210 is configured to acquire current visual information, task instruction information and current mechanical arm state information of the body robot for a target scene. A first determination module 220 is configured to input the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, and predict next time visual information of the body robot and an action block sequence of the body robot. A second determination module 230 is configured to, after the body robot performs the action block sequence, input the next time visual information and a mechanical arm state after performing the action block sequence into the global state intermediate supervision network model to continue the visual prediction processing and the action prediction processing until the body robot completes a target task corresponding to the task instruction information.

[0066] Further, the first determining module 220 is configured to input the current visual information, task instruction information and the current mechanical arm state information into the pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, and predict the next time visual information of the embodied robot and the action block sequence of the embodied robot. The global state intermediate supervision network model is used for visual prediction processing on the current visual information and the task instruction information to predict the next time visual information, wherein the current visual information is composed of a plurality of sub visual images. The global state intermediate supervision network model is used for action prediction processing on the current visual information, the task instruction information and the current mechanical arm state to predict the action block sequence.

[0067] Further, the first determining module 220 is configured to perform visual prediction processing on the current visual information and the task instruction information based on the global state intermediate supervision network model to predict the next time visual information. The global state intermediate supervision network model is used for depth image feature extraction, bidirectional cross-attention fusion processing and gate cycle fusion processing on the sub visual images obtained by the left mechanical arm, the right mechanical arm and the head of the embodied robot respectively, and outputs the first visual feature corresponding to each sub visual image. The global state intermediate supervision network model is used for multi-modal segmentation processing and full attention mechanism processing on the task instruction information, a plurality of sub visual images and the first visual feature corresponding to each sub visual image, and outputs the next time visual information.

[0068] Further, the first determining module 220 is configured to output the first visual feature corresponding to each sub visual image of the left mechanical arm. The features extracted from the sub visual images by the feature extraction network and the visual feature extraction network based on the global state intermediate supervision network model are fused to obtain an RGB visual feature. A residual depth visual feature extraction network layer composed of a plurality of convolutional layers is used for feature extraction on the RGB visual feature to generate a plurality of bottom features of the RGB visual feature, and a depth feature extraction network is used for depth feature extraction on the plurality of bottom features to obtain a depth feature corresponding to each bottom feature. Bidirectional cross-attention processing is performed on each bottom feature and the corresponding depth feature to obtain a plurality of first fusion features. Performing pyramid pooling processing on each of the underlying features and the corresponding deep features to generate multiple sub-features with different receptive fields, and performing gated cyclic fusion processing on each of the underlying features and the multiple sub-features of the corresponding deep features to obtain multiple second fused features; Performing upsampling and splicing processing on the plurality of the first fused features to generate a two-dimensional visual feature, and performing upsampling and splicing processing on the plurality of the second fused features to generate a deep visual feature; Dimension conversion is performed on the two-dimensional visual feature and the depth visual feature to generate a first visual feature corresponding to the sub-visual image.

[0069] Furthermore, the first determination module 220 is configured to perform multimodal segmentation processing and full attention mechanism processing on the task instruction information, the plurality of sub-visual images, and the first visual feature corresponding to each sub-visual image in the global state intermediate supervision network model, and output the visual information at the next moment: Performing text feature extraction on the task instruction information; For any of the sub-visual images, performing multimodal segmentation processing on the task instruction information and the sub-visual image to determine a segmented image, and performing cross-attention processing on features in the segmented image and the first visual features of the sub-visual image to obtain a second visual feature; Based on the transformer network layer, a full attention mechanism is performed on the multiple first visual features, the multiple second visual features and the text features to output the visual information at the next moment.

[0070] Furthermore, the first determination module 220 is configured to perform action prediction processing on the current visual information, the task instruction information, and the current robotic arm state based on the global state intermediate supervision network model to predict the action block sequence: Extract features of the current robotic arm state and determine multiple current robotic arm motion features; Based on the transformer network layer, a causal attention mechanism is performed on the multiple first visual features, the multiple second visual features, and the multiple current robotic arm motion features to predict the action block sequence.

[0071] Further, such as Figure 3 As shown, the processing device 200 of the embodied robot further includes a model training module 240. The model training module 240 determines the global state intermediate supervisory network model through the following steps: The sample visual information, sample task instruction information and sample mechanical arm state information of the embodied robot are input into an initial global state intermediate supervision network model to perform visual prediction processing and action prediction processing, and the predicted next time visual information of the embodied robot and the predicted action block sequence of the embodied robot are predicted. The network parameters of the initial global state intermediate supervision network model are iteratively trained based on the cross entropy loss value between the actual next time visual information, the actual action block sequence, the predicted next time visual information and the predicted action block sequence of the embodied robot, and the global state intermediate supervision network model is determined.

[0072] An embodiment of the present application provides a processing device of an embodied robot, the processing device comprising: an acquisition module configured to acquire current visual information, task instruction information and current mechanical arm state information of an embodied robot for a target scene; a first determination module configured to input the current visual information, the task instruction information and the current mechanical arm state information into a pre-trained global state intermediate supervision network model to perform visual prediction processing and action prediction processing, and predict next time visual information of the embodied robot and an action block sequence of the embodied robot; and a second determination module configured to input the next time visual information and the mechanical arm state after the action block sequence is executed into the global state intermediate supervision network model to continue the visual prediction processing and the action prediction processing until the embodied robot completes a target task corresponding to the task instruction information. The global state intermediate supervision network model can directly determine the next time visual information and the action block sequence, and the embodied robot can successfully complete the task, so that the adaptability and execution precision of the embodied robot in a complex operation scene are significantly improved.

[0073] Please refer to Figure 4 , Figure 4 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 4. Figure 4 As shown in FIG. 4, the electronic device 400 comprises a processor 410, a memory 420 and a bus 430.

[0074] The memory 420 stores machine readable instructions executable by the processor 410, and when the electronic device 400 is running, the processor 410 and the memory 420 communicate through the bus 430. The machine readable instructions executed by the processor 410 can perform the steps of the processing method of the embodied robot in the method embodiment shown in the above Figure 1 The specific implementation can be referred to the method embodiment, and will not be described here.

[0075] The embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the computer program can execute the method as described above. Figure 1 The steps of the processing method of the embodied robot in the method embodiment are described above, and the specific implementation manners can be referred to the method embodiment, which will not be described here.

[0076] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0077] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other manners. The device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, devices or units, which can be electrical, mechanical or other forms.

[0078] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment scheme.

[0079] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0080] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a nonvolatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present application essentially or the parts of the prior art that make contributions or parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0081] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, but not to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can make modifications or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacements to some technical features. The modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for processing an embodied robot, characterized in that: The processing method comprises: Obtain the current visual information, task instruction information, and current robotic arm status information of the embodied robot for the target scene; Inputting the current visual information, task instruction information, and current robotic arm state information into a pre-trained global state intermediate supervisory network model for visual prediction processing and action prediction processing, and predicting the next moment visual information of the embodied robot and the action block sequence of the embodied robot; When the embodied robot completes the action block sequence, the visual information at the next moment and the state of the robotic arm after executing the action block sequence are input into the global state intermediate supervision network model to continue visual prediction processing and action prediction processing until the embodied robot completes the target task corresponding to the task instruction information and stops prediction.

2. The processing method according to claim 1, characterized in that The current visual information, task instruction information, and current robotic arm state information are input into a pre-trained global state intermediate supervisory network model for visual prediction processing and action prediction processing, and the next moment visual information of the embodied robot and the action block sequence of the embodied robot are predicted, including: Performing visual prediction processing on the current visual information and task instruction information based on the global state intermediate supervision network model to predict the visual information at the next moment; wherein the current visual information is composed of multiple sub-visual images; Based on the global state intermediate supervision network model, action prediction processing is performed on the current visual information, the task instruction information and the current robot arm state to predict the action block sequence.

3. The processing method according to claim 2, characterized in that The performing visual prediction processing on the current visual information and the task instruction information based on the global state intermediate supervision network model to predict the visual information at the next moment includes: performing deep image feature extraction, bidirectional cross-attention fusion processing, and gated recurrent fusion processing on the sub-visual images respectively acquired from the left robotic arm, the right robotic arm, and the head of the embodied robot in the global state intermediate supervision network model, and outputting a first visual feature corresponding to each of the sub-visual images; In the global state intermediate supervision network model, the task instruction information, the multiple sub-visual images and the first visual feature corresponding to each sub-visual image are subjected to multimodal segmentation processing and full attention mechanism processing to output the visual information at the next moment.

4. The processing method according to claim 3, characterized in that For the sub-visual image of the left robotic arm, outputting the first visual feature corresponding to each sub-visual image includes: Fusing the features extracted from the sub-visual image by the feature extraction network based on the global state intermediate supervision network model and the visual feature extraction network to obtain RGB visual features; Performing feature extraction on the RGB visual features based on a residual deep visual feature extraction network layer composed of multiple convolutional layers to generate multiple underlying features of the RGB visual features, and performing deep feature extraction on the multiple underlying features based on a deep feature extraction network to obtain a depth feature corresponding to each of the underlying features; Performing bidirectional cross-attention processing on each of the underlying features and the corresponding deep features to obtain multiple first fusion features; Performing pyramid pooling processing on each of the underlying features and the corresponding deep features to generate multiple sub-features with different receptive fields, and performing gated cyclic fusion processing on each of the underlying features and the multiple sub-features of the corresponding deep features to obtain multiple second fused features; Performing upsampling and splicing processing on the plurality of the first fused features to generate a two-dimensional visual feature, and performing upsampling and splicing processing on the plurality of the second fused features to generate a deep visual feature; Dimension conversion is performed on the two-dimensional visual feature and the depth visual feature to generate a first visual feature corresponding to the sub-visual image.

5. The processing method according to claim 3, characterized in that: The step of performing multimodal segmentation processing and full attention mechanism processing on the task instruction information, the plurality of sub-visual images, and the first visual feature corresponding to each sub-visual image in the global state intermediate supervision network model, and outputting the visual information at the next moment, includes: Performing text feature extraction on the task instruction information; For any of the sub-visual images, performing multimodal segmentation processing on the task instruction information and the sub-visual image to determine a segmented image, and performing cross-attention processing on features in the segmented image and the first visual features of the sub-visual image to obtain a second visual feature; Based on the transformer network layer, a full attention mechanism is performed on the multiple first visual features, the multiple second visual features and the text features to output the visual information at the next moment.

6. The processing method according to claim 2, characterized in that The performing action prediction processing on the current visual information, the task instruction information, and the current state of the robotic arm based on the global state intermediate supervision network model to predict the action block sequence includes: Extract features of the current robotic arm state and determine multiple current robotic arm motion features; Based on the transformer network layer, a causal attention mechanism is performed on the multiple first visual features, the multiple second visual features, and the multiple current robotic arm motion features to predict the action block sequence.

7. The processing method according to claim 1, characterized in that The global state intermediate supervision network model is determined by the following steps: Inputting the sample visual information, sample task instruction information, and sample manipulator state information of the embodied robot into the initial global state intermediate supervisory network model for visual prediction processing and action prediction processing, and predicting the predicted visual information of the embodied robot at the next moment and the predicted action block sequence of the embodied robot; Based on the actual next-moment visual information of the embodied robot, the actual action block sequence, the predicted next-moment visual information, and the cross-entropy loss value between the predicted action block sequence, the network parameters of the initial global state intermediate supervision network model are iteratively trained to determine the global state intermediate supervision network model.

8. A processing device for an embodied robot, characterized in that: The processing device comprises: The acquisition module is used to obtain the current visual information, task instruction information and current robotic arm status information of the embodied robot for the target scene; A first determination module is configured to input the current visual information, task instruction information, and current robotic arm state information into a pre-trained global state intermediate supervisory network model for visual prediction processing and action prediction processing, thereby predicting the next moment visual information of the embodied robot and the action block sequence of the embodied robot; The second determination module is used to input the visual information at the next moment and the state of the robotic arm after executing the action block sequence into the global state intermediate supervision network model after the embodied robot completes the action block sequence, and continue to perform visual prediction processing and action prediction processing until the embodied robot completes the target task corresponding to the task instruction information and stops prediction.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and the machine-readable instructions are executed by the processor to execute the steps of the processing method of the embodied robot as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the processing method of the embodied robot according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Robot positioning method and device, electronic equipment and storage medium

    CN114511625A

  • Mechanical arm control method, device and equipment based on large visual model and storage medium

    CN118143940A

  • Robot control method and device, electronic equipment and storage medium

    CN118809592A

  • Multi-robot collaborative navigation method and system based on visual language large model

    CN119756375A

  • Artificial intelligence system and method serving electric power robot

    WO2022188379A1