Operation-scene-aligned generation method for robot foundation model, system, and medium

By combining a multimodal description method using visible and infrared images, and utilizing Mask2Former and visual Transformer models, robot operation videos aligned with real-world scenarios are generated. This solves the problem of insufficient alignment in existing video generation technologies and improves the reliability of robot teaching and planning.

WO2026102709A1PCT designated stage Publication Date: 2026-05-21SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
Filing Date
2024-11-15
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing technologies struggle to generate robot operation videos that align with real-world scenarios, lacking details related to physical interactions, which impacts teaching effectiveness and the reliability of subsequent planning.

Method used

A multimodal infrared image enhancement and description method is adopted, which combines visible and infrared images, performs panoramic segmentation through the Mask2Former model, generates object structure segmentation information and description information using a multimodal large language model, and performs video generation by combining a visual Transformer model to achieve denoising and stitching, generating a video aligned with the real scene.

Benefits of technology

It improves the effectiveness of robot teaching and the reliability of subsequent planning by enhancing the description of fine-grained structural information of the scene, and generates high-quality videos that are aligned with the real scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024132384_21052026_PF_FP_ABST
    Figure CN2024132384_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to an operation-scene-aligned generation method for a robot foundation model, a system, and a medium, for use in solving the problems in the prior art of an inability to generate videos aligned with real-world scenes and the lack of details related to physical interaction, thereby affecting a teaching effect and the reliability of subsequent planning. According to the solution of the present invention, an infrared image is used to enhance additional scene background information, so as to obtain fine-grained structural information in a scene, thereby obtaining a more comprehensive video generation context prompt. Subsequently, a diffusion model is used to accurately convert enhanced prompts into high-quality video content, such that a generated video is aligned with a real-world scene, thereby enhancing a robot teaching effect and the reliability of subsequent planning.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, systems, and media for aligning and generating operational scenarios for large robot models Technical Field

[0001] This disclosure relates to the field of robot teaching and planning, and in particular to methods, systems and media for generating work scene alignment for large robot models. Background Technology

[0002] Large robot models such as GR-2 can generate robot operation videos for teaching purposes, demonstrating strong generalization ability and multi-task versatility. However, in complex general scenarios, the generated videos often fail to accurately align with the real-world scene. Summary of the Invention

[0003] The purpose of this disclosure is to propose a method for generating work scene alignment for large robot models, in order to overcome the problems of existing technologies that cannot generate videos aligned with real-world scenes, lack details related to physical interactions, and affect the teaching effect and the reliability of subsequent planning. The specific technical solution is as follows.

[0004] In a first aspect, this disclosure proposes a method for aligning and generating work scenes for large robot models. The method includes the following steps: the robot acquires a visible image and an infrared image of the same work scene simultaneously; based on the visible image, a noisy visible video is obtained, and then visible video latent features are obtained; based on the infrared image, a noisy infrared video is obtained, and then infrared video latent features are obtained; the visible video latent features and infrared latent features are spliced ​​together to obtain hybrid video latent features; based on the visible image and the infrared image, object structure segmentation information and descriptive information about the work scene are acquired; based on the segmentation information, spatial condition features are obtained; based on the descriptive information, text features are obtained; the spatial condition features and the text features are used as multimodal conditions; under the guidance of the multimodal conditions, the hybrid video latent features are denoised to obtain denoised hybrid video latent features; based on the denoised hybrid video latent features, using the inverse operation of splicing, denoised visible video latent features and denoised infrared video latent features are obtained, thereby obtaining visible video and infrared video for robot work motion planning.

[0005] In one embodiment of the above technical solution, the spatial condition feature is constructed by the following formula: Where: m i For the i-th segmented region, t i Let T5 be the category of the i-th segmentation region, T5 be the text-to-text Transformer model, and N be the total number of segmentation regions.

[0006] In one embodiment of the above technical solution, the segmented region is obtained by constructing a panoramic segmentation model based on infrared and visible color images using the Mask2Former model, and the category corresponding to the segmented region is generated using a multimodal large language model prompt.

[0007] In one embodiment of the above technical solution, the descriptive information about the work scenario is generated using a multimodal large language model with prompts. The generated content includes an overview of the work scenario, its characteristics, several details, and an environmental description.

[0008] In one embodiment of the above technical solution, the visible video latent features are obtained using a first encoder based on the visible image and initial noise coding, and the infrared video latent features are obtained using a second encoder based on the infrared image and initial noise coding.

[0009] In one embodiment of the above technical solution, the steps for acquiring the visible video and infrared video used for robot operation motion planning include: using a first decoder to decode the latent features of the denoised visible video to obtain the visible video; and using a second decoder to decode the latent features of the denoised infrared video to obtain the infrared video.

[0010] In one embodiment of the above technical solution, under the guidance of multimodal conditions, a preset diffusion model is used to denoise the latent features of the hybrid video. In the diffusion model, a cross-attention mechanism is used to process the multimodal conditions.

[0011] In one embodiment of the above technical solution, the diffusion model is a visual Transformer model. In each block of the visual Transformer, for the features processed by the attention mechanism, a first cross-attention mechanism is first used to establish the association and fusion between the feature and the spatial condition feature to obtain a first feature, and then a second cross-attention mechanism is used to establish the association and fusion between the first feature and the text feature to obtain a second feature.

[0012] Secondly, this disclosure proposes a computer-readable storage medium storing a computer program that can be loaded by a processor and executed by any of the methods described above.

[0013] Thirdly, this disclosure proposes a robot, which includes a first camera, a second camera, and a video generation module. The video generation module includes a conditional unit, a hybrid video latent feature unit, a noise reduction unit, and a video unit. The first camera is used to acquire a visible image of the robot's working scene, and the second camera is used to acquire an infrared image of the same robot working scene. The hybrid video latent feature unit is configured to obtain a noisy visible video based on the visible image, thereby acquiring visible video latent features; to obtain a noisy infrared video based on the infrared image, thereby acquiring infrared video latent features; and to stitch the visible video latent features and the infrared latent features together to obtain hybrid video latent features. The conditional unit... The system is configured to acquire object structure segmentation information and descriptive information about the work scene based on the visible image and the infrared image. Based on the segmentation information, spatial condition features are obtained, and based on the descriptive information, text features are obtained. The spatial condition features and the text features are used as multimodal conditions. The denoising unit, guided by the multimodal conditions, denoises the hybrid video latent features to obtain denoised hybrid video latent features. The video unit, based on the denoised hybrid video latent features, uses the inverse operation of splicing to obtain denoised visible video latent features and denoised infrared video latent features, thereby obtaining visible video and infrared video for robot motion planning.

[0014] The beneficial technical effects of this disclosed solution are as follows: Infrared images are used to enhance scene information, thereby obtaining fine-grained structural information within the scene and thus acquiring more comprehensive contextual cues for video generation. Subsequently, by designing an efficient visual Transformer fine-tuning method, these enhanced cues are accurately converted into high-quality video content, ensuring that the generated video aligns with the real-world scene, enhancing the robot's teaching effectiveness and the reliability of subsequent planning. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1. Schematic diagram of the initial scene description of multimodal mode.

[0017] Figure 2. Schematic diagram of multimodal video generation.

[0018] In Figure 1, 100 represents panoramic segmentation, 200 represents dense description generation, a is a schematic of a visible image, b is a schematic of a region segmented by object structure, c is a schematic of an infrared image, d is a schematic of dense text annotation, 310 is the predicted segmented region category, and 320 is dense description annotation. In Figure 2, 510 is the initial video conditioned on a visible light image, 520 is the initial video conditioned on an infrared image, 530 is the hybrid video latent feature, 540 is the denoised hybrid video latent feature, 550 is the visible light video, 560 is the infrared video, 600 is the diffusion model, 610 is the first encoder, 620 is the second encoder, 630 is the first decoder, 640 is the second decoder, 650 represents the repeated iterative denoising operation, 311 is the spatial conditional feature, 321 is the text feature, and 300 is the multimodal condition. Detailed Implementation

[0019] Existing technologies often rely solely on image and text conditions for video generation, failing to accurately capture fine-grained structural information within the scene, especially details related to physical interactions. This results in video generation often using the current viewpoint image as an alignment condition. In complex real-world scenarios, relying solely on the current viewpoint image is insufficient to describe the scene's intricate physical properties, leading to videos that are difficult to accurately align with the real-world environment. This impacts teaching effectiveness and the reliability of subsequent planning.

[0020] Based on this, this paper proposes a scene alignment method based on multimodal infrared enhancement description to reduce the deviation between the generated video and the actual scene. This method is of great significance in the field of robot teaching and planning, as it can predict the physical interaction behavior of objects in the simulated scene, thereby optimizing subsequent action planning.

[0021] The following description, in conjunction with the accompanying drawings, clearly and completely describes how the technical solution of this case is implemented. Obviously, the described embodiments are only a part of the embodiments of this case, and not all of them. Based on the embodiments in this case, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0022] (I) Initial Scene Description

[0023] Compared to complex real-world scenarios, it's difficult to describe the complex physical properties of a scene using only the current viewpoint image. Currently, video generation often uses the current viewpoint image as an alignment condition. Considering the complementarity of infrared and visible image information, and the fact that the operational scenarios for large robot models are characterized by easily accessible semi-open scenes with infrared images, this approach uses both infrared and visible images for initial scene description. This initial scene description serves as supplementary background information, providing context for video generation and controlling its control.

[0024] It is evident that the image is preferably a color image (see Figure 1a, which illustrates a specific image). Figure 1c is the infrared image of the color image in Figure 1a.

[0025] Referring to Figure 1b, a panoramic segmentation model based on infrared and visible color images is constructed using the Mask2Former model. After panoramic segmentation (100), several object structure segmentation regions are obtained. In the segmentation, the infrared image can enhance scene information, which is beneficial for obtaining fine-grained structural information in the scene.

[0026] Constructing segmented images based on object structure region segmentation I c Fine-grained object segmentation information M = [(m i , t i )] i , where m i ∈R W×H×1 Let t be the i-th segmented region, and t i For the corresponding category description. R W×H×1 It is a real-valued domain with spatial dimensions of W×H×1. W is the width of the image, H is the height of the image, and 1 represents the number of category channels. Category descriptions are generated through prediction using a multimodal large language model (MLLM): t i =MLLM(m i prompt i )

[0027] among them prompt i = "Briefly describe the area" is a prompt for description generation. Taking the image visible in Figure 1 as an example, the predicted segmentation region category (310) with the assistance of infrared image is keyboard, metal cover and table.

[0028] The descriptive information about the work scenario is denoted as T as a dense description, and I is... c It is a visible image, I ir Visible image I c The corresponding infrared image: T = MLLM(I c I ir ,prompt ir = "Brief description of the area"

[0029] Among them, prompt ir Provide a detailed description of the diagram, prompt ir =“Summary, Features, Detail 1, Detail 2, Context Description” are prompts used for generating dense descriptions. d in Figure 1 represents dense text annotation, illustrating dense description generation (200) and dense description annotation (320), the contents of which are shown in Table 1. It should be noted that detailed descriptions can be added as needed, not limited to Detail 1 and Detail 2 above, but increasing the number of detailed descriptions will increase the difficulty of generating MLLMs.

[0030] Table 1

[0031] After completing the initial multimodal scene description process, fine-grained distribution information and multi-dimensional description information can be obtained. Multi-dimensional description information consists of category descriptions and dense descriptions. The specific content of the dense descriptions varies from person to person. The summary is an overview of the content displayed on the color image. Its characteristic is the description of the most prominent elements. Detail 1 and Detail 2 are feature descriptions based on different granularities; for example, Detail 1 describes large, relatively clear features, while Detail 2 describes even more detailed features. The environment description is a description of the background scene features.

[0032] (II) Fine-tuning

[0033] A video generation model is constructed using a neural network. In one implementation, the video generation model includes a first video encoder (610), a second video encoder (620), a diffusion model (600), a first decoder (630), and a second decoder (640), as shown in Figure 2.

[0034] During model initialization, the visible light image in the initial scene description is used as a condition to obtain an initialization video with the visible light image as a condition (510), and the infrared image in the initial scene description is used as a condition to obtain an initialization video with the infrared image as a condition (520).

[0035] A first video encoder is configured to acquire visible video latent features based on a given initial visible image and random Gaussian noise. A second video encoder is configured to acquire infrared video latent features based on a given initial infrared image and random Gaussian noise. The visible video latent features and the infrared latent features are concatenated to obtain hybrid video latent features (530). The hybrid video latent features consist of image features and noise features.

[0036] A preset diffusion model (600) is used to denoise the latent features of the hybrid video. The latent features of the hybrid video are used as input, and the generation of denoised latent features (540) is guided by additional multimodal conditions (300). The denoising operation (650) is repeated iteratively, using the denoised latent features of the hybrid video as input to the diffusion model, until the operation stops and the final hybrid latent features are obtained. The number of operations is set. The aforementioned operation process constitutes one denoising process. During each denoising operation, the diffusion model incorporates cross-attention into the internal latent features, thereby controlling the changes in the internal latent features during the denoising process.

[0037] The inverse operation of stitching together the final hybrid video latent features yields denoised visible light video latent features and denoised infrared video latent features. Based on the denoised visible light video latent features, a generated visible light video (550) is obtained using the configured first video decoder. Based on the denoised infrared video latent features, a generated infrared video (560) is obtained using the configured second video decoder.

[0038] The multimodal conditions involved in the above process include textual features (321) and spatial conditional features (311), which serve as cues for fine-tuning. The textual features are extracted based on dense descriptive text using model T5. T5 is a text-to-text Transformer model.

[0039] Spatial condition features are constructed using model T5 based on visible images from the initial scene description in the following manner:

[0040] Among them, T5(t) i )∈R 1×1×E Description of category t i Features, generating fine-grained spatial conditional features M cond ∈R W×H×E ∑ represents the summation symbol. N represents the total number of partitioned regions. M cond This helps to control and further enhance the alignment effect.

[0041] The aforementioned diffusion model can employ a visual Transformer. By designing an efficient visual Transformer fine-tuning method, these enhanced cues can be accurately converted into high-quality video content, making the generated video aligned with the real-world scene, thereby enhancing the robot's teaching effect and the reliability of subsequent planning.

[0042] For example, the designed visual Transformer has 30 blocks. In each block, the features processed by the attention mechanism are first fused with spatial condition features using a first cross-attention mechanism to obtain a first feature, and then fused with text features using a second cross-attention mechanism to obtain a second feature.

[0043] Specifically, the features processed by the attention mechanism in each block are denoted as x'. The first cross-attention mechanism is processed as follows: x1 = attention(x', down2d(M cond ), t)

[0044] Where t is the current iteration number. down2d will M condThe network is downsampled to a 4×4 size, and the internal depth features x′ are controlled by fine-grained spatial conditions fused using standard attention. Table 2 shows the downsampling network parameters for the down2d module.

[0045] Next, the second cross-attention mechanism is used to process the following: x2 = attention(x1, T5(T), t)

[0046] Table 2

[0047] The training method described above is a commonly used model training method. After the model is trained, an infrared camera can be deployed on the robot to obtain complementary modes to the visible light images. The text descriptions obtained from the visible and infrared images can be used to guide and predict the physical interaction behavior of objects in the work scene, which is then used for robot action planning.

[0048] In summary, this project enhances the expressive power of fine-grained spatial features through a multimodal large language model and supplements the shortcomings of the visible light modality using infrared modality, constructing a multi-layered scene background description that integrates multimodal information. Furthermore, a visual Transformer model is used for fine-tuning to more accurately align with real-world scenes and generate realistic robot operation videos, thereby effectively supporting the robot's automatic planning and teaching.

[0049] The visual Transformer in this case can be replaced by the Unet neural network.

[0050] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0051] Through the above description of the embodiments, those skilled in the art can clearly understand that the method of this invention can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, dedicated CPUs, dedicated memory, dedicated components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the purposes of this disclosure, software program implementation is more often a preferred implementation method.

[0052] Through the above description of the embodiments, those skilled in the art can clearly understand that a robot can be implemented based on the method of this invention. Exemplarily, the robot includes a first camera, a second camera, and a video generation module. The video generation module includes a condition unit, a hybrid video latent feature unit, a denoising unit, and a video unit. The first camera is used to acquire a visible image of the robot's working scene, and the second camera is used to acquire an infrared image of the same robot working scene. The hybrid video latent feature unit is configured to obtain a noisy visible video based on the visible image, and then obtain visible video latent features; obtain a noisy infrared video based on the infrared image, and then obtain infrared video latent features; and stitch the visible video latent features and the infrared latent features together to obtain hybrid video latent features. The condition unit is configured to acquire object structure segmentation information and descriptive information about the working scene based on the visible image and the infrared image; obtain spatial condition features based on the segmentation information; obtain text features based on the descriptive information; and use the spatial condition features and the text features as multimodal conditions. The denoising unit, guided by multimodal conditions, denoises the hybrid video latent features to obtain denoised hybrid video latent features. The video unit, based on the denoised hybrid video latent features, uses the inverse operation of splicing to obtain denoised visible video latent features and denoised infrared video latent features, thereby obtaining visible video and infrared video for robot operation motion planning.

[0053] The aforementioned robot can enhance additional scene background information using infrared images and supplement the deficiencies of the visible light modality with infrared modality, thereby obtaining fine-grained structural information in the scene and generating a more comprehensive video context for constructing multimodal conditions. By simultaneously initializing videos conditioned by visible light images and initializing videos conditioned by infrared images, it generates visible light and infrared videos aligned with the real scene and consistent with actual robot operations under the guidance of multimodal conditions. This provides effective support for the robot's automatic planning and teaching, improving reliability. Furthermore, the visual Transformer model is fine-tuned to accurately convert cues in multimodal conditions into high-quality video content.

[0054] Although the embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this disclosure is not limited to the specific embodiments and application fields described above. The specific embodiments described above are merely illustrative and instructive, and not restrictive. Those skilled in the art can make many other forms based on the guidance of this specification and without departing from the scope of protection of the claims of this disclosure, and all of these are within the scope of protection of this disclosure.

Claims

1. A robot-oriented large model job scene alignment generation method, characterized in that, The method includes the following steps: The robot simultaneously acquires visible and infrared images of the same work scene; Based on the visible image, a noisy visible video is obtained, and then the latent features of the visible video are obtained. Based on the infrared image, a noisy infrared video is obtained, and then the latent features of the infrared video are obtained. The latent features of the visible video and the latent features of the infrared video are spliced ​​together to obtain the hybrid video latent features. Based on the visible image and the infrared image, obtain object structure segmentation information and descriptive information about the work scene in the work scene, obtain spatial condition features based on the segmentation information, obtain text features based on the descriptive information, and use the spatial condition features and the text features as multimodal conditions; Guided by multimodal conditions, the hybrid video latent features are denoised to obtain denoised hybrid video latent features; Based on the denoised hybrid video latent features, the inverse operation of splicing is used to obtain denoised visible video latent features and denoised infrared video latent features, thereby obtaining visible video and infrared video for robot operation motion planning.

2. The method of claim 1, wherein, The spatial condition feature is constructed by the following formula: In the formula, m i is the i-th segmentation area, t i is the class of the i-th segmentation area, T5 is a text-to-text Transformer model, and N is the total number of segmentation areas.

3. The method of claim 2, wherein, The segmented regions are obtained by constructing a panoramic segmentation model based on infrared and visible color images using the Mask2Former model, and the categories corresponding to the segmented regions are generated using a multimodal large language model prompt.

4. The method of claim 1, wherein, The description of the work scenario is generated using a multimodal large language model with prompts. The generated content includes an overview of the work scenario, its characteristics, some details, and an environmental description.

5. The method of claim 1, wherein: The visible video latent features are obtained using a first encoder based on the visible image and initial noise encoding, and the infrared video latent features are obtained using a second encoder based on the infrared image and initial noise encoding.

6. The method of claim 1, wherein, The steps for acquiring the visible and infrared videos used for robot motion planning include: The first decoder is used to decode the latent features of the denoised visible video to obtain the visible video; the second decoder is used to decode the latent features of the denoised infrared video to obtain the infrared video.

7. The method of claim 1, wherein, Guided by multimodal conditions, a preset diffusion model is used to denoise the latent features of the hybrid video. In the diffusion model, a cross-attention mechanism is used to process the multimodal conditions.

8. The method of claim 7, wherein, The diffusion model is a visual Transformer model. In each block of the visual Transformer, for the features processed by the attention mechanism, the first feature is obtained by establishing the association between the feature and the spatial condition features using the first cross-attention mechanism, and then the second feature is obtained by establishing the association between the first feature and the text features using the second cross-attention mechanism.

9. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed according to any one of claims 1 to 8.

10. A robot, characterized in that The robot includes a first camera, a second camera, and a video generation module. The video generation module includes a condition unit, a hybrid video latent feature unit, a noise reduction unit, and a video unit. The first camera is used to acquire visible images of the robot's work scene, and the second camera is used to acquire infrared images of the same robot's work scene; The hybrid video latent feature unit is configured to obtain a noisy visible video based on the visible image, and then obtain visible video latent features; obtain a noisy infrared video based on the infrared image, and then obtain infrared video latent features; and stitch the visible video latent features and the infrared latent features together to obtain hybrid video latent features. The condition unit is configured to acquire object structure segmentation information and descriptive information about the work scene based on the visible image and the infrared image, obtain spatial condition features based on the segmentation information, obtain text features based on the descriptive information, and use the spatial condition features and the text features as multimodal conditions. The denoising unit is configured to denoise the hybrid video latent features under the guidance of multimodal conditions to obtain denoised hybrid video latent features. The video unit is configured to obtain denoised visible video latent features and denoised infrared video latent features based on the denoised hybrid video latent features using the inverse operation of splicing, thereby obtaining visible video and infrared video for robot operation motion planning.