Reference video segmentation method for complex action description of robot with body
By adopting the reference video segmentation framework under complex text description of the DINO-SAM model in the embodied robot vision system, the problem of difficulty in making full use of time context information in the prior art is solved, and the accurate understanding of complex action descriptions and efficient performance of video segmentation is achieved.
Patent Information
- Application Number
- CN202510510769.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In the prior art, when understanding complex action descriptions and performing reference video segmentation, it is difficult to fully utilize the time context information across frames, resulting in insufficient perception and aggregation capabilities of object actions.
The DINO-SAM model is used to construct a reference video segmentation framework under complex text descriptions, including the DINO-SAM segmentation module, action-aware aggregation module and text-mark matching module. Through the Transformer decoder and cross attention mechanism, the object action trajectory in the video time dimension is processed, and the precise matching between object markers and text descriptions is achieved.
It effectively improves segmentation performance, enhances the perception and aggregation ability of object actions, realizes accurate matching between language description and object, improves generalization ability and sample detection ability, and improves the navigation performance of embodied robots.
Smart Images

Figure CN120032302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robot vision processing, and in particular, relates to a reference video segmentation method for describing complex actions of an embodied robot. Background Art
[0002] An embodied robot is usually defined as an intelligent entity with a physical body that supports interaction with the environment. The core of this concept is "intelligence" rather than "form", that is, the robot can actively perceive the environment, understand task requirements, plan action paths, and complete tasks through interaction with the environment. For example, an embodied robot can obtain information through a variety of sensors such as vision, touch, and force, and use technologies such as reinforcement learning to make autonomous decisions. In robot vision systems, video segmentation can be used to identify and track target objects in the environment, helping robots to better navigate and perform tasks.
[0003] Reference video segmentation is to accurately segment and track specific targets in a video based on natural language descriptions. This technology can effectively use complex language descriptions to provide key information related to target motion, such as "flying away" and "circling", to understand video content. Recent datasets in this field, such as MeViS, have highlighted the key role of motion description in understanding multimodal information.
[0004] The task was first proposed by Gavrilyuk et al. in the paper "Actor and Action Video Segmentation from a Sentence". The introduction of datasets such as Ref-DAVIS17, Ref-Youtube-VOS and MeViS has promoted the continuous development of this field. Many early segmentation methods in the field mainly rely on object segmentation of videos frame by frame, focusing on the separate processing of image features and text features, while ignoring the association in the temporal dimension. SeongukSeo et al. proposed a large-scale reference video object segmentation benchmark and a unified segmentation framework URVOS in "URVOS: unified referring video object segmentation network with a large-scale benchmark". Early methods mainly rely on complex pipelines. In order to simplify the workflow, Jiannan Wu et al. adopted a query-based end-to-end framework to decode objects in multimodal features in the paper "Languages as queries for referring video object segmentation". However, in the decoding process, only the object information within each frame is focused on, and the temporal context information across frames is not fully utilized. The recently proposed MeViS dataset aims to highlight the importance of motion expression and point out the shortcomings of existing methods in understanding motion information in language and video.
[0005] With the development and rise of large models pre-trained on web-scale datasets, the Segment Anything Model (SAM) proposed by Alexander Kirillov et al. in the paper "Segment Anything" has attracted widespread attention as an important large-scale model in the field of computer vision. SAM can generate high-quality object masks based on a variety of prompts (such as points, boxes, and text), and has shown strong zero-shot capabilities in multiple segmentation tasks. However, the current version of SAM does not support complex text prompts for reference object segmentation and other advanced tasks that require semantic understanding. Summary of the invention
[0006] The purpose of the present invention is to provide a reference video segmentation method for complex action description of an embodied robot, which is used to segment objects that meet the text description in the video through natural language prompts, thereby helping the embodied robot to better navigate and perform tasks.
[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A reference video segmentation method for complex action description of an embodied robot comprises the following steps: S1, input text prompts and video frames into the DINO-SAM segmentation module; S2, based on a frame image in T-frame video Get the video features of the video with the corresponding video text description and its corresponding mask feature ,in H.W is the length and width of the single-frame image feature tensor, C is the number of channels; S3, based on video features and language features Get output action-aware aggregation tags ,in is the length of the text; S4, based on object action labeling and language features And video mask features Get the object segmentation mask.
[0008] Furthermore, the specific process of step S2 is as follows: S21, obtained through the Grounding DINO model Object detection boxes corresponding to text descriptions; S22, the object detection frame is aligned with the frame image Enter into SAM as a box prompt to get Image features and mask features ; S23, will Image features and mask features Get the overall features of the video and its corresponding mask feature .
[0009] Furthermore, the specific process of step S3 is as follows: video features and language features Get video object tags from Transformer decoder , and then obtain the object tag tracking results through Hungarian matching , and finally the tracking result based on the object tag and language features The output of the action-aware aggregation module is obtained through self-attention layers and multi-layer cross-attention layers .
[0010] Furthermore, in step S4, the language feature Acquired through the cross attention mechanism Learnable Queries , and then based on the action-aware aggregation module output Labeling Learnable Queries The filtered objects are obtained through the Transformer decoder and finally combined with the video mask features obtained in step S2 Multiply frame by frame to obtain the segmentation mask of the described object.
[0011] Furthermore, the method constructs a reference video segmentation framework under complex text description based on the DINO-SAM model; the framework contains three modules: a DINO-SAM segmentation module, an action-aware aggregation module, and a text-token matching module; the DINO-SAM segmentation module realizes the preliminary segmentation of video objects and generates video features With mask features ; Motion-aware aggregation module combined with video features and language features , to process the object motion trajectory in the temporal dimension of the video; the text-token matching module realizes the effective matching between object tags and text descriptions.
[0012] Compared with the prior art, the present invention has the following beneficial effects: The present invention obtains video object tags based on video features and language features, and obtains the object motion trajectory that is consistent with the text description in the candidate object tags in the decoder, which effectively improves the segmentation performance, enhances the perception and aggregation capabilities of object movements, and realizes accurate matching between language descriptions and objects, improves generalization capabilities and few-sample detection capabilities, and effectively improves the navigation performance of embodied robots. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0014] The present invention is further described below in conjunction with the accompanying drawings and embodiments. The embodiments of the present invention include but are not limited to the following embodiments. Example
[0015] The present invention discloses a reference video segmentation method for complex action description of an embodied robot. The method is based on the DINO-SAM model and constructs a reference video segmentation framework for complex text description. The framework contains three modules: DINO-SAM segmentation module, action perception aggregation module and text-token matching module. The DINO-SAM segmentation module realizes the preliminary segmentation of video objects and generates video features. With mask features ; Motion-aware aggregation module combined with video features and language features , to process the object motion trajectory in the temporal dimension of the video; the text-token matching module realizes the effective matching between object tags and text descriptions.
[0016] like Figure 1 The following is the structure diagram of the DINO-SAM segmentation module. The initial segmentation and feature generation process of the video object is as follows: (1) First, T A frame of image in a video Together with the corresponding video text description, it is input into the Grounding DINO model. Grounding DINO adopts a dual encoder-single decoder architecture, and performs feature extraction through the image backbone (Swin Transformer) and the text backbone (BERT) to generate image features and text features. Subsequently, these features are sent to the feature enhancer for cross-modal fusion. The enhancer includes a self-attention layer, a deformable self-attention layer, an image-to-text cross-attention layer, and a text-to-image cross-attention layer, which can optimize the alignment of features of different modalities. The enhanced features are then input into the language-guided query selection module to select fused features that are more relevant to the input text as the query of the decoder. Finally, the image features, text features, and the retrieved cross-modal query are input into the cross-modal decoder to generate the same Object detection boxes corresponding to text descriptions; (2) Associate the object detection frame with the frame image As a box hint, it is input into SAM. After being processed by the image encoder and the hint encoder, the mask decoder generates Image features and mask features .
[0017] (3) After inputting the video frames into DINO-SAM in sequence, all Image features and mask features Add in the time dimension to get the video features and video mask features .
[0018] In this embodiment, the process of capturing motion information of different time scales by the motion perception aggregation module is as follows: (1) Through the cross-attention mechanism, the text description clues Inject into Learnable static queries Get a query that extracts static and dynamic information from text descriptions : ; (2) Based on With video features Obtaining video object tags through the Transformer decoder Then, the object tags between adjacent frames are matched through Hungarian matching to obtain the object tag tracking results. : After matching Candidates Its motion trajectory ; (3) Label the video objects and language features Produce the final motion-aware aggregated markup through the motion aggregation block The action aggregation block consists of The cascaded layers consist of three core components: self-attention, multi-layer cross-attention, and feed-forward network (FFN) layers. Self-attention is used to calculate the long-range dependencies in the motion trajectory of each object in the previous layer stage: in, Indicates the number of stages of hierarchical operation of motion aggregation blocks, , Represents the object tag after combining global and local information through self-attention; (4) The similarity between the language features and the motion trajectory of each object after self-attention processing can be used to highlight the object described by the text description: in, represents the number of stages of multi-layer cross-attention layer operations, , , Represents the object's motion trajectory of T Frame and Attention map of language features. Combined and language features Richer object motion features can be obtained : in, for exist The sum is calculated over the dimension of Perform row-by-row summation, which represents language features For motion trajectory exist The sum of the effects of each frame on the frame; It can be regarded as a frame weight that represents the cross-domain The importance of the frame's motion trajectory relative to the long-term motion. The corresponding attention map is summed Perform a mark merge operation to merge two adjacent prompts into a single prompt: Among them, the merged object prompts . Multi-layer cross attention After rounds of iterations, the output of the multi-layer cross attention Combined with the input of the action aggregate: in, . Input to the feed-forward network (FFN) to get the output of the action aggregation block : in, As the next stage action aggregation block, the formula is: Input. After layer aggregation, we get the output of the action-aware aggregation module .
[0019] In this embodiment, the process of aligning the text description with the segmented object is as follows: (1) The same strategy as the action-aware aggregation module, using cross-attention to integrate language features Injected into a learnable query: in, yes Initialized learnable queries; (2) Queries injected via the cross-attention mechanism Action tags with objects Input into the Transformer decoder to obtain the object motion trajectory that matches the text description in the candidate object tag, so as to achieve accurate alignment and segmentation of the object; (3) By combining the object tags obtained by screening with the video mask features output by the DINO-SAM segmentation module Multiply frame by frame to get the object segmentation mask.
[0020] Through the above design, this method effectively improves the segmentation performance, enhances the perception and aggregation capabilities of object actions, achieves accurate matching between language descriptions and objects, and improves generalization and few-sample detection capabilities.
[0021] The above embodiment is only one of the preferred implementation modes of the present invention and should not be used to limit the protection scope of the present invention. Any changes or modifications that are made to the main design concept and spirit of the present invention and have no substantive significance, and the technical problems they solve are still consistent with the present invention, should be included in the protection scope of the present invention.
Claims
1. A reference video segmentation method for complex action description of an embodied robot, characterized in that: The following steps are involved: S1, input text prompts and video frames into the DINO-SAM segmentation module; S2, based on A frame of image in a video Get the video features of the video with the corresponding video text description and its corresponding mask feature ,in , is the length and width of the single-frame image feature tensor, is the number of channels; S3, based on video features and language features Get output action-aware aggregation tags ,in is the length of the text; S4, based on object action labeling and language features And video mask features Get the object segmentation mask.
2. A reference video segmentation method for complex action description of an embodied robot according to claim 1, characterized in that: The specific process of step S2 is as follows: S21, obtained through the Grounding DINO model Object detection boxes corresponding to text descriptions; S22, the object detection frame is aligned with the frame image Enter into SAM as a box prompt to get Image features and mask features ; S23, will Image features and mask features Get the overall features of the video and its corresponding mask feature .
3. A reference video segmentation method for complex action description of an embodied robot according to claim 2, characterized in that: The specific process of step S3 is: video features and language features Get video object tags from Transformer decoder , and then obtain the object tag tracking results through Hungarian matching , and finally the tracking result based on the object tag and language features The output of the action-aware aggregation module is obtained through self-attention layers and multi-layer cross-attention layers .
4. The reference video segmentation method for complex action description of an embodied robot according to claim 3, characterized in that: In step S4, the language feature Acquired through the cross attention mechanism Learnable Queries , and then based on the action-aware aggregation module output Labeling Learnable Queries The filtered objects are obtained through the Transformer decoder and finally combined with the video mask features obtained in step S2 Multiply frame by frame to obtain the segmentation mask of the described object.
5. A reference video segmentation method for complex action description of an embodied robot according to claim 4, characterized in that: Based on the DINO-SAM model, this method constructs a reference video segmentation framework under complex text description; the framework contains three modules: DINO-SAM segmentation module, action-aware aggregation module and text-token matching module; the DINO-SAM segmentation module realizes the preliminary segmentation of video objects and generates video features With mask features ; Action-aware aggregation module combined with video features and language features , to process the object motion trajectory in the temporal dimension of the video; the text-token matching module realizes the effective matching between object tags and text descriptions.
Citation Information
Patent Citations
Anaphora video segmentation method based on multi-modal query vector and confidence coefficient
CN116052040A
Automatic driving multi-mode perception decision-making method and device based on large language model
CN118115969A
Automatic prompt image segmentation method based on open context scene
CN118470321A
Marine work equipment identification method based on vision
CN119863700A
Cited By
Reference video object segmentation method and system based on motion modeling and multi-modal interaction
CN121121617A
A reference video object segmentation method and system based on motion modeling and multi-modal interaction
CN121121617B