A video object segmentation method and system based on multi-modal interactive cues

CN122799337APending Publication Date: 2026-09-22KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004451.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有技术中,提示形式多局限于单一点击或边界框,交互方式也偏向一次性

Benefits of technology

第一,通过层次化跨模态注意力融合机制,使文本语义提示能够直接调控视觉空间提示特征的响应分布,有效消除了复杂场景下单一模态提示的指代歧义,提高了目标对象指定的精确性。第二,通过在记忆读取过程中引入基于统一提示特征的空间偏置调制,使网络在检索历史帧信息时持续受当前提示约束,增强了分割网络在目标遮挡、形态剧变等挑战下的时序鲁棒性。第三,采用基于掩码残差的增量修正范式,用户追加提示后网络仅预测局部残差进行修补,在修正错误区域的同时完整保留用户已确认的正确分割区域,避免了传统全视频重新处理导致的操作反复。第四,修正传播过程仅需执行解码与残差预测操作,无需重新提取图像特征,计算开销大幅降低,实现了近即时的交互响应,显著提升了交互效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799337A_ABST
    Figure CN122799337A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video object segmentation, and discloses a video object segmentation method and system based on multi-modal interactive prompts. The method comprises the following steps: acquiring a video sequence and at least two modal prompts input by a user; encoding each modal prompt respectively, and generating unified prompt features through hierarchical cross-modal attention fusion; injecting the unified prompt features into a segmentation network based on space-time memory, generating a current frame segmentation mask by using a memory reading mechanism of a prompt condition, and storing the current frame segmentation mask in a memory bank; in response to additional prompts of the user, updating the unified prompt features through a prompt update network, driving a residual correction network to predict a mask residual, and correcting and propagating the current frame and subsequent frames. The application reduces the reference ambiguity in a multi-target scene through hierarchical fusion, enhances the time sequence robustness of long video segmentation through memory reading of a prompt condition, retains a confirmed correct area through incremental residual correction, and significantly improves the interactive correction efficiency and segmentation precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and video processing technology, and in particular to a video object segmentation method and system based on multimodal interactive prompts. Background Technology

[0002] Semi-automatic video object segmentation tasks require algorithms to continuously and accurately segment target objects throughout a video sequence based on prompts provided by the user in the initial frame. Existing technologies often limit prompts to a single click or bounding box, and the interaction methods tend to be one-off.

[0003] However, in practical applications, common challenges include: First, in complex scenes containing multiple similar-looking objects, single-modal cues (such as those with only a few click points) often lead to ambiguity due to insufficient spatial information, making it difficult for users to accurately convey their intentions. Second, in long videos, targets may undergo severe occlusion or significant deformation, causing segmentation drift. While existing spatiotemporal memory-based networks have made progress in temporal propagation, they typically only utilize cues in the initial frame, with the cues fading gradually during propagation, failing to effectively constrain memory retrieval and exhibiting insufficient robustness. Third, when users discover propagation errors, there is a lack of efficient local correction methods. Re-providing cues and reprocessing is not only time-consuming but also discards correctly segmented regions already recognized by the user, resulting in repetitive operations.

[0004] Therefore, there is an urgent need for a video object segmentation scheme that can deeply integrate multimodal cues, continuously implement cues constraints during spatiotemporal propagation, and support efficient local incremental correction. Summary of the Invention

[0005] To overcome the aforementioned deficiencies of the prior art, the present invention provides a video object segmentation method and system based on multimodal interactive prompts to solve the problems existing in the background art.

[0006] This invention provides the following technical solution: a video object segmentation method based on multimodal interactive prompts, comprising the following steps: Acquire a video sequence and multimodal interactive prompts input by the user; wherein the multimodal interactive prompts contain indication information in at least two different modalities; The indication information for each modality is feature-encoded separately to obtain the corresponding modality cue features; The modal cue features described above are fused through hierarchical cross-modal attention to generate a unified cue feature; The image features of the current video frame are extracted, and the unified cue features are combined with the image features. The combined features are then input into a video object segmentation network based on spatiotemporal memory. The memory retrieval mechanism of the cue conditions is used to generate an initial segmentation mask for the current frame. The memory retrieval mechanism of the cue conditions refers to the spatial modulation of the attention operation using the unified cue features when retrieving historical frame features from the memory bank. The image features, the initial segmentation mask, and the unified prompt features of the current frame are stored in the memory bank; In response to user-added interactive prompts, the added interactive prompts are fused with historical unified prompt features through a prompt update network to generate updated unified prompt features; Based on the updated unified cue features, the residual correction network is driven to predict the mask residual for the current frame, and the correction mask is obtained by combining it with the existing segmentation mask of the current frame. The memory is updated using the corrected mask and the updated unified cue feature, and then the mask residual is propagated and the mask is updated for subsequent frames based on the memory along the time axis.

[0007] Furthermore, the multimodal interactive prompts include at least two of the following modalities: spatial point prompts, bounding box prompts, text description prompts, and doodle mask prompts; The feature encoding of the indication information for each modality includes: The spatial point prompts are encoded as Gaussian heatmaps centered on the positive and negative point positions; The bounding box prompt is encoded as a binary mask image; The text description prompt is mapped into a text feature vector through a pre-trained text encoder; The graffiti mask prompts are extracted into mask feature maps using a convolutional neural network.

[0008] Furthermore, the generation of unified cue features through hierarchical cross-modal attention fusion specifically includes: One or more visual cue features existing in the spatial point cue, the bounding box cue, and the graffiti mask cue are concatenated in the channel dimension and then used to generate a visual cue fusion feature through the first convolutional layer. When the text description prompt exists, the text feature vector is used as the query, and the spatial rearrangement vector of the visual prompt fusion feature is used as the key and value. A multi-head cross attention operation is performed to obtain the text-conditional visual prompt features. When the text description prompt does not exist, a learnable semantic query vector is used as the query, or a pooled vector obtained by global pooling of the visual cue fusion features is used as the query. Multi-head cross attention operation is performed on the spatial rearranged vector of the visual cue fusion features to obtain conditional visual cue features. The conditional visual cue features are concatenated or fused with the visual cue fusion features, and then processed through a second convolutional layer and a feature modulation layer to output the unified cue features.

[0009] Furthermore, the video object segmentation network based on spatiotemporal memory includes an image encoder, a memory reader, and a decoder; The image encoder is used to extract multi-scale image features of the current video frame; The memory reader uses the multi-scale image features and the unified cue features as queries to perform cross-attention calculations with the key-value pairs of historical frames stored in the memory bank; wherein, when calculating the attention weights, the attention weights are weighted element-wise using the spatial bias map generated by the unified cue features to enhance memory retrieval of cue-related regions; The decoder is used to generate a segmentation mask based on the output of the memory reader and the multi-scale image features.

[0010] Furthermore, in response to user-added interactive prompts, the added interactive prompts are fused with historical unified prompt features through a prompt update network to generate updated unified prompt features, including: The added interactive prompts are feature-encoded to obtain the added prompt features; The additional prompt features are concatenated with the historical unified prompt features corresponding to the current frame obtained from the memory bank. The concatenated result is input into a cue update network consisting of several convolutional layers and gating mechanisms, and the output is an element-wise fusion weight map. The updated unified prompt features are obtained by weighting and combining the historical unified prompt features and the additional prompt features using the fusion weight graph.

[0011] Furthermore, the residual correction network shares the image encoder and part of the decoder with the video object segmentation network, and adds a residual prediction branch; The step of driving the residual correction network to predict the mask residual for the current frame based on the updated unified cue features includes: After combining the updated unified cue features with the image features of the current frame, the data is input into the decoder. The residual prediction branch outputs a residual map with a value range within a preset range, which serves as the mask residual.

[0012] Furthermore, the step of performing mask residual propagation and mask update on subsequent frames includes: For each subsequent frame after the current frame, based on the existing segmentation mask of that frame, the memory reader is used to read the memory bank in combination with the updated unified cueing features, and the residual correction network predicts the mask residual of that frame. The original segmentation mask of the frame is added to the predicted mask residual, and the updated segmentation mask is obtained after thresholding.

[0013] A video object segmentation system based on multimodal interactive cues, comprising: A multimodal cue acquisition module is used to acquire video sequences and user-inputted multimodal interactive cue messages; the multimodal interactive cue messages contain indication information in at least two different modalities; The feature encoding module is used to encode the indication information of each modality separately to obtain the corresponding modality cue features; The hierarchical cross-modal fusion module is used to fuse the modal cue features of each modality through hierarchical cross-modal attention to generate a unified cue feature; The spatiotemporal memory segmentation module is used to extract the image features of the current video frame, combine the unified prompt features with the image features, generate an initial segmentation mask for the current frame using the memory reading mechanism of the prompt conditions, and store the image features, the initial segmentation mask and the unified prompt features of the current frame into the memory bank. The interactive correction module is used to respond to user-added interactive prompts, generate updated unified prompt features through the prompt update network, drive the residual correction network to predict the mask residual for the current frame and obtain the corrected mask, and perform mask residual propagation and mask update for subsequent frames based on the memory.

[0014] Furthermore, the spatiotemporal memory segmentation module includes an image encoder, a memory reader, and a decoder; When calculating cross-attention with historical frame features in the memory bank, the memory reader uses the spatial bias map generated by the unified cue features to spatially modulate the attention weights.

[0015] A computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the method as described in any one of the preceding descriptions.

[0016] The technical effects and advantages of this invention are as follows: First, a hierarchical cross-modal attention fusion mechanism enables textual semantic cues to directly regulate the response distribution of visual spatial cue features, effectively eliminating the referential ambiguity of single-modal cues in complex scenes and improving the accuracy of target object designation. Second, by introducing spatial bias modulation based on unified cue features during memory retrieval, the network is continuously constrained by the current cue when retrieving historical frame information, enhancing the temporal robustness of the segmentation network under challenges such as target occlusion and drastic morphological changes. Third, an incremental correction paradigm based on mask residuals is adopted. After the user adds a cue, the network only predicts local residuals for repair, completely preserving the correctly segmented regions confirmed by the user while correcting erroneous regions, avoiding the repetitive operations caused by traditional full-video reprocessing. Fourth, the correction propagation process only requires decoding and residual prediction operations, without the need to re-extract image features, significantly reducing computational overhead and achieving near-instantaneous interactive response, thus significantly improving interaction efficiency. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a video object segmentation method based on multimodal interactive prompts, provided as an embodiment of the present invention.

[0018] Figure 2 This is a structural block diagram of a video object segmentation system based on multimodal interactive prompts, provided for an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of the structure of the multimodal cue feature encoding and hierarchical cross-modal fusion module in an embodiment of the present invention.

[0020] Figure 4 This is a schematic diagram of a spatiotemporal memory-based video object segmentation network structure that includes a prompting condition memory retrieval mechanism in an embodiment of the present invention.

[0021] Figure 5 This is a schematic diagram of the process of interactive correction and mask residual propagation after adding prompts in an embodiment of the present invention. Detailed Implementation

[0022] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.

[0023] Please see Figure 1 As shown, this invention provides a video object segmentation method based on multimodal interactive prompts. The method includes the following steps.

[0024] Step 1: Acquire video sequence and multimodal interactive prompts Users import a video sequence to be processed and input multimodal interactive prompts on a specific frame (usually the first frame) through a graphical user interface. These multimodal interactive prompts contain at least two different modalities of indication information. Specifically, selectable modalities include: spatial point prompts, bounding box prompts, text description prompts, and doodle mask prompts. In an exemplary scenario, the user clicks a positive dot on the target person, clicks a negative dot in the background area, and simultaneously drags the mouse to draw a bounding box to roughly select the target person, and inputs the text description "pedestrian in a red coat". The information from these four modalities together constitutes the multimodal prompt set for this interaction. If the user does not provide a prompt for a certain modality, the input for that modality is empty, and the corresponding feature map or vector can be set to a zero tensor in subsequent feature encoding steps.

[0025] Step 2: Perform feature encoding on the indication information of each modality. For different modalities of cue information, corresponding encoding methods are used to generate modal cue features, see [link to documentation]. Figure 3 The left side.

[0026] For spatial point cues, a single-channel Gaussian heatmap with the same resolution as the input video frame is generated. Specifically, at positive point locations, a positive response with an amplitude of 1 is generated using a two-dimensional Gaussian kernel (exemplarily a Gaussian radius of 4 pixels and a standard deviation σ=2) centered on the pixel coordinates; at negative point locations, a Gaussian response is also generated, but with an amplitude of -1. After linearly superimposing the Gaussian responses of all positive and negative points on the same heatmap, the pixel values ​​are cropped to the interval [-1,1] to obtain the point heatmap.

[0027] For bounding box hints, generate a single-channel binary mask with the same frame resolution, set the pixel values ​​of the inner region of the bounding box to 1, and set the pixel values ​​of the outer region to 0, to obtain the box mask.

[0028] For graffiti masking, if the user has made a rough graffiti annotation on the current frame, a lightweight network consisting of three convolutional layers is used for feature extraction. Each convolutional layer has a 3×3 kernel size, a stride of 1, padding of 1, and uses ReLU activation. The network output is a mask feature map whose spatial dimensions match those of a feature map in an intermediate layer of the subsequent image encoder. If there is no graffiti hint, this feature map is filled with zeros.

[0029] For text description prompts, a pre-trained text encoder is used to map them into fixed-length text feature vectors. For example, the text branch of the contrastive language-image pre-trained model CLIP is used as a frozen feature extractor to convert the text description into a 512-dimensional vector. This vector is then mapped to a 256-dimensional text feature vector through a two-layer fully connected network for subsequent fusion.

[0030] Step 3: Generate unified cue features through hierarchical cross-modal attention fusion. This step deeply fuses the aforementioned multimodal cue features to generate a unified cue feature with complete information. See the fusion structure below. Figure 3 .

[0031] First, the spatial visual cue features are concatenated along the channel dimension. Specifically, the point heatmap, bounding box mask map, and mask feature map (if present) are concatenated along the channel dimension and then fed into a first convolutional layer (kernel size 3×3, stride 1, padding 1) consisting of a convolutional layer, a batch normalization layer, and a ReLU activation function, outputting a visual cue fusion feature with 64 channels. The spatial dimensions of this feature are consistent with or appropriately downsampled from the input cue map.

[0032] Next, depending on whether the user provides a text description prompt, different query construction methods are used to perform cross-modal attention operations.

[0033] Scenario 1: When the user provides a text description suggestion. Using the text feature vector as the query and the spatial rearrangement vector of the visual cue fusion features as the key and value, a multi-head cross-attention operation is performed. Specifically, the text feature vector obtained in step two is mapped to the attention query vector through a linear layer. Integrating visual cues with features The bond is obtained after spatial rearrangement. Sum In the exemplary configuration, 4-head attention is used. For single-head attention, the calculation process is as follows: Multi-head attention maps Q, K, and V to multiple subspaces and computes the attention functions in parallel. The outputs of each head are then concatenated and linearly projected. In this way, textual semantics actively selects and strengthens spatial locations in visual cue features that are related to semantic description, resulting in text-conditional visual cue features.

[0034] Scenario 2: When the user does not provide a textual description prompt. That is, if the user only provides at least two visual modalities from spatial points, bounding boxes, and doodle masks, then an alternative query mechanism is used to perform the attention operation.

[0035] One exemplary approach is to set a learnable semantic query vector. This vector, as part of the network parameters, is automatically learned and implicitly encodes general target semantics during training. Using this learnable vector as a query, visual cues are fused with features. The spatially rearranged vectors are subjected to the same multi-head cross-attention operation described above to obtain conditional visual cue features.

[0036] Another exemplary approach is to fuse visual cues with features. Global average pooling is performed, and the pooling result is mapped through a linear layer to serve as the query vector. Multi-head cross-attention is also performed on this vector. This pooling vector aggregates global spatial information from all visual cues, providing adaptive attention guidance even without textual guidance.

[0037] Both alternatives can be used, and both are conventional choices that can be made by those skilled in the art based on this specification.

[0038] Then, the conditional visual cue features are spliced ​​or fused with the original visual cue fusion features. For example, the conditional visual cue features can be combined with... The concatenation is performed along the channel dimension. The concatenated result is fed into a second convolutional layer (1×1 kernel size) and a feature modulation layer for processing, outputting a unified cue feature with 128 channels. .

[0039] This fusion approach organically integrates multimodal information through gated splicing, cross-attention, and feature modulation, enabling unified cue features to simultaneously carry detailed spatial location information, shape priors, and high-level semantic attributes. Regardless of whether the user provides a combination of modalities containing text or purely visual information, effective unified cue features can be generated, eliminating the rigid dependence of the solution on the text modality.

[0040] Step 4: Initial Video Object Segmentation and Memory Storage The unified cue features generated in step three are injected into the spatiotemporal memory-based video object segmentation network to generate an initial segmentation mask for the current frame and update the memory database. See [link to network structure] for details. Figure 4 .

[0041] The network comprises three main components: an image encoder, a memory reader, and a decoder. The image encoder extracts multi-scale image features from the input video frames. For example, the image encoder uses a ResNet-50 pre-trained on the ImageNet dataset, taking the feature maps output from stage 2, stage 3, and stage 4 as multi-scale features. These feature maps have resolutions of 1 / 4, 1 / 8, and 1 / 16 relative to the input frame, respectively. (Uniform cue features...) During injection, the spatial resolution is first adjusted using bilinear interpolation to match the size of the stage 2 output feature map of the image encoder. Then, the feature map is concatenated with this feature map along the channel dimension. The concatenated joint features continue to propagate forward through subsequent layers of the encoder, allowing the cue information to be deeply embedded in subsequent feature representations.

[0042] The memory reader is responsible for retrieving historical information related to the current frame's spatiotemporal context from the memory. The memory stores key-value pairs of processed frames, where the memory key is the concatenated feature of the historical frame's stage4 feature and its binary mask downsampled feature, and the memory value is the stage4 feature of the current frame. The memory reader performs multi-head cross-attention computation with the key-value pairs in the memory, using a query composed of the current frame's stage4 feature and unified cue feature.

[0043] The key point is to introduce spatial modulation of the cue conditions when calculating attention weights. Let the original attention score matrix in the memory reader be... .in Query the position number for the current frame. This refers to the number of bank keys. Determined by the unified hint feature. After dimensionality reduction to a single channel by a 1×1 convolutional layer, and then upsampling or spatially tiling to the same size as S after the Sigmoid activation function, a spatial bias map is obtained. The final attention weight A is calculated as follows: in This is a learnable scalar parameter, and its initial value can be set to 0.1 for example. This mechanism allows the network to spontaneously pay more attention to memory regions that are related to the spatial location and semantics of the current prompt when matching historical information, effectively suppressing interference from background or similar objects.

[0044] The decoder receives the spatiotemporal aggregation features output from the memory reader and fuses them with the skip connection features from stage 3 and stage 2 of the image encoder. It then generates a segmentation mask for the current frame through progressive upsampling and convolution operations. The generated mask is a single-channel probability map with the same resolution as the input frame, which is then binarized with a threshold (e.g., 0.5) to obtain the initial segmentation mask.

[0045] After the current frame is processed, the stage4 features of that frame, the downsampled features of the segmentation mask, and the current unified cue features are stored as new memory items in the memory bank. The memory bank is implemented using a fixed-capacity first-in-first-out queue, with an exemplary capacity of 20 frames, to balance long temporal dependencies and video memory overhead. The entire video sequence is processed frame by frame in this manner to obtain the initial segmentation mask sequence.

[0046] Step 5: Respond to user-added interactive prompts If a user finds a segmentation error in a frame (e.g., frame t) while browsing the initial segmentation results, a new interactive prompt can be added to that frame via the user interface. For example, if the target object's foot is missed in segmentation, the user can click a dot in the missed area. Upon receiving this prompt, the system triggers a subsequent local correction process. See [link to relevant documentation]. Figure 5 .

[0047] Step Six: Update Prompts and Local Residual Correction First, the added interactive prompts are feature-encoded in the same way as in step two. For example, the newly added hourly rate is encoded as a Gaussian heatmap as a feature for the added prompt.

[0048] Then, the appended cue features are fused with the original historical unified cue features of the frame through a cue update network to generate updated unified cue features. Specifically, the historical unified cue features corresponding to the t-th frame are retrieved from the memory. The spatial tiling obtained after encoding the additional cue features Instead of concatenating along the channel dimension, the concatenated result is fed into a cue update network consisting of several convolutional layers and a gating mechanism. For example, the cue update network contains two 3×3 convolutional layers, where the output of the latter convolutional layer is activated by a sigmoid function to generate an element-wise fusion weight map W of the same dimension as the historical unified cue features. This fusion weight map is then used to weight and combine the historical unified cue features and the appended cue features to obtain the updated unified cue features: in This indicates element-wise multiplication. This adaptive gating fusion can selectively introduce new information while retaining the original valid prompts.

[0049] Next, based on the updated unified prompt features The residual correction network is driven to predict the mask residual for the current frame. The residual correction network shares the image encoder and part of the decoder with the aforementioned video object segmentation network, and adds a residual prediction branch at the end of the decoder. This residual prediction branch consists of a 3×3 convolutional layer and a Tanh activation function, outputting a single-channel residual map ΔM with a value range within a preset range (e.g., [-1, 1]), which serves as the mask residual. After being combined with the image features extracted in the current frame, the data is input into the decoder, and the residual prediction branch outputs the residual map.

[0050] Use the existing segmentation mask of the current frame The corrected mask is obtained by adding the mask residual ΔM (in probabilistic form) to the mask and then performing a clipping operation: The clamp operation sets values ​​less than 0 to 0 and values ​​greater than 1 to 1 element by element. Since the residuals are only significantly non-zero in the error regions and the residuals in the correct segmentation regions are close to zero, local corrections can be achieved while completely preserving the segmentation regions that the user has previously confirmed are correct.

[0051] Step 7: Memory Update and Mask Residual Propagation After obtaining the corrected mask for frame t, the memory entry corresponding to frame t in the memory bank is updated using this corrected mask and the updated unified cue features, replacing the original mask and cue features. Subsequently, mask residual propagation is performed forward along the time axis. For each subsequent frame after frame t (frame t+1, frame t+2, ...), the existing segmentation mask for that frame is used... Based on this, instead of re-extracting image features, the image features of the frame already stored in the memory bank are directly used as memory entries. The updated memory bank is read by a memory reader in conjunction with the updated unified cue features, and the residual correction network predicts the mask residual for the frame. The original segmentation mask of the frame is added to the predicted mask residual, and then thresholded to obtain the updated binary segmentation mask: in Here, τ is the indicator function, and τ is the segmentation threshold, exemplified as 0.5. This process only requires memory read, decoding, and residual prediction operations, without rerunning the image encoder. The computational load is extremely small, enabling the correction propagation of dozens of subsequent frames in a very short time, achieving near-instantaneous interactive response.

[0052] Users can add prompts to different frames multiple times and repeat steps five through seven to iterate and gradually optimize the segmentation results of the entire video sequence until they are satisfied.

[0053] Figure 2 A system architecture block diagram for implementing the above method is shown. The system includes a multimodal cue acquisition module, a feature encoding module, a hierarchical cross-modal fusion module, a spatiotemporal memory segmentation module, and an interaction correction module.

[0054] The multimodal cue acquisition module is used to receive multimodal interactive cue input by the user through the interactive interface and parse it into structured cue data. The cue contains indication information of at least two modalities among spatial points, bounding boxes, text descriptions, and graffiti masks.

[0055] The feature encoding module includes spatial point encoding units, bounding box encoding units, text encoding units, and mask encoding units, which respectively implement the various prompt encoding functions described in step two and output the corresponding modal prompt features.

[0056] The hierarchical cross-modal fusion module is used to fuse the cue features of various modalities through hierarchical cross-modal attention to generate unified cue features. Its internal structure includes a first convolutional layer for concatenation and convolutional fusion, a multi-head cross-attention unit for performing queries based on text features, and a second convolutional layer and a feature modulation layer for further fusion. The fusion process is as described in step three.

[0057] The spatiotemporal memory segmentation module is used to implement the initial segmentation function described in step four. This module internally includes an image encoder, a memory reader, a decoder, and a memory bank. Specifically, when calculating the cross-attention with historical frame features in the memory bank, the memory reader uses a spatial bias map generated from unified cue features to spatially modulate the attention weights. The formula for calculating the attention weights is as follows: The specific implementation method is as described in step four. After processing each frame, this module stores the image features, the initial segmentation mask, and the unified prompt features into the memory bank.

[0058] The interactive correction module responds to user-added interactive prompts, implementing the correction and propagation functions described in steps six and seven. This module includes a prompt update network and a residual correction network. The prompt update network adaptively fuses the added prompt features with historical unified prompt features through several convolutional layers and gating mechanisms to generate updated unified prompt features. The fusion formula is as follows: The residual correction network shares the image encoder and part of the decoder with the spatiotemporal memory segmentation module, and adds a residual prediction branch to predict the mask residual ΔM, which is then calculated using the formula... The corrected mask is obtained. The interactive correction module is also responsible for using the updated memory to perform mask residual propagation and mask update on subsequent frames, as described in step seven.

[0059] Exemplary training method The aforementioned network can be trained end-to-end using simulated multimodal interaction data. For example, based on a publicly available video object segmentation dataset, 1 to 3 frames are randomly selected from each video sequence to generate combined cues (random combinations of dots, boxes, text, and doodles), and cues are simulated and appended to random frames after propagation. The training loss function may include binary cross-entropy loss of the initial segmentation mask, binary cross-entropy loss of the corrected mask, and optional L1 regression loss on the mask residuals. The optimizer may be Adam, with an initial learning rate set to 1e-4. The specific network structures, parameters, and training strategies described above are exemplary implementations and are not intended to limit the present invention. Any equivalent changes or modifications made based on the concept of the present invention should be included within the scope of protection of the present invention.

[0060] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video object segmentation method based on multimodal interactive prompts, characterized in that, Includes the following steps: Acquire a video sequence and multimodal interactive prompts input by the user; wherein the multimodal interactive prompts contain indication information in at least two different modalities; The indication information for each modality is feature-encoded separately to obtain the corresponding modality cue features; The modal cue features described above are fused through hierarchical cross-modal attention to generate a unified cue feature; The image features of the current video frame are extracted, and the unified cue features are combined with the image features. The combined features are then input into a video object segmentation network based on spatiotemporal memory. The memory retrieval mechanism of the cue conditions is used to generate an initial segmentation mask for the current frame. The memory retrieval mechanism of the cue conditions refers to the spatial modulation of the attention operation using the unified cue features when retrieving historical frame features from the memory bank. The image features, the initial segmentation mask, and the unified prompt features of the current frame are stored in the memory bank; In response to user-added interactive prompts, the added interactive prompts are fused with historical unified prompt features through a prompt update network to generate updated unified prompt features; Based on the updated unified cue features, the residual correction network is driven to predict the mask residual for the current frame, and the correction mask is obtained by combining it with the existing segmentation mask of the current frame. The memory is updated using the corrected mask and the updated unified cue feature, and then the mask residual is propagated and the mask is updated for subsequent frames based on the memory along the time axis.

2. The video object segmentation method based on multimodal interactive prompts according to claim 1, characterized in that, The multimodal interactive prompts include at least two of the following modalities: spatial point prompts, bounding box prompts, text description prompts, and doodle mask prompts; The feature encoding of the indication information for each modality includes: The spatial point prompts are encoded as Gaussian heatmaps centered on the positive and negative point positions; The bounding box prompt is encoded as a binary mask image; The text description prompt is mapped into a text feature vector through a pre-trained text encoder; The graffiti mask prompts are extracted into mask feature maps using a convolutional neural network.

3. The video object segmentation method based on multimodal interactive prompts according to claim 2, characterized in that, The generation of unified cue features through hierarchical cross-modal attention fusion specifically includes: One or more visual cue features existing in the spatial point cue, the bounding box cue, and the graffiti mask cue are concatenated in the channel dimension and then used to generate a visual cue fusion feature through the first convolutional layer. When the text description prompt exists, the text feature vector is used as the query, and the spatial rearrangement vector of the visual prompt fusion feature is used as the key and value. A multi-head cross attention operation is performed to obtain the text-conditional visual prompt features. When the text description prompt does not exist, a learnable semantic query vector is used as the query, or a pooled vector obtained by global pooling of the visual cue fusion features is used as the query. Multi-head cross attention operation is performed on the spatial rearranged vector of the visual cue fusion features to obtain conditional visual cue features. The conditional visual cue features are concatenated or fused with the visual cue fusion features, and then processed through a second convolutional layer and a feature modulation layer to output the unified cue features.

4. The video object segmentation method based on multimodal interactive prompts according to claim 1, characterized in that, The video object segmentation network based on spatiotemporal memory includes an image encoder, a memory reader, and a decoder; The image encoder is used to extract multi-scale image features of the current video frame; The memory reader uses the multi-scale image features and the unified cue features as queries to perform cross-attention calculations with the key-value pairs of historical frames stored in the memory bank; wherein, when calculating the attention weights, the attention weights are weighted element-wise using the spatial bias map generated by the unified cue features to enhance memory retrieval of cue-related regions; The decoder is used to generate a segmentation mask based on the output of the memory reader and the multi-scale image features.

5. A video object segmentation method based on multimodal interactive prompts according to claim 1, characterized in that, The response to user-added interactive prompts involves fusing the added interactive prompts with historical unified prompt features through a prompt update network to generate updated unified prompt features, including: The added interactive prompts are feature-encoded to obtain the added prompt features; The additional prompt features are concatenated with the historical unified prompt features corresponding to the current frame obtained from the memory bank. The concatenated result is input into a cue update network consisting of several convolutional layers and gating mechanisms, and the output is an element-wise fusion weight map. The updated unified prompt features are obtained by weighting and combining the historical unified prompt features and the additional prompt features using the fusion weight graph.

6. The video object segmentation method based on multimodal interactive prompts according to claim 1, characterized in that, The residual correction network shares the image encoder and part of the decoder with the video object segmentation network, and adds a residual prediction branch; The step of driving the residual correction network to predict the mask residual for the current frame based on the updated unified cue features includes: After combining the updated unified cue features with the image features of the current frame, the data is input into the decoder. The residual prediction branch outputs a residual map with a value range within a preset range, which serves as the mask residual.

7. A video object segmentation method based on multimodal interactive prompts according to claim 1, characterized in that, The process of propagating mask residuals and updating masks for subsequent frames includes: For each subsequent frame after the current frame, based on the existing segmentation mask of that frame, the memory reader is used to read the memory bank in combination with the updated unified cueing features, and the residual correction network predicts the mask residual of that frame. The original segmentation mask of the frame is added to the predicted mask residual, and the updated segmentation mask is obtained after thresholding.

8. A video object segmentation system based on multimodal interactive prompts, characterized in that, include: A multimodal cue acquisition module is used to acquire video sequences and user-inputted multimodal interactive cue messages; the multimodal interactive cue messages contain indication information in at least two different modalities; The feature encoding module is used to encode the indication information of each modality separately to obtain the corresponding modality cue features; The hierarchical cross-modal fusion module is used to fuse the modal cue features of each modality through hierarchical cross-modal attention to generate a unified cue feature; The spatiotemporal memory segmentation module is used to extract the image features of the current video frame, combine the unified prompt features with the image features, generate an initial segmentation mask for the current frame using the memory reading mechanism of the prompt conditions, and store the image features, the initial segmentation mask and the unified prompt features of the current frame into the memory bank. The interactive correction module is used to respond to user-added interactive prompts, generate updated unified prompt features through the prompt update network, drive the residual correction network to predict the mask residual for the current frame and obtain the corrected mask, and perform mask residual propagation and mask update for subsequent frames based on the memory.

9. A video object segmentation system based on multimodal interactive prompts according to claim 8, characterized in that, The spatiotemporal memory segmentation module includes an image encoder, a memory reader, and a decoder; When calculating cross-attention with historical frame features in the memory bank, the memory reader uses the spatial bias map generated by the unified cue features to spatially modulate the attention weights.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.