Video full-object segmentation system and method based on 3D prior enhancement and double-branch structure
By adopting 3D prior enhancement and dual-branch structure in the video whole object segmentation system, combining the stereoscopic space perception and perceptual consistency regularization terms, the problem of poor application effect of basic models in video whole object segmentation in the existing technology is solved, and a more efficient and robust segmentation effect is achieved.
Patent Information
- Application Number
- CN202510196035.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, when using the prior knowledge of the basic model to segment the whole video object, it is difficult to ensure the segmentation effect under different objects and occlusion conditions.
A video full object segmentation system based on 3D prior enhancement and dual-branch structure is adopted, including visible area segmentation branches and whole object area segmentation branches. Features are extracted and prompt embedded through image encoder and prompt encoder are generated, and segmented in combination with memory attention and mask decoder, and the stereoscopic space perception and perceptual consistency regularization terms are introduced to improve the segmentation effect.
It significantly improves the segmentation ability of the target object in the video, improves the segmentation effect of the occlusion area, and improves the robustness and segmentation accuracy of the model under different objects and occlusion conditions.
Smart Images

Figure CN120198831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video segmentation, and in particular to a video full object segmentation system and method based on 3D prior enhancement and a dual-branch structure. Background Art
[0002] In video segmentation technology, full object segmentation aims to infer the complete shape of a partially occluded object, which is more complex than traditional instance segmentation. This task not only requires identifying the visible regions of the occluded object but also needs to completely identify and segment the entire object to provide more accurate object information. Full object segmentation is of great significance for many real-world applications, such as occlusion perception of vehicles and pedestrians in the autonomous driving scenario, assisting robotic arms to achieve precise operations, and accurate positioning of instruments in surgical operations.
[0003] To address this challenge, researchers at home and abroad have proposed various methods, which mainly focus on the utilization of shape priors and the exploration of spatio-temporal consistency. Shape priors, as the core technology of full object segmentation, are widely adopted. For example, the method of shape prior based on multi-level encoding was proposed by the authors of reference [1] et al., the segmentation performance was improved by learning shape information through variational autoencoders by the authors of reference [2] et al., and a coarse-to-fine framework based on the learned shape prior was designed by reference [3] to gradually optimize the segmentation results. In terms of technologies beyond shape priors, researchers have further combined information such as spatio-temporal consistency and global perspective knowledge. For example, reference [4] addressed the occlusion problem in dynamic scenes by integrating spatio-temporal consistency and dense object motion features in the video, and reference [5] enhanced the segmentation effect by using the global perspective information of the bird's-eye view. Other works such as reference [6] restored occluded objects through a pre-trained 3D reconstruction model, while references [7] and [8] significantly improved the segmentation performance based on the powerful prior knowledge learned by the generative diffusion model.
[0004] Although existing research has made significant progress in the field of full object segmentation, the application of basic models in occlusion segmentation is still in its infancy. Currently, there are few studies exploring how to utilize the prior knowledge contained in models such as references [9] and
[10] to solve the full object segmentation problem. These basic models are usually pre-trained on large-scale data and have powerful feature extraction and generalization capabilities, but how to effectively transfer this prior knowledge to the full object segmentation task remains an open problem, and it is impossible to ensure the segmentation effect under different objects and different occlusion conditions.
[0005] Literature of the Existing Technology:
[0006] [1] L. Qi, L. Jiang, S. Liu, X. Shen, and J. Jia, “Amodal instance segmentation with kins dataset,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 3014–3023, 2019.
[0007] [2] W.-D. Jang, D. Wei, X. Zhang, B. Leahy, H. Yang, J. Tompkin, D. Ben-Yosef, D. Needleman, and H. Pfister, “Learning vector quantized shape code for amodal blastomere instance segmentation,” arXiv preprint arXiv:2012.00985, 2020.
[0008] [3] J. Gao, X. Qian, Y. Wang, T. Xiao, T. He, Z. Zhang, and Y. Fu, “Coarse-to-fine amodal segmentation with shape prior,” Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 1262–1271, 2023.
[0009] [4] J. Yao, Y. Hong, C. Wang, T. Xiao, T. He, F. Locatello, D. Wipf, Y. Fu, and Z. Zhang, “Self-supervised amodal video object segmentation,” arXiv preprint arXiv:2210.12733, 2022.
[0010] [5]K.Fan, J.Lei, X.Qian, M.Yu, T.Xiao, T.He, Z.Zhang, and Y.Fu, “Rethinking amodal video segmentation from learning supervised signals with object-centric representation,” Proceedings of the IEEE / CVF International Conference on Computer Vision, pp.1272–1281, 2023.
[0011] [6]Z.Li, W.Ye, T.Jiang, and T.Huang, “2D amodal instance segmentation guided by 3D shape prior,” European Conference on Computer Vision (ECCV), pp.165–181, 2022.
[0012] [7]E.Ozguroglu, R.Liu, D.Surís, D.Chen, A.Dave, P.Tokmakov, and C.Vondrick, “pix2gestalt: Amodal segmentation by synthesizing wholes,” 2024 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.3931–3940, 2024.
[0013] [8]G.Zhan, C.Zheng, W.Xie, and A.Zisserman, “Amodal ground truth and completion in the wild,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp.28003–28013, 2024.
[0014] [9]A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, et al., “Segment anything,” Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 4015–4026, 2023.
[0015]
[10] N. Ravi, V. Gabeur, Y.-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. C. Rolland, L. Gustafson, et al., “Sam 2: Segment anything in images and videos,” arXiv preprint arXiv:2408.00714, 2024. Summary of the Invention
[0016] The objective of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a video full-object segmentation system and method based on 3D prior enhancement and a dual-branch structure, which can efficiently transfer the powerful prior knowledge of the basic segmentation model and improve the ability to segment the complete regions of target objects in videos.
[0017] The objective of the present invention can be achieved through the following technical solutions: A video full-object segmentation system based on 3D prior enhancement and a dual-branch structure, including a visible region segmentation branch and a full-object region segmentation branch. Both the visible region segmentation branch and the full-object region segmentation branch are connected to an image encoder and a prompt encoder. The visible region segmentation branch includes a first pixel-aware memory module, a first mask decoder, a first memory bank, and a first memory encoder. The full-object region segmentation branch includes a stereo spatial-aware memory module, a second pixel-aware memory module, a third pixel-aware memory module, a second mask decoder, a second memory bank, and a second memory encoder;
[0018] The image encoder is used to extract image features from the video and output them to the visible region segmentation branch and the full-object region segmentation branch respectively;
[0019] The prompt encoder is used to obtain the prompt embeddings of the visible region segmentation branch and the full-object segmentation branch and output them to the visible region segmentation branch and the full-object region segmentation branch correspondingly;
[0020] The visible region segmentation branch and the full object region segmentation branch use the predicted metrics to perform weighted integration on the forward and backward segmentation results, obtaining the final visible region mask and full object region mask.
[0021] A video full object segmentation method based on 3D prior enhancement and a dual-branch structure includes the following steps:
[0022] S1: Obtain a video sequence of the target object;
[0023] S2: Input each frame image in the video sequence into an image encoder to extract image features;
[0024] S3: Input the bounding box of the visible region of the target object and the bounding box enlarged by two times into a prompt encoder respectively to obtain the prompt embeddings of the visible region segmentation branch and the full object segmentation branch;
[0025] S4: For the visible region segmentation branch, the image features gather information from the first memory bank storing the prediction results of previous frames through an attention mechanism, and then predict the visible region mask of the current frame through a first mask decoder, and update the first memory bank through a first memory encoder;
[0026] For the full object segmentation branch, the second memory bank includes the prediction results of previous frames and spatial perception information, predicts the full object mask of the current frame through a second mask decoder, and updates the second memory bank through a second memory encoder;
[0027] During the training process, a perceptual consistency regularization term is introduced to the full object segmentation branch to constrain the consistency between the full object prediction mask under complete memory and the full object prediction mask under incomplete memory, and learn the different memories of the visible region segmentation task and the full object segmentation task;
[0028] S5: After training is completed, perform inference on the current video frame to be segmented, process the video sequence forward and backward, and perform weighted integration on the forward and backward segmentation results in combination with the prediction metrics to obtain the final visible region mask and full object region mask.
[0029] Further, the step S1 is specifically to select two subsets Movi-B and Movi-D of the Multi-Object Video (Movi) dataset as the video sequence of the target object. Among them, Movi-B consists of objects in the CLEVR dataset, and these objects have simple shapes and regular geometries; Movi-D consists of highly realistic objects in the Google Scanned Objects dataset, with richer and more complex shapes and textures, and more serious occlusion situations.
[0030] Further, in step S2, the image encoder adopts a hybrid adaptation mechanism. By attaching LoRA modules to each query / key / value projection layer and mixing multiple LoRA weights in the feed-forward network, the original prior knowledge is retained and the knowledge of different tasks is integrated;
[0031] For the image encoder Hiera, the attached LoRA modules are connected to the query / key / value projection layers of each transformer block; for the feed-forward network, multiple LoRA weights are mixed according to a sample-specific gating function. The input x of the adapted transformer block is specifically:
[0032]
[0033] where W and b represent the original projection weights and biases, A and B represent the LoRA weights, Attn and FFN represent the self-attention and feed-forward network modules respectively, and N e represents the number of LoRA weights attached to the feed-forward network, and δ represents a learnable gating module that maps to an N e -dimensional feature vector and normalizes it with the softmax function.
[0034] Further, in step S3, the prompt encoder shares the same structure for the visible region segmentation and full object segmentation tasks, but adopts different prompt strategies. For visible region segmentation, the true visible bounding box of the target object is used as the input prompt; for full object segmentation, the size of the visible bounding box is doubled as the input prompt. During training, random noise is added to the true visible bounding box for both the visible region segmentation and full object segmentation branches, and the random noise does not exceed 10% of the side length or at most 20 pixels.
[0035] Further, in step S4, for the full object segmentation branch, a method of enhancing stereo space perception is specifically adopted to supplement the second memory bank of the full object segmentation branch. For each frame I t , in addition to the image encoder, a monocular depth estimator D is also used to predict the corresponding depth image, and then the stereo space perception embedding is extracted from the intermediate layer of D to capture detailed depth-related features. In addition, a stereo space perception memory module X vol is designed and used to store the of all processed frames. The stereo space perception memory module X vol follows the design principle of SAM (Segment Anything Model) 2.
[0036] Furthermore, the memory attention process of the three-dimensional space perception memory module and the second pixel perception memory module in the full object segmentation branch is as follows:
[0037]
[0038] Among them, MA represents the same memory attention architecture as in the original SAM2, f represents a linearly layer initialized with zeros, and x t represents the image features extracted by the image encoder, and X pix represents the original pixel perception memory bank in SAM2. In this way, the volume perception information is first aligned with the latent space related to the original pixel perception information and aggregated into In addition, the linearly layer initialized with zeros ensures that improper embeddings at the beginning of training do not distort the latent space.
[0039] Furthermore, the specific process of introducing the perception consistency regularization term into the full object segmentation branch in step S4 is as follows:
[0040] First, establish the basic objective function for segmentation:
[0041]
[0042] Among them, includes DICE loss, Focal loss, L1-based IoU (Intersection-over-Union) regression loss, and cross-entropy-based occlusion classification loss, represents the predicted mask;
[0043] To bridge the prior knowledge gap between the visible region segmentation and the full object segmentation subtasks, an auxiliary consistency regularization term for the full object segmentation branch is introduced. For each input video V, in addition to using the complete pixel-related memory X pix to obtain the prediction at each time step i, j memory items are randomly discarded, where to obtain the new prediction Then, the regularization objective function is expressed as:
[0044]
[0045] where λ is a hyperparameter.
[0046] Furthermore, the specific process of inferring the current video frame to be segmented in step S5 is as follows:
[0047] For the current video frame to be segmented, first extract the image features through the image encoder, and then input the bounding boxes of the visible regions of the target object and the bounding boxes enlarged by two times into the prompt encoder respectively to obtain the prompt embeddings of the visible region segmentation branch and the full object segmentation branch;
[0048] Next, the image features are processed by the corresponding memory attention of the visible region segmentation branch and the full object segmentation branch, gather information from the first memory bank and the second memory bank, and then generate the visible region mask and the full object mask of the current frame through the first mask decoder and the second mask decoder respectively. Finally, update the first memory bank and the second memory bank through the first memory encoder and the second memory encoder respectively to ensure continuous learning and adaptation to new video sequences.
[0049] Further, the specific process of weighted integration of the forward and reverse segmentation results in step S5 in combination with the prediction metrics is as follows:
[0050] For video frame \(I_t\) t , first use the visible region segmentation branch and the full object segmentation branch for processing to generate the predicted full object mask and the estimated IoU value \(\eta\) t , then flip the video along the time dimension to obtain its reverse-order version \(V\) ′ , where the \((T - t + 1)\)-th frame is \(I\) t , repeat the inference process on \(V\) ′ to generate the new prediction and its estimated IoU value \(\eta\) ′ t , based on these results, the final prediction is calculated through the following weighted integration:
[0051]
[0052] where \(\omega\) t is the weight coefficient calculated based on the estimated IoU value.
[0053] Compared with the prior art, the present invention has the following advantages:
[0054] The present invention designs a visible region segmentation branch and a whole object region segmentation branch. The two branches share an image encoder and a prompt encoder, and each branch is equipped with an independent memory attention module, a memory encoder, a memory bank, and a mask decoder. Among them, the whole object region segmentation branch also introduces additional spatial perception enhancement and consistency regularization techniques, effectively solving the occlusion problem in whole object segmentation. By introducing depth information to supplement the original pixel perception features, the inference ability for occluded regions is enhanced; a perception consistency regularization term is introduced to constrain the consistency between the whole object prediction masks under complete memory and incomplete memory, which is beneficial to learning different memories for the visible region segmentation task and the whole object segmentation task. In the present invention, the visible region segmentation branch and the whole object region segmentation branch use the predicted metrics to perform weighted integration on the forward and backward segmentation results to obtain the final visible region mask and whole object region mask, which can significantly improve the problem of insufficient memory in early frames and improve the overall segmentation accuracy.
[0055] The present invention proposes a dual-branch structure based on SAM2, which improves the model's perception ability for occluded regions of objects while retaining the original visible region segmentation ability. On this basis, the present invention also proposes to introduce a stereo spatial perception memory module in the whole object segmentation branch, using depth information to supplement pixel-level features, enhancing the perception of the spatial arrangement and relative position of objects, and thus improving the segmentation effect of occluded regions. In addition, during training, two datasets containing common objects and different occlusion degrees are used for training, and a consistency regularization term is introduced into the whole object segmentation branch to enable it to learn different prior knowledge from the visible region segmentation branch. When inferring a video, the trained model processes the video sequence forward and backward respectively, and performs weighted integration through the predicted metrics to compensate for the insufficient memory in early frames, enhancing the robustness of the model prediction with a reliable design and achieving better segmentation effects under different objects and different occlusion conditions. Brief Description of the Drawings
[0056] Figure 1 is a schematic diagram of the system structure of the present invention;
[0057] Figure 2 is a schematic diagram of the structure of the image encoder in the present invention;
[0058] Figure 3 is a schematic diagram of the process of consistency regularization within the whole object segmentation branch of the present invention. Detailed Embodiments
[0059] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0060] Embodiment
[0061] As Figure 1As shown in the figure, a video full-object segmentation system based on 3D prior enhancement and a dual-branch structure is divided into a shared unit and dedicated units for two branches (visible region segmentation branch and full-object segmentation branch). Among them, the shared unit includes an image encoder and a prompt encoder;
[0062] The visible region segmentation branch includes a first pixel-aware memory module, a first mask decoder, a first memory bank, and a first memory encoder. The full-object region segmentation branch includes a stereo space-aware memory module, a second pixel-aware memory module, a third pixel-aware memory module, a second mask decoder, a second memory bank, and a second memory encoder.
[0063] During training, the full-object segmentation branch additionally predicts the segmentation mask of incomplete memories for calculating the consistency regularization term. For the visible region segmentation branch, the input to the prompt encoder is the true visible region bounding box of the target object, and the visible region mask is predicted by the first mask decoder. For the full-object segmentation branch, the input to the prompt encoder is the result of doubling the visible region bounding box, and the complete object mask is predicted by the second mask decoder.
[0064] Based on the above video full-object segmentation system, a video full-object segmentation method is implemented, including the following:
[0065] 1. Obtain a video sequence of the target object. In this embodiment, two subsets Movi-B and Movi-D of the Multi-Object Video dataset are used as the experimental datasets.
[0066] 2. Input each frame image of the video into the common image encoder to extract features.
[0067] 3. Input the bounding box of the visible region of the target object and the bounding box doubled respectively into the prompt encoder to obtain the prompt embeddings of the visible region segmentation branch and the full-object segmentation branch.
[0068] 4. For the visible region segmentation branch, the image features gather information from the first memory bank storing the prediction results of previous frames through the attention mechanism, then predict the visible region mask of the current frame through the first mask decoder, and finally update the first memory bank through the first memory encoder.
[0069] 5. For the full-object segmentation branch, the second memory bank includes not only the prediction results of previous frames but also specialized spatial perception information, predicts the full-object mask of the current frame through the second mask decoder, and finally updates the second memory bank through the second memory encoder.
[0070] 6. During the training process, a perceptual consistency regularization term is introduced to the full-object segmentation branch to constrain the consistency between the full-object prediction masks under complete memory and those under incomplete memory, encouraging the model to learn the different memories of the visible region segmentation task and the full-object segmentation task.
[0071] 7. After the model training is completed, the trained model is used to infer new video frames, and the results of visible region segmentation and complete object segmentation are obtained simultaneously.
[0072] 8. A two-way integration strategy is adopted to compensate for the insufficient memory of early frames. By processing the video sequence forward and backward and combining the estimated IoU metric for weighted integration, the overall segmentation accuracy is improved.
[0073] Specifically, in step 1, the Multi-Object Video (Movi) dataset is used as the data source for the evaluation benchmark test. The Movi dataset consists of synthetic data created by the Kubric platform, containing diverse scenes and objects. In this embodiment, two specific subsets are used: Movi-B and Movi-D. Movi-B consists of objects in the CLEVR dataset, which have simple shapes and regular geometries; Movi-D consists of highly realistic objects in the Google Scanned Objects dataset, providing richer and more complex shapes and textures, as well as more severe occlusion situations. Movi-B contains a total of 936 videos, of which 702 are used for training and 234 are used for testing. Movi-D contains a total of 838 videos, of which 630 are used for training and 208 are used for testing. The object complete shape mask information is obtained during the dataset creation process.
[0074] In step 2, as Figure 2 shown, the image encoder adopts a hybrid adaptation mechanism by attaching LoRA modules to each query / key / value projection layer and mixing multiple LoRA weights in the feed-forward network to retain the original prior knowledge and fuse the knowledge of different tasks. Specifically, for the image encoder Hiera, the attached LoRA modules are connected to the query / key / value projection layers of each transformer block. For the feed-forward network, multiple LoRA weights are mixed according to a sample-specific gating function, rather than using only a single LoRA. Formally, the input x of the adapted transformer block corresponds to the following functional expression:
[0075]
[0076] where w and b represent the original projection weights and biases, A and B represent the LoRA weights, Attn and FFN represent the self-attention and feed-forward network modules, respectively, and N edenotes the number of LoRA weights attached to the feed - forward network, and δ denotes a learnable gating module that maps to N e - dimensional feature vectors and normalizes them using the softmax function.
[0077] In step 3, the prompt encoder shares the same structure for visible region segmentation and full - object segmentation tasks, but adopts different prompt strategies. For visible region segmentation, the ground - truth visible bounding box of the target object is used as the input prompt. For full - object segmentation, since there is a lack of rough full - mask information, the size of the visible bounding box is designed to be doubled as the input prompt. During training, random noise not exceeding 10% of the side length or at most 20 pixels is added to the ground - truth visible bounding box, and then it is used for both the visible region segmentation and full - object segmentation branches.
[0078] In step 4, a dual - branch structure is adopted to handle visible region segmentation and full - object segmentation tasks. Specifically, after the input frame is processed by the image encoder, the memory attention modules of the two branches respectively gather information from the visible region memory bank (i.e., the first memory bank) and the full - object memory bank (i.e., the second memory bank) into the image embedding. These outputs are further processed by two independent mask decoders and memory encoders to obtain the prediction results and update the first memory bank and the second memory bank.
[0079] In step 5, a method of stereo - spatial perception enhancement is adopted to supplement the second memory bank of the full - object segmentation branch. For each frame I t , in addition to the image encoder, a off - the - shelf monocular depth estimator D is used to predict the corresponding depth image. Then, stereo - spatial perception embeddings are extracted from the intermediate layer of D to capture detailed depth - related features. To assist the full - object segmentation task, an additional memory module X vol (i.e., the stereo - spatial perception memory module) is designed to store the of all processed frames, following the design principle of SAM2. Based on this, the original memory attention process is modified as:
[0080]
[0081] where MA represents the same memory attention architecture as in the original SAM2, f represents a linearly - layer initialized with zeros, x t represents the image features extracted by the image encoder, and X pix represents the original pixel - aware memory bank in SAM2. In this way, the volume - perception information is first aligned with the latent space related to the original pixel - perception information and aggregated into . In addition, the zero - initialization of f ensures that inappropriate embeddings do not distort the latent space at the beginning of training.
[0082] In step 6, during the training process, a perceptual consistency regularization term is introduced to the full object segmentation branch to constrain the consistency between the full object prediction mask under complete memory and the full object prediction mask under incomplete memory. A basic segmentation objective function is constructed and an auxiliary consistency regularization term is introduced. First, the basic objective function for segmentation is established as follows:
[0083]
[0084] where includes DICE loss, Focal loss, L1-based IoU regression loss, and cross-entropy-based occlusion classification loss, represents the predicted mask. To bridge the prior knowledge gap between these two subtasks, an auxiliary consistency regularization term for the full object branch is further introduced. As Figure 3 shown, the consistency regularization technique calculates an additional loss between the mask predicted with complete memory and the mask predicted after randomly discarding a part of the memory. For each input video V, in addition to using the complete pixel-related memory X pix to obtain the prediction at each time step i, this scheme also randomly discards j memory items, where to obtain the new prediction The regularization objective function is expressed as:
[0085]
[0086] where λ is a hyperparameter.
[0087] In step 7, after the model training is completed, the trained model is used to infer new video frames, and at the same time, the results of visible region segmentation and full object segmentation are obtained. For a new video frame, first, image features are extracted through an image encoder, and then the visible region bounding box of the target object and the bounding box enlarged by two times are respectively input into the prompt encoder to obtain the prompt embeddings of the visible region segmentation branch and the full object segmentation branch. Then, the image features gather information from the memory bank through the corresponding memory attention module and generate the visible region mask and full object mask of the current frame through the mask decoder. Finally, the memory bank is updated through the memory encoder to ensure that the model can continuously learn and adapt to new video sequences.
[0088] In step 8, a two-way integration strategy is adopted to compensate for the memory shortage of early frames. By processing the video sequence forward and backward and combining the estimated IoU metric for weighted integration, the overall segmentation accuracy is improved. For the t-th frame I of the video t , first, the proposed model is used for processing to generate the predicted full object mask and the estimated IoU value η t The video is then flipped along the time dimension to obtain its reverse-order version V ′ , where the (T - t + 1)-th frame is I t On V ′ Repeat the inference process to generate new predictions and its estimated IoU value η ′ t Based on these results, the final prediction is calculated by the following weighted integration:
[0089]
[0090] where ω t is the weight coefficient calculated based on the estimated IoU value. By combining these two predictions, it is ensured that each frame is inferred under sufficient memory priors, thus integrating the advantages of the two methods and improving the overall segmentation accuracy.
[0091] In summary, this solution greatly improves the ability to segment the complete region of the target object in the video by efficiently transferring the powerful prior knowledge of the basic segmentation model. This solution proposes a dual-branch structure based on SAM2, which improves the model's perception ability of the occluded region of the object while retaining the original visible region segmentation ability. On this basis, the present invention also proposes to introduce a stereo spatial perception memory module into the full object segmentation branch, supplement pixel-level features with depth information, enhance the perception of the spatial arrangement and relative position of the object, and thus improve the segmentation effect of the occluded region. In addition, during training, a consistency regularization term is introduced into the full object segmentation branch to enable it to learn different prior knowledge from the visible region segmentation branch. When inferring the video, the predicted metrics are used to perform weighted integration on the forward and backward segmentation results, greatly improving the problem of insufficient memory in the early frames, improving the overall segmentation accuracy, and enhancing the robustness of the model prediction with a reliable design, and being able to achieve better segmentation effects under different objects and different occlusion conditions.
Claims
1. A video full object segmentation system based on 3D prior enhancement and dual-branch structure, characterized in that: It includes a visible area segmentation branch and a full object area segmentation branch, both of which are connected to an image encoder and a prompt encoder, the visible area segmentation branch includes a first pixel perception memory module, a first mask decoder, a first memory bank and a first memory encoder, and the full object area segmentation branch includes a stereoscopic space perception memory module, a second pixel perception memory module, a third pixel perception memory module, a second mask decoder, a second memory bank and a second memory encoder; The image encoder is used to extract image features from the video and output them to the visible area segmentation branch and the full object area segmentation branch respectively; The hint encoder is used to obtain hint embedding of the visible region segmentation branch and the full object segmentation branch, and output them to the visible region segmentation branch and the full object region segmentation branch respectively; The visible region segmentation branch and the full object region segmentation branch perform weighted integration on the forward and reverse segmentation results using the predicted index to obtain the final visible region mask and the full object region mask.
2. A video full object segmentation method based on 3D prior enhancement and dual-branch structure, applied to a video full object segmentation system based on 3D prior enhancement and dual-branch structure as claimed in claim 1, characterized in that: The following steps are involved: S1: Get a video sequence about the target object; S2: Input each frame of the video sequence into the image encoder to extract image features; S3: Input the bounding box of the visible area of the target object and the bounding box after being enlarged by two times into the prompt encoder respectively to obtain the prompt embedding of the visible area segmentation branch and the whole object segmentation branch; S4: For the visible area segmentation branch, the image features gather information from the first memory bank storing the prediction results of the previous frame through the attention mechanism, and then predict the visible area mask of the current frame through the first mask decoder, and then update the first memory bank through the first memory encoder; For the full object segmentation branch, the second memory bank includes the prediction results of the previous frame and the spatial perception information. The full object mask of the current frame is predicted by the second mask decoder, and then the second memory encoder is used to update the second memory bank. During the training process, a perceptual consistency regularization term is introduced into the full object segmentation branch to constrain the consistency between the full object prediction mask under complete memory and the full object prediction mask under incomplete memory, and learn different memories for the visible area segmentation task and the full object segmentation task; S5: After the training is completed, the current video frame to be segmented is inferred, the video sequence is processed forward and backward, and the forward and reverse segmentation results are weighted integrated in combination with the prediction index to obtain the final visible area mask and the full object area mask.
3. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 2 is characterized in that: The step S1 specifically selects two subsets of the Multi-Object Video dataset, Movi-B and Movi-D, as video sequences of target objects, wherein Movi-B consists of objects in the CLEVR dataset, which have simple shapes and regular geometry; Movi-D consists of highly realistic objects in the Google Scanned Objects dataset, which have richer and more complex shapes and textures, as well as more serious occlusion.
4. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 2, characterized in that: In step S2, the image encoder adopts a hybrid adaptation mechanism, by adding a LoRA module to each query / key / value projection layer and mixing multiple LoRA weights in the feedforward network to retain the original prior knowledge and integrate the knowledge of different tasks; For the image encoder Hiera, an additional LoRA module is connected to the query / key / value projection layer of each transformer block; for the feedforward network, multiple LoRA weights are mixed according to the sample-specific gating function, and the input x of the adapted transformer block is specifically: Among them, W and b represent the original projection weight and bias, A and B represent the LoRA weight, Attn and FFN represent the self-attention and feedforward network modules respectively, and N e represents the number of LoRA weights attached to the feedforward network, and δ represents the weights added to the feedforward network by the multi-layer perceptron MLP. Mapping to N e dimensional feature vector and normalized with a softmax function.
5. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 2, characterized in that: The prompt encoder in step S3 shares the same structure for visible region segmentation and full object segmentation tasks, but adopts different prompt strategies. For visible region segmentation, the real visible bounding box of the target object is used as the input prompt; For full object segmentation, the size of the visible bounding box is doubled as an input prompt. During the training process, random noise is added to the real visible bounding box for visible area segmentation and full object segmentation branches. The random noise does not exceed 10% of the side length or a maximum of 20 pixels.
6. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 4, characterized in that: In step S4, for the full object segmentation branch, a method of enhancing stereoscopic spatial perception is specifically used to supplement the second memory bank of the full object segmentation branch. t In addition to the image encoder, a monocular depth estimator D is used to predict the corresponding depth image, and then the stereo space-aware embedding is extracted from the intermediate layers of D. To capture detailed depth-related features, the design uses a stereoscopic spatial perception memory module X vol , which is used to store all processed frames The three-dimensional space perception memory module X vol Follow the design principles of SAM2.
7. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 6, characterized in that: The memory attention process of the three-dimensional space perception memory module and the second pixel perception memory module in the full object segmentation branch is: Where MA represents the same memory attention architecture as in the original SAM2, f represents a zero-initialized linear layer, and x t represents the image features extracted by the image encoder, X pix represents the raw pixel-aware memory in SAM2. In this way, the volume-aware information is first aligned with the latent space associated with the raw pixel-aware information and aggregated into In addition, the zero-initialized f ensures that inappropriate embeddings do not distort the latent space at the beginning of training.
8. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 7, characterized in that: The specific process of introducing the perceptual consistency regularization term to the full object segmentation branch in step S4 is: First, establish the basic objective function of segmentation: in, Including DICE loss, Focal loss, L1-based IoU regression loss, and cross-entropy-based occlusion classification loss. A mask representing the prediction; In order to bridge the prior knowledge gap between the visible region segmentation and full object segmentation subtasks, an auxiliary consistency regularization term is introduced for the full object segmentation branch. For each input video V, in addition to using the complete pixel-related memory X pix Get Forecast In addition, j memory items are randomly discarded at each time step i, where This leads to new predictions Then, the regularized objective function is expressed as: Here, λ is a hyperparameter.
9. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 8, characterized in that: The specific process of inferring the current video frame to be segmented in step S5 is as follows: For the current video frame to be segmented, firstly, the image features are extracted through the image encoder, and then the visible area bounding box of the target object and the bounding box after being enlarged twice are respectively input into the hint encoder to obtain the hint embedding of the visible area segmentation branch and the whole object segmentation branch; Next, the image features are processed through the corresponding memory attention of the visible area segmentation branch and the whole object segmentation branch, and information is gathered from the first memory bank and the second memory bank. The first mask decoder and the second mask decoder respectively generate the visible area mask and the whole object mask of the current frame. Finally, the first memory encoder and the second memory encoder are used to update the first memory bank and the second memory bank accordingly to ensure continuous learning and adaptation to new video sequences.
10. The video full object segmentation method based on 3D prior enhancement and dual-branch structure according to claim 9, characterized in that: The specific process of weighted integration of the forward and reverse segmentation results in combination with the prediction index in step S5 is as follows: For Video The tth frame I t First, the visible region segmentation branch and the full object segmentation branch are used to generate the predicted full object mask. And the estimated IoU value η t , and then flip the video along the time dimension to get its reverse version V′, where the (T-t+1)th frame is I t , repeat the inference process on V′ to generate new predictions and its estimated IoU value η′ t , based on these results, the final prediction is calculated by the following weighted ensemble: Among them, ω t It is the weight coefficient calculated based on the estimated IoU value.
Citation Information
Cited By
Mask-based multi-target tracking method and system, electronic equipment and storage medium
CN121482089A