Arbitrary modal assistance-based camouflage target segmentation method

By introducing the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner in the camouflage target segmentation task, the problem of insufficient camouflage target segmentation performance under a single modal input is solved, and effective fusion of multimodal data and efficient identification of camouflage targets are achieved.

CN120107975APending Publication Date: 2025-06-06TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510183100.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art relies on single modal input in camouflage target segmentation tasks, making it difficult to effectively fusion of multimodal data, especially in the absence of real multimodal data, segmentation performance is limited.

Method used

A camouflage target segmentation method based on arbitrary modal assistance is proposed. Through the combination of the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner, an effective transition from single mode to multimodal is achieved. UniSEG adopts a fusion mechanism between latent space and state space, and combines a cross-state space model to perform deep fusion; UniLearner uses cross-modal mapping to generate pseudo-modal images and semantic-rich latent vectors to improve the feature extraction capability of segmented networks.

Benefits of technology

It significantly improves the performance of camouflage target segmentation, improves the ability to identify camouflage targets in complex scenarios, enhances the model's resistance to noise, and shows excellent performance and wide applicability in multiple camouflage target segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107975A_ABST
    Figure CN120107975A_ABST
Patent Text Reader

Abstract

The invention discloses a camouflage target segmentation method based on arbitrary modal assistance. The detection precision of a camouflage target is improved through multi-modal data fusion. According to the method, a multi-modal divider UniSEG and a cross-modal knowledge learning network UniLearner are included. The UniSEG adopts a double-branch architecture, extracts the features of RGB images and other modal images, and carries out preliminary fusion through a potential space fusion module LSFM. The SSFM is combined with a cross-state space model (CSSM) to further fuse features in a unified state space. And the UniLearner learns a mapping relation between the RGB image and the target mode through an encoder-decoder structure, generates a pseudo-modal image and a semantic-rich potential vector, and injects the pseudo-modal image and the semantic-rich potential vector into a specific layer of the UniSEG to improve feature extraction and fusion effects. The method has plug-and-play flexibility, can be seamlessly integrated into an existing segmentation network, is widely applicable to the fields of ecology, medicine, surface monitoring and the like, and remarkably improves the performance and robustness of disguise target segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to computer vision technology, and in particular to a camouflage target segmentation method based on arbitrary modality assistance. Background Art

[0002] Camouflaged Object Segmentation (COS) aims to detect objects that are difficult to identify in a scene. This task is extremely challenging because the visual information difference between the camouflaged object and the surrounding background is minimal and lacks obvious visual features, making it difficult for both machines and humans to accurately identify these objects.

[0003] In recent years, COS research has made progress driven by technologies such as multi-scale, multi-space, multi-stage, and bionic strategies, which are mainly focused on improving the feature extraction capabilities of camouflaged images. Despite this, most methods still rely on single-modal input, which limits the potential of multimodal data, mainly because it is difficult to obtain multimodal data paired with camouflaged targets. The development of depth estimation technology has promoted the fusion of depth information, further demonstrating the advantages of multimodal methods. However, research on RGB-to-X conversion is still limited, which to some extent hinders the further development of additional modality-assisted COD tasks.

[0004] To overcome the limitations of single-image COS, a common strategy is to introduce auxiliary information from other modalities. For example, IPNet and PolarNet use polarization data to improve segmentation accuracy through 1,200 sets of RGB-polarization camouflaged target image pairs. However, these datasets are small in scale, and models trained on such sparse data usually only bring limited performance improvements.

[0005] With the development of passive depth estimation technology, the application of depth information in COS tasks is becoming more and more popular. For example, PopNet introduces depth maps into COS tasks through a dedicated network architecture and loss function to improve segmentation results. Similarly, DSAM combines the SAM framework to study the interaction between depth and RGB information in the COS field to more effectively fuse these modalities. However, when the target and background are in the same focal plane (such as Figure 1 As shown in Figure 2, or when there is high visual confusion, monocular depth estimation may fail, resulting in reduced depth differentiation ability, which significantly affects the effectiveness of these methods.

[0006] Infrared data is a modality with potential in target center segmentation tasks because it can capture the difference in thermal radiation of the target, thus providing effective clues to distinguish the camouflaged target from its surroundings. However, the introduction of infrared data in COS tasks faces significant challenges. It is extremely difficult to construct a paired dataset of infrared and camouflaged target images, and there is currently a lack of reliable methods to generate pseudo-infrared data for camouflaged target images. These problems hinder the effective integration of infrared data and other similar modalities in COS tasks.

[0007] Originated from classical control theory, state-space models (SSMs) are an important tool for analyzing long continuous sequence data. The structured state-space sequence model (S4) was originally used to model long-distance dependencies, while the recent Mamba model introduced a selection mechanism that enables it to extract relevant information from the input data. Mamba has been successfully applied to image restoration, segmentation and other fields, and achieved competitive performance.

[0008] In the image fusion task, methods such as MambaDFuse and FusionMamba have used Mamba to improve the performance. However, these methods only use SSM for feature extraction, while ignoring the cross-modal state space features and Mamba's ability in different modal feature selection.

[0009] It should be noted that the information disclosed in the above background technology section is only used for understanding the background of the present application, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention

[0010] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a camouflaged target segmentation method based on the assistance of any modality.

[0011] To achieve the above object, the present invention adopts the following technical solutions:

[0012] A camouflaged target segmentation method based on arbitrary modality assistance includes the following steps:

[0013] S1. Feature extraction: The features of RGB images and other modality images are extracted respectively through the two branches of the multimodal segmenter UniSEG;

[0014] S2, preliminary fusion: preliminary fusion of the features of the RGB image and the image of the other modality through the latent space fusion module LSFM;

[0015] S3, further fusion: the initially fused features are deeply fused in a unified state space through the state space fusion mechanism SSFM and the cross-state space model CSSM;

[0016] S4, generate segmentation results: convert the fused features into the final segmentation results through the decoder;

[0017] S5. Learning cross-modal knowledge: The cross-modal knowledge learning network UniLearner learns the relationship between RGB and the other modalities to generate a pseudo-modal image and a knowledge vector;

[0018] S6. Joint training: Inject the knowledge of the cross-modal knowledge learning network UniLearner into the multimodal segmenter UniSEG to improve the segmentation performance of UniSEG.

[0019] Furthermore, in step S1, when extracting features of RGB images and images of other modalities, a dual-branch encoder architecture is adopted, wherein the first branch is used to extract features of RGB images, and the second branch is used to extract features of images of other modalities, and the output features of the two branches have the same spatial resolution.

[0020] Furthermore, in step S2, the latent space fusion module LSFM performs weighted fusion on the features of the RGB image and the features of the images of other modalities to generate fused latent features, and enhances the expressiveness of the features through nonlinear activation functions and convolution operations.

[0021] Furthermore, in step S3, the state space fusion mechanism SSFM selectively integrates the features of different modalities in a unified state space, the cross-state space model CSSM captures the long-range dependencies between the features of different modalities, and balances the contribution of each modal feature through a gating mechanism.

[0022] Furthermore, in step S4, when generating the segmentation result, a multi-task decoder is used, which combines the fused features and the preliminary prediction results at each layer, gradually reconstructs the segmentation map, and provides additional supervision information through the edge reconstruction task to enhance the details and boundary accuracy of the segmentation result.

[0023] Furthermore, in step S5, when learning cross-modal knowledge, the cross-modal knowledge learning network UniLearner maps the RGB image to the target modal space through an encoder-decoder structure to generate a pseudo-modal image and a semantically rich latent vector, which is used to guide the feature extraction and fusion process of the multimodal segmenter UniSEG.

[0024] Furthermore, in step S5, the cross-modal knowledge learning network UniLearner optimizes its parameters through joint training, uses L1 norm loss to constrain the generation of pseudo-modal images, and injects the generated latent vector into the feature fusion layer of the multimodal segmenter UniSEG to enhance the segmentation network's utilization of cross-modal semantic information.

[0025] Furthermore, in step S6, during joint training, a weighted loss function is used to optimize the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner, wherein the loss function includes segmentation loss, edge reconstruction loss, and cross-modal generation loss, so as to simultaneously improve the segmentation performance and the learning effect of cross-modal knowledge.

[0026] Furthermore, the multimodal segmenter UniSEG specifically includes a feature feedback module FFM, a state space fusion mechanism SSFM and a cross-state space model CSSM, wherein: the feature feedback module FFM feeds back the initially fused features to the subsequent layers of other modality encoders, and dynamically adjusts the feature weights through a gating mechanism to guide other modality encoders to perform targeted feature extraction; the state space fusion mechanism SSFM selectively integrates RGB image features and other modality features in a unified state space, captures long-range dependencies through a state space model SSM, and uses a gating mechanism to balance the contributions of different modality features; the cross-state space model CSSM further fuses different modality features in the state space, enhances feature expression capabilities through deep convolution and nonlinear activation functions, and combines the channel attention mechanism to reduce feature redundancy, thereby improving the robustness and semantic richness of the fused features.

[0027] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the method for segmenting a camouflaged target based on the assistance of any modality is implemented.

[0028] The present invention has the following beneficial effects:

[0029] The present invention proposes a camouflaged target segmentation method based on arbitrary modality assistance, which significantly improves the performance of camouflaged target segmentation through an innovative multimodal fusion framework. The core of the method is to combine the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner to achieve an effective transition from single modality to multimodality. UniSEG combines the latent space fusion module (LSFM) and the state space fusion mechanism (SSFM) with the cross-state space model (CSSM) to efficiently fuse multimodal features in a unified state space, thereby enhancing the recognition ability of camouflaged targets in complex scenes. At the same time, UniLearner uses task-independent multimodal data to learn cross-modal mapping, generate pseudo-modal images and semantically rich latent vectors, which are embedded in UniSEG to enhance its feature extraction ability, thereby significantly improving the segmentation performance in the absence of real multimodal data. This fusion-feedback-fusion strategy not only improves the robustness of feature extraction, but also enhances the model's resistance to noise, so that the present invention shows excellent performance and wide applicability in multiple camouflaged target segmentation tasks.

[0030] In addition, the modular design of the present invention gives it a high degree of flexibility and scalability. Both UniSEG and UniLearner adopt a plug-and-play architecture, which can be seamlessly integrated into the existing segmentation network, and easily convert a single-modal network into a multi-modal network. This design not only simplifies the deployment process of the model, but also enables the method to be easily applied to other related fields, such as camouflaged animal detection in ecological research, lesion segmentation in medical images, and surface change monitoring. Through application in these fields, the present invention not only improves the detection accuracy of camouflaged targets, but also provides new technical means for research and practice in related fields, showing great application potential and value.

[0031] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Shown are an RGB image and its corresponding true segmentation value, a depth estimation map generated by PopNet, an infrared estimation image generated by a separately trained ResUNet, and an infrared image generated by UniLearner using the same network architecture.

[0033] Figure 2 It is a framework diagram of the UniCOS framework and FFM, LSFM, g_w and SSFM of an embodiment of the present invention.

[0034] Figure 3 This is a framework diagram of the CSSM (Cross State Space Module) of an embodiment of the present invention.

[0035] Figure 4 These are the qualitative results of UniCOS-I and other cutting-edge methods according to the embodiments of the present invention.

[0036] Figure 5 Visual comparison in the RGB-D COS task of the embodiment of the present invention

[0037] Figure 6 This is a visual comparison in the RGB-P COS task of an embodiment of the present invention.

[0038] Figure 7 The figure is an overall flow chart of the camouflaged target segmentation method based on arbitrary modality assistance of the present invention. DETAILED DESCRIPTION

[0039] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope and application of the present invention.

[0040] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0041] In recent years, COS research has made progress driven by technologies such as multi-scale, multi-space, multi-stage, and biomimetic strategies, which are mainly focused on improving the feature extraction capabilities of camouflaged images. Despite this, most methods still rely on single-modal input, which limits the potential of multimodal data, mainly because it is difficult to obtain multimodal data paired with camouflaged targets. The development of depth estimation technology has promoted the fusion of depth information, further demonstrating the advantages of multimodal methods. However, research on RGB-to-X conversion is still limited, which to some extent hinders the further development of COD tasks assisted by additional modalities. State-space models (SSMs) originated from classical control theory and are an important tool for analyzing continuous long sequence data. The structured state-space sequence model (S4) was originally used to model long-range dependencies, and the recent Mamba model introduced a selection mechanism that enables it to extract relevant information from the input data. Mamba has been successfully applied to image inpainting, segmentation and other fields, and has achieved competitive performance. In image fusion tasks, methods such as MambaDFuse and FusionMamba use Mamba to improve performance. However, these methods only use SSM for feature extraction, but ignore the cross-modal state space features and the ability of Mamba in feature selection across different modalities.

[0042] The present invention mainly solves the following four problems:

[0043] 1. Limitations of single-modal methods: Previous RGB single-modal methods have difficulty accurately segmenting camouflaged targets due to limited visual cues, especially when the target and background have similar colors and textures;

[0044] 2. The fusion problem of multimodal data: Although multimodal data (such as depth, infrared, polarization, etc.) can provide additional visual clues, how to effectively fuse these different modal data to improve segmentation performance is a key issue. For example, depth estimation may fail in some cases, and the acquisition and fusion of infrared data also face challenges;

[0045] 3. Lack of paired multimodal data: In practical applications, it is very difficult to obtain multimodal data pairs (such as RGB and infrared image pairs) related to the camouflaged target segmentation task, which limits the development and application of multimodal methods;

[0046] 4. Insufficient utilization of cross-modal knowledge: Even if there is multimodal data of non-camouflaged targets, how to use this data to improve the performance of camouflaged target segmentation models is an unsolved problem.

[0047] In order to make full use of effective features, the present invention proposes a state space fusion mechanism (SSFM) and combines it with a cross-state space model (CSSM) to unify multimodal features into a shared state space for efficient fusion. On this basis, the present invention further designs UniSEG, a unified network for MCOS tasks.

[0048] To avoid the interference of the uncertainty of pseudo-modal data on feature extraction, UniSEG uses the latent space fusion module (LSFM) to perform preliminary feature fusion in the latent space, and re-inputs the fusion result into the additional modality encoder through the feature feedback module (FFM) to provide targeted feature extraction guidance. Finally, SSFM is used to further fuse cross-modal information in the state space. Through this fusion-feedback-fusion strategy, UniSEG can effectively extract and integrate multi-modal key information and improve the segmentation performance of MCOS tasks.

[0049] In order to more effectively utilize additional modality information to improve the performance of COS tasks, this paper proposes UniLearner, a framework for acquiring cross-modal knowledge from an auxiliary RGB-X dataset. It is worth noting that the auxiliary dataset is irrelevant to the COS task itself. UniLearner generates pseudo-modal data and a semantically rich latent vector by learning cross-modal mapping, which maps RGB images to the auxiliary modality, thereby providing guidance for the segmentation network. By jointly optimizing UniLearner with the segmentation network, the framework can improve the quality of feature generation, thereby enhancing segmentation performance and achieving better results in cross-domain image conversion tasks.

[0050] UniSEG adopts a modular design, which enables it to be used as a plug-and-play enhancement component to existing segmentation networks. Its individual modules can seamlessly convert a unimodal segmentation network into a multimodal segmentation network. In addition, UniLearner can also work with a dual-branch multimodal segmentation network to further improve segmentation performance through efficient cross-modal knowledge fusion.

[0051] The embodiment of the present invention provides a camouflaged target segmentation method based on any modality assistance, see Figure 7 , including the following steps:

[0052] Step S1, feature extraction: The features of RGB images and other modality images are extracted respectively through the two branches of the multimodal segmenter UniSEG.

[0053] In a preferred embodiment, in step S1, when extracting features of RGB images and images of other modalities, a dual-branch encoder architecture is adopted, wherein the first branch is used to extract features of RGB images, and the second branch is used to extract features of images of other modalities, and the output features of the two branches have the same spatial resolution.

[0054] Step S2, preliminary fusion: Preliminary fusion of the features of the RGB image and the other modality images through the latent space fusion module LSFM.

[0055] In a preferred embodiment, in step S2, the latent space fusion module LSFM performs weighted fusion on the features of the RGB image and the features of the images of other modalities to generate fused latent features, and enhances the expressiveness of the features through nonlinear activation functions and convolution operations.

[0056] Step S3, further fusion: The initially fused features are deeply fused in a unified state space through the state space fusion mechanism SSFM and the cross-state space model CSSM.

[0057] In a preferred embodiment, in step S3, the state space fusion mechanism SSFM selectively integrates the features of different modalities in a unified state space, the cross-state space model CSSM captures the long-range dependencies between the features of different modalities, and balances the contribution of each modal feature through a gating mechanism.

[0058] Step S4, generating segmentation results: converting the fused features into the final segmentation results through the decoder.

[0059] In a preferred embodiment, in step S4, when generating the segmentation result, a multi-task decoder is used, which combines the fused features and the preliminary prediction results at each layer, gradually reconstructs the segmentation map, and provides additional supervision information through the edge reconstruction task to enhance the details and boundary accuracy of the segmentation result.

[0060] Step S5, learning cross-modal knowledge: the cross-modal knowledge learning network UniLearner learns the relationship between RGB and the other modalities to generate a pseudo-modal image and a knowledge vector.

[0061] In a preferred embodiment, in step S5, when learning cross-modal knowledge, the cross-modal knowledge learning network UniLearner maps the RGB image to the target modal space through an encoder-decoder structure, generates a pseudo-modal image and a semantically rich latent vector, and the latent vector is used to guide the feature extraction and fusion process of the multi-modal segmenter UniSEG. Further, the cross-modal knowledge learning network UniLearner optimizes its parameters through joint training, uses L1 norm loss to constrain the generation of pseudo-modal images, and injects the generated latent vector into the feature fusion layer of the multi-modal segmenter UniSEG to enhance the segmentation network's utilization of cross-modal semantic information.

[0062] Step S6, joint training: inject the knowledge of the cross-modal knowledge learning network UniLearner into the multimodal segmenter UniSEG, thereby improving the segmentation performance of UniSEG.

[0063] In a preferred embodiment, in step S6, during joint training, a weighted loss function is used to optimize the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner, wherein the loss function includes segmentation loss, edge reconstruction loss and cross-modal generation loss, so as to simultaneously improve the segmentation performance and the learning effect of cross-modal knowledge.

[0064] In a preferred embodiment, the multimodal segmenter UniSEG specifically includes a feature feedback module FFM, a state space fusion mechanism SSFM and a cross-state space model CSSM, wherein: the feature feedback module FFM feeds back the initially fused features to the subsequent layers of other modality encoders, and dynamically adjusts the feature weights through a gating mechanism to guide other modality encoders to perform targeted feature extraction; the state space fusion mechanism SSFM selectively integrates RGB image features and other modality features in a unified state space, captures long-range dependencies through a state space model SSM, and uses a gating mechanism to balance the contributions of different modality features; the cross-state space model CSSM further fuses different modality features in the state space, enhances feature expression capabilities through deep convolution and nonlinear activation functions, and combines the channel attention mechanism to reduce feature redundancy, thereby improving the robustness and semantic richness of the fused features.

[0065] The following further describes an algorithm implementation example and experimental verification of a specific embodiment of the present invention.

[0066] UniSEG: A unified multimodal segmenter

[0067] UniSEG fuses features from RGB images and additional modalities in both state and latent spaces. The framework uses Latent Space Fusion Module (LSFM) and State Space Fusion Mechanism (SSFM) to selectively combine features from RGB images and auxiliary modalities to improve the performance of camouflaged object segmentation. In addition, Feature Feedback Module (FFM) uses the output of LSFM at specific network layers to guide subsequent encoder layers for more effective feature extraction.

[0068] Figure 2 The UniCOS framework and FFM, LSFM, g w The detailed algorithm framework of SSFM. The modules marked with dashed boxes represent the modules introduced by UniLearner and can be omitted when using paired RGB-X data.

[0069] Multimodal Segmentation-Guided Encoder

[0070] UniSEG adopts a dual-branch encoder architecture to extract and utilize beneficial features of different modalities. i and X u , first interpolate them to a uniform size W × H. Then, use the base encoder ε i From X i Extracting deep feature sets Each of these The resolution is

[0071] To process features from additional modalities, an auxiliary encoder ε with similar architecture is introduced u , the encoder contains a custom embedding layer to adapt X u The output of the kth layer of the auxiliary encoder is recorded as Its resolution and same.

[0072] In order to fuse the features of different modalities in the latent space, the present invention introduces a latent space fusion module (LSFM) to fuse the features and Generate fused latent features at k = {1, 2, 3, 4}

[0073]

[0074] Among them, W c represents the convolution operation, Represents the convolution Conv+LeakyReLU (LReLU)+batch normalization (BN) combination block, ⊙ represents element-by-element multiplication. The final fused latent feature map It has rich semantic information and is enhanced through the expanded spatial pyramid pooling (ASPP) module A s Further processing to obtain preliminary predictions Its spatial resolution is The same as above and serves as the initial input to the decoder.

[0075] Different from The role of is to use existing features to guide ε u To achieve this goal, UniSEG introduces a feature feedback module (FFM) to gate the injection Generate updated features This feature will also be used as ε u The (k+1)th layer input and the state space fusion mechanism (SSFM) input after the kth layer:

[0076]

[0077] To achieve robust feature fusion, this paper proposes a state space fusion mechanism (SSFM) to selectively integrate features from different modalities in a unified state space representation:

[0078]

[0079] in, What you get Provide more complete contextual information, reduce redundancy, filter out noise, and capture the relationship between different modalities.

[0080] In the decoding stage, each layer of the decoder will k As a conditional input, and combined with A s Preliminary forecast of treatment As well as latent space fusion features, the reconstruction process is enhanced, enabling it to extract richer detail information and improve modality perception capabilities.

[0081] Details of SSFM

[0082] State Space Fusion Mechanism

[0083] In the visual state space model with a two-dimensional selective scanning module, features are flattened into sequences and scanned along four directions (from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right) to capture the long-range dependency of each sequence using discrete state space equations. The present invention proposes a cross state space model (CSSM) to promote information interaction between different sequences in the state space.

[0084] In formula (3), the present invention converts Reshape The present invention implements the Visual State Space Module (SSM) as a residual state space block as shown in MambaIR and uses it as a long-range self-attention mechanism to process and Compute homomodal correlation:

[0085]

[0086] Then, the cross-state space model (CSSM) proposed in the present invention is used to further integrate the bimodal features in the state space to handle the same-modal correlation and cross-modal correlation:

[0087]

[0088] The present invention utilizes a weighted gating mechanism g w Merge the transformed features as follows:

[0089]

[0090] This gating mechanism is based on Guidance, balance and Contribution of . Function and Generate intermediate signals, affecting the final fusion features. By using Sigmoid, ensure that g w Keep it between 0 and 1, thus adjusting each path to output F k relative contribution.

[0091] Cross State Space Model

[0092] Assume the input is in Can be or As shown in eq. (5), the present invention first applies linear projection to and The channel dimensions of are expanded to d×2, and they are split into two parts along the last dimension: as well as

[0093] Figure 3 The detailed algorithm framework of the CSSM (Cross State Space Module) of the embodiment of the present invention is shown.

[0094] Next, the present invention will Considered as shape And apply depthwise convolution with kernel size d conv , and then perform nonlinear activation:

[0095]

[0096] Here, the number of convolution groups is equal to the channel dimension d, SiLU is the activation function, and W c In order to fuse two modalities in the state space, the present invention constructs the following model:

[0097]

[0098] Among them, B n , C n , and Δ n is the matrix B, C, and Δ with the selectivity mechanism parameters and Corresponding.

[0099] After combining the four directional sequences, the present invention k The application layer is normalized and compared with Multiply the activation values ​​​​element-wise:

[0100]

[0101] Next, the present invention will y' k Mapping back to the desired output dimensions:

[0102] Y k =y' k W l +b l ,(10)

[0103] in and

[0104] Finally, in order to enhance the expressive power of different channels, the present invention introduces a channel attention mechanism (CA) in CSSM to reduce channel redundancy. In addition, the present invention uses two weighted residual connections, namely s and To improve the robustness of the network:

[0105]

[0106] Alternative segmentation decoder

[0107] Since the multimodal segmentation-oriented encoder of the present invention adopts a plug-and-play design, the decoder in UniSEG can be replaced by any decoder that uses coarse results or latent graphs and skip connections as input.

[0108] In the implementation of the present invention, a multi-task segmentation decoder such as ICEG is used by default. The decoder has separate task heads at each layer, one for segmentation and one for edge reconstruction, where edge reconstruction provides additional supervision information. The decoding process can be expressed by the following formula:

[0109]

[0110] in Represents a decoder, and Represent the segmentation results and reconstructed edges respectively.

[0111] optimization

[0112] As a unified plug-and-play method, the multimodal segmentation-oriented encoder and multi-space fusion of the present invention can be easily integrated into most non-dedicated input design decoders. Here, a multi-task segmentation decoder is taken as an example as a default decoder.

[0113] UniSEG uses a weighted intersection-over-union loss L I , weighted binary cross entropy loss L B To constrain the segmentation results And use the dice loss L D To supervise the edge reconstruction results Let the segmentation result y s and the marginal result y e is the true label, the total loss of UniSEG can be expressed as:

[0114]

[0115] UniLearner: Cross-modal knowledge learning

[0116] UniLearner It is a plug-and-play encoder-decoder network. When the COS dataset lacks corresponding multimodal data, UniLearner can learn the mapping relationship between images and modalities by introducing additional non-COS multimodal datasets, thereby helping to complete the COS task.

[0117] Specifically, the present invention represents the image of the introduced additional dataset as e i , the corresponding additional modal data is denoted as e u The present invention expects We can learn the mapping relationship between them and get:

[0118]

[0119] When working with UniSEG, UniLearner takes as input an image x i , through the encoding and decoding process, the corresponding pseudo-mode x is obtained u , and a latent vector z i→u , which embodies the mapping knowledge between images and modalities:

[0120]

[0121] in and They are The encoder and decoder of z i→u Indicates that it contains i to x u Latent vectors that map knowledge.

[0122] In order to z i→u Integrating it into the segmentation process, the present invention injects it into the 4th layer of UniSEG by replacing LSFM (Formula (1)) with a new formula:

[0123]

[0124] This operation fuses the mapping information between the image and the pseudo-modality, as well as the semantic information extracted from both modalities, into the latent space. This unified representation enhances the segmentation effect by leveraging complementary cross-modal knowledge.

[0125] optimization

[0126] When UniLearner is used, the present invention performs joint training of UniLearner and UniSEG, using a shared optimizer to optimize the parameters of the two networks. i and e u The mapping relationship between them, the present invention uses L1 norm loss, the formula is:

[0127]

[0128] The total loss L for this joint training setting t It is expressed as:

[0129]

[0130] Experiments and effects

[0131] Describe the technical effect of the invention, that is, what advantages it has over the background technology, and explain why such technical effect can be achieved by analyzing the innovative points. The description should be specific and realistic.

[0132] The performance of our method is evaluated on three multimodal COS tasks: RGB-infrared (RGB-I), RGB-depth (RGB-D), and RGB-polarization (RGB-P).

[0133] For the RGB-infrared task (UniCOS-I), a dataset unrelated to the COS task was used to demonstrate the ability of UniLearner in improving the COS task performance using unrelated data. In the RGB-depth task (UniCOS-D), pseudo depth data was used, while in the RGB-polarization task (UniCOS-P), real degree of linear polarization (DoLP) data was used. This experimental setup enables a comprehensive evaluation of the performance and robustness of UniSEG in dealing with pseudo-multimodal and real multimodal data.

[0134] For UniLearner, ResUNet with 9 residual blocks is used as the backbone network. For UniSEG, PVTv2 pre-trained on ImageNet is used as the backbone network by default, and experimental results based on ResNet50 are provided for fair comparison. All results are evaluated using consistent task-specific evaluation tools.

[0135] Quantitative and qualitative results

[0136] RGB and task-independent infrared data

[0137] Table 1 shows the quantitative comparison of UniCOS-I with 12 other SOTAs using two different types of backbone networks. Red indicates the best result.

[0138] Table 1

[0139]

[0140] Figure 4 Qualitative results of UniCOS-I and other cutting-edge methods are shown.

[0141] As shown in Table 1, the UniCOS-I method proposed in this paper outperforms all 12 latest advanced methods on multiple datasets. Figure 4 As shown in Figure 2, the segmentation maps generated by UniCOS-I are more complete and coherent than those of other leading methods, which further proves the effectiveness of the proposed method in multimodal data fusion. Figure 1 As shown in Figure 3, the joint training of UniSEG and UniLearner significantly improves the RGB to infrared reconstruction performance. This shows that UniLearner can effectively handle the semantic complexity inherent in RGB-infrared data, which is often difficult for traditional end-to-end image conversion methods to handle.

[0142] Paired RGB and pseudo-depth data

[0143] Figure 5 A visual comparison in the RGB-D COS task is shown.

[0144] Table 2 shows the results of RGB-depth COS. All methods are trained using the passive depth data provided by PopNet.

[0145] Table 2

[0146]

[0147] In the RGB-D task, the UniCOS-D model of the present invention effectively solves the challenge of camouflaged target segmentation by using pseudo-depth data paired with RGB images. The quantitative results in Table 2 show that UniCOS-D outperforms the competing methods in all evaluation indicators and achieves the highest scores. In addition, Figure 5 The visual comparison in shows that UniCOS-D can clearly distinguish foreground objects from the background. Figure 1 The RGB image and its corresponding segmentation truth value, the depth estimation map generated by PopNet, the infrared estimation image generated by the separately trained ResUNet, and the infrared image generated by UniLearner using the same network architecture are shown. The method of the present invention performs well in the RGB to infrared conversion task, and can more accurately present the structure and position information of the camouflaged target, thereby improving the segmentation performance. Even in the case of less depth information (such as Figure 1 As shown in the first row, UniCOS-D still maintains excellent segmentation performance. These results demonstrate the robustness of the proposed method and its effectiveness under complex conditions.

[0148] Paired RGB and true polarization data

[0149] Figure 6shows a visual comparison in the RGB-P COS task,

[0150] Table 3 Results of RGB-polarization COS

[0151]

[0152] In the RGB-P task, the UniCOS-P model of the present invention significantly improves the detection capability of camouflaged targets by combining real DoLP data with RGB images. As shown in Table 3, UniCOS-P achieves excellent results on the PCOD1200 dataset. With the help of polarization information, the model is able to reveal details that are difficult to perceive with traditional RGB sensors. These polarization cues are crucial for accurately depicting the boundaries of targets, such as Figure 6 As shown in the figure, UniCOS-P excels in segmenting tiny features and accurately outlining edges. The success of UniCOS-P in complex scenes shows that integrating true polarization data can provide significant advantages, making objects that are difficult to detect with traditional imaging systems clearly visible.

[0153] UniSEG's Impact

[0154] Table 4 shows the impact of UniSEG: ε u and ε i They represent the encoders for the additional modality and RGB image, respectively, and are each equipped with a corresponding fusion module.

[0155] Table 4

[0156]

[0157] As shown in Table 4, UniSEG significantly improves the segmentation performance by integrating multimodal data. u Or image encoder ε i When they are removed, the segmentation accuracy drops significantly, indicating their importance in the system. In addition, removing state-space based fusion mechanisms (such as SSFM or CSSM) or LSFM adversely affects the performance metrics, further validating the key role of these components in improving the robustness and accuracy of the model. At the same time, removing FFM also leads to a drop in performance, indicating the importance of FFM in optimizing feature fusion across stages.

[0158] The Impact of UniLearner

[0159] Table 5 shows the impact of UniLearner: Know-Inject refers to the integration of i→u to guide the segmentation process.

[0160] Table 5

[0161]

[0162] Referring to Table 5, UniLearner enhances the camouflaged target segmentation capability by leveraging cross-modal knowledge. If the “knowledge injection” process is disabled, that is, the integrated latent vector z is removed i→u This verifies the effectiveness of UniLearner in utilizing additional multimodal data to improve the performance of camouflaged target segmentation, improving the accuracy and consistency of segmentation results on multiple datasets.

[0163] Generalization capability of UniCOS

[0164] Table 6 shows the ablation study of applying the module of the present invention to other COS methods: The module proposed by UniSEG can easily convert the unimodal COS method into a multimodal method and improve the performance through UniLearner and multimodal data unrelated to COS.

[0165] Table 6

[0166]

[0167] As shown in Table 6, when the unimodal method FEDER is modified to a multimodal method using the UniCOS-D scheme of the present invention, the performance is improved. Furthermore, when the present invention applies the UniCOS-I scheme combined with UniLearner on the modified FEDER and the original multimodal method DaCOD, the performance is further improved. This result shows that the method of the present invention can effectively utilize multimodal data and exhibit good generalization ability in COS tasks. In addition, this also shows that the method of the present invention can be used as a plug-and-play framework to significantly improve the performance of COS tasks.

[0168] In summary, the technical highlights and innovative contributions of the present invention include:

[0169] 1. This paper proposes UniCOS, a unified multimodal camouflaged object segmentation (MCOS) framework, which integrates the multimodal segmenter UniSEG and the cross-modal knowledge learning plug-in UniLearner.

[0170] 2. UniSEG fuses the encoded multimodal and image features in the latent space and state space, and feeds the fused features back to the extra-modal encoder to guide further feature extraction. This iterative fusion-feedback mechanism enhances context understanding and noise robustness, thereby improving segmentation performance.

[0171] 3. UniLearner acquires cross-modal knowledge from task-independent multimodal data. It maps images to the target modal space, generates pseudo-modal content and mapping vectors. By embedding this vector into UniSEG, UniLearner establishes cross-modal semantic associations, thereby improving segmentation performance.

[0172] 4. Extensive experiments on multiple COS tasks show that the method of the present invention achieves the current optimal performance and has plug-and-play flexibility.

[0173] In view of the shortcomings of existing methods in the field of multimodal disguised target segmentation, the present invention proposes UniLearner, which is used to learn and utilize cross-modal information between images and different modalities to improve the performance of MCOD (multimodal disguised target detection). By embedding cross-modal semantic vectors into the segmenter and utilizing existing non-disguised multimodal data, the framework can still improve the performance of the COS task when there is a lack of multimodal data that truly contains disguised targets. The present invention proposes a universal state space fusion mechanism that utilizes cross-state space features and Mamba's ability to select features of different modalities in a unified state space. The mechanism integrates and selectively extracts cross-modal features in a unified state space, thereby improving the performance of multimodal disguised target segmentation (MCOS).

[0174] In order to make full use of effective features, the present invention proposes a state space fusion mechanism (SSFM) and combines it with a cross-state space model (CSSM) to unify multimodal features into a shared state space for efficient fusion. On this basis, the present invention further designs UniSEG, a unified network for MCOS tasks.

[0175] To avoid the interference of the uncertainty of pseudo-modal data on feature extraction, UniSEG uses the latent space fusion module (LSFM) to perform preliminary feature fusion in the latent space, and re-inputs the fusion result into the additional modality encoder through the feature feedback module (FFM) to provide targeted feature extraction guidance. Finally, SSFM is used to further fuse cross-modal information in the state space. Through this fusion-feedback-fusion strategy, UniSEG can effectively extract and integrate multi-modal key information and improve the segmentation performance of MCOS tasks.

[0176] In order to more effectively utilize additional modality information to improve the performance of COS tasks, this paper proposes UniLearner, a framework for acquiring cross-modal knowledge from an auxiliary RGB-X dataset. It is worth noting that the auxiliary dataset is irrelevant to the COS task itself. UniLearner generates pseudo-modal data and a semantically rich latent vector by learning cross-modal mapping, which maps RGB images to the auxiliary modality, thereby providing guidance for the segmentation network. By jointly optimizing UniLearner with the segmentation network, the framework can improve the quality of feature generation, thereby enhancing segmentation performance and achieving better results in cross-domain image conversion tasks.

[0177] UniSEG adopts a modular design, which enables it to be used as a plug-and-play enhancement component to existing segmentation networks. Its individual modules can seamlessly convert a unimodal segmentation network into a multimodal segmentation network. In addition, UniLearner can also work with a dual-branch multimodal segmentation network to further improve segmentation performance through efficient cross-modal knowledge fusion.

[0178] The methods proposed in the present invention are all plug-and-play methods, which can be easily applied to other methods and bring performance gains.

[0179] Specific application scenarios of the present invention include:

[0180] 1. This method can be used to detect camouflaged animals in natural environments and improve the efficiency of ecological research and biodiversity conservation. In addition, this method can also be used for steganography and watermark detection to identify information hidden in images or videos, such as digital watermarks or steganographic content.

[0181] 2. In the medical field, this method can be used to segment imperceptible lesions, such as tiny tumors, retinal lesions or skin lesions, to improve the early detection of diseases. In addition, this method can also be used for microscopic image analysis, such as detecting parasites or microorganisms, which are usually difficult to distinguish from the background.

[0182] 3. This method can be used to detect surface changes and identify geographic features disguised by vegetation or buildings, such as illegal mining or military bases. In addition, this method can also be used for forest fire monitoring, detecting early signals of smoke or fire sources in complex environments, and improving disaster warning capabilities.

[0183] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0184] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0185] An embodiment of the present invention further provides a processor, wherein the processor executes a computer program and at least executes the method described above.

[0186] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0187] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0188] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0189] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0190] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, etc. Various media that can store program codes.

[0191] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0192] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0193] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0194] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0195] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art of the present invention, several equivalent substitutions or obvious variations can be made without departing from the concept of the present invention, and the performance or use is the same, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A camouflaged target segmentation method based on arbitrary modality assistance, characterized in that: The following steps are involved: S1. Feature extraction: The features of RGB images and other modality images are extracted respectively through the two branches of the multimodal segmenter UniSEG; S2, preliminary fusion: preliminary fusion of the features of the RGB image and the image of the other modality through the latent space fusion module LSFM; S3, further fusion: the initially fused features are deeply fused in a unified state space through the state space fusion mechanism SSFM and the cross-state space model CSSM; S4, generate segmentation results: convert the fused features into the final segmentation results through the decoder; S5. Learning cross-modal knowledge: The cross-modal knowledge learning network UniLearner learns the relationship between RGB and the other modalities to generate a pseudo-modal image and a knowledge vector; S6. Joint training: Inject the knowledge of the cross-modal knowledge learning network UniLearner into the multimodal segmenter UniSEG to improve the segmentation performance of UniSEG.

2. The camouflaged target segmentation method based on arbitrary modality assistance according to claim 1, characterized in that: In step S1, when extracting features of RGB images and images of other modalities, a dual-branch encoder architecture is adopted, wherein the first branch is used to extract features of RGB images, and the second branch is used to extract features of images of other modalities, and the output features of the two branches have the same spatial resolution.

3. The camouflaged target segmentation method based on any modality assistance according to any one of claims 1 to 2, characterized in that: In step S2, the latent space fusion module LSFM performs weighted fusion on the features of the RGB image and the features of the images of other modalities to generate fused latent features, and enhances the expressiveness of the features through nonlinear activation functions and convolution operations.

4. The method for camouflaged target segmentation based on any modality assistance according to any one of claims 1 to 3, characterized in that: In step S3, the state space fusion mechanism SSFM selectively integrates the features of different modalities in a unified state space, the cross-state space model CSSM captures the long-range dependencies between the features of different modalities, and balances the contribution of each modal feature through a gating mechanism.

5. The camouflaged target segmentation method based on any modality assistance according to any one of claims 1 to 4, characterized in that: In step S4, when generating the segmentation result, a multi-task decoder is used, which combines the fused features and the preliminary prediction results at each layer, gradually reconstructs the segmentation map, and provides additional supervision information through the edge reconstruction task to enhance the details and boundary accuracy of the segmentation result.

6. The camouflaged target segmentation method based on any modality assistance according to any one of claims 1 to 5, characterized in that: In step S5, when learning cross-modal knowledge, the cross-modal knowledge learning network UniLearner maps the RGB image to the target modality space through an encoder-decoder structure to generate a pseudo-modal image and a semantically rich latent vector, which is used to guide the feature extraction and fusion process of the multimodal segmenter UniSEG.

7. The method for camouflaged target segmentation based on any modality assistance according to any one of claims 1 to 6, characterized in that: In step S5, the cross-modal knowledge learning network UniLearner optimizes its parameters through joint training, uses L1 norm loss to constrain the generation of pseudo-modal images, and injects the generated latent vector into the feature fusion layer of the multimodal segmenter UniSEG to enhance the segmentation network's utilization of cross-modal semantic information.

8. The camouflaged target segmentation method based on any modality assistance according to any one of claims 1 to 7, characterized in that: In step S6, during joint training, a weighted loss function is used to optimize the multimodal segmenter UniSEG and the cross-modal knowledge learning network UniLearner, where the loss function includes segmentation loss, edge reconstruction loss and cross-modal generation loss, so as to improve both the segmentation performance and the learning effect of cross-modal knowledge.

9. The camouflaged target segmentation method based on any modality assistance according to any one of claims 1 to 8, characterized in that: The multimodal segmenter UniSEG specifically includes a feature feedback module FFM, a state space fusion mechanism SSFM and a cross-state space model CSSM, wherein: the feature feedback module FFM feeds back the initially fused features to the subsequent layers of other modality encoders, and dynamically adjusts the feature weights through a gating mechanism to guide other modality encoders to perform targeted feature extraction; the state space fusion mechanism SSFM selectively integrates RGB image features and other modality features in a unified state space, captures long-range dependencies through a state space model SSM, and uses a gating mechanism to balance the contributions of different modality features; the cross-state space model CSSM further fuses different modality features in the state space, enhances feature expression capabilities through deep convolution and nonlinear activation functions, and combines the channel attention mechanism to reduce feature redundancy, thereby improving the robustness and semantic richness of the fused features.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for segmenting a disguised target based on assistance of any modality is implemented as claimed in any one of claims 1 to 9.

Citation Information

Cited By

  • Passenger abnormal behavior recognition method and device in elevator monitoring night vision mode

    CN120452068A