Camouflage object semantic segmentation method, system and device based on multi-view collaborative learning and medium

Through the multi-view collaborative learning method, combined with feature extraction and fusion of global and local information, the problem of perspective dependence and insufficient information in semantic segmentation of camouflage objects is solved, and a more efficient camouflage object segmentation effect is achieved.

CN120388175AActive Publication Date: 2025-07-29HENGYANG NORMAL UNIV

Patent Information

Application Number
CN202510472153.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has problems such as sensitivity to complex background interference, dependence on perspective, insufficient visual information and insufficient spatial information in semantic segmentation of camouflage objects, making it difficult to accurately segment camouflage targets.

Method used

Using a method based on multi-view collaborative learning, the image enhancement module, feature extraction module, global-local perception module, hybrid interaction module and dynamic pyramid shrinkage module are used to perform semantic segmentation of camouflage objects, and feature extraction and fusion are combined with global and local information.

Benefits of technology

It improves the accuracy and robustness of camouflage object segmentation, can better capture the details and spatial distribution information of camouflage objects, and enhances the model's expression ability and feature extraction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388175A_ABST
    Figure CN120388175A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image segmentation, and discloses a camouflage object semantic segmentation method, system and device based on multi-view collaborative learning, and a medium, and the method comprises the steps: obtaining a to-be-segmented multi-view image; inputting the multi-view image into an image segmentation processing model for prediction, and outputting a corresponding segmentation result; wherein the image segmentation processing model comprises an image enhancement module, a feature extraction module, a global-local sensing module, a hybrid interaction module and a dynamic pyramid contraction module which are connected in sequence. The invention provides a camouflage object semantic segmentation method based on multi-view collaborative learning, and more accurate camouflage object detection is realized through association and supplement of color, structure and position information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a method, system, device and medium for semantic segmentation of camouflaged objects based on multi-view collaborative learning. Background Art

[0002] The reason why the semantic segmentation task of camouflaged objects is more challenging than ordinary object detection is mainly reflected in "camouflage". At present, single-view detection is mostly used for the semantic segmentation task of camouflaged objects, which makes the model sensitive to background interference and unable to achieve accurate segmentation for some special scenarios.

[0003] Early traditional COD methods relied on handcrafted features to identify camouflaged targets. However, due to the limitations in the design of handcrafted features and the lack of adaptability to complex backgrounds, these methods only performed well in simple camouflage scenarios, but had limited effects in complex scenarios.

[0004] To overcome these limitations, more and more research has begun to adopt deep learning models, leveraging their powerful representation capabilities to perform camouflaged object detection in a data-driven manner. Current research focuses on single-view learning, and technologies such as uncertainty modeling and visual deformation are used to improve the learning effect and robustness of the model, and significant progress has been made. However, the variability, concealment, and scale non-fixedness of camouflaged objects pose significant challenges in this field. To comprehensively understand the image content, a large amount of context information is usually required, but the semantic information provided by single-view input is often insufficient. This weakness mainly stems from the high dependence on pixel-level information, resulting in the inability to effectively distinguish camouflaged targets from deceptive backgrounds. Single-view features usually have problems such as perspective dependence, insufficient visual information, and insufficient spatial information. Therefore, it is difficult to extract the complete form, discriminative features, and spatial distribution information of the target, and it is also impossible to accurately capture the details related to camouflaged objects. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, system, device and medium for semantic segmentation of camouflaged objects based on multi-view collaborative learning to solve the problems existing in the above-mentioned prior art.

[0006] To achieve the above purpose, the present invention provides a method for semantic segmentation of camouflaged objects based on multi-view collaborative learning, including:

[0007] Obtain multi-view images to be segmented;

[0008] Input the multi-view images into an image segmentation processing model for prediction, and output corresponding segmentation results; wherein, the image segmentation processing model includes an image enhancement module, a feature extraction module, a global-local perception module, a hybrid interaction module, and a dynamic pyramid shrinkage module connected in sequence.

[0009] Optionally, the training process of the image segmentation processing model specifically includes:

[0010] Obtain training data, where the training data includes multi-view training images and corresponding segmentation labels;

[0011] Construct an initial image segmentation processing model, input the training data into the initial image segmentation processing model for prediction, and take the minimum loss between the predicted initial training result and the segmentation label corresponding to the multi-view training image as the goal to perform training and obtain a trained image segmentation processing model.

[0012] Optionally, the processing process of the image segmentation processing model specifically includes:

[0013] Input the multi-view image into the image enhancement module for color space transformation, color jitter simulation, projection transformation, affine transformation simulation, and scale space theory simulation scaling processing to obtain multi-source input data;

[0014] Input the multi-source input data into the feature extraction module for multi-view feature extraction to obtain feature maps of multiple scales, where the feature maps include a first feature map, a second feature map, a third feature map, a fourth feature map, and a fifth feature map;

[0015] Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into the global-local perception module to extract global and local information to obtain global perception features and local perception features; the global-local perception module includes a parallel global perception module and a local perception module;

[0016] Input the global perception features and the local perception features into the hybrid interaction module for hybrid interaction to obtain multi-view complementary features;

[0017] Input the multi-view complementary features into the dynamic pyramid shrinking module for prediction and output the segmentation result.

[0018] Optionally, the processing process of the feature extraction module specifically includes:

[0019] Select a backbone network, input the multi-source input data into the backbone network, and the backbone network sequentially performs feature extraction through a shallow embedding block, a first middle embedding block, a second middle embedding block, a third middle embedding block, and a deep embedding block to output the first feature map, the second feature map, the third feature map, the fourth feature map, and the fifth feature map of different scales; among them, in each embedding block of the backbone network, semantic features of the image are gradually extracted through a convolutional layer, a batch normalization layer, and an activation function.

[0020] Optionally, the processing process of the global-local perception module specifically includes:

[0021] The processing process of the global perception module specifically includes:

[0022] Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into three side branches of the global perception module respectively:

[0023] The first group of side branches: Extract features through a convolutional layer with a kernel size of 3×3, and then through a fully connected layer and a multi-head attention mechanism operation to obtain the first global feature;

[0024] The second group of side branches: Extract features through a convolutional layer with a kernel size of 5×5, and then through a fully connected layer and a multi-head attention mechanism operation to obtain the second global feature;

[0025] The third group of side branches: Extract features through a convolutional layer with a kernel size of 7×7, and then through a fully connected layer and a multi-head attention mechanism operation to obtain the third global feature;

[0026] Input the three global features into the combination block of the global perception module, and fuse them through a convolutional layer with a kernel size of 3×3, a batch normalization layer, and an activation function to obtain the global perception feature;

[0027] The processing process of the local perception module specifically includes:

[0028] Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into the combination block of the local perception module to preliminarily fuse the local features to obtain the initial local features;

[0029] Perform the first grouping operation on the initial local features to obtain the first grouped features of the local area;

[0030] Perform the second grouping operation, further integrate and screen the first grouped features to obtain the local perception features.

[0031] Optionally, the processing process of the hybrid interaction module specifically includes:

[0032] Input the global perception feature and the local perception feature into the corresponding input ends of the hybrid interaction module respectively, use the adaptive weighting method to weight the global perception feature and the local perception feature, obtain the feature regions and dimensions worthy of attention, and through deep interaction iteration, analyze and fuse the multi-view features to obtain the multi-view complementary features.

[0033] Optionally, the processing process of the dynamic pyramid shrinking module specifically includes:

[0034] Input the multi-view complementary features into the dynamic pyramid shrinking module, and use the method of dynamic pyramid shrinking to iteratively decode layer by layer, progressively mine the multi-view complementary features, and output the prediction results of the model.

[0035] A camouflaged object semantic segmentation system based on multi-view collaborative learning, comprising:

[0036] A data acquisition module for acquiring multi-view images to be segmented;

[0037] A camouflaged object semantic segmentation module for inputting the multi-view images into an image segmentation processing model for prediction and outputting corresponding segmentation results; wherein, the image segmentation processing model includes an image enhancement module, a feature extraction module, a global-local perception module, a hybrid interaction module, and a dynamic pyramid shrinking module connected in sequence.

[0038] An electronic device, comprising a memory and a processor, the memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the described camouflaged object semantic segmentation method based on multi-view collaborative learning.

[0039] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the described camouflaged object semantic segmentation method based on multi-view collaborative learning.

[0040] The technical effects of the present invention are:

[0041] The camouflaged object semantic segmentation method based on multi-view collaborative learning provided by the present invention includes the following operations: input multi-view images to be segmented, use camouflaged object images with ground truth annotations to supervise and train the model, and obtain multi-view extraction features of the backbone network of the image to be segmented through supervised training; based on the multi-view extraction features, process the camouflaged objects in the multi-view images to be segmented by using cross-aggregation methods in different spaces and channels to obtain refined mask feature results; based on the mask feature results, utilize the spatial form and detail features of adaptively weighted fusion of global and local perception information; finally, use grouped iterative dynamic gated dilated convolution for layer-by-layer decoding, gradually fuse context information for each feature layer, the gating mechanism weakens the interference of useless information, and the pyramid shrinking strategy suppresses unimportant information and retains identifiable features, and outputs the predicted segmentation results of the model.

[0042] During the above operation process, when dealing with detecting camouflaged objects in images, first, the user provides an image. The algorithm generates a color jitter view, a perspective view, an oblique view, a magnified view, and a reduced view based on the input image, and inputs these six views into the model for training to segment the pixel regions containing camouflaged objects in the image. Specifically, the reason for using multi-view input is that multi-view feature extraction is beneficial for the model to better explore the information of camouflaged objects in the image. The variability, concealment, and scale non-fixedness of camouflaged objects pose significant challenges in this field. To comprehensively understand the image content, a large amount of context information is usually required, while the semantic information provided by single-view input is often insufficient. The reason for using the cross-aggregation method across different spaces and channels is that this operation helps the model retain more global and local information in different spaces and channels;

[0043] Specifically, the reason for adopting G-LPM is that GPM pays more attention to the structural features between modules, solves the problem of high dependence of single-view on pixel-level information, and mines more discriminative features. While LPM focuses more on small regions and can better capture the local morphology, texture, and spatial information of objects. Subsequently, HIM is adopted because this module learns which features are more important in the final prediction in the model through convolution and feature weighting reinforcement mechanisms of different branches, analyzes features from multiple perspectives, and adaptively weighting features is beneficial for the model to better explore the information of camouflaged objects in the image. The reason for adopting DSPM is that this module integrates visual information at different levels through an iterative aggregation strategy, thereby enhancing the decoder's ability to model multi-view features, and further improving the model's expression ability and feature extraction accuracy. Brief Description of the Drawings

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0045] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the accompanying drawings:

[0046] Figure 1 is the flowchart of the method implementation in the embodiment of the present invention;

[0047] Figure 2 is the overall structural schematic diagram of the segmentation model in the embodiment of the present invention;

[0048] Figure 3Schematic diagram of the hybrid interaction module structure in the embodiments of the present invention;

[0049] Figure 4 Schematic diagram of the dynamic gated dilated convolution module structure in the embodiments of the present invention;

[0050] Figure 5 Schematic diagram of the pyramid shrinking module structure in the embodiments of the present invention;

[0051] Figure 6 Schematic diagram of the input image in the embodiments of the present invention;

[0052] Figure 7 True value label in the embodiments of the present invention;

[0053] Figure 8 Effect diagram in the embodiments of the present invention;

[0054] Figure 9 Effect diagram of VSCode in the embodiments of the present invention;

[0055] Figure 10 Effect diagram of CamoDiffusion in the embodiments of the present invention;

[0056] Figure 11 Effect diagram of FEDER in the embodiments of the present invention. Detailed implementation manners

[0057] Now, various exemplary implementation manners of the present invention will be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, characteristics, and implementation schemes of the present invention.

[0058] It should be understood that the terms described in the present invention are only for describing specific implementation manners and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges can be independently included or excluded from the range.

[0059] Without departing from the scope or spirit of the present invention, various improvements and changes can be made to the specific implementation manners of the present invention's specification, which are obvious to those skilled in the art. Other implementation manners obtained from the present invention's specification are obvious to those skilled in the art. The present application's specification and embodiments are only exemplary.

[0060] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.

[0061] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0062] As Figure 1 - Figure 11 shown, in this embodiment, a method for semantic segmentation of camouflaged objects based on multi-view collaborative learning is provided, including: obtaining multi-view images to be segmented; inputting the multi-view images into an image segmentation processing model for prediction, and outputting corresponding segmentation results; wherein, the image segmentation processing model includes an image enhancement module, a feature extraction module, a global-local perception module, a hybrid interaction module, and a dynamic pyramid shrinkage module connected in sequence.

[0063] To solve the problems existing in the prior art, this embodiment proposes a method for semantic segmentation of camouflaged objects based on multi-view collaborative learning, inspired by the fact that when humans face unclear, blurred, or camouflaged targets, the visual system usually observes strategies such as different lighting and conditional environments (color jitter), viewing angles (tilt), scene perspective (projection), and scaling to find the target object more accurately. First, to accurately locate the target, this embodiment uses color space transformation and perception models to simulate color jitter, uses projective transformation and affine transformation to simulate tilt and projection effects, and simulates scaling strategies through scale space theory. This embodiment believes that by using information from different perspectives as supplementary inputs, it is possible to better obtain the complementarity and commonality between visual information and the dependence between different perspectives.

[0064] The method for semantic segmentation of camouflaged objects based on multi-view collaborative learning provided in this embodiment includes the following steps: S1 Input multi-view images to be segmented, and use camouflaged object images with ground truth annotations to supervise and train the model, and obtain the extraction results of multi-view features of the images to be segmented through the supervised training; S2 Based on the extraction results of the multi-view features, process the camouflaged objects in the images to be segmented in a global-local perception and hybrid interaction manner, and obtain multi-view complementary feature results by using the extracted diverse global information and local information in a hybrid interaction manner; S3 Based on the multi-view complementary feature results, use the method of dynamic pyramid shrinkage to iteratively decode layer by layer, and gradually mine the multi-view complementary features by decoding the features, and output the prediction results of the model.

[0065] The specific implementation process of this embodiment includes:

[0066] Input the multi-view image to be segmented, use the camouflaged object image with ground truth annotation to supervise and train the model, and obtain the multi-view extraction features of the backbone network of the image to be segmented through the supervised training; based on the multi-view extraction features, process the camouflaged objects in the multi-view image to be segmented by using the cross-aggregation method of different spaces and channels to obtain the high-resolution mask feature results; based on the mask feature results, use the spatial morphology and detail features of adaptive weighted fusion of global and local perception information; finally, use the grouped iterative dynamic gated dilated convolution for layer-by-layer decoding, and each feature layer progressively fuses the context information. The gating mechanism weakens the interference of useless information, and the pyramid shrinkage strategy suppresses unimportant information and retains the distinguishable features to output the predicted segmentation result of the model.

[0067] S101: Input the multi-view image to be segmented (the multi-view image to be segmented refers to the multi-view camouflaged image to be segmented), use the camouflaged object image with ground truth annotation to supervise and train the model (the model specifically refers to the image segmentation processing model), and obtain the multi-view extraction features of the backbone network of the image to be segmented through the supervised training;

[0068] S102: Based on the extraction result of the multi-view features, process the camouflaged objects in the image to be segmented by using the global-local perception and hybrid interaction method, and obtain the multi-view complementary feature results by using the hybrid interaction method for the extracted diverse global information and local information;

[0069] S103: Based on the multi-view complementary feature results, use the method of dynamic pyramid shrinkage for layer-by-layer iterative decoding, and progressively mine the multi-view complementary features by decoding the features to output the prediction result of the model.

[0070] As Figure 2 shown, the technical solution of this embodiment specifically divides the camouflaged object semantic segmentation task into the following three processing stages, namely the multi-view feature extraction stage, the global-local perception and hybrid interaction stage, and the dynamic pyramid shrinkage detection stage.

[0071] For the feature extraction stage, use the backbone network to perform multi-scale feature extraction on the multi-view image to be segmented;

[0072] For the global-local perception and hybrid interaction stage, first, use the global perception module to capture the global shape of the target and its relationship with the background from the multi-view input features (original view, magnified view, and reduced view) by means of the multi-head attention mechanism and linear units, obtaining high-resolution global mask features; then use the local perception module to extract fine-grained features of color, texture, deformation, and spatial relationships from the multi-view input features (color jitter view, tilted view, and projection view) by means of dilated convolution, depth convolution, batch processing, and activation layers, capturing local changes in different environments, obtaining high-resolution local mask features; then use the hybrid interaction module to effectively fuse the global spatial shape and local fine-grained features, and explore the correlation between global and local information as well as the commonalities and complementarities of multi-view feature information to gradually strengthen structural, visual, color, semantic, and location information from different perspectives, thereby improving the quality and discriminative ability of feature aggregation.

[0073] For the dynamic pyramid shrinking module, based on the multi-view complementary feature results, use the method of dynamic pyramid shrinking to iteratively decode layer by layer, and gradually mine the multi-view complementary features by decoding the features, and output the prediction results of the model.

[0074] As one or more embodiments, during the execution of the above step S101, use the camouflaged object images with ground truth annotations to supervise and train the model, and obtain the multi-scale feature extraction results of the multi-view images to be segmented through the supervised training, specifically including the following operation steps:

[0075] S10101: Use the backbone network and multi-view image processing to perform multi-scale feature extraction on the multi-view images to be segmented;

[0076] Among them, the multi-view image processing performs color jittering, projection, tilting, magnifying, and reducing operations, including: multi-source input of the original view, color jitter view, projection view, tilted view, magnified view, and reduced view.

[0077] The backbone network includes: a shallow embedding block, a middle embedding block 1, a middle embedding block 2, a middle embedding block 3, and a deep embedding block connected in sequence;

[0078] Among them, the backbone network includes, but is not limited to, any implementation: ResNet-50 backbone network or SwinTransformer backbone network. Therefore, the specific operations of each module included in the backbone network are not elaborated in the embodiments of the present application.

[0079] As one or more embodiments, during the execution of the above step S102, based on the extraction result of the multi-view features, the camouflaged objects in the image to be segmented are processed by means of global-local perception and hybrid interaction to obtain high-resolution mask features that are complementary from multiple perspectives. The specific operation steps are as follows:

[0080] S10201: Based on the second feature map, the third feature map, the fourth feature map, and the fifth feature map, they are respectively input into the global-local perception module (G-LPM) through a hierarchical operation to obtain the second multi-level mask map, the third multi-level mask map, the fourth multi-level mask map, and the fifth multi-level mask map of corresponding sizes;

[0081] Based on the second feature map, the third feature map, the fourth feature map, and the fifth feature map, they are input into the hybrid interaction module (HIM) through a hierarchical operation. The HIM uses adaptive weighting to obtain the feature regions and dimensions worthy of attention in the G-LPM, performs deep interaction iterations, and analyzes features from multiple perspectives.

[0082] The global-local perception module (G-LPM) includes:

[0083] A global perception module (GPM) and a local perception module (LPM);

[0084] Among them, the GPM and the LPM are independent of each other;

[0085] Among them, the GPM includes: a first set of side branches, a second set of side branches, a third set of side branches, and a combination block;

[0086] The first set of side branches includes: a convolutional layer with a convolution kernel size of 3*3, a fully connected layer, and then a first global feature is obtained through a multi-head attention mechanism operation;

[0087] The second set of side branches includes: a convolutional layer with a convolution kernel size of 5*5, a fully connected layer, and then a second global feature is obtained through a multi-head attention mechanism operation;

[0088] The third set of side branches includes: a convolutional layer with a convolution kernel size of 7*7, a fully connected layer, and then a third global feature is obtained through a multi-head attention mechanism operation;

[0089] The combination block includes: a convolutional layer with a convolution kernel size of 3*3, a batch normalization layer, and an activation function connected in sequence;

[0090] The LPM includes: a combination block, a first grouping, and a second grouping operation;

[0091] Among them, the second grouping operation is used to output the local second feature map, third feature map, fourth feature map, and fifth feature map;

[0092] The combined block includes: a convolutional layer with a convolutional kernel size of 3*3, a batch normalization layer, and an activation function connected in sequence;

[0093] The first grouping is the grouping operation of local features after the combined block fuses features;

[0094] The second grouping is to integrate and screen the information after the first grouping for secondary refinement grouping;

[0095] The HIM includes: a first GPM input end, a second GPM input end, a third GPM input end, a fourth GPM input end, a first LPM input end, a second LPM input end, a third LPM input end, and a fourth LPM input end;

[0096] Among them, the first GPM input end and the first LPM input end are jointly used to input the second feature map;

[0097] The second GPM input end and the second LPM input end are jointly used to input the third feature map;

[0098] The third GPM input end and the third LPM input end are jointly used to input the fourth feature map;

[0099] The fourth GPM input end and the fourth LPM input end are jointly used to input the fifth feature map;

[0100] As one or more embodiments, during the execution of the above step S103: Based on the multi-view complementary feature result, use the method of dynamic pyramid shrinkage to iteratively decode layer by layer, and through the decoded features, progressively mine the multi-view complementary features, and output the prediction result of the model (the model specifically refers to the image segmentation processing model), specifically including:

[0101] S10301: Use the method of dynamic pyramid shrinkage to gradually aggregate the multi-view complementary high-resolution mask features of the refined mask feature result, and finally output the predicted segmentation result map of the model (the model specifically refers to the image segmentation processing model).

[0102] The DPSM includes: a dynamic dilated gated convolution module (DDGC) and a pyramid shrinkage module (PSM);

[0103] Among them, DDGC and PSM are interrelated;

[0104] Among them, DDGC includes: By combining the receptive field weights of dynamic dilated gated convolution with the gating mechanism, adjust the weight coefficients of different branches and features, and focus on the most discriminative features in multi-level visual information. Specifically, split each feature in the feature along the channel dimension into two parts and Among them, after dynamic weight operation and then sigmoid activation operation, f is obtained G , successively through a series of cascaded operations such as a convolutional layer with a convolution kernel size of 3*3 and a dilation rate of 2, a convolutional layer with a convolution kernel size of 3*3 and a dilation rate of 4, a convolutional layer with a convolution kernel size of 3*3 and a dilation rate of 5, and a convolutional layer with a convolution kernel size of 3*3 and a dilation rate of 7, f is obtained D , f G and f D perform element-wise multiplication operation to obtain a feature map f with rich semantics G ' , f G ' after dynamic gating operation and then sigmoid activation operation, f is obtained S , and finally f G ' and f D perform element-wise multiplication operation to obtain the output;

[0105] The PSM includes: the interaction between features in the same layer and the interaction between cross-layer features. Different from the traditional pyramid shrinking module, here a progressive shrinking pyramid is adopted to capture the subtle but crucial features during the decoding process to the greatest extent and suppress the interference of background information. Specifically, assume G i and G i-1 are adjacent feature pairs in the current layer, where G i and G i-1 perform a series of cascaded operations such as convolution, batch normalization, and relu activation to aggregate valuable information and obtain the output.

[0106] The DPSM includes: dividing the feature block F into N groups (N represents the number of groups), and using a dynamic gating convolution that combines the dynamic receptive field weights of dilated convolutions between groups and a gating mechanism to control the weight coefficients, allowing the model to dynamically adjust different branches and features. Among them, the first group of features is copied into two copies (feature block 1 and feature block 2), and feature block 1 is divided into two copies (feature block 1' and feature block 1"), where feature block 1' is used for gating convolution with the next group of information, and feature block 1" is used for channel swapping. Feature block 2 is input into the PSM. The operation of the i-th group (1 < i < N) is similar to the above. The last group of feature blocks is different. After the gating convolution between the last group of feature blocks and the (N - 1)-th group of feature blocks, it does not need to be divided into two copies but is directly concatenated with the PSM, and finally, after feature fusion, it passes through the relu activation function to obtain the final detection result.

[0107] Deep neural network training, parameter initialization: For the shallow embedding block to the deep embedding block, the backbone network weight parameters pre-trained on the ImageNet1K dataset are used for network parameter initialization.

[0108] Training optimization details: The technical solution of this embodiment is implemented based on the PyTorch framework and trained using an NVIDIA RTX3090 GPU. This embodiment uses pre-trained basic encoders (ResNet-50, swin transformer) as the backbone network. The model parameters are updated using the Adam optimizer, with 120 training epochs, an initial learning rate of 1e-4, a batch size of 4, and a weight decay rate of 1e-4. It should be noted that during the training and testing processes of this embodiment, the size of the original view is adjusted to 384×384.

[0109] Dataset division: This embodiment is evaluated on three widely used public datasets, namely CAMO, COD10K, and NC4K. CAMO contains 1250 camouflage images, of which 1000 are used as training images and 250 are used as test images. CHAMELEON is a small dataset collected from the Internet through the Google search engine, consisting of 76 manually annotated images. COD10K is a more challenging large-scale dataset containing 10,000 images, including 5066 camouflage images, 3000 background images, and 1934 non-camouflaged images. NC4K is another COD dataset for testing, including 4121 test images.

[0110] Table 1 shows the quantitative comparison results between this embodiment and 18 state-of-the-art camouflaged object semantic segmentation models, including SegMaR, CamoFormer-R, ZoomNet, BSA-Net, BGNe, FPNet, FEDER, ICEG, MFFN, MSCAF-Net, HitNet, FSPNet, SARNet, PopNet, PRNet, VSCode-T, RISNet, CamoDiffusion.

[0111] Table 1 Quantitative comparison results between this embodiment and 18 state-of-the-art camouflaged object semantic segmentation models

[0112]

[0113]

[0114] Figures 6 - 11 The qualitative effect comparison diagram of the technical solution of this embodiment compared with the current three state-of-the-art models is shown.

[0115] The method for semantic segmentation of camouflaged objects based on multi-view collaborative learning in this embodiment includes the following operations: Input multi-view images to be segmented, use camouflaged object images with ground truth annotations to supervise the training of the model, and obtain multi-view extraction features of the backbone network of the images to be segmented through supervised training; Based on the multi-view extraction features, process the camouflaged objects in the multi-view images to be segmented by using cross-aggregation methods in different spaces and channels to obtain refined mask feature results; Based on the mask feature results, utilize the spatial morphology and detailed features of adaptive weighted fusion of global and local perception information; Finally, use grouped iterative dynamic gated dilated convolution for layer-by-layer decoding. Each feature layer progressively fuses context information. The gating mechanism weakens the interference of useless information, and the pyramid contraction strategy suppresses unimportant information, retains identifiable features, and outputs the predicted segmentation results of the model.

[0116] In the above operation process, when mining and processing camouflaged objects in images, first, the user provides an image. The algorithm generates color jitter views, projection views, tilt views, magnified views, and reduced views based on the image input by the user, and inputs the six views into the model for training to segment the pixel regions containing camouflaged objects in the image. Specifically, the reason for using multi-view input is that multi-view feature extraction is beneficial for the model to better explore the information of camouflaged objects in the image. The variability, concealment, and scale non-fixedness of camouflaged objects pose significant challenges in this field. To comprehensively understand the image content, a large amount of context information is usually required, while the semantic information provided by single-view input is often insufficient; The reason for using cross-aggregation methods in different spaces and channels is that this operation helps the model retain more global and local information in different spaces and channels.

[0117] Specifically, the reason for using G-LPM is that GPM pays more attention to the structural features between modules, solves the problem of high dependence on pixel-level information in single views, and mines more discriminative features, while LPM focuses more on small regions and can better capture the local morphology, texture, and spatial information of objects; Subsequently, HIM is used because this module learns which features in the model are more important in the final prediction through convolution and feature weighting reinforcement mechanisms in different branches, analyzes features from multiple perspectives, and adaptively weighting features is beneficial for the model to better explore the information of camouflaged objects in the image; The reason for using DSPM is that this module integrates visual information at different levels through an iterative aggregation strategy, thereby enhancing the decoder's ability to model multi-view features, and further improving the model's expression ability and feature extraction accuracy.

[0118] This embodiment also provides a semantic segmentation system for camouflaged objects based on multi-view collaborative learning, including:

[0119] The multi-view feature acquisition module is configured to: input multi-view images to be segmented, use camouflaged object images with ground-truth annotations to supervise and train the model, and obtain the extraction results of multi-scale features of the multi-view images to be segmented through the supervised training;

[0120] The global-local perception module is configured to: based on the extraction results of the multi-scale features of the multi-view images, use the global perception module to obtain the multi-source global and spatial information of the images for the multi-scale features of the original view, enlarged view, and reduced view, and use the local module to obtain the multi-source detailed information of the images for the multi-scale features of the color jitter view, tilted view, and projection view;

[0121] The hybrid interaction module is configured to: based on the extraction results of the multi-source global spatial information and multi-source detailed information, through an adaptive weighting mechanism, be divided into two stages to gradually strengthen the structure, vision, color, semantics, and position information from different perspectives. Through this process, HIM can achieve the interactive fusion of deep semantic features and spatial features, thereby improving the quality and discriminative ability of feature aggregation;

[0122] The dynamic pyramid shrinking module is configured to: based on the multi-view complementary feature results, use the method of dynamic pyramid shrinking to iteratively decode layer by layer, and gradually mine the multi-view complementary features through the decoded features to output the prediction results of the model.

[0123] It should be noted here that the above multi-view feature acquisition module, global-local perception module, hybrid interaction module, and dynamic pyramid shrinking module correspond to steps S101 to S103 in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0124] The system proposed in this embodiment can be implemented in various forms. For example, the foregoing system embodiment is only for illustrative purposes. The module division mentioned therein is only a possible solution based on logical functions. In actual implementation, there may be other division methods. For example, multiple modules can be merged or integrated into other systems, and some functions can also be selectively omitted or not executed according to actual needs.

[0125] This embodiment also proposes a system for semantic segmentation of camouflaged objects based on multi-view collaborative learning. The system consists of a processing unit and a storage module. Among them, the storage module is responsible for saving program instructions; the processing unit runs according to these instructions to implement the technical solutions described in Embodiment 1.

[0126] It should be noted that in this solution, the processing unit is not limited to a central processing unit (CPU), but can also be other general-purpose processors. Such general-purpose processors can be microprocessors or other common processing chips.

[0127] This embodiment also provides a computer-readable storage medium, which stores a computer program internally. When the program is run by the processing unit, it can execute the technical method described in the first embodiment.

[0128] As mentioned above, the above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A semantic segmentation method for camouflaged objects based on multi-view collaborative learning, characterized in that, Including: Obtain multi-view images to be segmented; Input the multi-view images into an image segmentation processing model for prediction, and output corresponding segmentation results; wherein, the image segmentation processing model includes an image enhancement module, a feature extraction module, a global-local perception module, a hybrid interaction module, and a dynamic pyramid shrinkage module connected in sequence.

2. The semantic segmentation method of camouflaged objects based on multi-view collaborative learning according to claim 1, characterized in that, The training process of the image segmentation processing model specifically includes: Obtain training data, where the training data includes multi-view training images and corresponding segmentation labels; Construct an initial image segmentation processing model, input the training data into the initial image segmentation processing model for prediction, and perform training with the goal of minimizing the loss between the predicted initial training results and the segmentation labels corresponding to the multi-view training images, to obtain a trained image segmentation processing model.

3. A semantic segmentation method of camouflaged objects based on multi-view collaborative learning according to claim 1, characterized in that The processing process of the image segmentation processing model specifically includes: Input the multi-view images into the image enhancement module for color space transformation, color jitter simulation, projection transformation, affine transformation simulation, and scale space theory simulation scaling processing, to obtain multi-source input data; Input the multi-source input data into the feature extraction module for multi-view feature extraction, to obtain feature maps of multiple scales, where the feature maps include a first feature map, a second feature map, a third feature map, a fourth feature map, and a fifth feature map; Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into the global-local perception module to extract global and local information, to obtain global perception features and local perception features; the global-local perception module includes a parallel global perception module and a local perception module; Input the global perception features and local perception features into the hybrid interaction module for hybrid interaction, to obtain multi-view complementary features; Input the multi-view complementary features into the dynamic pyramid shrinkage module for prediction, and output segmentation results.

4. A semantic segmentation method for camouflaged objects based on multi-view collaborative learning according to claim 3, characterized in that, The processing process of the feature extraction module specifically includes: Select a backbone network, input the multi-source input data into the backbone network, and the backbone network sequentially extracts features through a shallow embedding block, a first middle embedding block, a second middle embedding block, a third middle embedding block, and a deep embedding block, and outputs the first feature map, the second feature map, the third feature map, the fourth feature map, and the fifth feature map of different scales; wherein, in each embedding block of the backbone network, semantic features of the image are gradually extracted through a convolutional layer, a batch normalization layer, and an activation function.

5. A semantic segmentation method of camouflaged objects based on multi-view collaborative learning according to claim 3, characterized in that, The processing process of the global-local perception module specifically includes: The processing process of the global perception module specifically includes: Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into three side branches of the global perception module respectively: The first group of side branches: Extract features through a convolutional layer with a kernel size of 3×3, and then through a fully connected layer and a multi-head attention mechanism operation, to obtain a first global feature; The second group of side branches: Extract features through a convolutional layer with a kernel size of 5×5, and then through a fully connected layer and a multi-head attention mechanism operation, to obtain a second global feature; The third group of side branches: Extract features through a convolutional layer with a kernel size of 7×7, and then obtain the third global feature through a fully connected layer and a multi-head attention mechanism operation; Input the three global features into the combination block of the global perception module, and fuse them through a convolutional layer with a kernel size of 3×3, a batch normalization layer, and an activation function to obtain global perception features; The processing process of the local perception module specifically includes: Input the second feature map, the third feature map, the fourth feature map, and the fifth feature map into the combination block of the local perception module to preliminarily fuse the local features to obtain initial local features; Perform a first grouping operation on the initial local features to obtain the first grouped features of the local area; Perform a second grouping operation to further integrate and screen the first grouped features to obtain local perception features.

6. A semantic segmentation method for camouflaged objects based on multi-view collaborative learning according to claim 3, characterized in that The processing process of the hybrid interaction module specifically includes: Input the global perception features and the local perception features into the corresponding input ends of the hybrid interaction module respectively, use the adaptive weighting method to weight the global perception features and the local perception features, obtain the feature regions and dimensions worthy of attention, and through deep interaction iteration, analyze and fuse the multi-view features to obtain multi-view complementary features.

7. A semantic segmentation method of camouflaged objects based on multi-view collaborative learning according to claim 3, characterized in that The processing process of the dynamic pyramid shrinkage module specifically includes: Input the multi-view complementary features into the dynamic pyramid shrinkage module, and use the method of dynamic pyramid shrinkage to iteratively decode layer by layer, progressively mine the multi-view complementary features, and output the prediction results of the model.

8. A semantic segmentation system for camouflaged objects based on multi-view collaborative learning, characterized in that, It includes: A data acquisition module for acquiring multi-view images to be segmented; A camouflaged object semantic segmentation module for inputting the multi-view images into an image segmentation processing model for prediction and outputting corresponding segmentation results; wherein, the image segmentation processing model includes an image enhancement module, a feature extraction module, a global-local perception module, a hybrid interaction module, and a dynamic pyramid shrinkage module connected in sequence.

9. An electronic device, characterized in that, It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute a method for camouflaged object semantic segmentation based on multi-view collaborative learning according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor, it implements a method for camouflaged object semantic segmentation based on multi-view collaborative learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Camouflage target image segmentation method based on omnibearing perception

    CN114549567A

  • Camouflage object instance segmentation method, device and system based on multi-scale pooling modeling

    CN116433911A

  • Multi-modal feasible road segmentation method and system based on boundary perception

    CN117710667A

  • Camouflage object semantic segmentation method and system based on decision-level feature fusion modeling, medium and electronic equipment

    CN118470714A

  • Icon semantic segmentation method and system based on double-model prediction and automatic screening mechanism and storage medium

    CN118968072A

Cited By

  • Camouflage object semantic segmentation method and device based on adaptive candidate strategy, and medium

    CN121458978A