Detection method and device for camouflage object
By extracting and enhancing the target image, combining high and low frequency cross-frequency interactions, the image features are optimized to identify camouflage objects, and the problem of low detection accuracy of camouflage objects in the prior art is solved, achieving higher detection accuracy and detailed information mining.
Patent Information
- Application Number
- CN202510457339.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for detecting camouflage object are difficult to accurately identify target objects from the background, resulting in low detection accuracy.
By extracting the detected target image, multiple levels of image features are obtained, and image features are enhanced through densely connected hollow space pyramid pooling module, spatial attention module and channel attention module. Then, through feature complementation and fusion, combining high and low frequency cross-frequency interaction, the image features are optimized, and the camouflage object detection results of the target image are finally determined.
It improves the accuracy of camouflage object detection, can accurately identify the target object when the background environment is similar to the target object, and enhances the ability to mine image feature details.
Smart Images

Figure CN119992274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and in particular to a method and device for detecting a camouflaged object. Background Art
[0002] With the development of computer vision technology, image-based target detection technology has been widely used in various fields to complete specific tasks. Camouflaged object detection (COD) is one of the common tasks in target detection. Camouflaged objects refer to target objects that are easily integrated into the background environment, and the purpose of camouflaged object detection is to find targets hidden in the background environment.
[0003] At present, the detection method for camouflaged objects usually uses a sample data set of camouflaged objects to fine-tune the saliency detection algorithm, and then uses the adjusted algorithm to detect the camouflaged objects. The detection mechanism of the saliency detection algorithm is mainly to evaluate the saliency score of each pixel in the image, and identify the target in the image based on the saliency score of each pixel.
[0004] In existing detection methods, the recognition of camouflaged objects depends on the significant difference between the target object and the background environment. In the actual detection scenario of camouflaged objects, the camouflaged objects in the image usually have a high similarity with the background environment of the image. It is usually difficult to accurately distinguish the target object based on the existing methods, resulting in low accuracy of camouflaged object detection. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides a method for detecting a disguised object to solve the problem that it is difficult to accurately identify a target object from a background in the existing disguised object detection method, resulting in low detection accuracy of the disguised object.
[0006] The embodiment of the present invention further provides a detection device for a camouflaged object to ensure the practical implementation and application of the above method.
[0007] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0008] A first aspect of an embodiment of the present invention provides a method for detecting a camouflaged object, comprising:
[0009] Perform feature extraction on the target image to be detected to obtain image features at multiple levels;
[0010] Among the image features of each of the said levels, the image feature with the highest level is used as the first image feature, and each of the remaining image features is used as the second image feature;
[0011] Performing feature enhancement on the first image feature through a preset densely connected atrous spatial pyramid pooling module, a spatial attention module, and a channel attention module to obtain various enhanced features;
[0012] For each of the enhancement features, the enhancement feature is used as a dominant feature, and each enhancement feature other than the enhancement feature is used as a supplementary feature, so that the enhancement feature and the remaining enhancement features are complemented with each other to obtain a complementary optimization feature corresponding to the enhancement feature;
[0013] Performing feature fusion on each of the complementary optimization features to obtain enhanced image features;
[0014] In descending order of the hierarchical levels, based on the enhanced image features, sequentially performing high- and low-frequency cross-frequency interaction on each of the second image features to obtain a frequency enhanced image feature corresponding to each of the second image features;
[0015] The disguised object detection result of the target image is determined based on the frequency enhanced image feature corresponding to the second image feature with the lowest hierarchical level among the second image features.
[0016] A second aspect of an embodiment of the present invention provides a detection device for a disguised object, comprising:
[0017] A feature extraction unit, used to extract features of the target image to be detected and obtain image features at multiple levels;
[0018] A feature classification unit, configured to use, among the image features of each of the said levels, the image feature with the highest level as the first image feature, and use each of the remaining image features as the second image feature;
[0019] A feature enhancement unit, used to perform feature enhancement on the first image feature through a preset densely connected void space pyramid pooling module, a spatial attention module, and a channel attention module, respectively, to obtain various enhanced features;
[0020] A feature supplementation unit is used to, for each of the enhanced features, use the enhanced feature as a dominant feature and each enhanced feature other than the enhanced feature as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complemented with each other to obtain a complementary optimization feature corresponding to the enhanced feature;
[0021] A feature fusion unit, used for fusing the complementary optimization features to obtain enhanced image features;
[0022] a feature interaction unit, configured to perform high- and low-frequency cross-frequency interaction on each of the second image features in sequence based on the enhanced image features in descending order of hierarchical levels, so as to obtain a frequency enhanced image feature corresponding to each of the second image features;
[0023] The result determination unit is configured to determine a disguised object detection result of the target image based on a frequency enhanced image feature corresponding to a second image feature with the lowest hierarchical level among the second image features.
[0024] A camouflaged object detection method provided based on the above-mentioned embodiment of the present invention includes: extracting features from a target image to be detected to obtain image features of multiple levels; taking the image feature with the highest level in each level as the first image feature, and taking each of the remaining image features as the second image feature; performing feature enhancement on the first image feature through a preset densely connected void space pyramid pooling module, a spatial attention module, and a channel attention module to obtain each enhanced feature; for each enhanced feature, taking the enhanced feature as the dominant feature, taking each enhanced feature other than the enhanced feature as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complementary to each other to obtain a complementary optimized feature corresponding to the enhanced feature; performing feature fusion on each complementary optimized feature to obtain an enhanced image feature; performing high-low frequency cross-frequency interaction on each second image feature in order from high to low levels based on the enhanced image feature to obtain a frequency enhanced image feature corresponding to each second image feature; and determining a camouflaged object detection result of the target image based on the frequency enhanced image feature corresponding to the second image feature with the lowest level in each second image feature. By applying the method provided by the embodiment of the present invention, the target image to be detected can be feature extracted to obtain multi-level image features. By enhancing the information of the highest-level image features with richer semantic information, the enhanced image features are used to guide the cross-frequency interactive optimization of image features at other levels, and more accurate detail information can be mined from the image features at each level. Target detection can be performed based on the image features finally optimized, thereby identifying camouflaged objects in the target image. Multi-dimensional optimization based on image features can mine more accurate details in image features, which is conducive to accurate segmentation of camouflaged objects and background environments in images. When the background environment is similar to the target object, the target object can also be identified in the background environment, which is conducive to improving the accuracy of camouflaged object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0026] Figure 1 A method flow chart of a method for detecting a camouflaged object provided by an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of the structure of a cross-frequency interactive network provided by an embodiment of the present invention;
[0028] Figure 3 A schematic diagram of a feature discussion interaction process provided by an embodiment of the present invention;
[0029] Figure 4 A schematic diagram of a cross-frequency interaction module provided by an embodiment of the present invention;
[0030] Figure 5 A schematic diagram of the structure of a detection device for a camouflaged object provided by an embodiment of the present invention;
[0031] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element.
[0034] The embodiment of the present invention provides a method for detecting a disguised object. The method can be applied to a target detection system for detecting disguised objects. The execution subject can be a server of the system. The method flow chart of the method is as follows: Figure 1As shown, including:
[0035] S101: extracting features of the target image to be detected to obtain image features at multiple levels;
[0036] The method provided by the embodiment of the present invention can be applied to target detection scenarios of various camouflaged objects, that is, detection scenarios where the detection target is relatively similar to the background environment, such as polyp recognition and tumor recognition in medicine, defect detection in industry, locust monitoring in agriculture, etc. When it is necessary to detect a target object (i.e., a camouflaged object) in a certain image, the image can be used as a target image and loaded into a target detection system for target detection. The target detection system as a whole can also be regarded as a target detection model for detecting camouflaged objects.
[0037] In the method provided by the embodiment of the present invention, a deep convolutional network structure can be used to perform multi-scale feature extraction on the target image. For example, Res2Net-50 can be used as the backbone structure for feature extraction. Res2Net-50 is an improved residual network structure and an existing network structure, which will not be described in detail here. After feature extraction for the target image, multiple levels of image features can be obtained, and each image feature is a feature map. It can be understood that each level has a corresponding level, and the level of each level represents the height of the level, or the depth of the level. The higher the level of the level (that is, the deeper the level of the level), the richer the semantic information of the image features.
[0038] S102: Among the image features of each of the layers, the image feature with the highest level is used as the first image feature, and each of the remaining image features is used as the second image feature;
[0039] In the method provided by the embodiment of the present invention, in descending order of hierarchical levels, the image feature with the highest hierarchical level among the extracted image features is taken as the first image feature, which is the highest level feature among the image features, or the deepest feature. At the same time, each image feature except the first image feature among the image features is taken as the second image feature.
[0040] S103: performing feature enhancement on the first image feature through a preset densely connected atrous spatial pyramid pooling module, a spatial attention module, and a channel attention module, respectively, to obtain various enhanced features;
[0041] In the method provided by the embodiment of the present invention, a densely connected atrous spatial pyramid pooling module, a spatial attention module and a channel attention module are pre-set. The densely connected atrous spatial pyramid pooling module is a module constructed based on the dense atrous spatial pyramid pooling (Dense Atrous Spatial Pyramid Pooling, DenseASPP) structure, which can refine features from multi-scale dimensions. The spatial attention module is a module constructed based on the spatial attention (SpatialAttention, SA) mechanism, which can refine features in the spatial dimension. The channel attention module is a module constructed based on the channel attention (Channel Attention, CA) mechanism, which can refine features from the channel dimension. The dense atrous spatial pyramid pooling structure, the spatial attention mechanism and the channel attention mechanism are all existing network structures or mechanisms, which will not be described in detail here.
[0042] In the method provided by the embodiment of the present invention, multi-dimensional feature enhancement is performed on the first image feature through a densely connected atrous spatial pyramid pooling module, a spatial attention module, and a channel attention module, respectively, and the features can be refined from multi-scale dimensions, spatial dimensions, and channel dimensions to obtain various enhanced features. Specifically, a densely connected atrous spatial pyramid pooling module is applied to perform feature enhancement processing on the first image feature, and the image feature obtained after the processing is used as an enhanced feature. The spatial attention module is applied to perform feature enhancement processing on the first image feature, and the image feature obtained after the processing is used as an enhanced feature. The channel attention module is applied to perform feature enhancement processing on the first image feature, and the image feature obtained after the processing is used as an enhanced feature, thereby obtaining three enhanced features.
[0043] S104: for each enhanced feature, the enhanced feature is used as a dominant feature, and each enhanced feature other than the enhanced feature is used as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complemented with each other to obtain a complementary optimization feature corresponding to the enhanced feature;
[0044] In the method provided by the embodiment of the present invention, feature complementation processing is performed for each enhanced feature respectively to further refine each enhanced feature. When performing feature complementation processing on each enhanced feature, the enhanced feature currently undergoing feature complementation processing (i.e., the current enhanced feature) is used as the dominant feature, and each enhanced feature except the current enhanced feature among the enhanced features is used as a supplementary feature. With the current enhanced feature as a guide, the remaining enhanced features (each enhanced feature except the current enhanced feature) are feature-complemented with the current enhanced feature, and the features obtained after feature complementation between the current enhanced feature and the remaining enhanced features are used as the complementary optimized features corresponding to the current enhanced feature. Thus, the complementary optimized features corresponding to each enhanced feature are obtained. Feature complementation processing can be completed through operations such as feature concatenation, feature addition, convolution, regularization, and activation function processing.
[0045] S105: performing feature fusion on each of the complementary optimization features to obtain enhanced image features;
[0046] In the method provided in the embodiment of the present invention, each complementary optimization feature can be further fused through operations such as convolution, regularization or activation function processing, and the feature fusion result of each complementary optimization feature can be used as the enhanced image feature.
[0047] S106: performing high- and low-frequency cross-frequency interaction on each of the second image features in order from high to low hierarchical levels based on the enhanced image features, to obtain a frequency enhanced image feature corresponding to each of the second image features;
[0048] In the method provided by the embodiment of the present invention, high-low frequency cross-frequency interaction is performed on each second image feature in sequence from high to low hierarchical levels, that is, firstly, the second image feature with the highest hierarchical level is processed with high-low frequency cross-frequency interaction to obtain the frequency enhanced image feature corresponding to the second image feature, and then the second image feature of the next level is processed, and so on, and finally the second image feature with the lowest hierarchical level is processed.
[0049] Specifically, when processing each second image feature, an image feature at a higher level than the current image feature (i.e., the second image feature currently being processed) is subjected to cross-frequency interaction of high and low frequency features with the current image feature. For example, on the one hand, the enhanced image feature can be used as a global guiding feature, and on the other hand, the feature associated with the image feature at an upper level of the current image feature can be used as a local guiding feature. Based on the global guiding feature and the local guiding feature, a cross-frequency interaction of high and low frequency features is performed on the current image feature, and the final processed result is used as the frequency enhanced image feature corresponding to the current image feature. The feature associated with the image feature of the previous level of the current image feature refers to the image feature obtained by processing the image feature of the previous level of the current image feature. For example, for the second image feature with the highest hierarchical level, the image feature of the previous level is the first image feature, and the enhanced image feature can be used as the local guiding feature. For the remaining second image features (the second image feature not at the highest hierarchical level), the image feature of the previous level is the second image feature of the previous level. The feature obtained by processing the high-low frequency cross-frequency interaction of the second image feature of the previous level can be used as the current local guiding feature. For example, the intermediate feature obtained by processing the high-low frequency cross-frequency interaction of the second image feature of the previous level can be used as the local guiding feature. For example, the frequency enhanced image feature corresponding to the second image feature of the previous level can be used as the current local guiding feature. The selection method of the local guiding feature can be selected according to actual needs without affecting the function of the method provided in the embodiment of the present invention.
[0050] S107: Determine a disguised object detection result of the target image based on a frequency enhanced image feature corresponding to a second image feature with the lowest level among the second image features.
[0051] In the method provided by the embodiment of the present invention, each second image feature is sequentially subjected to high- and low-frequency cross-frequency interaction in order of hierarchical levels from high to low, so when processing the second image feature with the lowest hierarchical level, the detailed information of the high-level image feature has been integrated, that is, the frequency-enhanced image feature corresponding to the second image feature with the lowest hierarchical level is a feature that integrates the enhanced image feature and the high- and low-frequency cross-frequency interaction results of all second image features. Based on the frequency-enhanced image feature corresponding to the second image feature with the lowest hierarchical level, the camouflaged object in the target image is identified, thereby obtaining a camouflaged object detection result.
[0052] Based on the method provided by the embodiment of the present invention, when it is necessary to detect a camouflaged object in a target image, feature extraction is performed on the target image to be detected to obtain image features of multiple levels; the image feature with the highest level in each level of the image features is used as the first image feature, and each of the remaining image features is used as the second image feature; the first image feature is feature enhanced by a preset densely connected void space pyramid pooling module, a spatial attention module, and a channel attention module to obtain each enhanced feature; for each enhanced feature, the enhanced feature is used as the dominant feature, and each enhanced feature other than the enhanced feature is used as a supplementary feature, so that the enhanced feature and the remaining enhanced features are feature-complementary to obtain a complementary optimized feature corresponding to the enhanced feature; feature fusion is performed on each complementary optimized feature to obtain an enhanced image feature; in order from high to low levels of the levels, based on the enhanced image feature, each second image feature is sequentially subjected to high-low frequency cross-frequency interaction to obtain a frequency enhanced image feature corresponding to each second image feature; based on the frequency enhanced image feature corresponding to the second image feature with the lowest level in each second image feature, a camouflaged object detection result of the target image is determined. By applying the method provided by the embodiment of the present invention, the target image to be detected can be feature extracted to obtain multi-level image features. By enhancing the information of the highest-level image features with richer semantic information, the enhanced image features are used to guide the cross-frequency interactive optimization of image features at other levels, and more accurate detail information can be mined from the image features at each level. Target detection can be performed based on the image features finally optimized, thereby identifying camouflaged objects in the target image. Multi-dimensional optimization based on image features can mine more accurate details in image features, which is conducive to accurate segmentation of camouflaged objects and background environments in images. When the background environment is similar to the target object, the target object can also be identified in the background environment, which is conducive to improving the accuracy of camouflaged object detection.
[0053] exist Figure 1 On the basis of the method shown in the figure, in the method provided by the embodiment of the present invention, the process mentioned in step S103 of performing feature enhancement on the first image feature through a preset densely connected void space pyramid pooling module, a spatial attention module and a channel attention module to obtain each enhanced feature includes:
[0054] Perform channel splitting on the first image feature to obtain a first branch feature, a second branch feature, and a third branch feature;
[0055] In the method provided by the embodiment of the present invention, the convolution operation can be used to split the first image feature into three parts in the channel dimension to obtain three branch features, namely, the first branch feature, the second branch feature and the third branch feature.
[0056] Applying the densely connected atrous spatial pyramid pooling module, performing channel splitting on the first branch feature to obtain a plurality of pyramid sub-features, and performing multi-scale optimization on each of the pyramid sub-features through a densely connected splicing operation and a residual connection operation to obtain a multi-scale optimized feature;
[0057] In the method provided by the embodiment of the present invention, the first branch feature can be feature processed by a densely connected atrous space pyramid pooling module. The densely connected atrous space pyramid pooling structure can optimize feature information transmission through dense connections and atrous convolutions with different atrous rates to obtain features of different scales in image features. Specifically, when the densely connected atrous space pyramid pooling module is applied for processing, the module can further perform channel splitting on the first branch feature, and use the split features as pyramid sub-features. Through the densely connected splicing operation and residual connection operation in the module, multi-scale optimization is performed on each pyramid sub-feature, thereby extracting feature information of different scales, and using the processed features as multi-scale optimized features corresponding to the first branch feature.
[0058] Applying the spatial attention module, performing dimensionality reduction processing on the second branch feature to obtain a dimensionality reduction feature map, performing a pooling convolution operation on the dimensionality reduction feature map to obtain a spatial attention map, and performing element-by-element multiplication processing on the spatial attention map and the second branch feature to obtain a spatial optimization feature;
[0059] In the method provided by the embodiment of the present invention, the second branch feature can be feature processed by the spatial attention module. The spatial attention module can perform weighted processing on the feature data according to the importance of different areas in the feature through the spatial attention mechanism to highlight the feature areas with high contribution to target recognition. Specifically, when applying the spatial attention module to process the second branch feature, the module can perform dimensionality reduction processing on the second branch feature to obtain a dimensionality reduction feature map corresponding to the second branch feature, and perform a pooling convolution operation on the dimensionality reduction feature map to obtain a corresponding spatial attention map, perform element-wise multiplication processing on the spatial attention map and the second branch feature, and use the processing result as the spatial optimization feature corresponding to the second branch feature.
[0060] Applying the channel attention module to generate a channel attention map corresponding to the third branch feature, and using the channel attention map to dynamically adjust the weight of the channel in the third branch feature to obtain a channel optimization feature;
[0061] In the method provided by the embodiment of the present invention, the third branch feature can be feature processed by a channel attention module. The channel attention module can perform weighted processing on different channels of the feature through the channel attention mechanism to highlight the channels that are more important for target recognition. Specifically, when the channel attention module is applied to process the third branch feature, the module can generate a corresponding channel attention map for the third branch feature, and use the channel attention map to dynamically adjust the weight of the channel in the third branch feature, thereby optimizing the feature in the channel dimension, and using the adjusted feature as the channel optimization feature corresponding to the third branch feature.
[0062] The multi-scale optimization feature, the space optimization feature and the channel optimization feature are used as the respective enhanced features.
[0063] In the method provided by an embodiment of the present invention, the multi-scale optimization feature corresponding to the first branch feature, the spatial optimization feature corresponding to the second branch feature, and the channel optimization feature corresponding to the third branch feature are respectively used as enhanced features, thereby obtaining various enhanced features.
[0064] exist Figure 1 On the basis of the method shown, in the method provided by the embodiment of the present invention, the process mentioned in step S104 of taking the enhanced feature as the dominant feature, taking each enhanced feature other than the enhanced feature as the supplementary feature, making the enhanced feature complement each other with the remaining enhanced features, and obtaining the complementary optimization feature corresponding to the enhanced feature includes:
[0065] Based on the enhanced feature, each enhanced feature except the enhanced feature is optimized to obtain a first optimized feature and a second optimized feature;
[0066] In the method provided by the embodiment of the present invention, each enhanced feature is a feature obtained by processing the first image feature respectively through a densely connected hollow space pyramid pooling module, a spatial attention module and a channel attention module, so a total of three enhanced features will be obtained. When performing feature complementary processing on each enhanced feature respectively, the currently processed enhanced feature is used as the dominant feature, and each of the remaining enhanced features is used as a supplementary feature. During the processing, for each current supplementary feature, the supplementary feature is feature optimized with the current dominant feature, and the optimization result of each supplementary feature is used as an optimized feature, thereby obtaining a first optimized feature and a second optimized feature. For example, each enhanced feature obtained by feature enhancement of the first image feature is referred to as a first enhanced feature, a second enhanced feature and a third enhanced feature, respectively. When performing feature complementary processing on the first enhanced feature, the first enhanced feature is used as the dominant feature, the second enhanced feature and the third enhanced feature are used as supplementary features, the second enhanced feature is feature optimized based on the first enhanced feature, and the feature obtained after optimization is used as the first optimized feature, and the third enhanced feature is feature optimized based on the first enhanced feature, and the feature obtained after optimization is used as the second optimized feature. When the second enhanced feature is subjected to feature complementation processing, the second enhanced feature is used as the dominant feature, and the first enhanced feature and the third enhanced feature are used as supplementary features respectively. The third enhanced feature is processed in the same manner.
[0067] In the method provided by the embodiment of the present invention, the process of optimizing the remaining enhanced features based on the current enhanced feature can be implemented based on operations such as feature splicing, feature addition, convolution, regularization, and activation function processing. For example, in the feature complementation process with the first enhanced feature as the dominant feature, the feature optimization of the second enhanced feature based on the first enhanced feature can be performed by splicing the first enhanced feature with the second enhanced feature, and then performing operations such as convolution, regularization, and activation function processing on the feature splicing result, adding elements of the processed feature splicing result to the second enhanced feature to obtain an element addition result, and further performing operations such as convolution, regularization, and activation function processing on the element addition result, and using the processed result as the current first optimized feature. The principle of feature optimization of the third enhanced feature based on the first enhanced feature is the same as the principle of feature optimization of the second enhanced feature, and will not be repeated here.
[0068] Based on the first optimized feature, performing feature optimization on the second optimized feature to obtain a third optimized feature;
[0069] Based on the second optimized feature, performing feature optimization on the first optimized feature to obtain a fourth optimized feature;
[0070] In the method provided by the embodiment of the present invention, after obtaining the first optimization feature and the second optimization feature, the two optimization features are interactively optimized respectively. The second optimization feature is feature optimized with the first optimization feature, and the feature obtained after optimization is used as the third optimization feature, and the first optimization feature is feature optimized with the second optimization feature, and the feature obtained after optimization is used as the fourth optimization feature. Specifically, the process of feature optimization based on the second optimization feature / first optimization feature can be implemented by operations such as feature splicing, feature addition, convolution, regularization, and activation function processing. For example, the second optimization feature is feature optimized based on the first optimization feature, and the second optimization feature can be feature spliced with the first optimization feature, and then the feature splicing result is subjected to operations such as convolution, regularization, and activation function processing, and the processed feature splicing result is element-wise added with the second optimization feature, and the element addition result is subjected to operations such as convolution, regularization, and activation function processing, and the processed result is used as the third optimization feature. The principle of feature optimization of the first optimization feature based on the second optimization feature is the same as the principle of feature optimization of the second optimization feature based on the first optimization feature, and will not be repeated here.
[0071] The enhanced feature, the third optimized feature and the fourth optimized feature are subjected to feature fusion, and a result of the feature fusion is used as a complementary optimized feature corresponding to the enhanced feature.
[0072] In the method provided by the embodiment of the present invention, the enhanced feature, the third optimized feature and the fourth optimized feature that are currently undergoing feature complementary processing can be subjected to feature fusion processing through operations such as feature splicing, convolution, regularization and activation function processing, and the processed features can be used as complementary optimized features corresponding to the current enhanced feature. For example, the current enhanced feature can be subjected to feature splicing with the third optimized feature and the fourth optimized feature, and operations such as convolution, regularization and activation function processing can be performed on the feature splicing result, and the processed features can be used as complementary optimized features corresponding to the current enhanced feature.
[0073] Based on the method provided in the embodiment of the present invention, when performing feature complementation processing on the enhanced features, multiple feature interaction optimizations can be performed to complement the current enhanced features with the remaining enhanced features, so that the details of the remaining enhanced features can be better utilized to complement the current enhanced features, enrich the information of the current enhanced features, and further improve the feature optimization effect.
[0074] exist Figure 1 On the basis of the method shown, in the method provided by the embodiment of the present invention, the process of performing high-low frequency cross-frequency interaction on each second image feature mentioned in step S106 includes:
[0075] Based on the current upper-level image feature, determining the target image feature and the target prediction map corresponding to the current second image feature; the current upper-level image feature is an image feature whose hierarchical level is one level higher than the current second image feature among the image features of each of the hierarchical levels;
[0076] In the method provided by the embodiment of the present invention, when processing high-low frequency cross-frequency interaction for each second image feature, the image feature of the previous level of the current second image feature (i.e., the second image feature currently undergoing high-low frequency cross-frequency interaction processing) is used as the current previous level image feature, that is, among the image features of each level obtained by extracting features from the target image, in descending order of the hierarchical levels, the image feature one level previous to the current second image feature is the current previous level image feature. It can be understood that if the current second image feature is the second image feature with the highest hierarchical level, then the current previous level image feature is the first image feature, and if the current second image feature is not the second image feature with the highest hierarchical level, then the current previous level image feature is the second image feature of the previous level.
[0077] In the method provided by the embodiment of the present invention, the target image feature and the target prediction map corresponding to the current second image feature are determined according to the current upper image feature, so as to realize the feature cross-frequency interaction of the dual branches by using the target image feature and the target prediction map. The target image feature is the image feature obtained after the upper image feature is processed. For example, if the upper image feature is the first image feature, the target image feature may be an enhanced image feature. It can be known from the method provided by the above embodiment that the enhanced image feature is the image feature obtained after the first image feature is processed by feature enhancement, feature complementation and feature fusion. If the upper image feature is the second image feature of the previous level, the high-low frequency cross-frequency interaction process of the upper image feature has been completed, and the current target image feature may be the intermediate feature in the high-low frequency cross-frequency interaction process of the upper image feature. The target prediction map is the prediction map obtained by processing the associated features of the upper image feature. The associated features of the upper image feature refer to the features obtained after processing the upper image feature. For example, the target prediction map may be the prediction map obtained by processing the target image feature, or it may be the prediction map corresponding to other features obtained in the high-low frequency cross-frequency interaction process of the upper image feature.
[0078] Fusing the current second image feature with the target image feature to obtain a guiding branch feature;
[0079] Performing feature fusion on the current second image feature and the target prediction map to obtain prediction map branch features;
[0080] In the method provided by the embodiment of the present invention, the current second image feature can be fused with the target image feature and the target prediction map, respectively, so as to obtain the guide branch feature and the prediction map branch feature corresponding to the current second image feature. Specifically, in the feature fusion process, the current second image feature can be channel-reduced according to actual needs, and the second image feature after channel reduction is used as the compressed image feature.
[0081] In the method provided by an embodiment of the present invention, the compressed image feature corresponding to the current second image feature can be element-wise added to the target image feature to achieve feature fusion of the current second image feature and the target image feature, and the result of the element-wise addition of the compressed image feature and the target image feature is used as a guiding branch feature.
[0082] In the method provided by an embodiment of the present invention, the compressed image feature corresponding to the current second image feature can be element-wise multiplied with the target prediction map to achieve feature fusion of the current second image feature and the target prediction map, and the result of the element-wise multiplication of the compressed image feature and the target prediction map is used as the prediction map branch feature.
[0083] Performing cross-frequency interaction on the enhanced image feature, the guidance branch feature, and the prediction graph branch feature, respectively, to obtain a first cross-frequency interaction result and a second cross-frequency interaction result corresponding to the current second image feature;
[0084] In the method provided by the embodiment of the present invention, a cross-frequency interaction network can be constructed in advance according to actual needs. The pre-established cross-frequency interaction network is applied to perform feature cross-frequency interaction on the enhanced image features and the guided branch features, and the features obtained after processing are used as the first cross-frequency interaction results corresponding to the current second image features. Similarly, the pre-established cross-frequency interaction network is applied to perform feature cross-frequency interaction on the enhanced image features and the predicted graph branch features, and the features obtained after processing are used as the second cross-frequency interaction results corresponding to the current second image features.
[0085] Feature fusion is performed on the first cross-frequency interaction result and the second cross-frequency interaction result, and a result of the feature fusion is used as a frequency enhanced image feature corresponding to the current second image feature.
[0086] In the method provided by the embodiment of the present invention, the first cross-frequency interaction result and the second cross-frequency interaction result may be feature fused, and the fusion result of the first cross-frequency interaction result and the second cross-frequency interaction result may be used as the frequency enhanced image feature corresponding to the current second image feature.
[0087] Based on the method provided in the embodiment of the present invention, the global highest-level image features, the related features of the previous level of the current features and the prediction graph can be combined to perform cross-frequency interactive optimization on the current features, and guide the current features to perform pixel-level optimization from a global and local perspective, which is conducive to further mining feature information and improving the accuracy of target detection.
[0088] On the basis of the method provided in the above embodiment, in the method provided in the embodiment of the present invention, if the current second image feature is the second image feature with the highest hierarchical level among the second image features, the process of determining the target image feature and the target prediction map corresponding to the current second image feature based on the current upper-level image feature includes:
[0089] Using the enhanced image feature as the target image feature corresponding to the current second image feature;
[0090] In the method provided by an embodiment of the present invention, when high- and low-frequency cross-frequency interaction is performed on each second image feature, if the second image feature currently being processed is the second image feature with the highest hierarchical level, the current upper-level image feature is the first image feature, and the enhanced image feature can be used as the target image feature corresponding to the current second image feature.
[0091] Performing a convolution operation on the enhanced image feature to obtain a prediction map corresponding to the enhanced image feature, and using the prediction map corresponding to the enhanced image feature as a target prediction map corresponding to the current second image feature;
[0092] In the method provided by the embodiment of the present invention, if the current second image feature is the second image feature with the highest hierarchical level, a convolution operation can be performed on the enhanced image feature to obtain a prediction map corresponding to the enhanced image feature, and the prediction map is used as the current target prediction map. It can be understood that the prediction map is a result image of target detection obtained by processing the feature. Performing a convolution operation on a specified feature to obtain a prediction map corresponding to the specified feature refers to a prediction map obtained on the basis of performing a convolution operation on the specified feature using the specified feature as feature data for prediction. In the specific processing process, other operations can be further performed on the basis of the convolution operation according to actual needs to finally obtain a prediction map.
[0093] If the current second image feature is not the second image feature with the highest level in the hierarchy, the process of determining the target image feature and the target prediction map corresponding to the current second image feature based on the current upper-level image feature includes:
[0094] Determining a target image feature corresponding to the current second image feature based on a second cross-frequency interaction result corresponding to the current upper-level image feature;
[0095] In the method provided by the embodiment of the present invention, when performing high-low frequency cross-frequency interaction on each second image feature, if the current second image feature is not the second image feature with the highest hierarchical level, the current upper-level image feature is the second image feature of the upper level. It can be understood that the high-low frequency cross-frequency interaction processing of the upper-level image feature has been completed, and the first cross-frequency interaction result, the second cross-frequency interaction result, and the frequency enhanced image feature corresponding to the upper-level image feature can be obtained. At this time, the current target image feature can be determined based on the second cross-frequency interaction result corresponding to the upper-level image feature. For example, the second cross-frequency interaction result can be used as the target image feature, or the feature obtained by further processing the second cross-frequency interaction result of the upper-level image feature can be used as the target image feature, for example, the product of the second cross-frequency interaction result of the upper-level image feature and the preset scale parameter is used as the target image feature.
[0096] A convolution operation is performed on the frequency enhanced image feature corresponding to the current upper-level image feature to obtain an enhanced prediction map corresponding to the current upper-level image feature, and the enhanced prediction map is used as a target prediction map corresponding to the current second image feature.
[0097] In the method provided by an embodiment of the present invention, when the current second image feature is not the second image feature with the highest hierarchical level, the prediction map corresponding to the frequency enhanced image feature corresponding to the current upper-level image feature is used as the target prediction map, that is, a convolution operation is performed on the frequency enhanced image feature corresponding to the upper-level image feature, and the prediction map obtained thereby (the so-called enhanced prediction map) is used as the target prediction map corresponding to the current second image feature.
[0098] Based on the method provided by the embodiment of the present invention, when the current second image feature is not the second image feature with the highest hierarchical level, the second cross-frequency interaction result corresponding to the upper-level image feature can be used to determine the current target image feature, which is used to determine the guiding branch feature corresponding to the current second image feature. The current guiding branch feature can be determined in combination with the features processed by the branch features of the prediction graph of the previous level. This can better utilize feature information of different levels and dimensions, which is conducive to further refining the features and improving the accuracy of target recognition.
[0099] On the basis of the method provided in the above embodiment, in the method provided in the embodiment of the present invention, the process of performing cross-frequency interaction between the enhanced image feature and the guidance branch feature and the prediction graph branch feature respectively to obtain a first cross-frequency interaction result and a second cross-frequency interaction result corresponding to the current second image feature includes:
[0100] Applying a pre-trained local frequency interaction model to perform local feature interaction on the guide branch feature and the enhanced image feature to obtain a first high-frequency feature and a first low-frequency feature; the local frequency interaction model is a model constructed based on a convolutional neural network;
[0101] In the method provided by the embodiment of the present invention, a local frequency interaction model can be constructed in advance based on the structure of a convolutional neural network (CNN) for multi-mode interactive optimization of features in different frequency bands. The model parameters in the local frequency interaction model can be obtained by training based on sample data of camouflaged objects. In the high-low frequency cross-frequency interactive processing of the current second image feature, the local frequency interaction model can be applied to perform local feature interaction on the guided branch feature and the enhanced image feature. Specifically, the low-frequency features and high-frequency features of the two features can be extracted respectively, and the extracted low-frequency features and high-frequency features can be interactively optimized to obtain the high-frequency features (i.e., the first high-frequency features) and low-frequency features (i.e., the first low-frequency features) output after the local frequency interaction model is processed.
[0102] Applying a pre-trained global frequency interaction model to perform global feature interaction on the first high-frequency feature and the first low-frequency feature to obtain a first global optimization feature; the global frequency interaction model is a model built based on a Transformer network;
[0103] In the method provided by the embodiment of the present invention, a global frequency interaction model can be constructed in advance based on the network structure of the Transformer network to optimize the features at the pixel level. The model parameters in the global frequency interaction model can be obtained by training based on the sample data of the camouflaged object. The Transformer network is an existing network structure based on the self-attention mechanism, which is not described in detail here. In the high- and low-frequency cross-frequency interaction processing process of the current second image feature, the global frequency interaction model can be applied to perform global feature interaction on the first high-frequency feature and the first low-frequency feature output by the local frequency interaction model, and the feature output after processing by the global frequency interaction model is used as the first global optimization feature.
[0104] Performing feature fusion on the guiding branch feature, the first global optimization feature and the enhanced image feature, and using a result of the feature fusion as a first cross-frequency interaction result corresponding to the current second image feature;
[0105] In the method provided by the embodiment of the present invention, the guide branch feature, the first global optimization feature and the enhanced image feature can be fused by operations such as feature splicing, element addition, and convolution, and the fusion result of the three sets of features can be used as the first cross-frequency interaction result corresponding to the current second image feature. For example, the guide branch feature and the first global optimization feature can be element-wise added, and the element addition result can be subjected to operations such as convolution, regularization, and activation function processing, and the processed feature can be feature spliced with the enhanced image feature, and the feature splicing result can be further subjected to operations such as convolution, regularization, and activation function processing, and the feature obtained after processing can be used as the first cross-frequency interaction result.
[0106] In general, the local frequency interaction model and the global frequency interaction model are used to realize the cross-frequency interaction processing of enhanced image features and guided branch features.
[0107] Applying the local frequency interaction model, performing local feature interaction on the prediction graph branch feature and the enhanced image feature to obtain a second high-frequency feature and a second low-frequency feature;
[0108] In the method provided by the embodiment of the present invention, similar to the manner of cross-frequency interactive processing of enhanced image features and guided branch features, a local frequency interaction model and a global frequency interaction model are applied to realize cross-frequency interactive processing of enhanced image features and prediction graph branch features. Specifically, firstly, the local frequency interaction model is applied to extract low-frequency features and high-frequency features of the prediction graph branch features and enhanced image features, respectively, and the extracted low-frequency features and high-frequency features are interactively optimized to obtain high-frequency features (i.e., the second high-frequency features) and low-frequency features (i.e., the second low-frequency features) output by the local frequency interaction model after processing the prediction graph branch features and enhanced image features.
[0109] Applying the global frequency interaction model, performing global feature interaction on the second high-frequency feature and the second low-frequency feature, to obtain a second global optimization feature;
[0110] In the method provided by an embodiment of the present invention, a global frequency interaction model is applied to perform global feature interaction on the second high-frequency features and the second low-frequency features output by the local frequency interaction model. After the global frequency interaction model is processed on the second high-frequency features and the second low-frequency features, the output features are used as the second global optimization features.
[0111] The prediction graph branch feature, the second global optimization feature and the enhanced image feature are subjected to feature fusion, and a result of the feature fusion is used as a second cross-frequency interaction result corresponding to the current second image feature.
[0112] In the method provided by the embodiment of the present invention, the prediction graph branch features, the second global optimization features and the enhanced image features can be fused through operations such as feature concatenation, element addition, and convolution, and the fusion results of the three sets of features are used as the second cross-frequency interaction results corresponding to the current second image features. The processing method for fusing the prediction graph branch features, the second global optimization features and the enhanced image features is the same as the method for fusing the guide branch features, the first global optimization features and the enhanced image features in the previous text. Please refer to the previous examples and will not be repeated here.
[0113] Based on the method provided in the embodiment of the present invention, the features of different frequency bands can be optimized through multi-mode interaction by combining the CNN network and the Transformer network, so as to further mine the feature information, restore the feature details, and thus improve the accuracy of target detection.
[0114] On the basis of the method provided in the above embodiment, in the method provided in the embodiment of the present invention, the process of applying a pre-trained local frequency interaction model to perform local feature interaction on the guide branch feature and the enhanced image feature to obtain the first high-frequency feature and the first low-frequency feature includes:
[0115] Inputting the guiding branch features and the enhanced image features into the local frequency interaction model, so that the local frequency interaction model performs high- and low-frequency feature interaction processing on the guiding branch features and the enhanced image features;
[0116] In the method provided by an embodiment of the present invention, when a local frequency interaction model is applied to process the guided branch features and enhanced image features, the guided branch features and enhanced image features can be loaded into the local frequency interaction model to trigger the local frequency interaction model to perform high- and low-frequency feature interaction processing.
[0117] The process of the local frequency interaction model performing high- and low-frequency feature interaction processing on the guide branch feature and the enhanced image feature includes:
[0118] Based on the improved octave convolution, the high-frequency information in the guide branch feature is enhanced to obtain the guide branch high-frequency feature, and the low-frequency information in the enhanced image feature is extracted to obtain the enhanced branch high-frequency feature;
[0119] In the method provided by the embodiment of the present invention, based on the traditional CNN network, the convolution layer therein is improved, and the improved octave convolution is deployed to construct a local frequency interaction model. When the local frequency interaction model processes the input guide branch features and enhanced image features, the improved octave convolution can be used to further enhance the high-frequency features of the high-frequency information in the guide branch features, and the high-frequency features obtained after processing are used as the guide branch high-frequency features, and the improved octave convolution is used to extract the high-frequency features in the low-frequency information of the enhanced image features, and the extracted high-frequency features are used as the enhanced branch high-frequency features.
[0120] Adding the high-frequency feature of the guiding branch and the high-frequency feature of the enhancing branch, and using the result of the feature addition as the first high-frequency feature;
[0121] In the method provided by the embodiment of the present invention, the local frequency interaction model can add the high-frequency features of the guiding branch and the high-frequency features of the enhancing branch, and use the result of adding the high-frequency features of the guiding branch and the high-frequency features of the enhancing branch as the first high-frequency feature.
[0122] Based on the improved octave convolution, low-frequency features are extracted from the high-frequency information in the guide branch features to obtain the guide branch low-frequency features, and low-frequency features are enhanced on the low-frequency information in the enhanced image features to obtain the enhanced branch low-frequency features;
[0123] In the method provided by an embodiment of the present invention, the local frequency interaction model can extract low-frequency features in the high-frequency information of the guiding branch features based on the improved octave convolution, use the extracted low-frequency features as the guiding branch low-frequency features, and use the octave convolution to further enhance the low-frequency features of the low-frequency information in the enhanced image features, and use the low-frequency features obtained after processing as the enhanced branch low-frequency features.
[0124] The guiding branch low-frequency feature and the enhancing branch low-frequency feature are added, and a result of the feature addition is used as the first low-frequency feature.
[0125] In the method provided by the embodiment of the present invention, the local frequency interaction model adds the guiding branch low-frequency feature and the enhancing branch low-frequency feature, and uses the result of adding the guiding branch low-frequency feature and the enhancing branch low-frequency feature as the first low-frequency feature.
[0126] Based on the method provided in the embodiment of the present invention, the high-frequency features and low-frequency features in the guide branch features and enhanced image features can be extracted respectively through octave convolution to achieve local feature enhancement, which is conducive to mining feature details and further improving the accuracy of target detection.
[0127] Furthermore, in the method provided in an embodiment of the present invention, a local frequency interaction model is applied to perform local feature interaction on the prediction graph branch features and the enhanced image features to obtain the processing process of the second high-frequency features and the second low-frequency features. The principle is the same as the principle of applying the local frequency interaction model to process the guide branch features and the enhanced image features. Please refer to the description in the previous embodiment and will not be repeated here.
[0128] On the basis of the method provided in the above embodiment, in the method provided in the embodiment of the present invention, the process of applying a pre-trained global frequency interaction model to perform global feature interaction on the first high-frequency feature and the first low-frequency feature to obtain the first global optimization feature includes:
[0129] Inputting the first high-frequency feature and the first low-frequency feature into the global frequency interaction model, so that the global frequency interaction model performs global feature pixel-level optimization processing on the first high-frequency feature and the first low-frequency feature;
[0130] In the method provided by an embodiment of the present invention, when the global frequency interaction model is applied to process the first high-frequency feature and the first low-frequency feature, the first high-frequency feature and the first low-frequency feature can be loaded into the global frequency interaction model to trigger the global frequency interaction model to perform global feature pixel-level optimization processing.
[0131] The process of performing global feature pixel-level optimization processing on the first high-frequency feature and the first low-frequency feature by the global frequency interaction model includes:
[0132] Using a nested attention mechanism, extract global high-frequency information from the first high-frequency feature, and generate a high-frequency feature map corresponding to the first high-frequency feature based on the extracted high-frequency information;
[0133] In the method provided by the embodiment of the present invention, a nested attention mechanism is deployed on the basis of the traditional Transformer network to construct a global frequency interaction model. The global frequency interaction model can use a nested attention mechanism for the first high-frequency feature of the input to extract the global high-frequency information in the first high-frequency feature, and generate a corresponding high-frequency feature map based on the extracted high-frequency information, so as to optimize the global high-frequency information in the first high-frequency feature and obtain the optimized high-frequency feature. The nested attention mechanism can specifically adopt a cross-space attention mechanism.
[0134] The high-frequency feature map corresponding to the first high-frequency feature is concatenated with the first low-frequency feature, and a result of the feature concatenation is used as the first global optimization feature.
[0135] In the method provided by an embodiment of the present invention, the global frequency interaction model can splice the high-frequency feature map corresponding to the first high-frequency feature with the first low-frequency feature, and use the splicing result of the high-frequency feature map and the first low-frequency feature as the first global optimization feature.
[0136] Based on the method provided by the embodiment of the present invention, the global high-frequency information in the first high-frequency feature can be optimized through a nested attention mechanism, and the globally optimized high-frequency feature can be further fused with the low-frequency feature, which is conducive to further mining the detailed information of the high-frequency feature and improving the accuracy of target detection.
[0137] Furthermore, in the method provided in an embodiment of the present invention, a global frequency interaction model is applied to perform global feature interaction on the second high-frequency feature and the second low-frequency feature to obtain a processing process of the second global optimization feature. The principle is the same as the principle of applying the global frequency interaction model to process the first high-frequency feature and the first low-frequency feature. Please refer to the description in the previous embodiment and will not be repeated here.
[0138] On the basis of the method provided in the above embodiment, in the method provided in the embodiment of the present invention, the process of performing feature fusion on the first cross-frequency interaction result and the second cross-frequency interaction result includes:
[0139] Performing a product operation on the first cross-frequency interaction result and a pre-trained first scale parameter to obtain a first frequency enhancer feature;
[0140] Performing a product operation on the second cross-frequency interaction result and the pre-trained second scale parameter to obtain a second frequency enhancer feature;
[0141] In the method provided by the embodiment of the present invention, two scale parameters, namely, a first scale parameter and a second scale parameter, can be pre-trained by using sample data of the camouflaged object. The first scale parameter is a scale parameter associated with the cross-frequency interaction result of the guided branch feature and the enhanced image feature, and the second scale parameter is a scale parameter associated with the cross-frequency interaction result of the predicted image branch feature and the enhanced image feature.
[0142] In the method provided by the embodiment of the present invention, after obtaining the first cross-frequency interaction result and the second cross-frequency interaction result, the first cross-frequency interaction result can be multiplied by the trained first scale parameter, and the operation result is used as the first frequency enhancement sub-feature. The second cross-frequency interaction result can be multiplied by the trained second scale parameter, and the operation result is used as the second frequency enhancement sub-feature.
[0143] A feature addition process is performed on the first frequency enhancement sub-feature and the second frequency enhancement sub-feature, and a processing result is used as a feature fusion result of the first cross-frequency interaction result and the second cross-frequency interaction result.
[0144] In the method provided by the embodiment of the present invention, the first frequency enhancement sub-feature and the second frequency enhancement sub-feature can be fused by adding features, and the addition result of the first frequency enhancement sub-feature and the second frequency enhancement sub-feature is used as the feature fusion result of the two, that is, the frequency enhanced image feature corresponding to the current second image feature.
[0145] In order to better illustrate the method provided by the embodiment of the present invention, based on the methods provided by the previous embodiments and in combination with actual application scenarios, the embodiment of the present invention provides another method for detecting camouflaged objects. The method provided by the embodiment of the present invention can be specifically implemented by a target detection system, and the architecture of the system can be regarded as a cross-frequency interactive network for camouflaged object detection. Next, the construction process of the cross-frequency interactive network is described. The architecture of the cross-frequency interactive network provided by the embodiment of the present invention can be as follows: Figure 2 As shown in the figure, it mainly includes the backbone network (S1-S2-S3-S4) for feature extraction of the original image, the discussion module (DM) and the positioning guidance decoding structure (LGS) and other modules. The backbone network adopts the Res2Net-50 network, and the positioning guidance decoding structure (LGS) is provided with a cross-frequency interaction module (CFM). The cross-frequency interaction module contains two sub-modules, namely the CFM-C module and the CFM-T module. The CFM-C module is a module built based on the convolutional neural network CNN, and the CFM-T module is a module built based on the Transformer network. Figure 2 The symbol " " indicates the addition of elements, and the symbol " ” means element-wise multiplication.
[0146] Specifically, when the original image to be processed is loaded into the cross-frequency interaction network, the backbone network is first used to extract multi-level features from the image. For example, for an RGB image of size H×W×3, feature extraction is performed through the backbone network to obtain four-level features of sizes (H / 4)×(W / 4)×256, (H / 8)×(W / 8)×512, (H / 16)×(W / 16)×1024, and (H / 32)×(W / 32)×2048, respectively. Except for the features of the highest layer, the features of other layers are sent to three convolutional layers for channel reduction, and three-level features of sizes (H / 4)×(W / 4)×64, (H / 8)×(W / 8)×128, and (H / 16)×(W / 16)×256 are obtained, respectively. These three-level features are represented as F i (i=1,2,3), that is, in order from low to high levels, the three levels of features are F1, F2 and F3, and the highest level feature extracted by the backbone network is represented by F4.
[0147] The top-level feature F4 will be processed by the discussion module, which processes the features through three sub-modules, namely the densely connected atrous spatial pyramid pooling (DenseASPP) module, the spatial attention (SA) module, and the channel attention (CA) module. The discussion module first compresses and splits the top-level features in the channel dimension, and divides the top-level features into three branch features. These three branch features are represented by f1, f2, and f3. One branch feature is processed by one sub-module to enhance the positioning features from multi-scale perspectives, spatial perspectives, and channel perspectives. Specifically, f1 is processed by the densely connected atrous spatial pyramid pooling module, and the feature f is obtained after processing. a f2 is processed by the spatial attention module to obtain feature f b f3 is processed by the channel attention module, and the feature f is obtained after processing c For the features processed by each submodule, further discussion and interaction are carried out to obtain more accurate deep features L. It can be understood that f a 、f b and f c They are the various enhanced features in the previous embodiments, and L is the enhanced image feature in the previous embodiments.
[0148] Specifically, the discussion module obtains the feature f obtained by the above three submodules a 、f b and f c After that, we will further discuss these three features interactively. In the interactive discussion of each feature, the current feature is taken as the dominant feature, and the other two features are taken as supplementary features. The current feature guides the other two features to interactively optimize, and then the complementary optimization feature of the current feature is obtained. b Take interactive discussion as an example. The process of feature interactive discussion can be as follows: Figure 3 As shown, the symbol " " represents element addition. The symbol formed by three rectangles represents the CBR operation. The CBR operation is a combination of convolution operation, regularization operation and activation function ReLU processing. The symbol " " represents the Cat operation, i.e. the feature concatenation operation. The process can be described as follows:
[0149] f a '=CBR{CBR[Cat(f a ,f b )]+f a} (Formula 1)
[0150] f c '=CBR{CBR[Cat(f c,f b )]+f c} (Formula 2)
[0151] f a ''=CBR{CBR[Cat(f a ',f c ')]+f a '} (Formula 3)
[0152] f c ''=CBR{CBR[Cat(f a ',f c ')]+f c '} (Formula 4)
[0153] L b =CBR[Cat(f a '',f b ,f c '')] (Formula 5)
[0154] Similarly, we can get L a =CBR[Cat(f a ,f b '',f c '')],L c =CBR[Cat(f a '',f b '',f c )].
[0155] The feature L output by the discussion module (DM) after processing can be expressed as: L = CBR[Cat(L a ,L b ,L c )].
[0156] Among them, CBR represents the CBR operation, that is, the collection of convolution operation, regularization operation and activation function ReLU processing, and Cat represents the feature splicing operation.
[0157] Then, the localization guided decoding structure (LGS) is used to perform cross-frequency interaction optimization on each image feature (i.e., F1, F2, and F3) except the highest-level features. The cross-frequency interaction optimization process of each feature can be divided into two branches, namely, the guidance branch and the prediction graph branch. For example, Figure 2 In the structure shown, in the LGS1 structure that processes F1, the upper processing branch is the guidance branch, and the lower processing branch is the prediction graph branch. The same is true for other LGS structures. In the guidance branch, the features of the current layer and the upsampled features of the previous layer are added element-wise to obtain F. igIn the prediction graph branch, the prediction graph corresponding to the current layer feature and the previous layer feature is multiplied element-wise to obtain F ip Then, the cross-frequency interaction module CFM is used to ig and F ip Cross-frequency interaction optimization is performed to allow features of adjacent layers to fully interact. In the cross-frequency interaction process of the two branches, the highest-level features (L) after the discussion module optimization are used as guidance, and the features of the two branches of the current layer are interactively optimized from a local and global perspective using feature L. Taking a certain level of the guiding branch as an example, the brief schematic architecture of the cross-frequency interaction module CFM can be shown as follows: Figure 4 As shown, the module includes two modules, CFM-C and CFM-T. Figure 4 In the example, AP represents the average pooling operation, UP represents the upsampling operation, and the remaining characters are the same as Figure 3 The corresponding characters in have the same meaning, please refer to the previous description. The CFM-C module uses octave convolution to interact high- and low-frequency features locally to achieve local correction. The CFM-T module uses the Transformer network to optimize the pixel level from the global guidance features to achieve global association and finally achieve the progressive fusion of high-level features. Figure 4 The structure shown corresponds to F ig Taking L as input, the processing of the CFM-C module can be briefly described as:
[0158] F ig HF =CBR 3×3 [CBR 1×1 (F ig )] HF→HF +UP{CBR 3×3 {UP[CBR 1×1 (L)]}} LF→HF (Formula 6)
[0159] F ig LF =CBR 3×3 {Pool[CBR 1×1 (F ig )]} HF→LF +CBR 3×3 {UP[CBR 1×1 (L)]} LF→LF (Formula 7)
[0160] Among them, Pool represents the average pooling operation, UP represents the upsampling operation using the nearest neighbor interpolation, HF represents the high frequency, LF represents the low frequency, CBR represents the set of convolution operation, regularization operation and activation function ReLU processing, and the parameter in the CBR subscript represents the convolution size of the convolution operation. igHF is the high-frequency feature output by the CFM-C module, F ig LF It is the low-frequency feature output after processing by the CFM-C module. The high-frequency and low-frequency features output by the CFM-C module will be processed by the CFM-T module. First, the CFM-T module will globally optimize the high-frequency features based on the cross-space attention mechanism CSA (i.e., the nested attention mechanism mentioned in the previous embodiment), and then perform feature splicing on the optimized high-frequency features and the input low-frequency features, and use the feature splicing results as the features optimized by the CFM-T module.
[0161] The process of optimizing the high-frequency characteristics of the CFM-C module output by the CFM-T module can be briefly described as follows:
[0162] (Formula 8)
[0163] (Formula 9)
[0164] Among them, F ig HF ' indicates F ig HF The low-frequency information in F ig HF The optimized features of high-frequency information query in F ig HF '' is the CFM-T module to F ig HF The final feature obtained after optimization is F ig HF The high frequency information in F ig HF The low-frequency information in ' is used to query the optimized features, which contains the spatial dependency between high-frequency and low-frequency information. H represents height, W represents width, and N represents H×W. Q, K, and V are the features based on F in the processing process. ig HF The three new feature maps generated, K', V' are based on F ig HF 'Generate two new feature maps. Q i: represents the i-th row of matrix Q, Q i: ' represents the i-th row of the matrix Q'. K j: represents the j-th column of matrix K, V j: represents the j-th column of the matrix V, K j: ' represents the jth column of matrix K', V j: ' represents the jth column of the matrix V'. x is the spatial attention map generated based on Q' and K during the processing, x ijrepresents the jth high-frequency spatial feature in x (K j: ) for the i-th low-frequency spatial feature (Q i: '), and x' is the spatial attention map generated based on Q and K' during the processing, x' ij represents the jth low-frequency spatial feature in x' (K j: ') for the i-th high-frequency spatial feature (Q i: ). Pool is the average pooling operation, and exp represents the exponential function operation.
[0165] In general, after cross-frequency interaction optimization for each feature outside the highest level, the corresponding optimized feature F of each level can be obtained. i '(i=1,2,3), that is, the frequency enhanced image feature corresponding to each second image feature in the previous embodiment. The process of determining the optimized features corresponding to each level based on the features output by the cross-frequency interaction module in the two branches can be expressed as follows:
[0166] F i ' = α × CFM i α [F i +UP(L)]+β×CFM i β [F i ×UP(P i+1 )], i=(1,2,3) (Formula 10)
[0167] Among them, α and β are learnable scale parameters initialized to 1 (i.e., the first scale parameter and the second scale parameter in the previous embodiment), CFM i α represents the CFM module on the guiding branch in level i, CFM i β represents the CFM module on the prediction graph branch in level i, P i+1 It represents a prediction map for cross-frequency interaction at level i, which is a prediction map obtained based on feature processing in the previous level of level i.
[0168] In the process of constructing a cross-frequency interaction network, the sample image can be used as the original image to train each trainable parameter in the network, and the true value (that is, the actual image) corresponding to the original image can be used as supervision. Based on each predicted image and the true value obtained in the cross-frequency interaction network, gradient conduction is performed under the supervision of the loss function to optimize each parameter in the network, such as the trainable parameters of the corresponding network structure in the CFM module and the learnable scale parameters α and β. In an embodiment of the present invention, the loss function of the cross-frequency interaction network is calculated using weighted binary intersection loss (wBCE) and IoU loss (wIoU), which is conducive to calculating the difference between the center pixel and its surrounding environment, and paying more attention to hard pixels to enhance the generalization ability of the network.
[0169] Based on the positioning guidance decoding structure in the cross-frequency interaction network provided by the embodiment of the present invention, it is possible to effectively integrate feature information at different levels, use the deepest features (i.e., the highest level features) for global guidance, and gradually upsample to obtain the final segmentation result. The discussion module is used to obtain accurate deepest low-frequency features, which is conducive to using the deepest features to effectively guide features at other levels. The cross-frequency interaction module can extract high- and low-frequency information in the features and restore feature details through high- and low-frequency interactions.
[0170] and Figure 1 Corresponding to the method for detecting a disguised object shown in FIG. 1 , an embodiment of the present invention further provides a device for detecting a disguised object, Figure 1 The specific implementation of the method shown in is shown in the structural diagram Figure 5 As shown, including:
[0171] A feature extraction unit 201 is used to extract features of a target image to be detected and obtain image features at multiple levels;
[0172] A feature classification unit 202, configured to use the image feature with the highest level among the image features of each level as the first image feature, and use each of the remaining image features as the second image feature;
[0173] A feature enhancement unit 203 is used to perform feature enhancement on the first image feature through a preset densely connected void space pyramid pooling module, a spatial attention module and a channel attention module to obtain various enhanced features;
[0174] The feature supplementation unit 204 is used to take each enhanced feature as a dominant feature and each enhanced feature other than the enhanced feature as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complemented to obtain a complementary optimized feature corresponding to the enhanced feature;
[0175] A feature fusion unit 205 is used to fuse the complementary optimization features to obtain enhanced image features;
[0176] A feature interaction unit 206 is used to perform high-low frequency cross-frequency interaction on each of the second image features in sequence based on the enhanced image features in descending order of hierarchical levels, so as to obtain a frequency enhanced image feature corresponding to each of the second image features;
[0177] The result determination unit 207 is configured to determine the disguised object detection result of the target image based on the frequency enhanced image feature corresponding to the second image feature with the lowest level among the second image features.
[0178] By using the device provided by the embodiment of the present invention, the target image to be detected can be extracted to obtain multi-level image features. By enhancing the information of the highest-level image features with richer semantic information, the enhanced image features are used to guide the cross-frequency interactive optimization of image features at other levels, and more accurate detail information can be mined from the image features at each level. Target detection can be performed based on the image features finally optimized, thereby identifying camouflaged objects in the target image. Multi-dimensional optimization based on image features can mine more accurate details in image features, which is conducive to accurate segmentation of camouflaged objects and background environments in images. When the background environment is similar to the target object, the target object can also be identified in the background environment, which is conducive to improving the accuracy of camouflaged object detection.
[0179] exist Figure 5 Based on the device shown, the device provided by the embodiment of the present invention can be further expanded to include multiple units. The function of each unit can refer to the description of each embodiment provided for the detection method of camouflaged objects in the foregoing text, and no further examples will be given here.
[0180] An embodiment of the present invention further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned method for detecting a camouflaged object.
[0181] The embodiment of the present invention further provides an electronic device, the structural diagram of which is shown in FIG. Figure 6 As shown, it specifically includes a memory 301 and one or more instructions 302, wherein the one or more instructions 302 are stored in the memory 301 and are configured to be executed by one or more processors 303 to perform the following operations:
[0182] Perform feature extraction on the target image to be detected to obtain image features at multiple levels;
[0183] Among the image features of each of the said levels, the image feature with the highest level is used as the first image feature, and each of the remaining image features is used as the second image feature;
[0184] Performing feature enhancement on the first image feature through a preset densely connected atrous spatial pyramid pooling module, a spatial attention module, and a channel attention module to obtain various enhanced features;
[0185] For each of the enhancement features, the enhancement feature is used as a dominant feature, and each enhancement feature other than the enhancement feature is used as a supplementary feature, so that the enhancement feature and the remaining enhancement features are complemented with each other to obtain a complementary optimization feature corresponding to the enhancement feature;
[0186] Performing feature fusion on each of the complementary optimization features to obtain enhanced image features;
[0187] In descending order of the hierarchical levels, based on the enhanced image features, sequentially performing high- and low-frequency cross-frequency interaction on each of the second image features to obtain a frequency enhanced image feature corresponding to each of the second image features;
[0188] The disguised object detection result of the target image is determined based on the frequency enhanced image feature corresponding to the second image feature with the lowest hierarchical level among the second image features.
[0189] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0190] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0191] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting a camouflaged object, characterized in that: include: Perform feature extraction on the target image to be detected to obtain image features at multiple levels; Among the image features of each of the said levels, the image feature with the highest level is used as the first image feature, and each of the remaining image features is used as the second image feature; The first image feature is enhanced by respectively using a preset densely connected atrous spatial pyramid pooling module, a spatial attention module, and a channel attention module to obtain various enhanced features; For each of the enhanced features, the enhanced feature is used as a dominant feature, and each enhanced feature other than the enhanced feature is used as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complemented with each other to obtain a complementary optimization feature corresponding to the enhanced feature; Performing feature fusion on each of the complementary optimization features to obtain enhanced image features; In descending order of the hierarchical levels, based on the enhanced image features, sequentially performing high- and low-frequency cross-frequency interaction on each of the second image features to obtain a frequency enhanced image feature corresponding to each of the second image features; The disguised object detection result of the target image is determined based on the frequency enhanced image feature corresponding to the second image feature with the lowest hierarchical level among the second image features.
2. The method for detecting a disguised object according to claim 1, characterized in that: The method of performing feature enhancement on the first image feature by respectively using a preset densely connected void space pyramid pooling module, a spatial attention module, and a channel attention module to obtain various enhanced features includes: Perform channel splitting on the first image feature to obtain a first branch feature, a second branch feature, and a third branch feature; Applying the densely connected atrous spatial pyramid pooling module, performing channel splitting on the first branch feature to obtain a plurality of pyramid sub-features, and performing multi-scale optimization on each of the pyramid sub-features through a densely connected splicing operation and a residual connection operation to obtain a multi-scale optimized feature; Applying the spatial attention module, performing dimensionality reduction processing on the second branch feature to obtain a dimensionality reduction feature map, performing a pooling convolution operation on the dimensionality reduction feature map to obtain a spatial attention map, and performing element-by-element multiplication processing on the spatial attention map and the second branch feature to obtain a spatial optimization feature; Applying the channel attention module to generate a channel attention map corresponding to the third branch feature, and using the channel attention map to dynamically adjust the weight of the channel in the third branch feature to obtain a channel optimization feature; The multi-scale optimization feature, the space optimization feature and the channel optimization feature are used as the respective enhanced features.
3. The method for detecting a disguised object according to claim 1, characterized in that: The enhancement feature is used as the dominant feature, and each enhancement feature other than the enhancement feature is used as a supplementary feature, so that the enhancement feature and the other enhancement features complement each other to obtain the complementary optimization feature corresponding to the enhancement feature, including: Based on the enhanced feature, each enhanced feature except the enhanced feature is optimized to obtain a first optimized feature and a second optimized feature; Based on the first optimized feature, performing feature optimization on the second optimized feature to obtain a third optimized feature; Based on the second optimized feature, performing feature optimization on the first optimized feature to obtain a fourth optimized feature; The enhanced feature, the third optimized feature and the fourth optimized feature are subjected to feature fusion, and a result of the feature fusion is used as a complementary optimized feature corresponding to the enhanced feature.
4. The method for detecting a disguised object according to claim 1, characterized in that: The process of performing high-low frequency cross-frequency interaction on each of the second image features includes: Based on the current upper-level image feature, determining the target image feature and the target prediction map corresponding to the current second image feature; the current upper-level image feature is an image feature whose hierarchical level is one level higher than the current second image feature among the image features of each of the hierarchical levels; Fusing the current second image feature with the target image feature to obtain a guiding branch feature; Performing feature fusion on the current second image feature and the target prediction map to obtain prediction map branch features; Performing cross-frequency interaction on the enhanced image feature, the guidance branch feature, and the prediction graph branch feature, respectively, to obtain a first cross-frequency interaction result and a second cross-frequency interaction result corresponding to the current second image feature; Feature fusion is performed on the first cross-frequency interaction result and the second cross-frequency interaction result, and a result of the feature fusion is used as a frequency enhanced image feature corresponding to the current second image feature.
5. The method for detecting a disguised object according to claim 4, characterized in that: If the current second image feature is the second image feature with the highest hierarchical level among the second image features, determining the target image feature and the target prediction map corresponding to the current second image feature based on the current upper-level image feature includes: Using the enhanced image feature as the target image feature corresponding to the current second image feature; Performing a convolution operation on the enhanced image feature to obtain a prediction map corresponding to the enhanced image feature, and using the prediction map corresponding to the enhanced image feature as a target prediction map corresponding to the current second image feature; If the current second image feature is not the second image feature with the highest level in the hierarchy, determining the target image feature and the target prediction map corresponding to the current second image feature based on the current upper-level image feature includes: Determining a target image feature corresponding to the current second image feature based on a second cross-frequency interaction result corresponding to the current upper-level image feature; A convolution operation is performed on the frequency enhanced image feature corresponding to the current upper-level image feature to obtain an enhanced prediction map corresponding to the current upper-level image feature, and the enhanced prediction map is used as a target prediction map corresponding to the current second image feature.
6. The method for detecting a disguised object according to claim 4, characterized in that: The step of performing cross-frequency interaction with the enhanced image feature, the guidance branch feature, and the prediction graph branch feature to obtain a first cross-frequency interaction result and a second cross-frequency interaction result corresponding to the current second image feature includes: Applying a pre-trained local frequency interaction model to perform local feature interaction on the guide branch feature and the enhanced image feature to obtain a first high-frequency feature and a first low-frequency feature; the local frequency interaction model is a model constructed based on a convolutional neural network; Applying a pre-trained global frequency interaction model to perform global feature interaction on the first high-frequency feature and the first low-frequency feature to obtain a first global optimization feature; the global frequency interaction model is a model built based on a Transformer network; Performing feature fusion on the guiding branch feature, the first global optimization feature and the enhanced image feature, and using a result of the feature fusion as a first cross-frequency interaction result corresponding to the current second image feature; Applying the local frequency interaction model, performing local feature interaction on the prediction graph branch feature and the enhanced image feature to obtain a second high-frequency feature and a second low-frequency feature; Applying the global frequency interaction model, performing global feature interaction on the second high-frequency feature and the second low-frequency feature, to obtain a second global optimization feature; The prediction graph branch feature, the second global optimization feature and the enhanced image feature are subjected to feature fusion, and a result of the feature fusion is used as a second cross-frequency interaction result corresponding to the current second image feature.
7. The method for detecting a disguised object according to claim 6, characterized in that: The method of applying a pre-trained local frequency interaction model to perform local feature interaction on the guide branch feature and the enhanced image feature to obtain a first high-frequency feature and a first low-frequency feature includes: Inputting the guiding branch features and the enhanced image features into the local frequency interaction model, so that the local frequency interaction model performs high- and low-frequency feature interaction processing on the guiding branch features and the enhanced image features; The process of the local frequency interaction model performing high- and low-frequency feature interaction processing on the guide branch feature and the enhanced image feature includes: Based on the improved octave convolution, the high-frequency information in the guide branch feature is enhanced to obtain the guide branch high-frequency feature, and the low-frequency information in the enhanced image feature is extracted to obtain the enhanced branch high-frequency feature; Adding the high-frequency feature of the guiding branch and the high-frequency feature of the enhancing branch, and using the result of the feature addition as the first high-frequency feature; Based on the improved octave convolution, low-frequency features are extracted from the high-frequency information in the guide branch features to obtain the guide branch low-frequency features, and low-frequency features are enhanced on the low-frequency information in the enhanced image features to obtain the enhanced branch low-frequency features; The guiding branch low-frequency feature and the enhancing branch low-frequency feature are added, and a result of the feature addition is used as the first low-frequency feature.
8. The method for detecting a disguised object according to claim 6, characterized in that: The applying a pre-trained global frequency interaction model to perform global feature interaction on the first high-frequency feature and the first low-frequency feature to obtain a first global optimization feature includes: Inputting the first high-frequency feature and the first low-frequency feature into the global frequency interaction model, so that the global frequency interaction model performs global feature pixel-level optimization processing on the first high-frequency feature and the first low-frequency feature; The process of performing global feature pixel-level optimization processing on the first high-frequency feature and the first low-frequency feature by the global frequency interaction model includes: Using a nested attention mechanism, extract global high-frequency information from the first high-frequency feature, and generate a high-frequency feature map corresponding to the first high-frequency feature based on the extracted high-frequency information; The high-frequency feature map corresponding to the first high-frequency feature is concatenated with the first low-frequency feature, and a result of the concatenated features is used as the first global optimization feature.
9. The method for detecting a disguised object according to claim 4, characterized in that: The performing feature fusion on the first cross-frequency interaction result and the second cross-frequency interaction result includes: Performing a product operation on the first cross-frequency interaction result and a pre-trained first scale parameter to obtain a first frequency enhancer feature; Performing a product operation on the second cross-frequency interaction result and the pre-trained second scale parameter to obtain a second frequency enhancer feature; A feature addition process is performed on the first frequency enhancement sub-feature and the second frequency enhancement sub-feature, and a processing result is used as a feature fusion result of the first cross-frequency interaction result and the second cross-frequency interaction result.
10. A detection device for a disguised object, characterized in that: include: A feature extraction unit, used to extract features of the target image to be detected and obtain image features at multiple levels; A feature classification unit, configured to use, among the image features of each of the said levels, the image feature with the highest level as the first image feature, and use each of the remaining image features as the second image feature; A feature enhancement unit, used to perform feature enhancement on the first image feature through a preset densely connected void space pyramid pooling module, a spatial attention module, and a channel attention module, respectively, to obtain various enhanced features; A feature supplementation unit is used to, for each of the enhanced features, use the enhanced feature as a dominant feature and each enhanced feature other than the enhanced feature as a supplementary feature, so that the enhanced feature and the remaining enhanced features are complemented with each other to obtain a complementary optimization feature corresponding to the enhanced feature; A feature fusion unit, used for fusing the complementary optimization features to obtain enhanced image features; a feature interaction unit, configured to perform high- and low-frequency cross-frequency interaction on each of the second image features in sequence based on the enhanced image features in descending order of hierarchical levels, so as to obtain a frequency enhanced image feature corresponding to each of the second image features; The result determination unit is configured to determine a disguised object detection result of the target image based on a frequency enhanced image feature corresponding to a second image feature with the lowest hierarchical level among the second image features.
Citation Information
Patent Citations
Camouflage target image segmentation method and system based on multilevel feature fusion
CN116703950A
Camouflage target detection method based on Gaussian attention
CN117541815A
Frequency decomposition and attention guidance-based camouflage target detection method
CN118397292A
Camouflage target detection method for hourglass vision Transform double-channel coding and decoding network
CN118747842A
Camouflage target detection method, system and device and medium
CN119723044A
Cited By
Camouflage target identification method and system
CN120451517A