Camouflage target detection method and device, terminal and computer storage medium

By building a feature pyramid and performing multiple feature extraction and enhancement processing, the problem of poor detection of camouflage targets in the prior art is solved, and the accurate identification and positioning of camouflage targets is achieved.

CN120259618APending Publication Date: 2025-07-04SHANGHAI ADVANCED RES INST CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317707.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing object detection methods deal with complex scenarios, especially when camouflage targets are similar to backgrounds, it is difficult to fully capture multi-level and multi-scale feature information, resulting in poor detection results.

Method used

By building a feature pyramid, feature extraction is performed with preset times, combined with hierarchical feature enhancement and global context-aware processing, the features are gradually optimized to achieve more accurate object detection.

Benefits of technology

Accurate identification and positioning of camouflage targets is achieved, and the effect of target detection is improved, especially in complex backgrounds such as camouflage camouflage personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259618A_ABST
    Figure CN120259618A_ABST
Patent Text Reader

Abstract

The invention provides a camouflage target detection method and device, a terminal and a computer storage medium, and the method comprises the steps: carrying out the construction of an image pyramid based on a to-be-detected image, and obtaining a feature pyramid which comprises a preset number of layers of feature maps; on the basis of each feature map, executing feature extraction for a preset number of times, and obtaining a highest-layer prediction map and a bottommost-layer prediction map; and adding and fusing the highest-layer prediction map and the lowest-layer prediction map to obtain a final prediction map, and carrying out identification detection based on the final prediction map to obtain a detection result. According to the method, the features are gradually optimized and enhanced through the preset number of times of feature extraction and multiple iterations, so that detail information and semantic information of the detection target are captured more accurately, accurate recognition and positioning of the camouflage target are achieved, and a good target detection effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision and relates to an object detection technology, in particular to a method, device, terminal, and computer storage medium for detecting camouflaged objects. Background Art

[0002] Object detection is a computer vision task for accurately identifying and locating objects in images, and has extensive applications in fields such as military, medicine, ecology, industry, and security.

[0003] Existing object detection methods usually extract features from images to capture the features of the detection object and achieve the recognition and location of the object. However, existing object detection methods usually perform one-time feature extraction, and the processing effect for complex scenes is not good. Especially when the detection object is similar to the background, for example, when identifying and locating camouflaged personnel in military reconnaissance, due to the highly similar color and texture features between the camouflaged personnel and the natural background, the correlation relationship between features is relatively complex, and one-time feature extraction often fails to comprehensively capture the multi-level and multi-scale feature information of the detection object, resulting in poor object detection results.

[0004] Therefore, how to accurately identify and detect camouflaged objects is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, terminal, and computer storage medium for detecting camouflaged objects, which are used to solve the problem that existing object detection methods usually adopt one-time feature extraction, are difficult to comprehensively capture the multi-level and multi-scale feature information of the detection object, and have poor effects when detecting camouflaged objects.

[0006] In a first aspect, this application provides a method for detecting camouflaged objects, including: constructing an image pyramid based on the image to be detected to obtain a feature pyramid, where the feature pyramid includes feature maps with a preset number of layers; performing feature extraction a preset number of times based on each of the feature maps to obtain a top-layer prediction map and a bottom-layer prediction map; adding and fusing the top-layer prediction map and the bottom-layer prediction map to obtain a final prediction map, and performing recognition and detection based on the final prediction map to obtain a detection result.

[0007] In an embodiment of the present application, the feature map corresponding to the bottom layer of the feature pyramid is used as the bottom layer feature map; a single feature extraction process includes: obtaining a current optimization item; based on the current optimization item, updating the bottom layer feature map to obtain a current feature pyramid; performing hierarchical feature enhancement processing on each feature layer in the current feature pyramid to obtain corresponding enhanced feature maps for each layer at the current time; performing global context awareness processing based on the current feature pyramid to obtain a current global feature map; performing channel fusion on the current global feature map and the enhanced feature map corresponding to the top layer feature map to obtain a top layer optimized feature map; based on the top layer optimized feature map, adding and fusing the enhanced feature maps in sequence from high to low to obtain a bottom layer optimized feature map; determining whether the current extraction times is the preset number of times, if so, performing convolution on the top layer optimized feature map to obtain a top layer prediction map, and performing convolution on the bottom layer optimized feature map to obtain a bottom layer prediction map, otherwise, using the bottom layer optimized feature map as a new optimization item and continuing to execute the next feature extraction.

[0008] In an embodiment of the present application, the hierarchical feature enhancement processing includes: based on each feature map, obtaining all adjacent layers and performing adjacent layer feature fusion processing to obtain corresponding fused feature maps; based on each fused feature map, respectively performing gated convolution enhancement processing to obtain corresponding enhanced feature maps.

[0009] In an embodiment of the present application, the performing global context awareness processing based on the current feature pyramid to obtain a current global feature map includes:

[0010] Based on the current feature pyramid, performing channel fusion on the feature maps of each level to obtain an intermediate global feature map;

[0011] Based on the intermediate global feature map, obtaining a global feature map through a self-attention mechanism.

[0012] In one embodiment of the present application, for each of the fused feature maps, the gated convolution enhancement process includes: based on the fused feature map, through convolutional attention mechanism processing, obtaining a corresponding weighted feature map; adding and fusing the weighted feature map and the corresponding fused feature map, and performing layer normalization processing to obtain an intermediate feature map; performing channel expansion and channel splitting processing on the intermediate feature map to obtain a first expanded feature map and a second expanded feature map; performing depthwise separable convolution operations on the first expanded feature map to obtain a first convolutional feature map, and performing depthwise separable convolution operations on the second expanded feature map to obtain a second convolutional feature map; based on the first convolutional feature map, combining with a non-linear activation function to obtain a gated weight matrix, and multiplying the gated weight matrix element-wise with the second convolutional feature map to obtain a gated feature map; performing channel alignment based on the gated feature map, and adding and fusing it with the fused feature map to obtain an enhanced feature.

[0013] In one embodiment of the present application, the obtaining all adjacent layers based on each of the feature maps and performing adjacent layer feature fusion processing to obtain corresponding fused feature maps includes: if the feature map corresponds to the bottom layer or the top layer of the feature pyramid, obtaining its corresponding adjacent layer feature map, and performing dual-scale feature fusion on the feature map and the adjacent layer feature map to obtain the corresponding fused feature map; otherwise, obtaining all its corresponding adjacent layer feature maps, and performing triple-scale feature fusion on the feature map and all the adjacent layer feature maps to obtain the corresponding fused feature map.

[0014] In one embodiment of the present application, the obtaining a corresponding weighted feature map by convolutional attention mechanism processing based on the fused feature map includes: performing layer normalization processing on the fused feature map, and performing convolution and depthwise separable convolution to obtain a group of attention matrices; splitting based on the attention matrix to obtain a query matrix, a reference matrix, and an information matrix; performing dot product calculation on the query matrix and the reference matrix, and performing normalization processing to obtain an attention weight matrix; performing dot product calculation on the attention weight matrix and the information matrix to obtain the weighted feature map.

[0015] In a second aspect, the present application provides a camouflaged target detection device, including a pyramid construction module, a feature extraction module, and a detection and recognition module; the pyramid construction module is configured to perform image pyramid construction on a to-be-detected image to obtain a feature pyramid, and the feature pyramid includes feature maps with a preset number of layers; the feature extraction module is configured to perform feature extraction a preset number of times based on each of the feature maps to obtain a top layer prediction map and a bottom layer prediction map; the detection and recognition module is configured to add and fuse the top layer prediction map and the bottom layer prediction map to obtain a final prediction map, and perform recognition and detection based on the final prediction map to obtain a detection result.

[0016] In a third aspect, the present application provides a terminal, including: a processor and a memory, where the memory is communicatively connected to the processor;

[0017] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes the camouflaged target detection method as described above.

[0018] In a fourth aspect, the present application provides a computer storage medium, where the computer storage medium stores a computer program, and when the computer program is executed by a processor, the camouflaged target detection method as described above is implemented.

[0019] As described above, the present application provides a camouflaged target detection method, device, terminal, and computer storage medium. By performing feature extraction multiple times to gradually optimize and enhance the features, the interference caused by the high similarity between the background and the detection target is eliminated, and the accurate identification and positioning of the camouflaged target are achieved, resulting in a good target detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It shows a schematic flowchart of a camouflaged target detection method according to an embodiment of the present application.

[0021] Figure 2 It shows a schematic framework diagram of a camouflaged target detection method according to an embodiment of the present application.

[0022] Figure 3 It shows a schematic flowchart of a single feature extraction process according to an embodiment of the present application.

[0023] Figure 4 It shows a schematic flowchart of a hierarchical feature enhancement process according to an embodiment of the present application.

[0024] Figure 5 It shows a schematic diagram of the fusion situation of each layer of a feature pyramid according to an embodiment of the present application.

[0025] Figure 6 It shows a schematic flowchart of a process of fusing features of adjacent layers according to an embodiment of the present application.

[0026] Figure 7 It shows a schematic framework diagram of a process of fusing features of adjacent layers according to an embodiment of the present application.

[0027] Figure 8 It shows a schematic flowchart of a method for obtaining a global feature map according to an embodiment of the present application.

[0028] Figure 9It shows a schematic flowchart of a gating convolution enhancement processing procedure described in an embodiment of the present application.

[0029] Figure 10 It shows a schematic flowchart of a method for implementing a convolutional attention mechanism described in an embodiment of the present application.

[0030] Figure 11 It shows a schematic structural diagram of a camouflage target detection device described in an embodiment of the present application.

[0031] Figure 12 It shows a terminal described in an embodiment of the present application.

[0032] Explanation of reference numerals

[0033] 41: Pyramid construction module; 42: Feature extraction module; 43: Detection and recognition module; 50: Terminal; 51: Processor; 52: Memory; 521: Operating system; 522: Application program; 53: User interface; 54: Network interface; 55: Bus system. Detailed implementation manners

[0034] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0035] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0036] Existing target detection methods usually process the image to be detected by means of one-time feature extraction to achieve the recognition and positioning of the target. However, in complex scenarios such as military reconnaissance, for example, the recognition and positioning of camouflaged personnel, the correlation between the detected target features is relatively complex, and one-time feature extraction is difficult to fully capture effective features, thus resulting in poor detection effects of these target detection methods.

[0037] In view of the technical problems existing in the prior art, the following embodiments of the present application provide a method, apparatus, terminal and computer storage medium for detecting camouflaged targets. By performing feature extraction a preset number of times and iteratively optimizing and enhancing the features step by step, it is convenient to more accurately capture the detailed information and semantic information of the detection target, so as to realize the accurate identification and positioning of the camouflaged target, achieving a good target detection effect.

[0038] The following embodiments of the present application provide a method, apparatus, terminal and computer storage medium for detecting camouflaged targets, including but not limited to being applied to target detection scenarios where the detection target is highly similar to the background. The following will take the identification and detection of camouflaged personnel in military reconnaissance as an example for description.

[0039] It should be noted that the method, apparatus, terminal and computer storage medium for detecting camouflaged targets provided in the following embodiments of the present application can also be applied to other target detection scenarios, including but not limited to identifying and tracking animals with camouflage ability in the natural environment, identifying pests or weeds in farmland that are close in color and shape to crops, or detecting hidden weapons or dangerous items in a complex background. The present application does not make specific limitations here.

[0040] Next, the technical solutions in the embodiments of the present application will be described in detail with reference to the accompanying drawings in the embodiments of the present application.

[0041] As Figure 1 and Figure 2 shown, the present embodiment provides a method for detecting camouflaged targets, including:

[0042] S100, constructing an image pyramid based on the image to be detected to obtain a feature pyramid, where the feature pyramid includes feature maps of a preset number of layers.

[0043] Exemplarily, an image pyramid is constructed through a preset pyramid feature extraction network, such as pyramid feature extraction networks such as FPN, PANet, BiFPN, ASPP, PSPNet, HRNet, TridentNet, and NAS-FPN.

[0044] It should be noted that the feature pyramid constructed through the pyramid feature extraction network has a preset number of layers. As Figure 2 shown, for example, 4 layers, and each layer corresponds to one of the feature maps. These feature maps have different resolutions and semantic information. Among them, the feature map corresponding to the highest layer of the feature pyramid, that is, Figure 2 F4 in Figure 2 has the lowest resolution, and the feature map corresponding to the bottom layer, that is,

[0045] Further, in this embodiment, after obtaining the feature pyramid, channel alignment is also performed on the feature maps of each layer of the feature pyramid so that each of the feature maps has the same number of channels, facilitating subsequent feature fusion. Specifically, among the feature maps of each layer of the feature pyramid, the feature map corresponding to the highest layer has the fewest channels. Based on this, the number of channels of each of the feature maps is reduced to be equal to the number of channels of the feature map corresponding to the highest layer. Exemplarily, the number of channels of each of the feature maps is reduced to 64 through a convolution operation.

[0046] S200. Based on each of the feature maps, perform feature extraction a preset number of times to obtain the highest-layer prediction map and the lowest-layer prediction map.

[0047] Among them, the feature extraction refers to the process of extracting key information from the image to be inspected that can effectively represent the detection target. Through the feature extraction, the features of the detection target can be captured, thereby realizing the recognition and positioning of the target to be inspected.

[0048] Specifically, the feature extraction process includes subjecting each of the feature maps to hierarchical feature enhancement processing to obtain corresponding enhanced feature maps for enhanced fusion of the features corresponding to the detection target, and subjecting each of the feature maps to global context awareness processing to obtain global feature maps. Combining each of the enhanced feature maps with the global feature maps further optimizes the effect of feature fusion, facilitating subsequent object detection.

[0049] When the detection target is highly similar to the background, for example, in military reconnaissance for the recognition and positioning of camouflaged personnel, single-time feature extraction often cannot accurately capture features. Based on this, in this embodiment, feature extraction is performed a preset number of times to gradually enhance and optimize the features, thereby realizing accurate and effective feature capture.

[0050] Specifically, for the feature extraction process of the preset number of times, an optimization item is obtained based on the previous feature extraction process, and the current feature extraction process is performed based on the optimization item, thereby realizing multiple iterations to gradually optimize and enhance the features, and further facilitating more accurate capture of the features of the detection target to achieve a better detection effect.

[0051] Among them, the optimization item is used to characterize the feature optimization situation after feature fusion and feature enhancement through the feature extraction process. Through this optimization item, the feature optimization effect after the previous feature extraction can be superimposed on the subsequent feature extraction process to gradually optimize and enhance the features.

[0052] Exemplarily, the preset number of times is 3 times.

[0053] In some alternative embodiments, as Figure 3 shown, a single feature extraction process includes:

[0054] S210. Obtain the current optimization item; based on the current optimization item, update the bottommost feature map to obtain the current feature pyramid.

[0055] Wherein, the bottommost feature map is the feature map corresponding to the bottommost layer of the feature pyramid.

[0056] Specifically, as Figure 2 shown, the plus sign in the figure is used to represent addition fusion. Additively fuse the optimization item and the bottommost optimized feature map to update the bottommost feature map.

[0057] Since the bottommost feature map is the feature map with the highest resolution in the feature pyramid and has richer detail information, by additively fusing the bottommost feature map and the optimization item, higher-quality image details can be generated, thereby further improving the optimization effect of the features and facilitating better capture of the detected target features.

[0058] It should be noted that when the current feature extraction process is the first feature extraction, the corresponding optimization item is a preset value, for example, the optimization item is 0; otherwise, the optimization item is the optimization item obtained based on the previous feature extraction process.

[0059] It should be noted that those skilled in the art should be aware of the specific implementation methods and principles of the additive fusion between the bottommost feature map and the optimization item, and this embodiment does not specifically describe them here.

[0060] S220. Perform hierarchical feature enhancement processing on each feature layer in the current feature pyramid to obtain the corresponding enhanced feature map for each layer.

[0061] Wherein, the hierarchical feature enhancement processing includes: adjacent layer feature fusion for fusing features of each feature map to obtain each fused feature map to facilitate capture of features at different scales; and gated convolution enhancement processing for enhancing features of each fused feature map to obtain each enhanced feature map, thereby realizing enhanced fusion of features to facilitate capture of the detected target features.

[0062] In some optional embodiments, as Figure 4 shown, the hierarchical feature enhancement processing includes:

[0063] S221. Based on each feature map, perform adjacent layer feature fusion to obtain the corresponding fused feature maps.

[0064] Since the characteristic scales corresponding to the camouflaged personnel in the image to be inspected vary greatly, it is usually necessary to capture features at different scales for the recognition and positioning of the camouflaged personnel. Specifically, the feature maps corresponding to the upper layers of the feature pyramid usually have relatively rich semantic information, which is beneficial to capturing large-scale features; while the feature maps corresponding to the lower layers usually have rich detailed information, which is beneficial to capturing small-scale features. By fusing the feature maps with those of other levels to generate the corresponding fused feature maps, the fused feature maps can integrate the advantages of feature maps at different levels, which is beneficial to capturing features at different scales, thereby improving the detection effect of the camouflaged targets.

[0065] It should be noted that since the semantic information between the layers of the feature pyramid decreases from top to bottom, while the detailed information increases from top to bottom, there are usually large semantic differences and detailed differences between the feature maps of different levels. Fusion may lead to unstable information propagation, and thus the obtained fused feature maps cannot effectively express the key information of the target.

[0066] Based on this, in this embodiment, each of the fused feature maps is obtained through adjacent-layer feature fusion, that is, feature fusion is performed between the feature maps of adjacent layers. This is because the feature maps of adjacent layers have good continuity in semantic information and detailed information, and the information propagation is more stable, that is, each of the obtained fused feature maps can better express the key information of the target.

[0067] Exemplarily, as Figure 5 shown, since there is only one adjacent layer for the feature map corresponding to the bottom layer or the top layer of the feature pyramid, while the feature maps corresponding to other levels each have two adjacent layers (upper and lower), therefore, for the feature map corresponding to the bottom layer or the top layer of the feature pyramid, dual-scale feature fusion is performed, as Figure 5 shown in the first layer and the fourth layer in Figure 5 and triple-scale feature fusion is performed on the feature maps corresponding to other levels to obtain the corresponding fused feature maps for each level, as

[0068] shown in the second layer and the third layer in Figure 6 and Figure 7 shown. In some optional implementation manners, as

[0069] S2211, if the feature map corresponds to the bottom layer or the top layer of the feature pyramid, obtain the corresponding adjacent-layer feature map, and perform dual-scale feature fusion on the feature map and the adjacent-layer feature map to obtain the corresponding fused feature map.

[0070] Specifically, as Figure 7As shown on the left, the convolution in the figure represents the convolution operation, the symbol C represents the channel fusion operation, and the symbol CBR (Convolution - Batch Normalization - ReLU) represents performing convolution, batch normalization, and activation operations in sequence. Specifically, the two feature maps to be fused are convolved to extract features. Exemplarily, the convolution kernel is 3*3. For each of the convolved feature maps, channel fusion is performed, that is, the channels of each feature map are concatenated, and then convolution, batch normalization, and ReLU linear correction are performed on the concatenated image in sequence to achieve feature fusion of the feature maps of two adjacent layers.

[0071] S2212, otherwise, obtain all its corresponding adjacent - layer feature maps, perform three - scale feature fusion on the feature map and all its adjacent - layer feature maps, and obtain the corresponding fused feature map.

[0072] Similarly, as Figure 7 As shown on the right, the convolution in the figure represents the convolution operation, the symbol C represents the channel fusion operation, and the symbol CBR (Convolution - Batch Normalization - ReLU) represents performing convolution, batch normalization, and activation operations in sequence. Specifically, after convolving the three feature maps to be fused, channel fusion is performed, and then convolution, batch normalization, and ReLU linear correction are performed in sequence to achieve feature fusion of the feature maps of three adjacent layers. Among them, the specific implementation method and principle of this three - scale feature fusion please refer to the above - mentioned two - scale feature fusion, and this embodiment will not elaborate here.

[0073] S222, based on each of the fused feature maps, perform gated convolution enhancement processing respectively to obtain the corresponding enhanced feature maps.

[0074] Specifically, by enhancing the features of each of the fused feature maps, the detection target and the features corresponding to the camouflaged personnel are highlighted to facilitate capturing these features.

[0075] It should be noted that during the process of identifying and detecting camouflaged personnel, since the camouflaged personnel in the image to be detected may be at different distances, angles, and lighting conditions, the feature scale and form of the detection target vary greatly. It is often difficult to effectively highlight the target features by using a fixed feature enhancement strategy.

[0076] Based on this, in this embodiment, the gated convolution enhancement process is used to enhance the features of each of the fused feature maps to achieve adaptive adjustment of feature enhancement. Specifically, the gated convolution enhancement process includes: based on the fused feature maps, capturing the correlation relationships between the features corresponding to the detection targets through a convolutional attention mechanism to calculate attention weights, and enhancing the corresponding features through these attention weights, thereby realizing dynamic enhancement of the features; and introducing a non-linear activation function through a gating mechanism for element-wise multiplication to further adaptively adjust the features, so as to achieve accurate recognition and positioning of detection targets with large variations in scale and shape.

[0077] Among them, the convolutional attention mechanism refers to an attention mechanism combined with convolutional operations. By performing convolutional operations on the fused feature maps, an attention weight map is obtained, and dynamic feature enhancement is performed based on this attention weight; the gating mechanism refers to a mechanism that obtains a gating weight matrix through non-linear transformation and adaptively adjusts the features through element-wise multiplication.

[0078] By performing the adjacent layer feature fusion process and the gated convolution enhancement process, the features of each of the feature maps are fused and enhanced, so as to improve the accuracy of feature capture.

[0079] S230, perform global context awareness processing based on the current feature pyramid to obtain the current global feature map.

[0080] Among them, the global feature map is used to represent the global semantic information of the image to be inspected.

[0081] In some alternative embodiments, as Figure 8 shown, the method for obtaining the global feature map includes:

[0082] S231, perform channel fusion on each of the feature maps to obtain an intermediate global feature map.

[0083] Specifically, through channel fusion, the integration and superposition of semantic information and detail information in each of the feature maps are realized, so that the intermediate global feature map can contain multi-scale semantic information and multi-scale detail information from the global perspective.

[0084] It should be noted that before performing channel fusion on each of the feature maps, resolution alignment is also performed on each of the feature maps so that each of the feature maps has the same resolution, which is convenient for performing channel fusion.

[0085] S232, based on the intermediate global feature map, obtain the global feature map through a self-attention mechanism.

[0086] Specifically, the self-attention mechanism is to dynamically adjust features by calculating the correlation between each position in the feature map and other positions, so as to better capture global context information, so that the obtained global feature map has relatively comprehensive global semantic information and global detail information.

[0087] S240, perform channel fusion on the current global feature map and the enhanced feature map corresponding to the highest-level feature map to obtain the highest-level optimized feature map; based on the highest-level optimized feature map, sequentially add and fuse each of the enhanced feature maps from high to low to obtain the bottom-most optimized feature map.

[0088] Fuse the enhanced feature map with the global feature map to add the multi-scale feature information enhanced by the enhanced feature map to the global information in the feature map, further optimizing the effect of feature fusion, enabling features to achieve a better balance between global and local, which is beneficial to achieving a better detection effect.

[0089] Furthermore, as Figure 2 shown, in order to fuse each level of the enhanced feature map with the global feature map, perform channel fusion on the global feature map and the enhanced feature map corresponding to the highest level of the feature pyramid. Please refer to Figure 2 the symbol C in to obtain the optimized feature map corresponding to the highest level of the feature pyramid, that is, the highest-level optimized feature map, and perform addition fusion on the highest-level optimized feature map and the enhanced feature map corresponding to the second-highest level of the feature pyramid to obtain the optimized feature map corresponding to the second-highest level of the feature pyramid, that is, the second-highest-level optimized feature map, and so on. Sequentially perform addition fusion on each level of the enhanced feature map from high to low until finally obtaining the optimized feature map corresponding to the bottom-most level of the feature pyramid, that is, the bottom-most optimized feature map. Based on this, the fusion of each level of the enhanced feature map with the global feature map can be realized.

[0090] It should be noted that before adding and fusing each optimized feature map with the enhanced feature map of the next level, resolution alignment is performed. Please refer to Figure 2 the symbol U in to make the optimized feature map and the enhanced feature map have the same resolution, so as to facilitate addition fusion. Specifically, perform an upsampling operation on the optimized feature map so that the optimized feature map has the same resolution as the enhanced feature map.

[0091] Furthermore, as Figure 2 shown, perform CBR operations on the image obtained after performing channel fusion or addition fusion, that is, perform convolution, batch normalization, and ReLU linear correction in sequence to realize feature fusion of the feature maps of two adjacent layers, and use the image after performing CBR operations as the optimized feature map.

[0092] S250, determine whether the current extraction times is the preset times. If so, perform convolution on the highest-level optimized feature map to obtain the highest-level prediction map, and perform convolution on the lowest-level optimized feature map to obtain the lowest-level prediction map. Otherwise, use the lowest-level optimized feature map as a new optimization item and continue to perform the next feature extraction.

[0093] Specifically, when the current extraction times is the preset times and the feature extraction times have reached the expectation, it means that the feature enhancement and fusion effect of the image to be detected is good. At this time, perform convolution on the highest-level optimized feature map to obtain the highest-level prediction map, and perform convolution on the lowest-level optimized feature map to obtain the lowest-level prediction map as the result of feature extraction.

[0094] It should be noted that since the lowest-level optimized feature map is obtained by fusing the enhanced feature maps and the global feature maps of each level, it contains relatively comprehensive detailed information and semantic information. Based on this, as Figure 2 shown, the dotted arrow in the figure is used to indicate that when the current extraction times is not the preset times, the lowest-level optimized feature map is used as the optimization item for feedback, added and fused with the current lowest-level feature map to obtain a new lowest-level feature map, and based on this new lowest-level feature map, the next feature extraction is performed, so as to improve the feature capture effect of the next feature extraction and achieve a better object detection effect.

[0095] S300, add and fuse the highest-level prediction map and the lowest-level prediction map to obtain the final prediction map, and perform recognition detection based on the final prediction map to obtain the detection result.

[0096] It should be noted that since the highest-level prediction map is actually obtained by convolving the highest-level optimized feature map and corresponds to the highest level of the feature pyramid, it has relatively rich global semantic information; the lowest-level prediction map is actually obtained by convolving the lowest-level optimized feature map and corresponds to the lowest level of the feature pyramid, it has relatively rich local detailed information. By fusing the highest-level prediction map and the lowest-level prediction map, the final prediction map can make use of the advantages of both global semantics and local details, thereby improving the accuracy and precision of object detection.

[0097] It should be noted that those skilled in the art should know the specific method of performing recognition detection based on the final prediction map, and this embodiment does not make a specific explanation here.

[0098] Based on this, through the final prediction map, the detection result can be obtained, that is, the detection of the camouflaged target is realized.

[0099] It should be noted that, in order to facilitate the addition and fusion of the top - level prediction map and the bottom - level prediction map, before the addition and fusion, the resolutions of the top - level prediction map and the bottom - level prediction map are aligned. Exemplarily, after performing the convolution operation in step S250 on the top - level prediction map and the bottom - level prediction map, the resolutions of the top - level prediction map and the bottom - level prediction map are aligned. Exemplarily, as Figure 2 shown, the symbol U represents the up - sampling operation, that is, the top - level prediction map and the bottom - level prediction map after the convolution operation are respectively subjected to the up - sampling operation to achieve resolution alignment.

[0100] Furthermore, the addition and fusion of the top - level prediction map and the bottom - level prediction map is a weighted fusion, that is, the product of the top - level prediction map and the corresponding weight is added to the product of the bottom - level prediction map and the corresponding weight to achieve a better fusion effect. Among them, the specific setting of the weight can be set according to actual needs, and this embodiment does not make specific limitations here. Exemplarily, the weights of both the top - level prediction map and the bottom - level prediction map are 0.5.

[0101] In some alternative embodiments, as Figure 9 shown, the execution method of the gated convolution enhancement processing includes:

[0102] S2221, based on the fused feature map, through the convolutional attention mechanism processing, obtain the corresponding weighted feature map.

[0103] Specifically, the weighted feature map refers to the image after dynamic feature enhancement through the convolutional attention mechanism.

[0104] Exemplarily, as Figure 10 shown, the execution method of the convolutional attention mechanism includes:

[0105] S22211, perform layer normalization processing on the fused feature map, and perform convolution and depth - separable convolution to obtain a group of attention matrices.

[0106] Among them, the group of attention matrices is used to obtain and enhance the features corresponding to the detection target in the fused feature map based on the attention mechanism.

[0107] S22212, based on the group of attention matrices, perform splitting to obtain a query matrix, a reference matrix, and an information matrix.

[0108] Exemplarily, the splitting of the group of attention matrices is realized through the Split function.

[0109] Among them, the query matrix, the reference matrix, and the information matrix are all obtained by splitting the attention matrix group to achieve feature enhancement through the interaction among the three.

[0110] The query matrix is used to represent the feature positions for which attention weights need to be calculated; the reference matrix is used to represent the feature values at each position in the fused feature map. By performing matching calculations between the query matrix and the reference matrix to obtain the similarity, the weight distribution of the features in the fused feature map can be obtained.

[0111] The information matrix is used as the object to be weighted and summed to perform feature enhancement based on the weight distribution.

[0112] S22213, perform a dot product calculation on the query matrix and the reference matrix, and perform normalization processing to obtain the attention weight matrix; perform a dot product calculation on the attention weight matrix and the information matrix to obtain the weighted feature map.

[0113] Specifically, perform a dot product calculation on the query matrix and the reference matrix to calculate the similarity between the query matrix and the reference matrix, thereby obtaining the attention weight matrix based on this similarity, and then adaptively adjust the features based on the attention weight matrix.

[0114] Based on this, the dynamic and effective enhancement of features can be achieved through the combination of convolution operations and attention mechanisms.

[0115] S2222, add and fuse the weighted feature map and the corresponding fused feature map, and perform layer normalization processing to obtain the intermediate feature map.

[0116] Specifically, through the addition and fusion of the weighted feature map and the fused feature map, the dynamically enhanced features are superimposed on the fused feature map, thereby achieving the feature enhancement of the fused feature map.

[0117] Furthermore, before adding and fusing the weighted feature map and the fused feature map, perform resolution alignment so that the weighted feature map and the fused feature map have the same resolution to facilitate addition and fusion. Exemplarily, the resolution alignment is achieved through upsampling operations or downsampling operations.

[0118] Even further, in order to improve the stability of the intermediate feature map and improve the accuracy of subsequent steps, after adding and fusing the weighted feature map and the corresponding fused feature map, perform layer normalization processing on the fused image to obtain the intermediate feature map.

[0119] S2223. Perform channel expansion and channel splitting on the intermediate feature map to obtain a first expanded feature map and a second expanded feature map.

[0120] It should be noted that the gating mechanism realizes feature enhancement by introducing a non - linear transformation and performing element - by - element multiplication with the feature map. Among them, the non - linear transformation is obtained by transforming the feature map through a non - linear activation function. Based on this, in this embodiment, the intermediate feature map is expanded in channels and split into two expanded feature maps, thereby introducing a non - linear transformation.

[0121] Exemplarily, the channel expansion of the intermediate feature map is achieved by applying a fully - connected layer to the intermediate feature map for linear transformation; the splitting of the intermediate feature map after channel expansion is achieved through the Split function.

[0122] S2224. Perform depth - separable convolution operations on the first expanded feature map to obtain a first convolutional feature map, and perform depth - separable convolution operations on the second expanded feature map to obtain a second convolutional feature map.

[0123] Specifically, feature extraction is performed on the first expanded feature map and the second expanded feature map through convolution operations to facilitate subsequent feature enhancement. Among them, in order to reduce the computational amount and achieve an efficient object detection process, this convolution operation is a depth - separable convolution operation.

[0124] S2225. Based on the first convolutional feature map, in combination with a non - linear activation function, obtain a gating weight matrix, and multiply the gating weight matrix and the second convolutional feature map element - by - element to obtain a gated feature map.

[0125] Through the element - by - element multiplication of the gating weight matrix and the second convolutional feature map, feature enhancement is realized. Among them, since the gating weight matrix is obtained through non - linear transformation based on the first convolutional feature map, it can adaptively adjust the weights, that is, an adaptive feature enhancement process is realized.

[0126] Exemplarily, the non - linear activation function is the GELU function.

[0127] S2226. Perform channel alignment based on the gated feature map, and add and fuse it with the fused feature map to obtain an enhanced feature map.

[0128] Specifically, similar to step S2222, through the addition and fusion of the weighted feature map and the fused feature map, the dynamically enhanced features are superimposed on the fused feature map, and through the addition and fusion of the gated feature map and the fused feature map, the features dynamically enhanced based on the gating mechanism are superimposed on the fused feature map, thereby realizing the feature enhancement of the fused feature map.

[0129] Further, to facilitate the execution of the addition fusion process, before the gated feature map and the fusion feature map are added and fused, channel alignment is also required. Specifically, the resolution and the number of channels of the gated feature map are adjusted to be the same as those of the fusion feature map. Exemplarily, the channel alignment is achieved by applying a fully connected layer for linear transformation.

[0130] As Figure 11 shown, a camouflage target detection device provided in this embodiment includes a pyramid construction module 41, a feature extraction module 42, and a detection and recognition module 43.

[0131] Among them, the pyramid construction module 41 is configured to construct an image pyramid based on the image to be detected, and obtain a feature pyramid, where the feature pyramid includes feature maps of a preset number of layers;

[0132] The feature extraction module 42 is configured to perform feature extraction a preset number of times based on each of the feature maps, and obtain a top-layer prediction map and a bottom-layer prediction map;

[0133] The detection and recognition module 43 is configured to perform addition fusion on the top-layer prediction map and the bottom-layer prediction map to obtain a detection result.

[0134] Based on the same inventive concept, the camouflage target detection method provided in the embodiments of the present invention can be implemented on the terminal side or the server side.

[0135] As Figure 12 shown, it is a schematic diagram of an optional hardware structure of a terminal provided in an embodiment of the present invention. The terminal 50 may be a mobile phone, a computer device, a tablet device, a personal digital processing device, a factory background processing device, etc. The terminal 50 includes: at least one processor 51, a memory 52, at least one network interface 54, and a user interface 53. Each component in the device is coupled together through a bus system 55. It can be understood that the bus system 55 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 55 also includes a power bus, a control bus, and a status signal bus.

[0136] Among them, the user interface 53 may include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.

[0137] It can be understood that the memory 52 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory characterized in the embodiments of the present invention is intended to include but not be limited to these and any other suitable categories of memory.

[0138] The memory 52 in the embodiments of the present invention is used to store various categories of data to support the operation of the terminal. Examples of these data include: any executable programs for operating on the terminal 50, such as the operating system 521 and application programs 522; the operating system 521 contains various system programs, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application program 522 can contain various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. Implementing the camouflage target detection method provided by the embodiments of the present invention can be included in the application program 522.

[0139] The method disclosed in the above embodiments of the present invention can be applied to the processor 51 or implemented by the processor 51. The processor 51 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 51 or by instructions in software form. The above processor can be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 51 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present invention. The processor 51 can be a microprocessor or any conventional processor, etc. Combining the steps of the accessory optimization method provided by the embodiments of the present invention can be directly embodied as being completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module can be located in the storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.

[0140] In an exemplary embodiment, the terminal 50 may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs) to execute the foregoing method.

[0141] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the program is called by a processor, the camouflage target detection method provided by the present invention is implemented.

[0142] Among them, the computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example (but not limited to), an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), static random access memories (SRAMs), portable compact disk read-only memories (CD-ROMs), digital versatile disks (DVDs), memory sticks, floppy disks, and mechanical encoding devices.

[0143] The computer-readable program characterized herein may be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0144] In summary, through feature extraction for a preset number of times, the present application gradually improves the optimization and enhancement effect of features, thereby facilitating the capture of camouflaged personnel, that is, the features corresponding to the detection target. In the case where the detection target is highly similar to the background, the accuracy of feature capture is ensured, and a good target detection effect is achieved, thus realizing the accurate identification and positioning of the camouflage target, and having high industrial application value.

[0145] The descriptions of the processes or structures corresponding to the above respective drawings have their own emphases. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0146] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed in the present application should still be covered by the claims of the present application.

Claims

1. A method for detecting camouflaged targets, comprising: Constructing an image pyramid based on the image to be detected to obtain a feature pyramid, where the feature pyramid includes feature maps of a preset number of layers; Performing feature extraction a preset number of times based on each of the feature maps to obtain a top-layer prediction map and a bottom-layer prediction map; Adding and fusing the top-layer prediction map and the bottom-layer prediction map to obtain a final prediction map, and performing recognition and detection based on the final prediction map to obtain a detection result.

2. The camouflage target detection method according to claim 1, wherein Taking the feature map corresponding to the bottom layer of the feature pyramid as the bottom-layer feature map; The single feature extraction process includes: Obtaining the current optimization item; Based on the current optimization item, updating the bottom-layer feature map to obtain the current feature pyramid; Performing hierarchical feature enhancement processing on each of the feature layers in the current feature pyramid to obtain corresponding enhanced feature maps for each layer at the current time; Performing global context awareness processing based on the current feature pyramid to obtain the current global feature map; Performing channel fusion on the current global feature map and the enhanced feature map corresponding to the top-layer feature map to obtain a top-layer optimized feature map; Based on the top-layer optimized feature map, adding and fusing the enhanced feature maps in sequence from high to low to obtain a bottom-layer optimized feature map; Determining whether the current extraction times is the preset number of times. If so, performing convolution on the top-layer optimized feature map to obtain a top-layer prediction map, and performing convolution on the bottom-layer optimized feature map to obtain a bottom-layer prediction map. Otherwise, taking the bottom-layer optimized feature map as a new optimization item and continuing to perform the next feature extraction.

3. The camouflage target detection method according to claim 2, characterized in that, The hierarchical feature enhancement processing includes: Based on each of the feature maps, obtaining all adjacent layers and performing adjacent-layer feature fusion processing to obtain corresponding fused feature maps; Performing gated convolution enhancement processing on each of the fused feature maps to obtain corresponding enhanced feature maps.

4. The camouflage target detection method according to claim 2, wherein The performing global context awareness processing based on the current feature pyramid to obtain the current global feature map includes: Performing channel fusion on the feature maps of each level based on the current feature pyramid to obtain an intermediate global feature map; Based on the intermediate global feature map, obtaining a global feature map through a self-attention mechanism.

5. The camouflage target detection method according to claim 3, wherein For each of the fused feature maps, the gated convolution enhancement processing includes: Based on the fused feature map, performing processing through a convolutional attention mechanism to obtain a corresponding weighted feature map; Adding and fusing the weighted feature map and the corresponding fused feature map, and performing layer normalization processing to obtain an intermediate feature map; Performing channel expansion and channel splitting processing on the intermediate feature map to obtain a first expanded feature map and a second expanded feature map; Performing a depthwise separable convolution operation on the first expanded feature map to obtain a first convolutional feature map, and performing a depthwise separable convolution operation on the second expanded feature map to obtain a second convolutional feature map; Based on the first convolutional feature map, combining a non-linear activation function to obtain a gated weight matrix, and multiplying the gated weight matrix element-wise with the second convolutional feature map to obtain a gated feature map; Channel alignment is performed based on the gated feature map, and added and fused with the fusion feature map to obtain an enhanced feature map.

6. The camouflage target detection method according to claim 3, characterized in that Based on each of the feature maps, all adjacent layers are obtained, and adjacent layer feature fusion processing is performed to obtain corresponding fusion feature maps, including: If the feature map corresponds to the bottom layer or the top layer of the feature pyramid, the corresponding adjacent layer feature map is obtained, and the feature map and the adjacent layer feature map are subjected to dual-scale feature fusion to obtain the corresponding fusion feature map; Otherwise, all corresponding adjacent layer feature maps are obtained, and the feature map and all the adjacent layer feature maps are subjected to triple-scale feature fusion to obtain the corresponding fusion feature map.

7. The camouflage target detection method according to claim 5, wherein Based on the fusion feature map, through convolutional attention mechanism processing, the corresponding weighted feature map is obtained, including: Performing layer normalization processing on the fusion feature map, and performing convolution and depthwise separable convolution to obtain a group of attention matrices; Based on the splitting of the attention matrix, a query matrix, a reference matrix, and an information matrix are obtained; Performing dot product calculation on the query matrix and the reference matrix, and performing normalization processing to obtain an attention weight matrix; performing dot product calculation on the attention weight matrix and the information matrix to obtain the weighted feature map.

8. A camouflaged target detection device, characterized in that, Including a pyramid construction module, a feature extraction module, and a detection and recognition module; The pyramid construction module is used to construct an image pyramid based on the image to be detected to obtain a feature pyramid, and the feature pyramid includes feature maps with a preset number of layers; The feature extraction module is used to perform feature extraction a preset number of times based on each of the feature maps to obtain a top layer prediction map and a bottom layer prediction map; The detection and recognition module is used to add and fuse the top layer prediction map and the bottom layer prediction map to obtain a final prediction map, and perform recognition detection based on the final prediction map to obtain a detection result.

9. A terminal, characterized in that, Including: A processor and a memory, and the memory is communicatively connected to the processor; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes the camouflage target detection method according to any one of claims 1 to 7.

10. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the camouflage target detection method according to any one of claims 1 to 7 is implemented.