A method and system for detecting camouflaged targets based on simulated hyperspectral images to assist deep learning
By combining deep learning technology that simulates hyperspectral images and RGB images, the accuracy and robustness of camouflage object detection in complex backgrounds are solved, and the accurate detection and recognition of camouflage object is achieved.
Patent Information
- Application Number
- CN202410789372.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-06-18
AI Technical Summary
The existing camouflage object detection method based on deep learning is difficult to achieve high-precision and robust camouflage object detection in complex contexts, and traditional visual feature methods are prone to missing or misidentifying camouflage objects.
Combining simulated hyperspectral images and RGB images as input data, deep learning technology is used to generate a camouflage object detection model through feature extraction, shallow feature fusion and deep feature correction to achieve accurate detection and recognition of camouflage object.
The accuracy and robustness of camouflage object detection are improved in complex backgrounds, and are suitable for camouflage object detection tasks in various complex scenarios.
Smart Images

Figure CN118823305B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a camouflaged target detection method and system based on simulated hyperspectral images to assist deep learning. Background Art
[0002] Camouflaged target detection is an important technology aimed at identifying and locating targets concealed by camouflage techniques. In this field, the development of background technologies is crucial for improving the accuracy and robustness of detection.
[0003] First of all, traditional target detection methods usually rely on visual features or the saliency of image regions to achieve. However, the emergence of camouflaged targets makes these methods perform poorly in complex backgrounds because camouflage techniques blur the boundaries of targets, making them blend in with the background. Therefore, methods based on visual features may miss camouflaged targets or misidentify camouflaged targets in the background as the background.
[0004] In recent years, with the development of deep learning technology, deep learning-based camouflaged target detection methods have gradually become mainstream. These methods use deep neural networks to perform end-to-end learning on images, thereby learning more discriminative feature representations. For example, deep models such as convolutional neural networks (CNNs) can automatically learn the abstract features of targets, thus improving the robustness and accuracy of detection.
[0005] However, with the gradual increase in detection requirements, ordinary deep models are also difficult to meet the detection accuracy. Summary of the Invention
[0006] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a camouflaged target detection method and system based on simulated hyperspectral images to assist deep learning. By using simulated hyperspectral images and the original RGB images as input data, making full use of the spectral information in the hyperspectral images, and combining deep learning technology, accurate detection and identification of camouflaged targets are achieved in complex backgrounds, improving the detection accuracy and robustness.
[0007] According to one aspect of the present invention, the present invention provides a camouflaged target detection method based on simulated hyperspectral images to assist deep learning, and the method includes the following steps:
[0008] S1: Simulate hyperspectral image data according to the original image, and the hyperspectral image data includes spectral information of different bands;
[0009] S2: Process the original image and the hyperspectral image data;
[0010] S3: Input the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model;
[0011] S4: Apply the trained camouflage target detection model to achieve the detection and recognition of camouflage targets.
[0012] Preferably, the simulation of the hyperspectral image data from the original image includes:
[0013] Perform hyperspectral reconstruction on the original image to obtain simulated hyperspectral image data.
[0014] Preferably, the processing of the original image and the hyperspectral image data includes:
[0015] Input the original image and the hyperspectral image into an encoder for feature extraction, extract shallow features and deep features, fuse the shallow features to obtain shared spatial features, and generate an initial prediction map; adopt an attention mechanism to assign weights to the deep features and perform deep feature correction.
[0016] Preferably, the feature extraction, shallow feature fusion, and deep feature correction specifically include:
[0017] (1) Feature extraction:
[0018] Input the original image I R and the hyperspectral image I H into the encoder for extracting shallow features respectively to obtain shallow features S R , S H ; input the original image I R and the hyperspectral image I H into the residual neural network model for extracting deep features respectively to obtain deep features and where R represents the features of the original image, H represents the features of the hyperspectral image, and 1, 2, 3 represent the layers for extracting features from top to bottom;
[0019] (2) Shallow feature fusion:
[0020] For the shallow features S R , S H , perform global average pooling to obtain average encodings m R , m H , and the two pass through a multi-layer perceptron MLP to obtain m S , then calculate the corresponding weights a and b of the two according to cosine similarity, and the formula is as follows:
[0021] m R = GAP(S R )
[0022] m H = GAP(S H )
[0023]
[0024] a = COS(m S , m R ), b = 1 - a
[0025] where represents matrix multiplication, and the shallow features are fused according to the calculated weights to obtain the initial feature F initial and the initial prediction P initial :
[0026] F initial = a·m R + b·m H
[0027] (3) Deep feature correction:
[0028] Connect zhi to obtain D 1 :
[0029]
[0030] Apply channel attention to D 1 to obtain D' with assigned weights 1 :
[0031] D' 1 = CA(D 1 )
[0032] Use D' 1 to correct the initial prediction P initial The specific correction measures are as follows: D' 1 is upsampled to obtain D' up1 , and sigmoid is used for normalization. The normalized result and the reverse result are respectively multiplied by the feature map of the current scale to obtain the foreground attention feature F fa and the background attention feature F ba :
[0033]
[0034]
[0035] Calculate the foreground interference and background interference based on the foreground attention feature F fa , the background attention feature F ba and the initial prediction P initial . Eliminate the foreground interference through element addition and the background interference through element subtraction. Finally, obtain the prediction map P of the prediction feature of the next layer through a convolutional layer1 , enter the correction link of the next layer, and repeat this process until the final prediction P 3 , which is the output of the final prediction map.
[0036] Preferably, the detection and recognition of the camouflaged target includes:
[0037] According to the size set during model training, process the original image I R and the simulated hyperspectral data I H to meet the model input requirements;
[0038] Load the trained model, set the model to the test mode, and enter the computing environment;
[0039] Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map F of the camouflaged target existing in the predicted image final :
[0040] F fianl = model(I R , I H );
[0041] Resize F final to the size of the original image, and then obtain the final camouflaged target boundary map P through the Sigmoid activation function fianl :
[0042] P final = sigmoid(F fianl ).
[0043] According to another aspect of the present invention, the present invention also provides a camouflaged target detection system based on simulated hyperspectral image-assisted deep learning, and the system includes:
[0044] A reconstruction module for simulating hyperspectral image data according to the original image, and the hyperspectral image data includes spectral information of different bands;
[0045] A processing module for processing the original image and the hyperspectral image data;
[0046] A training module for inputting the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model;
[0047] A detection module for applying the trained camouflaged target detection model to realize the detection and recognition of the camouflaged target.
[0048] Preferably, the reconstruction module simulates hyperspectral image data according to the original image, including:
[0049] Perform hyperspectral reconstruction on the original image to obtain simulated hyperspectral image data.
[0050] Preferably, the processing module processes the original image and the hyperspectral image data, including:
[0051] Input the original image and the hyperspectral image into an encoder for feature extraction, extract shallow features and deep features, fuse the shallow features to obtain shared spatial features, and generate an initial prediction map; adopt an attention mechanism to assign weights to the deep features and perform deep feature correction.
[0052] Preferably, the feature extraction, shallow feature fusion, and deep feature correction specifically include:
[0053] (1) Feature extraction:
[0054] Input the original image I R and the hyperspectral image I H into an encoder for extracting shallow features respectively to obtain shallow features S R , S H ; input the original image I R and the hyperspectral image I H into a residual neural network model for extracting deep features respectively to obtain deep features and where R represents the features of the original image, H represents the features of the hyperspectral image, and 1, 2, 3 represent the layers of extracting features from top to bottom;
[0055] (2) Shallow feature fusion:
[0056] For the shallow features S R , S H , perform global average pooling to obtain average encodings m R , m H , and the two pass through a multi-layer perceptron MLP to obtain m s , and then calculate the corresponding weights a and b of the two according to the cosine similarity. The formula is as follows:
[0057] m R = GAP(S R )
[0058] m H = GAP(S H )
[0059]
[0060] a = COS(m s , m R ), b = 1 - a
[0061] Among them represents matrix multiplication, and the shallow features are fused according to the calculated weights to obtain the initial feature F initial and the initial prediction P initial :
[0062] F initial = a·m R + b·m H
[0063] (3) Deep feature correction:
[0064] Connect known to obtain D 1 :
[0065]
[0066] Connect D 1 through channel attention to obtain D' after weight assignment 1 :
[0067] D' 1 = CA(D 1 )
[0068] Use D' 1 to correct the initial prediction P initial , and the specific correction measures are as follows: D' 1 is upsampled to obtain D' up , and sigmoid is used for normalization. The normalized result and the reverse result are multiplied by the feature map of the current scale respectively to obtain the foreground attention feature F fa and the background attention feature F ba :
[0069]
[0070]
[0071] According to the foreground attention feature F fa , the background attention feature F ba and the initial prediction P initial calculate the foreground interference and the background interference. Eliminate the foreground interference through element addition, and eliminate the background interference through element subtraction. Finally, obtain the prediction map P 1 of the prediction feature of the next layer through a convolutional layer, and enter the correction link of the next layer. Repeat this process until the final prediction P 3 , which is the output of the final prediction map.
[0072] Preferably, the detection module realizes the detection and recognition of the camouflage target, including:
[0073] Process the original image I according to the size set during model training R and the simulated hyperspectral data I H to meet the model input requirements;
[0074] Load the trained model, set the model to test mode, and enter the computing environment;
[0075] Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map F of the camouflage targets existing in the predicted image final :
[0076] F fiant = model(I R , I H );
[0077] Resize Ff inal to the size of the original image, and then obtain the final camouflage target boundary map Pf through the Sigmoid activation function ianl :
[0078] P final = sigmoid(F fianl ).
[0079] Beneficial effects: This method uses simulated hyperspectral images and original RGB images as input data, combines deep learning techniques, and realizes accurate detection and recognition of camouflage targets in complex backgrounds. By making full use of the spectral information in hyperspectral images, the detection accuracy and robustness are improved, and it is applicable to camouflage target detection tasks in various complex scenarios.
[0080] The features and advantages of the present invention will become clear by referring to the following drawings and the detailed description of the specific embodiments of the present invention. Description of the Drawings
[0081] Figure 1 is a flowchart of a camouflage target detection method based on simulated hyperspectral image assisted deep learning;
[0082] Figure 2 is a schematic diagram of a camouflage target detection model based on simulated hyperspectral image assisted deep learning;
[0083] Figure 3 is a schematic diagram of a camouflage target detection system based on simulated hyperspectral image assisted deep learning. Detailed Description of the Specific Embodiments
[0084] Combined with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0085] Embodiment 1
[0086] Reference Figure 1 And Figure 2 , this embodiment provides a method for detecting camouflaged targets based on simulated hyperspectral images to assist deep learning. The method includes the following steps:
[0087] S1: Simulate hyperspectral image data according to the original image, and the hyperspectral image data includes spectral information of different bands;
[0088] S2: Process the original image and the hyperspectral image data;
[0089] S3: Input the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model;
[0090] S4: Apply the trained camouflaged target detection model to realize the detection and recognition of camouflaged targets.
[0091] In this embodiment, the original image includes the original RGB image. By using the spectral information in the hyperspectral image, this embodiment improves the detection accuracy and robustness, and is applicable to the camouflaged target detection tasks in various complex scenarios.
[0092] Preferably, the simulating the hyperspectral image data according to the original image includes:
[0093] Perform hyperspectral reconstruction on the original image to obtain simulated hyperspectral image data.
[0094] Preferably, the processing the original image and the hyperspectral image data includes:
[0095] Input the original image and the hyperspectral image into an encoder for feature extraction, extract shallow features and deep features, fuse the shallow features to obtain shared spatial features, and generate an initial prediction map; adopt an attention mechanism to assign weights to the deep features for deep feature correction.
[0096] Preferably, taking the original image as the original RGB image as an example, the feature extraction, shallow feature fusion, and deep feature correction specifically include:
[0097] (1) Feature extraction:
[0098] Input the original image I R and the hyperspectral image I H into the encoder for extracting shallow features respectively to obtain the shallow feature S R , S H ; Input the original image I R and the hyperspectral image I H into the residual neural network model for extracting deep features respectively to obtain the deep feature known wherein, R represents the feature of the original image, H represents the feature of the hyperspectral image, and 1, 2, 3 represent the layers of extracting features from top to bottom; the encoder can be the virtual converter VIT (Visual Transformer).
[0099] (2) Shallow feature fusion:
[0100] For the shallow features S R , S H , perform global average pooling to obtain the average encoding m R , m H , and the two pass through the multi-layer perceptron MLP to obtain m S , then calculate the corresponding weights a and b of the two according to the cosine similarity, and the formula is as follows:
[0101] m R = GAP(S R )
[0102] m H = GAP(S H )
[0103]
[0104] a = COS(m S , m R ), b = 1 - a
[0105] where represents matrix multiplication, and fuse the shallow features according to the calculated weights to obtain the initial feature F initial and the initial prediction P initial :
[0106] F initial = a·m R + b·m H
[0107] (3) Correct the deep features:
[0108] Connect known to obtain D 1 :
[0109]
[0110] Apply D 1 Through channel attention, obtain D' with the assigned weights 1 :
[0111] D' 1 = CA(D 1 )
[0112] Utilize D' 1 To correct the initial prediction P initial The specific correction measures are as follows: D' 1 After upsampling, obtain D' up1 , and use sigmoid for normalization. Multiply the normalization result and the reverse result by the feature map of the current scale respectively to obtain the foreground attention feature F fa And the background attention feature F ba :
[0113]
[0114]
[0115] According to the foreground attention feature F fa , the background attention feature F ba And the initial prediction P initial Calculate the foreground interference and background interference. Eliminate the foreground interference through element-wise addition and the background interference through element-wise subtraction. Finally, obtain the prediction map P of the prediction feature of the next layer through a convolutional layer 1 , enter the correction link of the next layer, and repeat this process until the final prediction P 3 , which is the output of the final prediction map.
[0116] Preferably, the detection and recognition of the camouflage target include:
[0117] Process the original image I R And the simulated hyperspectral data I H According to the size set during model training to meet the model input requirements;
[0118] Load the trained model, set the model to the test mode, and enter the computing environment;
[0119] Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map F of the camouflage target existing in the predicted image final :
[0120] F fianl = model(IR , I H );
[0121] Resize F final to the size of the original image, and then obtain the final camouflaged target boundary map P through the Sigmoid activation function fianl :
[0122] P final = sigmoid(F fianl ).
[0123] Specifically, the modified size of the original image can be 224 * 224.
[0124] This embodiment uses the simulated hyperspectral image and the original RGB image as input data, combines deep learning technology, and realizes the accurate detection and recognition of camouflaged targets in complex backgrounds. By making full use of the spectral information in the hyperspectral image, the detection accuracy and robustness are improved, and it is applicable to the camouflaged target detection tasks under various complex scenarios.
[0125] Embodiment 2
[0126] Figure 3 is a schematic diagram of a camouflaged target detection system based on simulated hyperspectral image assisted deep learning. As Figure 3 shown, this embodiment provides a camouflaged target detection system based on simulated hyperspectral image assisted deep learning, and the system includes:
[0127] A reconstruction module 301 for simulating hyperspectral image data according to the original image, where the hyperspectral image data includes spectral information of different bands;
[0128] A processing module 302 for processing the original image and the hyperspectral image data;
[0129] A training module 303 for inputting the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model;
[0130] A detection module 304 for applying the trained camouflaged target detection model to realize the detection and recognition of camouflaged targets.
[0131] Preferably, the reconstruction module 301 simulates hyperspectral image data according to the original image, including:
[0132] Performing hyperspectral reconstruction on the original image to obtain simulated hyperspectral image data.
[0133] Preferably, the processing module 302 processes the original image and the hyperspectral image data, including:
[0134] The original image and the hyperspectral image are input into an encoder for feature extraction to extract shallow features and deep features. The shallow features are fused to obtain shared spatial features, and an initial prediction map is generated. The deep features are weighted by an attention mechanism for deep feature correction.
[0135] Preferably, the feature extraction, shallow feature fusion, and deep feature correction specifically include:
[0136] (1) Feature extraction:
[0137] The original image I R and the hyperspectral image I H are respectively input into an encoder for extracting shallow features to obtain shallow features S R , S H ; The original image I R and the hyperspectral image I H are respectively input into a residual neural network model for extracting deep features to obtain deep features and where R represents the features of the original image, H represents the features of the hyperspectral image, and 1, 2, 3 represent the layers of feature extraction from top to bottom;
[0138] (2) Shallow feature fusion:
[0139] For the shallow features S R , S H , global average pooling is performed to obtain average encodings m R , m H , and the two are passed through a multi-layer perceptron MLP to obtain m S , and then the corresponding weights a and b are calculated according to the cosine similarity. The formula is as follows:
[0140] m R = GAP(S R )
[0141] m H = GAP(S H )
[0142]
[0143] a = COS(m S , m R ), b = 1 - a
[0144] where represents matrix multiplication, and the shallow features are fused according to the calculated weights to obtain the initial feature F initial and the initial prediction P initial :
[0145] Finitial = a·m R + b·m H
[0146] (3) Deep feature correction:
[0147] Connect known to obtain D 1 :
[0148]
[0149] Connect D 1 through channel attention to obtain D' with the assigned weights 1 :
[0150] D' 1 = CA(D 1 )
[0151] Use D' 1 to correct the initial prediction P initial , and the specific correction measures are as follows: D' 1 is upsampled to obtain D' up1 , and sigmoid is used for normalization. The normalized result and the reverse result are respectively multiplied by the feature map of the current scale to obtain the foreground attention feature F fa and the background attention feature F ba :
[0152]
[0153]
[0154] According to the foreground attention feature F fa , the background attention feature F ba and the initial prediction P initial calculate the foreground interference and the background interference, eliminate the foreground interference through element addition, eliminate the background interference through element subtraction, and finally obtain the prediction map P 1 of the prediction feature of the next layer through a convolutional layer, and enter the correction link of the next layer, and so on until the final prediction P 3 , which is the output of the final prediction map.
[0155] Preferably, the detection module 304 realizes the detection and recognition of the camouflage target, including:
[0156] Process the original image I R and the simulated hyperspectral data I H according to the size set during model training to meet the model input requirements;
[0157] Load the trained model, set the model to test mode, and enter the computing environment;
[0158] Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map F of the camouflage targets existing in the predicted image final :
[0159] F fianl = model(I R , I H );
[0160] Resize F to the size of the original image, and then obtain the final camouflage target boundary map P through the Sigmoid activation function final : fianl :
[0161] P final = sigmoid(F fianl ).
[0162] The specific implementation processes of the functions implemented by each module in Embodiment 2 of the present invention are the same as the implementation processes of the steps in Embodiment 1, and will not be elaborated here.
[0163] Embodiment 3
[0164] This embodiment provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the method steps in Embodiment 1 are implemented. The specific implementation process can refer to the implementation process of the method steps in Embodiment 1, and will not be elaborated here.
[0165] It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes.
[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the steps of the functions specified in Figure 1 one process or more processes and / or boxes Figure 1 the functions specified in one box or more boxes.
[0168] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A camouflaged target detection method based on simulated hyperspectral image assisted deep learning, characterized in that: The method comprises the following steps: S1: simulating hyperspectral image data according to the original image, wherein the hyperspectral image data includes spectral information of different bands; S2: Processing the original image and hyperspectral image data; S3: Input the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model; S4: applying the trained camouflage target detection model to detect and identify the camouflage target; simulating the hyperspectral image data based on the original image, including: The original image is hyperspectral reconstructed to obtain simulated hyperspectral image data; the processing of the original image and the hyperspectral image data includes: The original image and the hyperspectral image are input into the encoder for feature extraction to extract shallow features and deep features. The shallow features are fused to obtain shared spatial features and generate an initial prediction map. The attention mechanism is used to assign weights to the deep features and perform deep feature correction. The feature extraction, shallow feature fusion, and deep feature correction specifically include: (1) Feature extraction: The original image and hyperspectral images Input into the encoder used to extract shallow features to obtain shallow features , ; The original image and hyperspectral images The deep features are respectively input into the residual neural network model used to extract deep features , , and , , , where R represents the features of the original image, H represents the features of the hyperspectral image, and 1, 2, and 3 represent the number of layers for extracting features from top to bottom; (2) Shallow feature fusion: For shallow features , , perform global average pooling to get the average encoding , , both are obtained through multi-layer perceptron MLP , and then calculate the corresponding weights of the two according to the cosine similarity , , the formula is as follows: in It represents matrix multiplication, which fuses shallow features according to the calculated weights to obtain the initial features. and initial prediction : (3) Correction of deep features: Will Connect, get : Will Through channel attention, we get the weighted : use Initial prediction The specific correction measures are as follows: After upsampling , and use sigmoid for normalization, multiply the normalized result and the reverse result by the feature map of the current scale, and obtain the foreground attention features respectively and background attention features : According to the foreground attention feature , Background attention features and initial prediction Calculate the foreground interference and background interference, eliminate the foreground interference by element addition, eliminate the background interference by element subtraction, and finally obtain the prediction map of the prediction features of the next layer through a convolution layer , enter the next level of correction, and repeat this process until the final prediction , which is the final prediction graph output.
2. The method according to claim 1, characterized in that The detection and identification of the camouflaged target includes: According to the size set during model training, the original image And simulated hyperspectral data Processing to meet model input requirements; Load the trained model, switch the model to test mode, and enter the computing environment; Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map of the camouflaged targets in the predicted image. : ; Will The final camouflaged target boundary map is obtained by resizing it to the size of the original image and then using the Sigmoid activation function. : 。 3. A camouflaged target detection system based on simulated hyperspectral image assisted deep learning, characterized in that: The system comprises: A reconstruction module, used for simulating hyperspectral image data according to the original image, wherein the hyperspectral image data includes spectral information of different bands; A processing module, used for processing raw image and hyperspectral image data; A training module, used for inputting the processed hyperspectral image data and the original image into a preset model for training to obtain a camouflaged target detection model; The detection module is used to apply the trained camouflaged target detection model to detect and identify the camouflaged target; the reconstruction module simulates the hyperspectral image data according to the original image, including: The original image is hyperspectrally reconstructed to obtain simulated hyperspectral image data; the processing module processes the original image and the hyperspectral image data, including: The original image and the hyperspectral image are input into an encoder for feature extraction, shallow features and deep features are extracted, the shallow features are fused to obtain shared spatial features, and an initial prediction map is generated; the deep features are assigned weights using an attention mechanism to correct the deep features; the feature extraction, shallow feature fusion, and deep feature correction specifically include: (1) Feature extraction: The original image and hyperspectral images Input into the encoder used to extract shallow features to obtain shallow features , ; The original image and hyperspectral images The deep features are respectively input into the residual neural network model used to extract deep features , , and , , , where R represents the features of the original image, H represents the features of the hyperspectral image, and 1, 2, and 3 represent the number of layers for extracting features from top to bottom; (2) Shallow feature fusion: For shallow features , , perform global average pooling to get the average encoding , , both are obtained through multi-layer perceptron MLP , and then calculate the corresponding weights of the two according to the cosine similarity , , the formula is as follows: in It represents matrix multiplication, which fuses shallow features according to the calculated weights to obtain the initial features. and initial prediction : (3) Correction of deep features: Will Connect, get : Will Through channel attention, we get the weighted : use Initial prediction The specific correction measures are as follows: After upsampling , and use sigmoid for normalization, multiply the normalized result and the reverse result by the feature map of the current scale, and obtain the foreground attention features respectively and background attention features : According to the foreground attention feature , Background attention features and initial prediction Calculate the foreground interference and background interference, eliminate the foreground interference by element addition, eliminate the background interference by element subtraction, and finally obtain the prediction map of the prediction features of the next layer through a convolution layer , enter the next level of correction, and repeat this process until the final prediction , which is the final prediction graph output.
4. The system according to claim 3, characterized in that The detection module realizes the detection and recognition of the camouflaged target, including: And simulated hyperspectral data Processing to meet model input requirements; Load the trained model, switch the model to test mode, and enter the computing environment; Input the original image to be predicted and the corresponding hyperspectral data into the loaded model to obtain the distribution probability map of the camouflaged targets in the predicted image. : ; Will The final camouflaged target boundary map is obtained by resizing it to the size of the original image and then using the Sigmoid activation function. : 。
Citation Information
Patent Citations
Target identification method based on visible light and hyperspectral image information fusion
CN117994624A