Optical remote sensing image target detection method based on semantic invariant feature learning
By using semantic invariant feature learning and frequency domain transformation techniques in remote sensing image processing, the problem of insufficient generalization ability of deep learning object detection methods in the remote sensing field is solved, and more efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510313052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
The existing deep learning object detection methods in the remote sensing field have reduced generalization capabilities due to inter-domain differences, and the detection performance is significantly attenuated.
The optical remote sensing image object detection method based on semantic invariant feature learning is adopted, and the image is converted from the spatial domain to the frequency domain through discrete cosine transformation, and the frequency domain amplitude is enhanced, and the semantic invariant features is adaptively enhanced in combination with the cross attention mechanism.
The detailed and effective transformation of image style is realized, the semantic invariant feature representation in feature maps is enhanced, and the generalization ability and detection accuracy of object detection are improved.
Smart Images

Figure CN120147889A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optical image imaging, and in particular relates to an optical remote sensing image target detection method based on semantically invariant feature learning. Background Art
[0002] With the rapid development of artificial intelligence technology, target detection methods based on deep learning have been widely used in the field of remote sensing, such as target reconnaissance and battlefield monitoring in the military field, and urban monitoring, environmental monitoring and disaster assessment in the civilian field. Deep learning models have become the main means of intelligent remote sensing image processing with their powerful feature extraction and classification capabilities. However, most existing target detection methods rely on data-driven architectures, requiring that the statistical distribution of training data and test data is highly consistent to ensure the detection performance of the model. However, in actual remote sensing application scenarios, they are often affected by multiple factors such as scene changes, meteorological conditions, and seasonal changes, resulting in significant inter-domain differences between test data and training data. This inter-domain difference will lead to a decrease in the generalization ability of the model in a new environment, which in turn will cause a significant attenuation of target detection performance, limiting the actual application effect of deep learning methods in the field of remote sensing. Summary of the invention
[0003] In order to solve the problem that the detection network is sensitive to semantically changing features, resulting in weak generalization detection ability, the present invention provides an optical remote sensing image target detection method based on semantically invariant feature learning, which can more accurately control the degree of feature enhancement at different frequencies, thereby achieving more detailed and effective style transformation.
[0004] An optical remote sensing image target detection method based on semantic invariant feature learning comprises the following steps:
[0005] S1: Use discrete cosine transform to transform the optical remote sensing image I from the spatial domain to the frequency domain, and enhance the frequency domain amplitude to obtain the enhanced frequency domain remote sensing image I s ;
[0006] S2: Use the feature extraction network Resnet-50 to extract the optical remote sensing image I and the frequency domain remote sensing image I s Perform feature extraction to obtain a set of semantic feature maps F and semantic feature maps F s ;
[0007] S3: Adopt cross-attention mechanism to adaptively enhance semantic feature map F and semantic feature map F s The semantically invariant features in , and irrelevant features are suppressed to obtain the final enhanced feature map
[0008] S4: Enhance the feature map Input it into the target detection network to complete the target detection in the optical remote sensing image I.
[0009] Furthermore, for the frequency-domain remote sensing image I s the acquisition method is as follows:
[0010] S11: Perform a discrete cosine transform on the input optical remote sensing image I to obtain its frequency-domain representation. The specific operation is:
[0011] I f = DCT(I)
[0012] where DCT(·) represents the standard two-dimensional discrete cosine transform operation; I f is the frequency-domain representation after performing the two-dimensional discrete cosine transform operation on the input optical remote sensing image I;
[0013] S12: Use the set amplitude enhancement coefficient α to enhance the amplitude of I f The specific operation is:
[0014]
[0015] where the amplitude enhancement coefficient α is a random number in the range of 1 - 2, and × represents element-wise multiplication, is the enhanced representation obtained by enhancing the amplitude of the frequency-domain representation I f ;
[0016] S13: Perform an inverse discrete cosine transform on the enhanced representation to convert it into a spatial-domain representation to obtain the enhanced frequency-domain remote sensing image I s , and the specific operation is:
[0017] I s = IDCT(I f )
[0018] where IDCT(·) represents the standard two-dimensional inverse discrete cosine transform operation; I s is the spatial-domain representation after performing the two-dimensional inverse discrete cosine transform operation on the frequency-domain representation I f .
[0019] Furthermore, for the semantic feature map F and the semantic feature map F s the acquisition method is as follows:
[0020] F = Resnet(I)
[0021] F s = Resnet(I s )
[0022] Among them, Resnet(·) represents the feature extraction network of the Resnet-50 structure.
[0023] Further, the method for obtaining the enhanced feature map is as follows:
[0024] S31: Use the linear mapping layer to perform linear mapping on the semantic feature map F and the semantic feature map F s The specific operation is:
[0025] Q = Linear Q (F s )
[0026] K = Linear K (F)
[0027] V = Linear V (F)
[0028] Among them, Linear Q is the standard linear mapping layer that converts the semantic feature map F s into the query feature Q, Linear K is the standard linear mapping layer that converts the semantic feature map F into the key feature K, and Linear V is the standard linear mapping layer that converts the semantic feature map F into the value feature V;
[0029] S32: Calculate the correlation between the query feature Q and the key feature K to obtain the correlation weight matrix A. The specific operation is:
[0030] A = Softmax(Q T ·K)
[0031] Among them, Softmax(·) represents the standard Softmax function, which is used to normalize the value obtained after calculating Q T ·K, and T represents the transpose;
[0032] S33: Use the correlation weight matrix A to perform a weighting operation on the value feature V to obtain the final enhanced feature map The specific operation is:
[0033]
[0034] Among them, × represents the element-wise multiplication operation, and Norm(·) is used to normalize the value obtained after calculating V×A + V.
[0035] Further, the method for obtaining the optical remote sensing image I is:
[0036] Obtained by using a satellite to observe and photograph the area of interest where the target exists.
[0037] Beneficial effects:
[0038] 1. The present invention provides an optical remote sensing image target detection method based on semantic invariant feature learning. First, the input optical remote sensing image is subjected to style transformation to generate a set of image pairs. Then, a feature extraction network is used to extract the features of the image pairs respectively, and through the proposed image-level semantic invariant feature enhancement module, the semantic invariant feature representation in the feature map is strengthened, and finally, target detection with strong generalization ability is achieved. That is to say, the present invention transfers the input optical remote sensing image to the frequency domain by applying the discrete cosine transform (DCT) and enhances its amplitude to achieve an effective transformation of the image style representation. The way of directly operating on the frequency components of the image in the frequency domain in the present invention can more precisely control the enhancement degree of features at different frequencies, thereby achieving a more detailed and effective style transformation.
[0039] 2. The present invention provides an optical remote sensing image target detection method based on semantic invariant feature learning. By analyzing the style differences between groups of feature maps, a cross-attention mechanism is adopted to adaptively enhance highly relevant features and suppress irrelevant features, so as to effectively enhance the semantic invariant features in the feature map. Compared with the existing methods, the present invention dynamically adjusts the attention weights during the feature fusion process, can more accurately identify and strengthen the key semantic information, and at the same time effectively suppress the interference features, thereby achieving a more detailed and efficient feature enhancement. Brief description of the drawings
[0040] Figure 1 It is a schematic diagram of the overall process of an optical remote sensing image target detection method based on semantic invariant feature learning provided by the present invention;
[0041] Figure 2 It is a flowchart of the frequency domain image enhancement module;
[0042] Figure 3 It is a flowchart of extracting the features of the input image pairs based on the feature extraction network Resnet-50;
[0043] Figure 4 It is a flowchart of the image-level semantic invariant feature enhancement module. Detailed implementation manners
[0044] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.
[0045] The present invention discloses an optical remote sensing image target detection method based on semantic invariant feature learning, aiming to further enhance the generalization ability of the detection network, so as to effectively address the problem of decreased detection accuracy caused by factors such as style changes. The core of the present invention lies in realizing effective transformation of the image style through a frequency-domain image transformation module and a feature enhancement module, and further using the transformed image to enhance the network's ability to extract semantic invariant features. Specifically, first, a frequency-domain image transformation module is introduced, and the input optical remote sensing image is transformed from the spatial domain to the frequency domain using the discrete cosine transform, and the frequency-domain amplitude is enhanced. This process can not only achieve effective transformation of the image style, but also precisely control the enhancement degree of different frequency components in the frequency domain, thereby finely adjusting the overall style features of the image. Then, an image-level semantic invariant feature enhancement module is designed, which adopts a cross-attention mechanism to effectively enhance the semantic invariant features by adaptively enhancing highly relevant semantic invariant features and suppressing irrelevant features. Finally, the feature map after feature enhancement is input into the target detection network to complete the precise detection of targets in the optical remote sensing image.
[0046] Specifically, as Figure 1 shown, an optical remote sensing image target detection method based on semantic invariant feature learning includes the following steps:
[0047] Step 0: Acquisition of optical remote sensing image
[0048] Use a satellite to observe and photograph a region of interest with targets to obtain a high-resolution optical remote sensing image.
[0049] Step 1: Use the discrete cosine transform to transform the optical remote sensing image I from the spatial domain to the frequency domain, and enhance the frequency-domain amplitude to obtain the enhanced frequency-domain remote sensing image I s ;
[0050] As Figure 2 shown, given an input optical remote sensing image The specific method for first using the frequency-domain image enhancement module to perform style transformation on it is:
[0051] First, perform a discrete cosine transform on the input optical remote sensing image I to obtain its frequency-domain representation, and the specific operation can be expressed as:
[0052] I f = DCT(I)
[0053] where DCT(·) represents the standard two-dimensional discrete cosine transform operation; I f is the frequency-domain representation after performing the two-dimensional discrete cosine transform operation on the input image I, and thus the dimension is the same as that of the input image and is
[0054] Next, set the amplitude enhancement coefficient α to enhance the amplitude of I f , and the specific operation can be expressed as:
[0055]
[0056] where the amplitude enhancement coefficient α is a random number in the range of 1-2, and × represents element-wise multiplication, is the enhancement operation on the input I f , so the dimension is the same as the input image and is
[0057] Finally, perform an inverse discrete cosine transform on the amplitude-enhanced to convert it into a spatial domain representation to obtain the finally style-transformed optical remote sensing image, and the specific operation can be expressed as:
[0058] I s = IDCT(I f )
[0059] where IDCT(·) represents the standard two-dimensional inverse discrete cosine transform operation; I s is the spatial domain representation after performing the two-dimensional inverse discrete cosine transform operation on the input I f , so the dimension is the same as the input image and is
[0060] Step 2: Use the feature extraction network Resnet-50 to extract features from the optical remote sensing image I and the frequency domain remote sensing image I s respectively to obtain a set of semantic feature maps F and semantic feature maps F s ;
[0061] As Figure 3 shown, use the feature extraction network Resnet-50 to perform feature extraction operations on the images I and I s obtained in step one respectively to obtain the feature maps F and F s , and the specific operation can be expressed as:
[0062] F = Resnet(I), F s = Resnet(I s )
[0063] where Resnet(·) represents the feature extraction network with the Resnet-50 structure; I and I s represent the pair of optical remote sensing images obtained in step one. F and F s are the semantic feature maps extracted from the corresponding input images respectively,
[0064] Step 3: Use the cross-attention mechanism to adaptively enhance the semantic feature map F and the semantic feature map F s in the semantic invariant features and suppress irrelevant features to obtain the final enhanced feature map
[0065] Specifically, since the main difference between the I and I s images lies in the style difference, which belongs to the semantic variation features. Therefore, the proposed image-level semantic invariant feature enhancement module uses the cross-attention mechanism to calculate the feature correlation between the feature map groups F and F s and enhances the semantic invariant features in the extracted feature map by enhancing highly correlated features and suppressing irrelevant features. As Figure 4 shown, the specific method is as follows:
[0066] First, use the linear mapping layer to perform a linear mapping on the input feature map groups F and F s , and the specific operation can be expressed as:
[0067] Q = Linear Q (F s )
[0068] K = Linear K (F)
[0069] V = Linear V (F)
[0070] where Linear Q is the standard linear mapping layer that converts the semantic feature map F s into the query feature Q, Linear K is the standard linear mapping layer that converts the semantic feature map F into the key feature K, and Linear V is the standard linear mapping layer that converts the semantic feature map F into the value feature V; in the present invention, F s is used as an auxiliary feature map and is converted into the query feature Q by the linear mapping layer Linear Q , and F is used as the main feature map and is respectively converted into the key feature K and the value feature V by the linear mapping layers Linear K , Linear V ;
[0071] Next, calculate the correlation between the query feature Q and the key feature K to obtain the correlation weight matrix A, and the specific operation can be expressed as:
[0072] A = Softmax(Q T ·K)
[0073] Among them, "·" represents the matrix inner product, and Softmax(·) represents the standard Softmax function, which is used to normalize the calculated values.
[0074] Finally, the obtained correlation weight matrix A is used to weight the value feature V to obtain the final enhanced feature map. The specific operation can be expressed as:
[0075]
[0076] Among them, × represents the element-wise multiplication operation, and Norm(·) is used to normalize the data of the feature map. represents the input feature map after weighted processing, whose semantic invariant features are enhanced and semantic related features are suppressed. The dimension is consistent with the input.
[0077] Step 4: Input the enhanced feature map into the object detection network to complete the object detection in the optical remote sensing image I.
[0078] That is to say, the present invention finally uses the enhanced feature map obtained in Step 3 as the input of the detection result prediction network to obtain the prediction result for the input image.
[0079] Of course, the present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for target detection in optical remote sensing images based on semantically invariant feature learning, characterized in that: The following steps are involved: S1: Use discrete cosine transform to transform the optical remote sensing image I from the spatial domain to the frequency domain, and enhance the frequency domain amplitude to obtain the enhanced frequency domain remote sensing image I s ; S2: Use the feature extraction network Resnet-50 to extract the optical remote sensing image I and the frequency domain remote sensing image I s Perform feature extraction to obtain a set of semantic feature maps F and semantic feature maps F s ; S3: Adopt cross attention mechanism to adaptively enhance semantic feature map F and semantic feature map F s The semantically invariant features in , and irrelevant features are suppressed to obtain the final enhanced feature map S4: Enhance the feature map Input into the target detection network to complete the target detection in the optical remote sensing image I.
2. The method for optical remote sensing image target detection based on semantically invariant feature learning as claimed in claim 1, characterized in that: Frequency Domain Remote Sensing Image I s The method to obtain is: S11: Perform discrete cosine transform on the input optical remote sensing image I to obtain its frequency domain representation. The specific operation is: Where DCT(·) represents the standard two-dimensional discrete cosine transform operation; f is the frequency domain representation of the input optical remote sensing image I after performing a two-dimensional discrete cosine transform operation; S12: Use the set amplitude enhancement coefficient α to I f The amplitude is enhanced, and the specific operation is: Among them, the amplitude enhancement coefficient α is a random number ranging from 1 to 2, and × represents the multiplication of corresponding elements. I is the frequency domain representation f An enhanced representation obtained by performing amplitude enhancement; S13: Enhanced representation Perform inverse discrete cosine transform to convert it into spatial domain representation to obtain the enhanced frequency domain remote sensing image I s , the specific operations are: Where IDCT(·) represents the standard two-dimensional inverse discrete cosine transform operation; s To represent I in the frequency domain f Spatial domain representation after a two-dimensional inverse discrete cosine transform operation.
3. The method for optical remote sensing image target detection based on semantically invariant feature learning as claimed in claim 1, characterized in that: Semantic feature map F and semantic feature map F s The method to obtain is: F=Resnet(I) F s =Resnet(I s ) Among them, Resnet(·) represents the feature extraction network of Resnet-50 structure.
4. The method for optical remote sensing image target detection based on semantically invariant feature learning as claimed in claim 1, characterized in that: Enhanced feature map The method to obtain is: S31: Use the linear mapping layer to map the semantic feature map F and the semantic feature map F s Perform linear mapping, the specific operations are: Q=Linear Q (F s ) K=Linear K (F) V=Linear V (F) Among them, Linear Q To transform the semantic feature map F s Converted to the standard linear mapping layer of query feature Q, Linear K To transform the semantic feature map F into the key feature K, Linear V A standard linear mapping layer that converts the semantic feature map F into the value feature V; S32: Calculate the correlation between the query feature Q and the key feature K to obtain a correlation weight matrix A. The specific operation is: A=Softmax(Q T ·K) Among them, Softmax(·) represents the standard Softmax function, which is used to calculate Q T The values obtained after K are normalized, and T represents transposition; S33: Use the correlation weight matrix A to perform a weighted operation on the value feature V to obtain the final enhanced feature map The specific operations are: Here, × represents an element-by-element multiplication operation, and Norm(·) is used to normalize the value obtained after calculating V×A+V.
5. The method for optical remote sensing image target detection based on semantically invariant feature learning as claimed in claim 1, characterized in that: The method for obtaining the optical remote sensing image I is: The satellite is used to observe and photograph the area of interest where the target exists.