Remote sensing image scene target interpretation method and device

Through the combined method of ResNet-50 and Transformer encoder, multi-scale convolutional features of remote sensing image scene targets are extracted and feature fusion is performed, which solves the problems of insufficient feature utilization and large calculation amount in the prior art, and achieves high-precision remote sensing image interpretation.

CN120259722APending Publication Date: 2025-07-04PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510173911.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has problems of insufficient feature utilization rate and large calculation amount in the interpretation of remote sensing image scene targets, resulting in low interpretation accuracy.

Method used

The ResNet-50 network is used to combine the Transformer encoder and attention mechanism to extract global semantics and local fine features through multi-scale convolutional features, and perform feature fusion, and design upper and lower branch feature extraction modules to achieve comprehensive utilization of remote sensing image scene goals.

Benefits of technology

The accuracy of interpretation of remote sensing image scene targets is improved, the calculation process is simplified, and the calculation complexity and parameter quantity are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259722A_ABST
    Figure CN120259722A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image scene target interpretation method and device, and the method comprises the steps: obtaining a remote sensing image scene target image, and dividing the image into training data and verification data according to a proportion; determining a multi-scale convolution feature corresponding to the current training data by using a ResNet-50 network; sequentially determining upper branch enhancement features and lower branch fusion features corresponding to the multi-scale convolution features; and fusing the upper branch enhanced feature and the lower branch fusion feature to obtain a final fusion feature, thereby realizing comprehensive utilization of global semantics and local fine features of a remote sensing image scene target. And further determining an optimal interpretation model of the remote sensing image scene target. According to the method, an attention mechanism and a Transform encoder are combined, an upper branch feature enhancement module is designed to extract global semantic features, a lower branch feature fusion module is designed to extract local fine features, through fusion of the two features, rich feature information of a remote sensing image scene target is obtained, and the method has high interpretation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a method and device for interpreting remote sensing image scene targets. Background Art

[0002] The practical application scope of remote sensing image scene target interpretation is extensive. It provides important technical support and scientific basis in the utilization and management of land resources, the investigation and assessment of geological disasters, the monitoring and protection of the ecological environment, and the reconnaissance and surveillance activities in the military field. The core of its technology lies in the in-depth analysis and understanding of multiple features in remote sensing images, such as tone, color, shape, size, shadow, texture, pattern, position, and layout. Through the executed target interpretation tasks, technicians can extract the high-level semantic features jointly contained in different ground objects in remote sensing images, thereby realizing the analysis of the scene content and the overall semantic interpretation of remote sensing images.

[0003] Currently, the mainstream method systems mainly include technologies based on traditional machine learning and strategies based on deep learning frameworks. Under the traditional machine learning path, such as Bayesian classifiers, support vector machines, decision tree models, and random forest algorithms, are widely used in the interpretation of remote sensing image scene targets. However, these methods largely rely on the manually designed feature extraction process and use learning algorithms to construct interpretation models. Although they are applied in practice, due to the limited low-level feature expression ability, they often lead to low classification accuracy and require increased intervention of manual operations.

[0004] In contrast, deep learning technology trains and interprets remote sensing image scene targets through deep neural networks, showing significant advantages. The core competitiveness of such methods lies in their ability to autonomously extract high-level feature information from remote sensing images, and after sufficient training, the neural network shows excellent interpretation performance. Nevertheless, the efficient operation of deep learning models depends on large-scale datasets for training, which is a time-consuming process and also has relatively high requirements for computing resources and storage space.

[0005] Currently, there are mainly two implementation methods for remote sensing image scene target interpretation methods based on ResNet and Transformer: The first method uses ResNet and Transformer as backbone networks in a "parallel" manner to respectively extract the deep features of remote sensing image scene targets, and uses the fused features of the two networks at the end for the training and interpretation of remote sensing scene targets; The second method uses ResNet and Transformer as backbone networks in a "series" manner to extract the semantic features of remote sensing image scene targets for training and interpretation.

[0006] However, the above methods still have some deficiencies in the extraction of scene target features in remote sensing images. Specifically, for the first type of method, only the depth features of the last layer of ResNet and Transformer are concatenated and fused. The method is simple and prone to feature redundancy, and the depth features of scene targets learned by other network layers of the two neural networks are not utilized. The insufficient feature utilization leads to insufficient interpretation accuracy. For the second type of method, adding a Transformer after the convolutional layer only enhances the features of the connection layer, without considering the role of other adjacent convolutional layers. The feature utilization is insufficient, and the network becomes deeper, resulting in a larger computational amount during training and affecting the feature learning efficiency. Summary of the Invention

[0007] The technical solution adopted by the present invention is to design a remote sensing image scene target interpretation network based on ResNet-50 and Transformer to solve the above problems and improve the interpretation effect of remote sensing image scene targets. In view of this, the present invention provides a remote sensing image scene target interpretation method and device.

[0008] The technical solution of the present invention proposes a remote sensing image scene target interpretation method, including: Step S1, obtain a remote sensing image scene target image, divide the image into training data and validation data according to a ratio, and perform preprocessing; Step S2, use the configured ResNet-50 network to determine the multi-scale convolutional features corresponding to the current training data, where the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset; Step S3, use the configured Transformer encoder module to determine the upper-branch enhanced features corresponding to the multi-scale convolutional features, and realize the extraction of the global semantic features of the remote sensing image scene target; Step S4, use the configured adjacent layer guidance enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolutional features, and realize the extraction of the local fine features of the remote sensing image scene target; Step S5, fuse the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features, and realize the comprehensive utilization of the global semantic and local fine features of the remote sensing image scene target; Step S6, input the final fusion features into the interpretation unit to determine the optimal interpretation model of the remote sensing image scene target.

[0009] In one embodiment, the step S1 includes: Step S101, randomly crop the training data, where the crop covers at least 80% of the training data picture; Step S102: Randomly rotate the training data, where the rotation range is between -45 degrees and 45 degrees; Step S103: Horizontally flip the current training data with a probability of 0.5; Step S104: Crop from the center of the current training data towards both sides; Step S105: Convert the current training data into tensor data of a preset shape; Step S106: Normalize the tensor data channel by channel, changing the mean to 0 and the standard deviation to 1 to obtain the preprocessed training data.

[0010] In one embodiment, the step S2 includes: Step S201: Build a ResNet-50 network and import the pre-trained parameters "resnet50-11ad3fa6.pth" on the ImageNet dataset; Step S202: Input the current training data; Step S203: Use the ResNet-50 network to perform convolution operations on the training data in sequence, and extract four layers of convolution features with different scales in sequence.

[0011] In one embodiment, the step S3 includes: Step S301: Use the squeeze-and-excitation attention mechanism to enhance the features of the four layers of convolution features with different scales; Step S302: Perform pooling operations on the features obtained after feature enhancement to obtain the maximum pooling feature and the average pooling feature respectively; Step S303: Perform convolution operations on the maximum pooling feature and the average pooling feature to obtain the guiding information of the four different scale features; Step S304: Perform convolution operations on the fourth layer of convolution features to obtain the image patch embedding features; Step S305: Based on the image patch embedding features, input the guiding features of the four different scale features into the configured Transformer encoder for processing in sequence to obtain the upper branch enhanced features, and realize the extraction of the global semantic features of the remote sensing image scene target.

[0012] In one embodiment, the step S4 includes: Step S401: Use the squeeze-and-excitation attention mechanism to enhance the features of the four different scale convolution features; Step S402: Use the pre-configured convolution kernel to perform dimensionality reduction processing on the enhanced current features; Step S403: Fuse the adjacent layers of the four dimensionality-reduced current features to obtain the primary fusion features; Step S404: Perform feature-guided enhancement on adjacent layers of the 4 primary fusion features to obtain the lower-branch fusion features, thereby achieving the extraction of local fine features of the remote sensing image scene target.

[0013] In one embodiment, step S5 includes: Step S501: Extract the upper-branch enhanced features from the upper-branch feature enhancement module. Step S502: Sequentially splice, perform global average pooling on, and flatten the lower-branch fusion features using the Flatten function in PyTorch to obtain the lower-branch fusion features. Step S503: Use the Concat function in PyTorch to splice the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features.

[0014] In one embodiment, step S6 includes: Step S601: Determine the features of the fully-connected layer based on the final fusion features. Step S602: Input the features of the fully-connected layer into the Softmax function to obtain the discrimination probability for each remote sensing image scene target category. Step S603: Use the cross-entropy loss function to determine the distance loss between the true probability and the predicted probability. Step S604: Determine the model corresponding to the minimum value of the distance loss as the best interpretation model for the remote sensing image scene target.

[0015] On the other hand, the present invention also provides a remote sensing image scene target interpretation device, including: A remote sensing image scene preprocessing unit, configured to obtain training data, divide the data into training data and validation data according to a ratio, and perform preprocessing. A multi-scale feature extraction unit, configured to use the configured ResNet-50 network to determine 4 multi-scale convolutional features corresponding to the current training data, where the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset. An upper-branch feature extraction unit, configured to use the configured Transformer encoder module to determine the upper-branch enhanced features corresponding to the multi-scale convolutional features, thereby achieving the extraction of global semantic features of the remote sensing image scene target. A lower-branch feature extraction unit, configured to use the configured adjacent-layer guidance enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolutional features, thereby achieving the extraction of local fine features of the remote sensing image scene target. The up-down feature fusion generation unit is configured to fuse the enhanced features of the upper branch and the fused features of the lower branch to obtain the final fused features, so as to realize the comprehensive utilization of the global semantics and local fine features of the remote sensing image scene target; The image scene target interpretation unit is configured to determine the optimal interpretation model of the remote sensing image scene target based on the final fused features through global average pooling, Softmax function and cross-entropy loss function.

[0016] Adopting the above technical solution, the present invention has at least the following advantages: 1) The present invention uses the four multi-scale convolutional features of the ResNet-50 network for interpretation, and further designs an upper branch feature enhancement module to extract global semantic features and a lower branch feature fusion module to extract local fine features by combining the attention mechanism and the Transformer encoder. By fusing and utilizing the two features, richer feature information of the remote sensing image scene target is obtained. Therefore, it has a high interpretation accuracy.

[0017] 2) The method proposed by the present invention selects the ResNet-50 network, combines the Transformer encoder and the attention mechanism, does not use a more complex image feature extraction neural network, and only selects a 50-layer convolutional network to extract the deep semantic features of the remote sensing image scene target. In addition, at the upper branch, lower branch and double-branch fusion stages of the network, the number of parameters is small, and the calculations are mostly multiplication and addition operations. Therefore, it is more convenient. Description of the Drawings

[0018] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 It is a schematic flowchart of the method for interpreting the remote sensing image scene target according to an embodiment of the present invention; Figure 2 It is a schematic block diagram of the principle of the method for interpreting the remote sensing image scene target according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the "multi-head attention guidance" module according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the "multi-head attention-guided feature fusion" module according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the calculation logic of the low-level feature guidance coefficient according to an embodiment of the present invention. Detailed Embodiments

[0019] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the present invention will be described in detail as follows in conjunction with the accompanying drawings and preferred embodiments.

[0020] In the accompanying drawings, although exemplary embodiments of the present invention are shown, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art. The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.

[0021] In the first embodiment of the present invention, a method and device for interpreting remote sensing image scene targets are as Figure 1 shown, and include the following steps: Step S1, obtain a remote sensing image scene target image, divide the image into training data and validation data according to a ratio, and perform preprocessing; Step S2, use the configured ResNet-50 network to determine the multi-scale convolutional features corresponding to the current training data, wherein the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset; Step S3, use the configured Transformer module to determine the upper-branch enhanced features corresponding to the multi-scale convolutional features, and realize the extraction of the global semantic features of the remote sensing image scene target; Step S4, use the configured adjacent-layer guided enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolutional features, and realize the extraction of the local fine features of the remote sensing image scene target; Step S5, fuse the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features; Step S6, input the final fusion features into the interpretation unit to determine the best interpretation model for the remote sensing image scene target.

[0022] Referring to Figure 2 , the method provided in this embodiment will be described in detail step by step below.

[0023] Step S1, obtain a remote sensing image scene target image, divide the image into training data and validation data according to a ratio, and perform preprocessing. Data preprocessing is to increase the data. By transforming the training set, the training set is made more abundant and the model has stronger generalization ability. The following processing is performed on the training set data using the transforms function in the torchvison module in PyTorch: Step S101, randomly crop the training data to 256×256 pixels, and cover at least 80% of the picture during the cropping process; Step S102: Randomly rotate the training data within the range of -45 degrees to 45 degrees; Step S103: Horizontally flip the training data with a probability of 0.5; Step S104: Crop from the center of the training data along both sides, and the size of the cropped image is 224×224 pixels; Step S105: Convert the training data into tensor format; Step S106: Normalize the tensor data channel by channel, changing the mean to 0 and the standard deviation to 1; Step S2: Use the configured ResNet-50 network to determine 4 multi-scale convolutional features corresponding to the current training data, where the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset.

[0024] In this embodiment, the ResNet-50 network is built, the pre-trained parameter "resnet50-11ad3fa6.pth" on the ImageNet dataset is imported, and then the training data is input , and the ResNet-50 network is used to perform convolutional operations on the training data in sequence, and the output convolutional features of the Stage1 layer, Stage2 layer, Stage3 layer, and Stage4 layer are extracted in sequence 、 、 and .

[0025] Step S3: Use the configured Transformer encoder module to determine the upper-branch enhanced features corresponding to the multi-scale convolutional features, and realize the extraction of the global semantic features of the remote sensing image scene target.

[0026] Refer to Figure 3 , the upper-branch Transformer encoder module is mainly used to calculate the global semantic information of the remote sensing image scene target of four convolutional layer features 、 、 and .

[0027] Specifically, it includes feature focusing, feature pooling, image patch embedding features, trainable parameter embedding, and multi-head attention-guided feature calculation. The specific process is as follows: (I) Feature focusing Use the squeeze-and-excitation attention mechanism to perform feature focusing calculations on the 4 convolutional layer features at different scales 、 、 、 respectively, and obtain the focused features 、 , and .

[0028] (II) Feature Pooling Perform pooling operations on the focused features respectively to obtain the maximum pooling feature and the average pooling feature.

[0029] (1) Maximum Pooling Extract local information from the focused features, find the maximum value of the feature points in the neighborhood, and obtain the maximum pooling feature.

[0030] (2) Average Pooling Extract global information from the focused features, find the average value of the feature points in the neighborhood, and obtain the average pooling feature.

[0031] (III) Image Patch Embedding Features Perform convolution operations on the local features and global features simultaneously to obtain the image patch embedding information of the features.

[0032] First, perform convolution operations on the maximum pooling feature and the average pooling feature of the first-layer and second-layer features to obtain the upsampled features, and perform convolution operations on the maximum pooling feature and the average pooling feature of the third-layer and fourth-layer features to obtain the downsampled features. Then, "flatten" the upsampled features and the downsampled features respectively using the Flatten function, and use the Transforse function to swap the first and second dimensions of the features.

[0033] Finally, obtain the image patch embedding features after maximum pooling , , and .

[0034] Similarly, obtain the image patch embedding features after average pooling , , and .

[0035] Finally, perform distribution standardization processing on the image patch embedding features after maximum pooling and the image patch embedding features after average pooling respectively to obtain new features.

[0036] (IV) Trainable Parameter Embedding This part is mainly divided into three parts: embedding the target token, adding position encoding, and regularization processing. By initializing a trainable parameter and concatenating it with the token sequence, the features of the embedded target token are obtained. Since there is no processing of the information on the positions of the image patches, position encoding information is added, and the information that different positions of the image patches may generate different semantics is added to the image patch embedding tensor to make up for the lack of position information. To make the model more generalizable, a Dropout layer is introduced, and the multi-head attention-guided features under four different scales of convolutional features are obtained as follows: The guided feature of the first layer and ; The guided feature of the second layer and ; The guided feature of the third layer and ; The guided feature of the fourth layer and .

[0037] (V) Multi-head attention-guided features (1) Input features First, perform "image patch embedding" and "trainable parameter embedding" operations on the features of the fourth layer of the ResNet-50 network . That is, through image patch embedding, concatenating the target token, and position information embedding, the input features of the encoder block are finally obtained .

[0038] (2) Encoder block Refer to Figure 4 , in order to let the network capture richer information, the processing of the encoder block is divided into two steps: multi-head self-attention feature calculation and multi-head attention-guided feature calculation

[0039] The first step: Calculate the multi-head self-attention feature First, normalize the input features , and then perform multi-head self-attention calculation to obtain the multi-head self-attention information . Add the multi-head self-attention information and the input features element-wise, and then perform normalization processing to obtain the multi-head self-attention feature .

[0040] The second step: Calculate the multi-head attention-guided feature Multiply the multi-head self-attention feature with the fourth guided feature and They are input into the multi-head self-attention module for calculation to obtain multi-head attention-guided features The multi-head attention-guided features and the multi-head self-attention features are added element-wise and then normalized to obtain the multi-head attention-guided features .

[0041] (5) Transformer Encoder Similarly, the third-layer guided features and , the second-layer guided features and , and the first-layer guided features and are input into the encoder block to obtain the final output features of the Transformer encoder .

[0042] Step S4: Use the configured adjacent-layer guidance enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolution features. This part mainly calculates the local semantic information of the features , , and of the remote sensing image scene target at four different scales.

[0043] (I) Feature Enhancement Use the squeeze-and-excitation attention mechanism to perform feature enhancement on the four initial features , , , respectively. It mainly includes the following steps: First, perform global average pooling to obtain the pooled features, then use a fully connected network to perform dimensionality reduction and dimensionality increase calculations in sequence, input the results into the Sigmoid function to calculate the channel weight coefficients, and finally multiply the channel weight coefficients element-wise with the features to be enhanced to obtain the enhanced features.

[0044] (II) Feature Dimensionality Reduction Use convolutional kernels to perform convolutional dimensionality reduction processing on the enhanced features to obtain the dimensionality-reduced features , , and .

[0045] (III) Primary Fusion of Adjacent Features Fuse the dimensionality-reduced features of adjacent layers in the way of "+". First, expand the spatial scale of the dimensionality-reduced features of the higher layer of adjacent layers by 2 times, and then fuse the feature information in the way of element-wise addition. The specific steps are as follows: (1) Spatial Dimension Expansion Using the bilinear interpolation algorithm for and and perform scale amplification, and successively obtain and and .

[0046] (2) Primary feature fusion Sum the elements of the features of adjacent layers feature by feature to fuse into primary features and and and .

[0047] (4) Advanced fusion of adjacent features (1) Amplification of high-level feature size Using the bilinear interpolation algorithm for the primary fusion features and and perform spatial scale amplification to make their scale sizes equal to , and obtain the features at the high-level fusion stage and and and .

[0048] (2) Guiding coefficients of low-level features Referring to Figure 5 , calculate the guiding coefficients of the low-level features of adjacent layers, mainly including calculating the global attention feature, local attention feature, and guiding coefficient.

[0049] In the calculation of the global attention feature, the features , and are respectively subjected to two-dimensional average pooling operations. When processing , first perform "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, and padding numbers of 0 on both sides. Then add a rectified linear unit to increase the non-linearity of the network. Finally, use a fully connected layer for calculation to obtain the average value weight .

[0050] Similarly, processing the feature obtains the average value weight , processing obtains the average value weight .

[0051] In the calculation of the local attention feature, the primary features and and are respectively subjected to two-dimensional maximum pooling operations. When processing At this time, first perform an "inner product operation" using a two-dimensional convolution with a convolution kernel of 1×1, a stride of 1, and padding of 0 on both sides. Then add a rectified linear unit to increase the non-linearity of the network. Finally, use a fully connected layer for calculation to obtain the maximum value weight 。

[0052] Similarly, process the feature to obtain the maximum value weight 、process to obtain the maximum value weight 。

[0053] In the calculation of the guiding coefficient, add the average value weight corresponding to the primary feature and the maximum value weight element-wise and input the result into the Sigmoid function to obtain the guiding coefficient ; Similarly, use the average value weight corresponding to the feature and the maximum value weight to obtain the guiding coefficient 。

[0054] Similarly, use the average value weight corresponding to the feature and the maximum value weight to obtain the guiding coefficient 。

[0055] (3) High-level fusion feature enhancement Multiply the low-level guiding coefficients of adjacent layers with the corresponding elements of the high-level features to obtain high-level fusion features.

[0056] Multiply the guiding coefficient and the high-level feature element-wise to obtain the high-level fusion feature ; Multiply the guiding coefficient and the high-level feature element-wise to obtain the high-level fusion feature ; Multiply the guiding coefficient and the high-level feature element-wise to obtain the high-level fusion feature 。

[0057] Thus, four high-level fusion features , , , 。

[0058] Step S5: Fuse the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features, realizing the comprehensive utilization of the global semantics and local fine features of the remote sensing image scene target.

[0059] For the upper branch, the obtained output features Extract the target tokens, and the result is used as the upper-branch enhanced features, denoted as .

[0060] For the lower branch, after obtaining the four high-level fusion features , , , , first perform splicing to obtain the feature , then perform global average pooling to obtain the feature , and finally use the Flatten function in PyTorch to "flatten" to obtain the feature .

[0061] The upper-branch enhanced features and the lower-branch features Are concatenated using the Concat function in PyTorch to obtain the final fusion features .

[0062] Step S6: Input the final fusion features into the interpretation unit to determine the best interpretation model for the remote sensing image scene target.

[0063] First, input the final fusion features Into the global average pooling layer to calculate the global average pooling features; then, input the global average pooling features into the Softmax function to obtain the discrimination probability of each remote sensing image scene target category; secondly, use the cross-entropy loss function to determine the distance loss between the true probability and the predicted probability; finally, after training, the model corresponding to the minimum value of the distance loss is determined as the best interpretation model for the remote sensing image scene target.

[0064] Compared with the prior art, this embodiment has at least the following advantages: 1) The present invention uses the 4 multi-scale convolution features of the ResNet-50 network for interpretation, and further designs an upper-branch feature enhancement module combining the attention mechanism and the Transformer encoder to extract global semantic features and a lower-branch feature fusion module to extract local fine features. By fusing and utilizing the two features, richer feature information of the remote sensing image scene target is obtained. Therefore, it has a high interpretation accuracy.

[0065] 2) The method proposed by the present invention selects the ResNet-50 network, combines the Transformer encoder and the attention mechanism, does not use a more complex neural network for extracting image features, and only selects a 50-layer convolutional network to extract the deep semantic features of the remote sensing image scene targets. Additionally, in the upper branch, lower branch, and dual-branch fusion stages of the network, the number of parameters is small, and the calculations are mostly multiplication and addition operations, so it is more convenient.

[0066] The second embodiment of the present invention corresponds to the first embodiment. This embodiment introduces a device for interpreting remote sensing image scene targets, which includes the following components: A remote sensing image scene preprocessing unit, configured to obtain training data, divide the data into training data and validation data according to a ratio, and perform preprocessing; A multi-scale feature extraction unit, configured to use the configured ResNet-50 network to determine 4 multi-scale convolutional features corresponding to the current training data, wherein the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset; An upper branch feature extraction unit, configured to use the configured Transformer encoder module to determine the upper branch enhanced features corresponding to the multi-scale convolutional features, and realize the extraction of the global semantic features of the remote sensing image scene targets; A lower branch feature extraction unit, configured to use the configured adjacent layer guidance enhancement module to determine the lower branch fusion features corresponding to the multi-scale convolutional features, and realize the extraction of the local fine features of the remote sensing image scene targets; An upper and lower feature fusion generation unit, configured to fuse the upper branch enhanced features and the lower branch fusion features to obtain the final fusion features, and realize the comprehensive utilization of the global semantic and local fine features of the remote sensing image scene targets; An image scene target interpretation unit, configured to determine the best interpretation model of the remote sensing image scene targets based on the final fusion features through global average pooling, the Softmax function, and the cross-entropy loss function.

[0067] It should be noted that in the embodiments of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0068] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0070] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.

Claims

1. A method for interpreting scene targets in remote sensing images, characterized in that, Including: Step S1: Obtain the remote sensing image scene target image, divide the image into training data and validation data according to a ratio, and perform preprocessing. Step S2: Use the configured ResNet-50 network to determine the multi-scale convolution features corresponding to the current training data, where the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset. Step S3: Use the configured Transformer encoder module to determine the upper-branch enhanced features corresponding to the multi-scale convolution features, and extract the global semantic features of the remote sensing image scene target. Step S4: Use the configured adjacent-layer guidance enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolution features, and extract the local fine features of the remote sensing image scene target. Step S5: Fuse the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features, and comprehensively utilize the global semantic and local fine features of the remote sensing image scene target. Step S6: Input the final fusion features into the interpretation unit to determine the best interpretation model for the remote sensing image scene target.

2. The method for interpreting remote sensing image scene targets according to claim 1, characterized in that, The step S1 includes: Step S101: Randomly crop the training data, where the crop covers at least 80% of the training data picture. Step S102: Randomly rotate the training data, where the rotation range is between -45 degrees and 45 degrees. Step S103: Horizontally flip the current training data with a probability of 0.

5. Step S104: Crop from the center of the current training data to both sides. Step S105: Convert the current training data into tensor data of a preset shape. Step S106: Normalize the tensor data channel by channel, change the mean to 0, and the standard deviation to 1 to obtain the preprocessed training data.

3. The method for interpreting remote sensing image scene targets according to claim 2, wherein, The step S2 includes: Step S201: Build the ResNet-50 network and import the pre-trained parameter "resnet50-11ad3fa6.pth" on the ImageNet dataset. Step S202: Input the current training data. Step S203: Use the ResNet-50 network to perform convolution operations on the training data in sequence, and extract four layers of convolution features with different scales in sequence.

4. The method for interpreting remote sensing image scene targets according to claim 3, characterized in that, The step S3 includes: Step S301: Use the squeeze-and-excitation attention mechanism to enhance the features of the four layers of convolution features with different scales. Step S302: Perform pooling operations on the features obtained after feature enhancement to obtain the maximum pooling feature and the average pooling feature respectively. Step S303: Perform convolution operations on the maximum pooling feature and the average pooling feature to obtain the guiding information of four different-scale features. Step S304: Perform convolution operations on the fourth-layer convolution feature to obtain the image patch embedding feature. Step S305: Based on the image patch embedding feature, sequentially input the guiding features of the four different-scale features into the configured Transformer encoder for processing to obtain the upper-branch enhanced features, and extract the global semantic features of the remote sensing image scene target.

5. The method for interpreting remote sensing image scene targets according to claim 4, wherein The step S4 includes: Step S401: Using a compression and excitation attention mechanism to enhance the features of four convolution features with different scales; Step S402: Using a pre-configured convolution kernel to perform dimensionality reduction on the enhanced current features; Step S403: Fusing the adjacent layers of the four current features after dimensionality reduction to obtain primary fusion features; Step S404: Performing feature-guided enhancement on the adjacent layers of the four primary fusion features to obtain lower-branch fusion features, realizing the extraction of local fine features of remote sensing image scene targets.

6. The method for interpreting remote sensing image scene targets according to claim 5, characterized in that The step S5 includes: Step S501: Extracting the upper-branch enhanced features from the upper-branch feature enhancement module; Step S502: Sequentially splicing, globally average pooling the lower-branch fusion features, and flattening them using the Flatten function in PyTorch to obtain lower-branch fusion features; Step S503: Using the Concat function in PyTorch to splice the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features.

7. The method for interpreting remote sensing image scene targets according to claim 6, characterized in that The step S6 includes: Step S601: Determining the globally average pooled features based on the final fusion features; Step S602: Inputting the globally average pooled features into the Softmax function to obtain the discrimination probability of each remote sensing image scene target category; Step S603: Using the cross-entropy loss function to determine the distance loss between the true probability and the predicted probability; Step S604: Determining the model corresponding to the minimum value of the distance loss as the best interpretation model for remote sensing image scene targets.

8. A device for interpreting scene targets in remote sensing images, characterized in that, It includes: A remote sensing image scene preprocessing unit, configured to obtain training data, divide the data into training data and validation data according to a ratio, and perform preprocessing; A multi-scale feature extraction unit, configured to use the configured ResNet-50 network to determine four multi-scale convolution features corresponding to the current training data, where the ResNet-50 network includes pre-trained parameters imported on the ImageNet dataset; An upper-branch feature extraction unit, configured to use the configured Transformer encoder module to determine the upper-branch enhanced features corresponding to the multi-scale convolution features, realizing the extraction of global semantic features of remote sensing image scene targets; A lower-branch feature extraction unit, configured to use the configured adjacent layer guidance enhancement module to determine the lower-branch fusion features corresponding to the multi-scale convolution features, realizing the extraction of local fine features of remote sensing image scene targets; An upper and lower feature fusion generation unit, configured to fuse the upper-branch enhanced features and the lower-branch fusion features to obtain the final fusion features, realizing the comprehensive utilization of global semantic and local fine features of remote sensing image scene targets; An image scene target interpretation unit, configured to determine the best interpretation model for remote sensing image scene targets based on the final fusion features through a fully connected network, the Softmax function, and the cross-entropy loss function.

Citation Information

Cited By

  • Dense pedestrian detection method and system based on improved D-FINE-N network

    CN122493497A