Complex scene information prior-oriented remote sensing image natural language description generation method

Through a method of complex scene information prior, combined with deep learning and natural language processing technology, multi-scale features and text prior information of remote sensing images are extracted and fused, the problem that the existing technology is not comprehensive and accurate enough in complex scene descriptions is solved, and a more accurate and detailed natural language description generation is achieved.

CN120014288AActive Publication Date: 2025-05-16UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411963075.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing natural language description generation method for remote sensing images is difficult to accurately capture and express all key information when facing complex and changeable remote sensing scenarios, resulting in the generated description being incomplete, accurate and even misunderstandings.

Method used

Using a priori method for complex scene information, a data set that combines Chinese and English is constructed, a pre-trained image recognition neural network is used to extract shallow and deep coding features of multiple scales, and combined with a priori information construction module, text priori features are extracted from visual global features, and multi-feature cross-fusion is performed, and natural language description is finally generated through a pre-trained natural language model.

Benefits of technology

It effectively improves the accuracy and meticulousness of natural language descriptions of remote sensing images, can better understand the spatial and semantic relationships between land objects in complex scenes, and generates a more comprehensive and accurate description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014288A_ABST
    Figure CN120014288A_ABST
Patent Text Reader

Abstract

The invention discloses a complex scene information prior-oriented remote sensing image natural language description generation method, and belongs to the field of image processing and analysis. The method provided by the invention comprises the following steps: constructing a Chinese and English combined data set; constructing global features and local features of vision; constructing information prior features; strengthening the global features and the local features; carrying out multi-feature cross fusion; and performing natural language description generation on the crossed and fused features. According to the technical scheme, the accuracy of remote sensing image description related to a large number of complex scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of deep learning and image processing, and in particular to a method for generating natural language descriptions of remote sensing images based on prior information of complex scenes. Background Art

[0002] With the rapid development of science and technology, remote sensing technology has become an indispensable tool in the fields of earth observation, environmental monitoring, urban planning and disaster assessment. As the core link of this technology, remote sensing image analysis has undergone a profound transformation from traditional manual interpretation to automated processing based on machine learning and deep learning. This transformation not only significantly improved the speed and accuracy of image processing, but also greatly broadened the application boundaries of remote sensing technology. As an emerging direction in the field of remote sensing image analysis, natural language description generation of remote sensing images aims to convert complex image information into natural language descriptions that can be understood by humans, thereby greatly improving the convenience and intuitiveness of information acquisition. The realization of this technology is of great significance to promoting the popularization and application of remote sensing data, enhancing decision-making support capabilities, and promoting interdisciplinary integration. It is an important bridge connecting the technological world and human cognition.

[0003] The development of natural language description generation for remote sensing images can be traced back to early rule-based methods, which attempt to map specific features in images to corresponding natural language descriptions through preset templates and rules. However, as the complexity and diversity of image content increase, the limitations of such methods become increasingly prominent. Subsequently, with the cross-integration of natural language processing and computer vision technologies, deep learning-based models have begun to emerge, which can learn the potential associations between images and texts and generate richer and more accurate descriptions. The progress at this stage marks the entry of remote sensing image natural language description generation technology into a new era of intelligence, promoting the transition from simple feature description to complex scene narration.

[0004] Although existing methods have made certain achievements in generating natural language descriptions of remote sensing images, they still face many challenges when facing complex and ever-changing remote sensing scenes. Complex scenes often contain multi-level and multi-category ground object information, as well as complex spatial relationships, which places higher demands on the model's understanding ability, generalization ability, and detail capture ability. When dealing with such scenes, existing methods often find it difficult to accurately capture and express all key information, resulting in the generated description being incomplete, inaccurate, or even misleading. In view of this, we propose an innovative method for generating natural language descriptions of remote sensing images based on prior information of complex scenes. This method makes full use of the advantages of deep learning models in processing large-scale data, and combines prior knowledge. By enhancing the model's ability to understand and parse complex scenes, it aims to generate more accurate, detailed, and insightful natural language descriptions, thereby effectively overcoming the limitations of existing technologies and promoting the further development of natural language description generation technology for remote sensing images. Summary of the invention

[0005] The present application provides a method for generating natural language descriptions of remote sensing images based on prior information of complex scenes, so as to improve the generation quality of current methods for generating natural language descriptions of remote sensing images.

[0006] In order to solve the above technical problems, the technical solutions adopted in this application are as follows:

[0007] A method for generating natural language description of remote sensing images based on prior information of complex scenes, the method comprising the following steps:

[0008] Step 1: Build a joint Chinese and English remote sensing image natural language description dataset;

[0009] Step 2: Use the backbone network of the pre-trained image recognition neural network to extract shallow coding features of multiple scales and deep coding features of multiple scales of the input image;

[0010] The shallow coding features of multiple scales are integrated to obtain the global visual features, and the deep coding features of multiple scales are integrated to obtain the local visual features.

[0011] Step 3, constructing a module based on prior information to extract text prior features from visual global features;

[0012] Step 4, respectively enhancing the visual global features and the visual local features to obtain enhanced visual global features and visual local features;

[0013] Step 5, feature fusion of the text prior features and the enhanced visual global features is performed through a multi-feature cross-fuser to obtain global fusion features, and feature fusion of the text prior features and the enhanced visual local features is performed to obtain local fusion features; and feature fusion of the global fusion features and the local fusion features is performed to obtain the final coding features;

[0014] Step 6: Decode the final encoded features based on the pre-trained natural language model to generate a natural language description of the remote sensing image.

[0015] Furthermore, step 1 includes: obtaining a public natural language description dataset of remote sensing images as an initial dataset, and obtaining an English dataset based on the English descriptions of the remote sensing images in the initial dataset; based on the English dataset, re-annotating with matching Chinese, and constructing a corresponding Chinese dataset based on the Chinese annotations.

[0016] Furthermore, in step 2, the shallow coding features and the deep coding features include two scales respectively, corresponding to two stages.

[0017] Furthermore, in step 2, the pre-trained backbone network is a Transformer-based network structure or a convolutional neural network-based network structure.

[0018] Furthermore, in step 2, the shallow coding features of multiple scales and the deep coding features of multiple scales extracted have their channel numbers gradually increasing and their height and width gradually decreasing in the forward propagation direction.

[0019] Furthermore, step 2 also includes fine-tuning and learning some network parameters in the backbone network, wherein the fine-tuning and learned part of the network parameters are network parameters of the network layer used to extract deep coding features.

[0020] Furthermore, in step 2, for shallow coding features of multiple scales, a layer-by-layer fusion method is used to fuse multi-scale shallow coding features. According to the forward propagation direction of the backbone network, the shallow / deep coding features of each scale are traversed in turn, and the feature dimension of the shallow / deep coding feature of the current scale is adjusted to the dimension of the next shallow / deep coding feature through convolution and downsampling operations, and then the features are superimposed. The current superimposed shallow / deep coding feature is used as the new next shallow / deep coding feature, and the feature adjustment and superposition of the new next shallow / deep coding feature and the next shallow / deep coding feature are continued until the shallow / deep coding feature of the last scale is obtained.

[0021] For example, for shallow coding features including two scales, the dimension of the first scale is adjusted to the second scale and then the features are superimposed to obtain the visual global feature.

[0022] Furthermore, in step 3, the processing of the prior information construction module includes:

[0023] Step 3-1: Define the global visual feature as G∈R C×H×W ;

[0024] The visual global feature G is processed by global maximum pooling and global average pooling respectively to obtain global vectors g1 and g2;

[0025] Then the global vectors g1 and g2 are respectively passed through a shared weight parameter and pre-trained multi-attribute predictor to obtain the attribute prediction probability p1∈R of the scene attribute of the remote sensing image d and p2∈R d , where d represents the number of scene attributes of the remote sensing image;

[0026] Step 3-2: Sort the elements in the comprehensive multi-attribute prediction probabilities p1 and p2, select the top K largest image attribute prediction probabilities, and convert them into one-hot vector representations to obtain the matrix U∈R d×k ;

[0027] Step 3-3: Construct a learnable matrix M∈R t×d , so that it and the matrix U∈R d×k Perform matrix multiplication to obtain a new matrix E∈R t×k , where t is the set coding length;

[0028] Step 3-4: Expand the matrix E and then replicate it in space to obtain the feature map E f ∈R B×H×W ; Among them, the number of feature map channels B = t × k;

[0029] The feature map E f Through a convolution layer with a convolution kernel of 1×1, the dimension is adjusted to obtain the text prior feature P∈R C ×H×W .

[0030] Informative prior features can provide more information about scene categories, thus playing an auxiliary role in understanding the semantic and spatial relationships of ground objects in complex scenes.

[0031] Furthermore, in step 4, the visual global features are enhanced by the set global feature enhancement module, including:

[0032] Step 4-1-A: Define the global visual feature as G∈R C×H×W , where C is the number of channels and H×W is the height and width of the feature;

[0033] The global feature G is convolved by parallel depth-wise separable convolutions with kernel sizes of 3x3, 5x5, and 7x7 respectively;

[0034] The results of the three-way convolution operation are superimposed according to the channel dimension, and then the number of feature channels is adjusted to be consistent with the visual global feature G through a convolution layer with a convolution kernel of 1x1, and the feature S∈R is obtained. C×H×W ;

[0035] Split the feature S into two new features S1∈R according to the channel C / 2×H×W and S2∈R C / 2×H×W ;

[0036] The feature S1 is pooled horizontally and globally to obtain the feature S _h ∈R C / 2×H×1 ; Feature S2 is pooled vertically and globally to obtain feature S _w ∈R C / 2×1×W ;

[0037] For feature S _h Multiply by point with S1 to get feature T1∈R C / 2×H×W , for S _w Multiply by point with S2 to get feature T2∈R C / 2×H×W ;

[0038] Superimpose features T1 and T2 to obtain the enhanced visual global feature T∈R C×H×W .

[0039] Furthermore, in step 4, the local visual features are enhanced by a local feature enhancement module, including:

[0040] Step 4-2-A: Define the local visual feature as L∈R C×H×W , where C is the number of channels and H×W is the height and width of the feature;

[0041] The local visual feature L is passed through the convolution layer with a convolution kernel of 3x3 and the ReLU activation function to obtain the feature L / ∈R C×H×W ;

[0042] Step 4-2-B: Set feature L / ∈R C×H×W Divide into four features L1∈R according to the channel order C / 4×H×W , L2∈R C / 4×H×W , L3∈R C / 4×H×W and L4∈R C / 4×H×W ;

[0043] Features L1 and L3 are superimposed to obtain L 1-3 ∈R C / 2×H×W, features L2 and L4 are superimposed to obtain L 2-4 ∈R C / 2×H×W , features L1 and L4 are superimposed to obtain L 1-4 ∈R C / 2×H×W , features L2 and L3 are superimposed to obtain L 2-3 ∈R C / 2×H×W ; This process cross-superimposes the features in the channel dimension of the local features, which can enhance the perception of inter-channel features and thus enhance the local features;

[0044] Step 4-2-C: For feature L 1-3 , L 2-4 , L 1-4 and L 2-3 Channel attention will be performed separately to obtain enhanced features and and

[0045] Step 4-2-D: Enhanced features and According to the channel superposition, the feature U∈R is obtained C×H×W , the enhanced features and According to the channel superposition, we get the feature V∈R C×H×W ;

[0046] Then add the features U and V pixel by pixel to get the enhanced visual local feature O∈R C×H×W .

[0047] Furthermore, the specific execution process of the multi-feature cross-fusion device in step 5 includes:

[0048] Set the text prior feature to P∈R C×H×W , the enhanced visual global feature is T∈R C×H×W , the enhanced local visual feature is O∈R C×H×W ;

[0049] Perform a joint spatial attention operation on the enhanced visual global feature T and the text prior feature P to obtain the feature

[0050] Perform a joint spatial attention operation on the enhanced local visual feature O and the text prior feature P to obtain the feature

[0051] feature and Features According to the channel dimension superposition, the first fusion coding feature Y∈R is obtained by sequentially passing through the convolution layer with a convolution kernel of 3x3 and the activation function C×H×W; Through the cross-fusion of multiple features, we can effectively achieve a deep understanding of the semantic relationship of ground objects in complex scenes;

[0052] The first fused encoding feature Y is passed through an atrous spatial convolutional pooling pyramid (ASPP) module to obtain the final encoding feature Q∈R C×H×W Further mining multi-scale information based on fusion features can achieve a more comprehensive understanding of complex scenes.

[0053] Furthermore, the enhanced visual global feature T and the text prior feature P are subjected to a joint spatial attention operation to obtain the feature include:

[0054] The visual global feature T is obtained through two convolution layers with a convolution kernel of 1×1 to obtain features T1 and T2;

[0055] Adjust the dimension of feature T2 to H×W×C, and then perform matrix multiplication with feature P to obtain the spatial attention relationship map map1 = (H×W)×(H×W); and use the softmax function to weight the spatial attention relationship map map1;

[0056] Then add feature T1 and the weight-activated spatial attention relationship map map2 pixel by pixel to get the feature

[0057] Furthermore, the enhanced local visual feature O and the text prior feature P are subjected to a joint spatial attention operation to obtain the feature

[0058] The local visual feature O is passed through two convolution layers with a convolution kernel of 1×1 to obtain features O1 and O2;

[0059] Adjust the dimension of O2 to H×W×C, and then perform matrix multiplication with feature P to obtain the spatial attention relationship map map2 = (H×W)×(H×W), and perform weight activation on the spatial attention relationship map map2 through the softmax function;

[0060] Then add feature O1 and the weight-activated spatial attention relationship map map2 pixel by pixel to get the feature

[0061] The technical solution provided by this application brings at least the following beneficial effects:

[0062] This application uses a pre-trained neural network to construct global features and local features, and constructs information prior features based on global features. Information prior features can be used to achieve a better understanding of the spatial and semantic relationships between objects in complex scenes. In addition, the cross-fusion of global features and local features with information prior features can effectively improve the accuracy of image description of remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0064] Figure 1 This is a structural diagram of a global feature enhancement module according to an embodiment of the present application;

[0065] Figure 2 A structural diagram of a priori information building module in an embodiment of the present application;

[0066] Figure 3 This is a structural diagram of a multi-cross fuser according to an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions of the embodiments of the present application will be described in detail and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the embodiments described with reference to the drawings are exemplary and are intended to be used to explain the present application, and cannot be understood as limiting the present application.

[0068] This application provides a method for generating natural language descriptions of remote sensing images based on prior information of complex scenes, which aims to improve the effect of current methods for generating natural language descriptions of remote sensing images.

[0069] In one embodiment, the method for generating natural language description of remote sensing images provided in the present application comprises the following steps:

[0070] Step 1: Build a joint Chinese and English dataset.

[0071] Obtain open source natural language description datasets of remote sensing images from home and abroad, retain the original English dataset, refer to the English description, and re-annotate it using appropriate Chinese.

[0072] In this embodiment, in order to further ensure the quality of annotation, the following requirements need to be noted when constructing a joint Chinese and English dataset:

[0073] 1) In the process of remote sensing image annotation, it is necessary not only to accurately identify and annotate the target objects, scene backgrounds and their subordinate relationships in the image, but also to use deep learning technology to deeply mine the implicit information in the image, such as the shape characteristics, texture details and color attributes of the target objects. This information will be structured and integrated into the annotation text to generate a more detailed and comprehensive image description.

[0074] 2) First, natural language generation technology is used to automatically generate preliminary annotation text. Subsequently, professionals conduct manual review and optimization to ensure that the annotation text meets language standards, can accurately and clearly convey the key information in the remote sensing image, and maintain fluent and natural written expression.

[0075] 3) In the final stage of annotating the text, each text will be carefully checked to remove redundant information and ambiguous expressions to ensure that each annotation is closely related to the image content and can accurately reflect the target object and scene characteristics in the image.

[0076] 4) To ensure the quality of annotation, a complete annotation quality assessment system will be established to regularly assess and inspect the completed annotation results. Based on the assessment results, targeted feedback and improvement suggestions will be provided to continuously improve the accuracy and overall quality of the annotation work and ensure the high-level output of the annotation results.

[0077] Step 2: Construction of global and local visual features.

[0078] In this embodiment, a neural network based on a Transformer structure or CNN pre-trained on the ImageNet dataset is used to obtain encoding features of multiple scales of the input image, and based on this, global and local visual features are constructed.

[0079] The specific steps of constructing the global and local features of vision in the above process are:

[0080] Step 2-1: remove the last average pooling layer and fully connected layer of the pre-trained neural network constructed based on Transfromer or CNN; that is, the backbone network based on the pre-trained image recognition neural network extracts the coding features of multiple scales of the input image, and the backbone network constitutes the pre-trained encoder Encoder of the present application; in this embodiment, the coding features of the four stages are extracted from the backbone network in the forward propagation direction, and the number of channels of the extracted coding features of the four stages gradually increases, and the feature resolution of the coding features gradually decreases, that is, the height and width will gradually decrease.

[0081] Step 2-2: Freeze the parameters of the first and third stages in the pre-trained model, and fine-tune the parameters of the second and fourth stages during the training process;

[0082] Step 2-3: Obtain the remote sensing image encoding features F of the four stages in the encoder n Then, based on the encoding feature F n Construct global and local features of vision.

[0083] Given a remote sensing image input I∈R C×H×W , using the pre-trained encoder Encoder to extract visual features of remote sensing images of different scales F n :

[0084] F n =Encoder(I;W1,W2)

[0085] Among them, n represents the extraction stage of different features, W1 represents the weight parameters frozen in the pre-training model, and W2 represents the weight parameters that need to be fine-tuned in the pre-training model.

[0086] Visual Features n It contains the encoding features F1, F2, F3 and F4 of the four stages of the pre-training model. The number of channels of the encoding features of the four stages will gradually increase, while the height and width will gradually decrease. The superposition of F1 and F2 is used as the local feature, and the superposition of F3 and F4 is used as the global feature. Considering the inconsistency of the dimensions of the feature map, before superimposing the features, it is necessary to adjust the size of F1 to be consistent with F2 through convolution and downsampling operations, and adjust the size of F3 to be consistent with F4.

[0087] Step 3: Construct informative prior features.

[0088] Based on the constructed visual global features, information prior features are constructed through the information prior construction module.

[0089] See also Figure 1 The specific execution process of the information priori construction module in this embodiment includes:

[0090] Step 3-A: Set the global feature to G∈R C×H×W , where C is the number of channels and H×W is the resolution of the feature, i.e., height and width. First, the global features are processed by global maximum pooling and global average pooling to obtain global vectors g1 and g2. Then, the global vectors g1 and g2 are respectively processed by a shared weight parameter and pre-trained multi-attribute predictor to obtain the multi-attribute prediction probability p1∈R of the image scene. d and p2∈R d . Where d represents the number of scene attributes, such as airport, forest, school, city, etc.

[0091] Step 3-B: Multi-attribute prediction probability p1∈R for the above step predictiond and p2∈R d The elements in are sorted in descending order, and then the first k (preset value) attribute prediction probabilities are selected based on the two probability prediction results, and converted into a one-hot vector representation to obtain the matrix U∈R d×k .

[0092] Step 3-C: Construct a learnable matrix M∈R t×d , so that it and the matrix U∈R d×k Perform matrix multiplication to obtain a new matrix E∈R t×k . Where t is the set coding length.

[0093] Step 3-D: Expand the matrix E and then replicate it in space to obtain the feature map E f ∈R B×H×W . Where B = t × k. Next, the feature map E f Through a 1×1 convolution, the feature map with dimension size C×H×W is adjusted to obtain the text prior feature P∈R C×H×W .

[0094] Step 4: Strengthening of global and local features.

[0095] The global feature enhancement module and the local feature enhancement module are used to enhance the constructed global features and local features respectively.

[0096] See also Figure 2 , the specific execution process of the global feature enhancement module in this embodiment is:

[0097] Step 4-1-A: Set the global feature to G∈R C×H×W The global features are processed by parallel depth-wise separable convolutions with kernel sizes of 3x3, 5x5, and 7x7, respectively. The above process is expressed as follows:

[0098] G1=DWConv3×3(G)

[0099] G2=DWConv5×5(G)

[0100] G3=DWConv7×7(G)

[0101] Among them, G1, G2 and G3 represent the features processed by depthwise separable convolution with kernel sizes of 3x3, 5x5 and 7x7 respectively, and all dimensions are consistent with G. DWConv3×3, DWConv5×5 and DWconv7×7 represent depthwise separable convolution with kernel sizes of 3x3, 5x5 and 7x7 respectively.

[0102] Step 4-1-B: Superimpose the features G1, G2 and G3 obtained in the above steps according to the channel dimension, and adjust the number of channels to be consistent with feature G through 1x1 convolution. The expression is as follows:

[0103] S=Conv1×1(ConaCat(G1,G2,G3))

[0104] Step 4-1-C: The feature S∈R obtained in the above step C×H×W Split into two new features S1∈R according to the channel C / 2×H×W and S2∈R C / 2×H×W .

[0105] Next, feature S1 is pooled horizontally to obtain feature S _h ∈R C / 2×H×1 , feature S2 will be pooled vertically to obtain feature S _w ∈R C / 2×1×W Then, S _h Multiply by point with S1 to get feature T1∈R C / 2×H×W , S _w Multiply by point with S2 to get feature T2∈R C / 2×H×W Finally, the features are superimposed to obtain the output feature T∈R C×H×W .

[0106] In one embodiment, the specific execution process of the local feature enhancement module in step 4 is:

[0107] Step 4-2-A: Set the local feature to L∈R C×H×W Feature L will first be re-integrated through 3x3 convolution and ReLU activation function to obtain feature L / ∈R C×H×W .

[0108] Step 4-2-B: Feature L obtained in the above steps / ∈R C×H×W Divide it into four features L1∈R according to the channel order C / 4×H×W , L2∈R C / 4×H×W , L3∈R C / 4×H×W and L4∈R C / 4×H×W Then, L1 and L3 are superimposed to obtain L 1-3 ∈R C / 2×H×W , L2 and L4 are superimposed to obtain L 2-4 ∈R C / 2×H×W , features L1 and L4 are superimposed to obtain L 1-4 ∈R C / 2×H×W , L2 and L3 are superimposed to obtain L 2-3 ∈R C / 2×H×WThe above process cross-superimposes the features in the channel dimension in the local features, which can enhance the perception ability of inter-channel features and further enhance the local features.

[0109] Step 4-2-C: Feature L obtained in the above steps 1-3 , L 2-4 , L 1-4 and L 2-3 Channel attention will be performed separately to obtain enhanced features and and

[0110] Step 4-2-D: Then take the features obtained in the above steps and According to the channel superposition, we can get the feature U∈R C ×H×W ,feature and According to the channel superposition, we get the feature V∈R C×H×W The feature U and feature V are superimposed to obtain the final output feature O∈R C×H×W .

[0111] Step 5: Multi-feature fusion: The fusion of global features, local features and text prior features is achieved through a multi-feature cross-fuser.

[0112] See also Figure 3 The specific execution process of the multi-cross fuser in this embodiment is as follows:

[0113] Step 5-A: Set the enhanced global feature to T∈R C×H×W , the text prior feature is P∈R C×H×W . Next, features T and P will undergo a joint spatial attention operation. Specifically, first, the global feature T will be passed through two 1x1 convolutions to obtain features T1 and T2. Then, feature T2 will change its dimension to HxWxC and perform matrix multiplication with feature P to obtain the spatial attention relationship map map = (H×W)×(H×W). The spatial attention heat map map will be weighted activated by the softmax function. Next, the features obtained by matrix multiplication of feature T1 and the spatial attention relationship heat map are added pixel by pixel to obtain the output

[0114] Step 5-B: Set the enhanced local feature to O∈R C×H×W , the text prior feature is P∈R C×H×W. Next, features O and P will undergo a joint spatial attention operation. Specifically, first, the local feature O will be convolved with two 1x1 convolutions to obtain features O1 and O2. Then, feature O2 will change its dimension to H×W×C and perform matrix multiplication with feature P to obtain the spatial attention relationship map map = (H×W)×(H×W). The spatial attention heat map map will be weighted activated by the softmax function. Next, the features obtained by matrix multiplication of feature O1 and the spatial attention relationship map are added pixel by pixel to O to obtain the output Through the above-mentioned cross-fusion of multiple features, we can effectively achieve a deep understanding of the semantic relationship of ground objects in complex scenes.

[0115] Step 5-C: Features obtained from the above steps and Features It will be superimposed according to the channel dimension and then through 3x3 convolution and activation function to obtain the output fused encoding feature Y∈R C×H×W .

[0116] Step 5-D: The feature Y obtained in the above step will pass through an atrous spatial convolutional pooling pyramid (ASPP) module to obtain the final encoding feature Q∈R C×H×W Further mining multi-scale information based on fusion features can achieve a more comprehensive understanding of complex scenes.

[0117] Step 6: Generate natural language description. Based on the current pre-trained natural language model, decode the features fused in step 5 to generate the final natural language description.

[0118] Example

[0119] In one embodiment, a method for generating natural language description of remote sensing images based on complex scene information prior provided in the embodiment of the present application specifically comprises the following steps:

[0120] Step S1: Construct a joint Chinese and English dataset. Obtain open source natural language description datasets of remote sensing images from home and abroad, retain the original English dataset, refer to the English description, and annotate it again with appropriate Chinese.

[0121] Step S2: Construction of global and local visual features. Use a neural network based on the Transformer or CNN structure pre-trained on the ImageNet dataset to obtain the encoding features of the input image, and use this as a basis to construct global and local visual features. In the specific implementation process, the pre-trained ResNe50 can be selected as the feature extraction network. The size of the input remote sensing image is I∈R 3×512×512 , the input image will get 4 stages of encoding features after ResNe50, which are: F1∈R256×128×128 , F2∈R 512×64×64 , F3∈R 1024×32×32 and F4∈R 2048×16×16 Next, F1 will be downsampled and convolved by maximum pooling to make the feature dimension consistent with F2, and then superimposed with F2 along the channel dimension to obtain the local feature L∈R 1024×64×64 Similarly, F3 will downsample and convolve through maximum pooling to keep the feature dimension consistent with F4, and then superimpose it with F4 along the channel dimension to obtain the global feature G∈R 4096×16×16 .

[0122] Step S3: Construction of information prior features. First, the global feature G∈R 4096×16×16 After global maximum pooling and global average pooling, the global vector g1∈R is obtained. 4096×1×1 and g2∈R 4096×1×1 Then, the global vectors g1 and g2 are respectively passed through a shared weight parameter and pre-trained multi-attribute predictor to obtain the multi-attribute prediction probability p1∈R d and p2∈R d . Where d represents the number of attributes. The multi-attribute prediction probability p1∈R predicted in the above steps d and p2∈R d The elements in are sorted in descending order, and then the first k attribute prediction probabilities are selected by combining the two probability prediction results, and converted into a one-hot vector representation to obtain the matrix U∈R d×k Next, we construct a learnable matrix M∈R t×d , so that it and the matrix U∈R d×k Perform matrix multiplication to obtain a new matrix E∈R t×k . Where t is the set encoding length. Next, the matrix E is expanded and then replicated in space to obtain the feature map E∈R B×H×W . Among them, B = t × k. Finally, the feature map E is resized to P ∈ R by a 1 × 1 convolution. 4096×16×16 .

[0123] Step S4: Strengthening of global features and local features. The global features are strengthened by the global feature strengthening module. Specifically, the global feature is G∈R 4096×16×16 The feature G1∈R is processed by parallel depth-wise separable convolution with kernel sizes of 3x3, 5x5, and 7x7 respectively. 4096×16×16 , G2∈R 4096×16×16 and G3∈R 4096×16×16 The results are superimposed according to the channel dimension, and the number of channels is adjusted to be consistent with G through 1x1 convolution to obtain the feature S∈R 4096×16×16 Next, the feature S is split into two new features S1∈R according to the channel2048×16×16 and S2∈R 2048×16×16 Next, feature S1 is pooled horizontally to obtain feature S _h ∈R 2048×16×1 , feature S2 will be pooled vertically to obtain feature S _w ∈R 2048×1×16 Then, S _h The dot product with S1 gives the same features as S _w The feature superposition obtained by point multiplication with S2 is obtained to obtain the output feature T∈R 4096×16×16 .

[0124] The local features will be enhanced through the local feature enhancement module. Local feature L∈R 1024×64×64 The spatial features will first be reintegrated through 3x3 convolution and ReLU activation function to obtain feature L / ∈R 1024×64×64 Next, feature L / Divide it into four features L1∈R according to the channel order 256×64×64 , L2∈R 256×64×64 , L3∈R 256×64×64 and L4∈R 256×64×64 Then, L1 and L3 are superimposed to obtain L 1-3 ∈R 512×64×64 , L2 and L4 are superimposed to obtain L 2-4 ∈R 512×64×64 , features L1 and L4 are superimposed to obtain L 1-4 ∈R 512×64×64 , L2 and L3 are superimposed to obtain L 2-3 ∈R 512×64×64 Next, feature L 1-3 , L 2-4 , L 1-4 and L 2-3 Channel attention will be performed separately to obtain enhanced features and and Finally, the features and According to the channel superposition, the feature U∈R is obtained 512×64×64 ,feature and According to the channel superposition, we get the feature V∈R 512×64×64 , feature U and feature V are superimposed to obtain the final output feature O∈R 1024×64×64 .

[0125] Step S5: Multi-feature cross fusion. The enhanced global feature is T∈R 4096×16×16 And the enhanced local feature is O∈R 1024×64×64 will be respectively and the information prior features are P∈R4096×16×16 Take the cross fusion of global features and text prior features as an example. First, the global feature T is obtained through two 1x1 convolutions to obtain the feature T1∈R 4096×16×16 and T2∈R 4096×16×16 . Then, feature T2 will change its dimension to HxWxC and perform matrix multiplication with feature P to obtain the spatial attention relationship map map = (16×16)×(16×16). The spatial attention heat map map will be weighted activated by the softmax function. Next, the feature obtained by matrix multiplication of feature T1 and the spatial attention relationship heat map is added pixel by pixel to obtain the output The cross-fusion process of the enhanced local features and text prior features is consistent with the above description, and the output features are obtained However, before the local feature O is cross-fused, it needs to be adjusted to the same dimension as P through convolution and upsampling. and Features The final encoding feature Y is obtained by adding pixel by pixel. The feature Y obtained in the above steps will pass through an atrous spatial convolutional pooling pyramid (ASPP) module to obtain the final encoding feature Q∈R 4096×16×16 .

[0126] Step S6: Based on the current pre-trained natural language model, decode the features fused in step 5 to generate the final natural language description.

[0127] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0128] In addition, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include at least one of the features.

[0129] Any process or method description described in a flowchart or otherwise in this specification may be understood to represent a module, fragment or portion of code that includes one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0130] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0131] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating natural language description of remote sensing images based on prior information of complex scenes, characterized in that: The following steps are involved: Step 1: Build a joint Chinese and English remote sensing image natural language description dataset; Step 2: Use the backbone network of the pre-trained image recognition neural network to extract shallow coding features of multiple scales and deep coding features of multiple scales of the input image; The shallow coding features of multiple scales are integrated to obtain the global visual features, and the deep coding features of multiple scales are integrated to obtain the local visual features. Step 3, constructing a module based on prior information to extract text prior features from visual global features; Step 4, respectively enhancing the visual global features and the visual local features to obtain enhanced visual global features and visual local features; Step 5, feature fusion of the text prior features and the enhanced visual global features is performed through a multi-feature cross-fuser to obtain global fusion features, and feature fusion of the text prior features and the enhanced visual local features is performed to obtain local fusion features; and feature fusion of the global fusion features and the local fusion features is performed to obtain the final coding features; Step 6: Decode the final encoded features based on the pre-trained natural language model to generate a natural language description of the remote sensing image.

2. The method according to claim 1, characterized in that Step 1 includes: obtaining a public natural language description dataset of remote sensing images as an initial dataset, and obtaining an English dataset based on the English descriptions of the remote sensing images in the initial dataset; based on the English dataset, re-annotating with matching Chinese, and constructing a corresponding Chinese dataset based on the Chinese annotations.

3. The method according to claim 1, characterized in that In step 2, the pre-trained backbone network is a Transformer-based network structure or a convolutional neural network-based network structure.

4. The method according to claim 1, characterized in that Step 2 also includes fine-tuning and learning some network parameters in the backbone network, wherein the fine-tuned and learned part of the network parameters are network parameters of the network layer used to extract deep coding features.

5. The method according to claim 1, characterized in that In step 2, for shallow coding features of multiple scales, a layer-by-layer fusion method is used to fuse multi-scale shallow coding features. According to the forward propagation direction of the backbone network, the shallow / deep coding features of each scale are traversed in turn, and the feature dimension of the shallow / deep coding feature of the current scale is adjusted to the dimension of the next shallow / deep coding feature through convolution and downsampling operations, and then the features are superimposed. The current superimposed shallow / deep coding feature is used as the new next shallow / deep coding feature, and the feature adjustment and superposition of the new next shallow / deep coding feature and the next shallow / deep coding feature are continued until the shallow / deep coding feature of the last scale is obtained.

6. The method according to claim 1, characterized in that In step 3, the processing of the prior information construction module includes: Step 3-1: Define the global visual feature as G∈R C×H×W ; The visual global feature G is processed by global maximum pooling and global average pooling respectively to obtain global vectors g1 and g2; Then the global vectors g1 and g2 are respectively passed through a shared weight parameter and pre-trained multi-attribute predictor to obtain the attribute prediction probability p1∈R of the scene attribute of the remote sensing image d and p2∈R d , where d represents the number of scene attributes of the remote sensing image; Step 3-2: Sort the elements in the comprehensive multi-attribute prediction probabilities p1 and p2, select the top K largest image attribute prediction probabilities, and convert them into one-hot vector representations to obtain the matrix U∈R d×k ; Step 3-3: Construct a learnable matrix M∈R t×d , so that it and the matrix U∈R d×k Perform matrix multiplication to obtain a new matrix E∈R t×k , where t is the set coding length; Step 3-4: Expand the matrix E and then replicate it in space to obtain the feature map E f ∈R B×H×W ; Among them, the number of feature map channels B = t × k; The feature map E f Through a convolution layer with a convolution kernel of 1×1, the dimension is adjusted to obtain the text prior feature P∈R C×H×W . Informative prior features can provide more information about scene categories, thus playing an auxiliary role in understanding the semantic and spatial relationships of ground objects in complex scenes.

7. The method according to claim 1, characterized in that In step 4, the visual global features are enhanced by the set global feature enhancement module, including: Step 4-1-A: Define the global visual feature as G∈R C×H×W , where C is the number of channels and H×W is the height and width of the feature; The global feature G is convolved by parallel depth-wise separable convolutions with kernel sizes of 3x3, 5x5, and 7x7 respectively; The results of the three-way convolution operation are superimposed according to the channel dimension, and then the number of feature channels is adjusted to be consistent with the visual global feature G through a convolution layer with a convolution kernel of 1x1, and the feature S∈R is obtained. C×H×W ; Split the feature S into two new features S1∈R according to the channel C / 2×H×W and S2∈R C / 2×H×W ; The feature S1 is pooled horizontally and globally to obtain the feature S _h ∈R C / 2×H×1 ; Feature S2 is pooled vertically and globally to obtain feature S _w ∈R C / 2×1×W ; For feature S _h Multiply by point with S1 to get feature T1∈R C / 2×H×W , for S _w Multiply by point with S2 to get feature T2∈R C / 2×H×W ; Superimpose features T1 and T2 to obtain the enhanced visual global feature T∈R C×H×W .

8. The method according to claim 1, characterized in that In step 4, the local visual features are enhanced by the local feature enhancement module, including: Step 4-2-A: Define the local visual feature as L∈R C×H×W , where C is the number of channels and H×W is the height and width of the feature; The local visual feature L is passed through the convolution layer with a convolution kernel of 3x3 and the ReLU activation function to obtain the feature L / ∈R C ×H×W ; Step 4-2-B: Set feature L / ∈R C×H×W Divide into four features L1∈R according to the channel order C / 4×H×W , L2∈R C / 4×H×W , L3∈R C / 4×H×W and L4∈R C / 4×H×W ; Features L1 and L3 are superimposed to obtain L 1-3 ∈R C / 2×H×W , features L2 and L4 are superimposed to obtain L 2-4 ∈R C / 2×H×W , features L1 and L4 are superimposed to obtain L 1-4 ∈R C / 2×H×W , features L2 and L3 are superimposed to obtain L 2-3 ∈R C / 2×H×W ; Step 4-2-C: For feature L 1-3 , L 2-4 , L 1-4 and L 2-3 Channel attention will be performed separately to obtain enhanced features and and Step 4-2-D: Enhanced features and According to the channel superposition, the feature U∈R is obtained C×H×W , the enhanced features and According to the channel superposition, we get the feature V∈R C×H×W ; Then add the features U and V pixel by pixel to get the enhanced visual local feature O∈R C×H×W .

9. The method according to claim 1, characterized in that The specific execution process of the multi-feature cross-fusion device in step 5 includes: Set the text prior feature to P∈R C×H×W , the enhanced visual global feature is T∈R C×H×W , the enhanced local visual feature is O∈R C×H×W ; Perform a joint spatial attention operation on the enhanced visual global feature T and the text prior feature P to obtain the feature Perform a joint spatial attention operation on the enhanced local visual feature O and the text prior feature P to obtain the feature feature and Features According to the channel dimension superposition, the first fusion coding feature Y∈R is obtained by sequentially passing through the convolution layer with a convolution kernel of 3x3 and the activation function C×H×W ; The first fused encoding feature Y is passed through a dilated spatial convolutional pooling pyramid module to obtain the final encoding feature Q∈R C×H×W .

10. The method according to claim 9, characterized in that Perform a joint spatial attention operation on the enhanced visual global feature T and the text prior feature P to obtain the feature include: The visual global feature T is obtained through two convolution layers with a convolution kernel of 1×1 to obtain features T1 and T2; Adjust the dimension of feature T2 to H×W×C, and then perform matrix multiplication with feature P to obtain the spatial attention relationship map map1 = (H×W)×(H×W); and use the softmax function to weight the spatial attention relationship map map1; Then add feature T1 and the weight-activated spatial attention relationship map map2 pixel by pixel to get the feature Perform a joint spatial attention operation on the enhanced local visual feature O and the text prior feature P to obtain the feature The local visual feature O is passed through two convolution layers with a convolution kernel of 1×1 to obtain features O1 and O2; Adjust the dimension of O2 to H×W×C, and then perform matrix multiplication with feature P to obtain the spatial attention relationship map map2 = (H×W)×(H×W), and perform weight activation on the spatial attention relationship map map2 through the softmax function; Then add feature O1 and the weight-activated spatial attention relationship map map2 pixel by pixel to get the feature

Citation Information

Patent Citations

  • Method and system for generating attention remote sensing image description based on high-low layer feature fusion

    CN111860235A

  • Image description method and device, equipment and medium

    CN115186061A

  • Method for fusing self-adaptive localized language model and visual generation technology

    CN116958772A

  • Remote sensing image change detection method based on language guidance

    CN119169449A