Semantic alignment method and device of remote sensing image, storage medium, electronic equipment and computer program product

Through feature extraction embedding module, cross-modal feature alignment module and multi-scale semantic alignment module, the fine-grained understanding of large scene aerial images in remote sensing image processing is solved, and the precise alignment and feature fusion of remote sensing images and text description data is realized, which improves the accuracy and accuracy of semantic alignment.

CN120544199APending Publication Date: 2025-08-26启元实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510626302.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing remote sensing image processing technologies cannot effectively process the fine-grained understanding of large-scene aerial images, especially in the cross-modal text and remote sensing image matching tasks, which cannot accurately align visual features and text semantic vectors, and lack an understanding of the spatial relationship between multiple targets.

Method used

The feature extraction embedding module, a cross-modal feature alignment module and a multi-scale semantic alignment module are adopted to achieve accurate information alignment and feature fusion of remote sensing images and text description data through multi-scale feature extraction and cross-modal feature alignment, and the self-attention mechanism and cross-modal attention mechanism are used to achieve accurate information alignment and feature fusion of remote sensing images and text description data.

Benefits of technology

It realizes the precise positioning of mesoscale semantic entities in large-scene aerial images, overcomes the limitations of traditional methods, can handle multi-hop references and text descriptions of complex spatial relationships, and improves the accuracy and accuracy of semantic alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544199A_ABST
    Figure CN120544199A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic alignment method and device of a remote sensing image, a storage medium, electronic equipment and a computer program product. The semantic alignment method comprises the following steps: responding to a remote sensing image input by a user, and determining a first image embedding vector according to the remote sensing image; responding to text description data input by a user, and determining a first text embedding vector according to the text description data; performing cross-modal feature alignment processing on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector; determining a correlation output graph of the text description data and the remote sensing image according to the second text embedding vector and the second image embedding vector; and determining an area corresponding to the text description data in the remote sensing image according to the correlation output graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of semantic alignment of remote sensing images, and in particular to a method, device, storage medium, electronic device, and computer program product for semantic alignment of remote sensing images. Background Art

[0002] With the continuous improvement of remote sensing image resolution and dataset size, existing image-level description, object detection and segmentation technologies are gradually unable to meet the requirements of fine-grained understanding of large-scene aerial images.

[0003] For example, in the task of matching cross-modal text with remote sensing images, existing methods, when dealing with text descriptions with complex semantic references, typically only obtain the general semantics of the image by matching the entire image with the text, but are unable to accurately locate large-scale semantic entities in the image. When the text description involves multiple objects or complex spatial relationships, there is a problem of being unable to accurately align visual features and text semantic vectors.

[0004] Moreover, existing object-level detection and segmentation methods mainly focus on identifying and segmenting single objects. Although they can identify targets such as buildings and roads, they lack the understanding of the spatial relationships between multiple targets, and therefore have limited ability to understand mesoscale areas in large-scale scenes.

[0005] The contents of the background technology are merely technologies known to the public and do not necessarily represent existing technologies in this field. Summary of the Invention

[0006] According to one aspect of the present application, the present application provides a semantic alignment method for remote sensing images, and the semantic alignment method includes: responding to a remote sensing image input by a user to determine a first image embedding vector based on the remote sensing image; responding to text description data input by the user to determine a first text embedding vector based on the text description data; performing cross-modal feature alignment processing on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector; determining a correlation output map of the text description data and the remote sensing image based on the second text embedding vector and the second image embedding vector; and determining an area in the remote sensing image corresponding to the text description data based on the correlation output map.

[0007] According to some embodiments of the present application, determining a first image embedding vector based on a remote sensing image includes: determining a standard remote sensing image based on the remote sensing image; and determining the first image embedding vector based on the standard remote sensing image. Determining a first text embedding vector based on text description data includes: determining standard text description data based on the text description data; and determining the first text embedding vector based on the standard text description data.

[0008] According to some embodiments of the present application, determining a correlation output graph of text description data and remote sensing images based on a second text embedding vector and a second image embedding vector includes: determining a text semantic vector based on the second text embedding vector; performing reshape processing on the second image embedding vector to obtain a third image embedding vector with a preset scale; determining a correlation graph between the text semantic vector and the third image embedding vector; and determining a correlation output graph based on the correlation graph.

[0009] According to some embodiments of the present application, determining a correlation graph between a text semantic vector and a third image embedding vector includes: determining a first pixel of a first scale of the third image embedding vector; determining a first correlation between the first pixel and the text semantic vector; traversing all pixels of all scales of the third image embedding vector to obtain a correlation set between all pixels of all scales and the text embedding vector; and determining a correlation graph based on the correlation set.

[0010] According to some embodiments of the present application, after determining the area in the remote sensing image corresponding to the text description data based on the correlation output map, the semantic alignment method also includes: determining a multi-scale semantic alignment image based on the correlation output map and the remote sensing image; and determining the area in the remote sensing image corresponding to the text description data based on the multi-scale semantic alignment image.

[0011] According to one aspect of the present application, the present application provides a semantic alignment device for remote sensing images, characterized in that the semantic alignment device includes a feature extraction and embedding module, a cross-modal feature alignment module, a multi-scale semantic alignment module, and an output module. The feature extraction and embedding module responds to the remote sensing image input by the user to determine a first image embedding vector based on the remote sensing image; the feature extraction and embedding module responds to the text description data input by the user to determine a first text embedding vector based on the text description data; the cross-modal feature alignment module performs cross-modal feature alignment processing on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector; the multi-scale semantic alignment module determines a correlation output map of the text description data and the remote sensing image based on the second text embedding vector and the second image embedding vector; the output module determines the area in the remote sensing image corresponding to the text description data based on the correlation output map.

[0012] According to some embodiments of the present application, a multi-scale semantic alignment module includes a preprocessing unit, a relevance assignment unit, and a multi-scale integration unit. The preprocessing unit determines a text semantic vector based on the second text embedding vector, and reshapes the second image embedding vector to obtain a third image embedding vector; the relevance assignment unit determines a relevance graph between the text semantic vector and the third image embedding vector; and the multi-scale integration unit determines a relevance output graph based on the relevance graph.

[0013] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the semantic alignment method of remote sensing images as described above.

[0014] According to another aspect of the present application, the present application also provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when the one or more programs are executed by one or more processors, enables the one or more processors to implement the semantic alignment method of remote sensing images as described above.

[0015] According to another aspect of the present application, the present application also provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by the computer, the computer executes the semantic alignment method of remote sensing images as described above.

[0016] Beneficial effects

[0017] The present application can determine a first text embedding vector from text description data, and can determine a first image embedding vector from a remote sensing image. The present application can obtain a second text embedding vector and a second image embedding vector by performing a cross-modal attention mechanism on the first text embedding vector and the first image embedding vector. The present application can determine a correlation output graph between the text description data and the remote sensing image using the second text embedding vector and the second image embedding vector. The present application can determine the area in the remote sensing image corresponding to the text description data using the correlation output graph.

[0018] The semantic alignment method provided in this application can be used for medium-scale semantic entities (such as buildings, industrial areas, urban green spaces, etc.) in large-scene aerial images through multi-scale feature extraction and cross-modal feature alignment, overcoming the limitation of traditional methods that can only process object-level targets.

[0019] This application can achieve more accurate information alignment and feature fusion between remote sensing images and text description data through the self-attention mechanism and cross-modal attention mechanism in cross-modal feature alignment, providing technical support for cross-modal tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A schematic structural diagram of a semantic alignment device according to an embodiment of the present application is shown;

[0022] Figure 2 A schematic structural diagram of a multi-scale semantic alignment module according to an embodiment of the present application is shown;

[0023] Figure 3 A schematic structural diagram of an output module according to an embodiment of the present application is shown;

[0024] Figure 4 A flowchart of a semantic alignment method 1000 according to an embodiment of the present application is shown;

[0025] Figure 5 A schematic diagram of a process of step S100 according to an embodiment of the present application is shown;

[0026] Figure 6 A schematic diagram of a process of step S200 according to an embodiment of the present application is shown;

[0027] Figure 7 A schematic diagram of a process of step S400 according to an embodiment of the present application is shown;

[0028] Figure 8 A schematic diagram of a process of step S430 according to an embodiment of the present application is shown;

[0029] Figure 9 Another flowchart of the semantic alignment method 1000 according to an embodiment of the present application is shown.

[0030] Description of reference numerals:

[0031] Semantic alignment device 200.

[0032] Feature extraction and embedding module 210; cross-modal feature alignment module 220; multi-scale semantic alignment module 230; output module 240.

[0033] Preprocessing unit 231; correlation allocation unit 232; multi-scale integration unit 233. DETAILED DESCRIPTION

[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. Like reference numerals in the drawings represent like or similar parts, and thus repetitive description thereof will be omitted.

[0035] The described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of these specific details, or other methods, components, materials, devices, etc. may be employed. In these cases, well-known structures, methods, devices, implementations, materials or operations will not be shown or described in detail.

[0036] Furthermore, the terms "include," "comprise," and "have," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0037] The terms "first", "second" and the like in the specification, claims and drawings of this application are used to distinguish different objects rather than to describe a specific order.

[0038] The following is a clear and complete description of the technical solution of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of this application.

[0039] The English terms and their full English names and corresponding Chinese meanings involved in this application are:

[0040] NLP: Natural Language Processing, natural language processing;

[0041] SWIN: SWIN Vision Transformer, SWIN transformer;

[0042] VIT: Vision Transformer, visual transformer;

[0043] BERT: Bidirectional Encoder Representations from Transformers, based on Transformer’s bidirectional encoder representation;

[0044] GPT: Generative Pretrained Transformer, generative pretrained transformer;

[0045] CNN: Convolutional Neural Network, convolutional neural network.

[0046] Existing image-level description tasks and image retrieval methods often rely on matching the entire image with text description data to obtain the general semantics of the image. Such methods can achieve good results when processing small-scale images or single objects, but when faced with large-scale aerial imagery, image-level tasks often fail to meet the following requirements:

[0047] (1) Lack of fine-grained localization capabilities: Existing image-level description methods are usually based on matching the overall image with text descriptions. They tend to ignore the presence of mesoscale semantic entities in the image (e.g., building complexes, urban areas, industrial zones, etc.). These mesoscale semantic entities are usually composed of multiple targets, and their spatial layout and environmental background information are crucial for understanding remote sensing images. Existing image-level tasks cannot effectively extract the detailed information of these areas and cannot accurately perform semantic region localization.

[0048] (2) Insufficient understanding of complex spatial relationships: Existing image-level methods rely more on a macroscopic understanding of the image, but they often ignore the spatial relationships between regions in the image, especially when the description involves multiple objects or there are complex spatial relationships between objects. For example, the description may involve "the park next to the building complex." Understanding such spatial relationships is crucial for accurate semantic localization in the image, but existing image-level methods cannot accurately handle these complex contexts.

[0049] (3) Low cross-modal matching accuracy: Existing image-level description and text matching methods are generally unable to accurately align the visual features of the image with the complex semantic information in the text. For example, in remote sensing images, the text description data contains relatively complex geographic information and the relationship between objects, and existing image-level description methods cannot effectively map this semantic information to the image, resulting in insufficient semantic alignment accuracy.

[0050] According to one aspect of the present application, the present application provides a remote sensing image semantic alignment device 200. Figure 1 The semantic alignment device 200 includes a feature extraction and embedding module 210, a cross-modal feature alignment module 220, a multi-scale semantic alignment module 230, and an output module 240. Figure 2 The multi-scale semantic alignment module 230 includes a pre-processing unit 231 , a correlation allocation unit 232 and a multi-scale integration unit 233 .

[0051] The following combination Figure 1 、 Figure 2 and Figure 3 The present invention describes a method for semantic alignment of remote sensing images.

[0052] See also Figure 4 , the semantic alignment method 1000 may include steps S100 to S500.

[0053] In step S100 , the feature extraction and embedding module 210 responds to the remote sensing image input by the user to determine a first image embedding vector according to the remote sensing image.

[0054] According to example embodiments, the remote sensing image may be image data for which the semantic alignment task is to be performed. The remote sensing image may include a multispectral image or a high-resolution remote sensing image. The first image embedding vector may be an image vector obtained by performing feature extraction on the remote sensing image. The first image embedding vector may be a multi-scale vector.

[0055] The feature pre-embedding module can process the remote sensing image into a standard remote sensing image by a standard method. The feature extraction and embedding module 210 can extract the standard remote sensing image by a multi-scale feature extractor, thereby obtaining a multi-scale first image embedding vector.

[0056] In step S200 , the feature extraction and embedding module 210 responds to the text description data input by the user to determine a first text embedding vector according to the text description data.

[0057] According to an example embodiment, the text description data may be task language description data corresponding to the remote sensing image. For example, the text description data may be “find the orange buildings around the parking lot in the remote sensing image”.

[0058] The feature pre-embedding module can process the text description data into standard text description data through natural language processing (NLP). The feature extraction and embedding module 210 can extract the standard text description data through a pre-trained semantic model to obtain a first text embedding vector.

[0059] In step S300 , the cross-modal feature alignment module 220 performs cross-modal feature alignment processing on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector.

[0060] According to an example embodiment, cross-modal feature alignment may be a processing mechanism for deeply fusing and aligning remote sensing images and text description data. Cross-modal feature alignment may include a self-attention mechanism and a cross-modal attention mechanism.

[0061] The second text embedding vector may be a text embedding vector obtained by processing the first text embedding vector through the self-attention mechanism and the cross-modal attention mechanism. The second image embedding vector may be an image embedding vector obtained by processing the first image embedding vector through the self-attention mechanism and the cross-modal attention mechanism.

[0062] The cross-modal feature alignment module 220 can perform a self-attention calculation on the first text embedding vector and the first image embedding vector. The cross-modal feature alignment module 220 performs a cross-modal attention process on the first text embedding vector and the first image embedding vector calculated by the self-attention mechanism to obtain a second text embedding vector and a second image embedding vector.

[0063] For example, the cross-modal feature alignment module 220 performs a cross-modal attention mechanism on each text vector in the first text embedding vector calculated by the self-attention mechanism and the visual feature vector of each pixel in the first image embedding vector calculated by the self-attention mechanism to obtain a second text embedding vector and a second image embedding vector.

[0064] The second text embedding vector is a text embedding vector obtained by fusion enhancement of the first text embedding vector. The second image embedding vector is an image embedding vector obtained by fusion enhancement of the first image embedding vector.

[0065] In step S400 , the multi-scale semantic alignment module 230 determines a correlation output graph between the text description data and the remote sensing image based on the second text embedding vector and the second image embedding vector.

[0066] According to an example embodiment, the correlation may be data on the degree of association between the text description data and each pixel in the remote sensing image. The correlation output graph may be a visualization graph of a scale determined according to all the correlations.

[0067] The multi-scale semantic alignment module 230 may determine a text semantic vector based on the second text embedding vector. The multi-scale semantic alignment module 230 may reshape the second image embedding vector to obtain a third image embedding vector having a preset scale. The multi-scale semantic alignment module 230 may determine a correlation graph between the text semantic vector and the third image embedding vector. The multi-scale semantic alignment module 230 may determine a correlation output graph based on the correlation graph.

[0068] The correlation output graph can use different colors or brightness levels to represent the correlation values ​​between the text description data and the pixels in the remote sensing image data. The correlation output graph shows the degree of match between the text description data and each pixel in the remote sensing image, thereby achieving pixel-level precision in locating the target area in the remote sensing image.

[0069] In step S500 , the output module 240 determines the area in the remote sensing image corresponding to the text description data based on the correlation output map.

[0070] According to an exemplary embodiment, the output module 240 may determine the degree of match between the text description data and the pixels in the remote sensing image based on the different colors of the correlation output graph (i.e., the magnitude of the correlation value). The closer the correlation value is to 1, the higher the degree of match between the text description data and the pixels in the remote sensing image, which can be understood as indicating that the pixels in the remote sensing image conform to the text description in the text description data.

[0071] The output module 240 can accurately identify pixels in the remote sensing image that conform to the text description data by analyzing and processing the correlation output graph, thereby determining the area in the remote sensing image that conforms to the text description data.

[0072] Through the above embodiments, the present application can determine a first text embedding vector from text description data, and the present application can determine a first image embedding vector based on a remote sensing image. The present application can obtain a second text embedding vector and a second image embedding vector by performing a cross-modal attention mechanism on the first text embedding vector and the first image embedding vector. The present application can determine a correlation output graph between the text description data and the remote sensing image using the second text embedding vector and the second image embedding vector. The present application can determine the area in the remote sensing image corresponding to the text description data using the correlation output graph.

[0073] The semantic alignment method provided in this application can be used for medium-scale semantic entities (such as buildings, industrial areas, urban green spaces, etc.) in large-scene aerial images through multi-scale feature extraction and cross-modal feature alignment, overcoming the limitation of traditional methods that can only process object-level targets.

[0074] This application can achieve more accurate information alignment and feature fusion between remote sensing images and text description data through the self-attention mechanism and cross-modal attention mechanism in cross-modal feature alignment, providing technical support for cross-modal tasks.

[0075] This application can effectively process text descriptions of multi-hop references and complex spatial relationships through cross-modal feature alignment, thereby solving the difficulties of existing methods in processing complex contexts containing relative positions, environmental backgrounds and multi-target descriptions.

[0076] Alternatively, see Figure 5 , step S100 may include step S110 and step S120.

[0077] In step S110 , the feature extraction and embedding module 210 determines a standard remote sensing image based on the remote sensing image.

[0078] According to an example embodiment, the standard remote sensing image may be a remote sensing image that has undergone standard processing. The feature extraction and embedding module 210 may process the remote sensing image through a labeling process (e.g., remote sensing image loading and standardization method) to obtain a standard remote sensing image with consistency (e.g., consistency in resolution and number of channels).

[0079] In step S120 , the feature extraction and embedding module 210 determines a first image embedding vector based on the standard remote sensing image.

[0080] The feature extraction and embedding module 210 can perform multi-scale processing on the standard remote sensing image through an attention mechanism network (e.g., SWIN and VIT). The attention mechanism network decomposes the standard remote sensing image into feature maps of different resolutions. The low-scale feature maps can capture the detailed information in the standard remote sensing image. The high-scale feature maps can capture the global background in the standard remote sensing image.

[0081] The feature extraction and embedding module 210 can use the attention mechanism network to learn the visual features of the standard remote sensing image at different scales and generate the corresponding first image embedding vector, ensuring that the information of the standard remote sensing image at multiple scales can be fully utilized.

[0082] See also Figure 6 , step S200 may include step S210 and step S220.

[0083] In step S210 , the feature extraction and embedding module 210 determines standard text description data according to the text description data.

[0084] According to an example embodiment, the standard text description data may be text description data that has been standardized. The text description data may also be descriptive text of a remote sensing image. The text description data may include a simple target description (such as "building complex") or a more complex context description (such as "park next to the building complex"). The feature extraction and embedding module 210 may obtain the standard text description data through NLP processing. NLP processing may include steps such as word embedding, word segmentation, and syntactic analysis.

[0085] In step S220 , the feature extraction and embedding module 210 determines a first text embedding vector based on the standard text description data.

[0086] The feature extraction and embedding module 210 can convert the standard text description data into a first text embedding vector through a pre-trained semantic model (for example, BERT and GPT, etc.).

[0087] Through the above embodiments, the present application can determine a standard remote sensing image from remote sensing image data, and can determine a first image embedding vector from a standard remote sensing image. The present application can determine standard text description data from text description data, and can determine a first text embedding vector from the standard text description data.

[0088] The semantic alignment method provided in the present application can determine the first image embedding vector and the first text embedding vector through annotation processing and feature extraction, so that the present application can effectively extract the features of the remote sensing image and the features of the text description data.

[0089] Alternatively, see Figure 7 , step S400 may include steps S410 to S440.

[0090] In step S410 , the preprocessing unit 231 determines a text semantic vector based on the second text embedding vector.

[0091] According to an example embodiment, the text semantic vector may be a text vector obtained by unifying the second text embedding vector. For example, see Figure 2 The pre-processing unit 231 may average the second text embedding vectors to generate a unified text semantic vector. The text semantic vector can carry semantic information from the text description data and the remote sensing image in the context of the remote sensing image.

[0092] In step S420 , the pre-processing unit 231 performs a reshape process on the second image embedding vector to obtain a third image embedding vector with a preset scale.

[0093] According to an example embodiment, the reshaping may be a process of re-adjusting and changing the dimension and structure of the second image embedding vector. The third image embedding vector may be an image embedding vector after the second image embedding vector has been reshaped. The preset scale may be a preset scale of the third image embedding vector. The preset scale may be a plurality of different scales. For example, see Figure 2 The pre-processing unit 231 may perform a reshape process on the second image embedding vector to obtain a third image embedding vector with different preset scales (preset scale 1, preset scale 2, ... preset scale L).

[0094] In step S430 , the correlation assigning unit 232 determines a correlation graph of the text semantic vector and the third image embedding vector.

[0095] According to an example embodiment, the correlation graph may be a multi-scale visualization graph determined based on correlations. The correlation assignment unit 232 may calculate the correlations between all pixels at all scales of the third image embedding vector and the text semantic vector, thereby obtaining a multi-scale correlation set. The text semantic vector may serve as a query vector, and all pixels at all scales of the third image embedding vector may serve as a key vector. The correlation assignment unit 232 may perform format conversion, storage, and export processing on the data of the multi-scale correlation set, thereby obtaining a multi-scale correlation graph.

[0096] For example, see Figure 2 , the correlation allocation unit 232 can determine a correlation graph between the third image embedding vector at the preset scale 1 and the text semantic vector at the preset scale 1. The correlation allocation unit 232 can also determine a correlation graph between the third image embedding vector at the preset scale 2 and the text semantic vector at the preset scale 2. Similarly, the correlation allocation unit 232 can determine a correlation graph between the third image embedding vector at the preset scale L and the text semantic vector at the preset scale L.

[0097] Alternatively, see Figure 8 , step S430 may include steps S431 to S434.

[0098] In step S431 , the correlation assigning unit 232 determines a first pixel of a first scale of the third image embedding vector.

[0099] According to an example embodiment, the first scale may be a scale level of the third image embedding vector. The first pixel may be a pixel in the first scale of the third image embedding vector. The correlation allocation unit 232 may determine the first pixel by pixel coordinates.

[0100] In step S432 , the correlation assigning unit 232 determines a first correlation between the first pixel and the text semantic vector.

[0101] According to an example embodiment, the first correlation may be association degree data between the first pixel and the text semantic vector. The correlation allocating unit 232 may calculate the first correlation by using a cosine similarity method, a CNN network, or the like.

[0102] For example, the correlation assignment unit 232 can calculate the first correlation using a cosine similarity method. The correlation assignment unit 232 can normalize the third image embedding vector and the text semantic vector so that the norm of each third image embedding vector and text semantic vector is 1, thereby avoiding calculation bias caused by scale differences or different data distributions. The correlation assignment unit 232 then calculates the first correlation (i.e., cosine similarity) between the first pixel and the text semantic vector.

[0103] In step S433 , the correlation assigning unit 232 traverses all pixels of all scales of the third image embedding vector to obtain a correlation set between all pixels of all scales and the text embedding vector.

[0104] According to an example embodiment, the correlation set may be a set of correlations between all pixels at all preset scales of the third image embedding vector and the text semantic vector. The correlation assignment unit 232 may traverse all pixels at all scales of the third image embedding vector using a cosine similarity method, a CNN network, or the like, to obtain a multi-scale correlation set between all pixels at all scales and the text embedding vector.

[0105] In step S434 , the correlation allocation unit 232 determines a correlation graph based on the correlation set.

[0106] According to an example embodiment, the correlation allocation unit 232 may perform format conversion, storage, and export processing on the data of the multi-scale correlation set, thereby obtaining a multi-scale correlation graph.

[0107] In step S440 , the multi-scale integration unit 233 determines a correlation output graph according to the correlation graph.

[0108] According to an example embodiment, the multi-scale integration unit 233 may perform weighted summation on the correlation graphs from different scales and average the correlation graphs from different scales to generate a correlation output graph of the same scale, thereby better fusing the features of the correlation graphs from different scales.

[0109] For example, see Figure 2 The multi-scale integration unit 233 may perform a weighted summation of the correlation maps from different scales (preset scale 1, preset scale 2, ..., preset scale L) and average the correlation maps from different scales (preset scale 1, preset scale 2, ..., preset scale L) to generate a correlation output map at the same scale. The scale of the correlation output map may be the same as that of the remote sensing image.

[0110] Through the above embodiments, the present application can determine a text semantic vector based on the second text embedding vector, and the present application can reshape the second image embedding vector to obtain a third image embedding vector with a preset scale. The present application can determine a correlation output graph by determining a correlation graph between the text semantic vector and the third image embedding vector.

[0111] The semantic alignment method provided in this application can achieve multi-scale hierarchical semantic alignment of remote sensing images and text description data.

[0112] Alternatively, see Figure 9, the semantic alignment method 1000 may further include step S600 and step S700.

[0113] In step S600 , the output module 240 determines a multi-scale semantically aligned image according to the correlation output map and the remote sensing image.

[0114] According to an example embodiment, the multi-scale semantic alignment image may be an image obtained by superimposing the correlation output map and the remote sensing image. Figure 3 The output module 240 can generate a multi-scale semantic alignment image by overlaying the correlation output map and the remote sensing image.

[0115] The multi-scale semantic alignment image includes the matching value of the correlation output map and the original image data in the remote sensing image.

[0116] In step S700 , the output module 240 determines the area in the remote sensing image corresponding to the text description data based on the multi-scale semantic alignment image.

[0117] According to an example embodiment, the output module 240 can accurately identify pixels in the remote sensing image that conform to the text description data by analyzing and processing the multi-scale semantically aligned image, thereby more clearly displaying the area in the remote sensing image that conforms to the text description data.

[0118] Through the above embodiments, the present application can determine a multi-scale semantically aligned image through a correlation output graph and a remote sensing image, and the present application can determine an area in a remote sensing image corresponding to text description data through the multi-scale semantically aligned image.

[0119] The semantic alignment method provided in this application can more clearly display the areas in the remote sensing image that conform to the text description data through multi-scale semantic alignment of images.

[0120] According to another aspect of the present application, the present application further provides a non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the semantic alignment method of remote sensing images as described above.

[0121] According to another aspect of the present application, the present application also provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when the one or more programs are executed by one or more processors, enables the one or more processors to implement the semantic alignment method of remote sensing images as described above.

[0122] According to another aspect of the present application, the present application also provides a computer program product, including: a computer program stored on a computer-readable storage medium; the computer program includes program instructions, and when the program instructions are executed by the computer, the computer executes the semantic alignment method of remote sensing images as described above.

[0123] Finally, it should be noted that the above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Although the present application is described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions of the aforementioned embodiments or replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A semantic alignment method for remote sensing images, characterized in that: The semantic alignment method comprises: In response to a remote sensing image input by a user, determining a first image embedding vector based on the remote sensing image; In response to text description data input by a user, determining a first text embedding vector according to the text description data; Performing cross-modal feature alignment on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector; Determining a correlation output graph between the text description data and the remote sensing image according to the second text embedding vector and the second image embedding vector; According to the correlation output graph, an area in the remote sensing image corresponding to the text description data is determined.

2. The semantic alignment method according to claim 1, characterized in that Determining a first image embedding vector according to the remote sensing image includes: determining a standard remote sensing image based on the remote sensing image; Determining the first image embedding vector according to the standard remote sensing image; Determining a first text embedding vector according to the text description data includes: Determining standard text description data according to the text description data; The first text embedding vector is determined according to the standard text description data.

3. The semantic alignment method according to claim 1, characterized in that Determining a correlation output graph between the text description data and the remote sensing image based on the second text embedding vector and the second image embedding vector includes: Determine a text semantic vector based on the second text embedding vector; performing a reshape process on the second image embedding vector to obtain a third image embedding vector having a preset scale; determining a correlation graph between the text semantic vector and the third image embedding vector; The correlation output graph is determined according to the correlation graph.

4. The semantic alignment method according to claim 3, characterized in that Determining a correlation graph between the text semantic vector and the third image embedding vector includes: determining a first pixel of a first scale of the third image embedding vector; determining a first correlation between the first pixel and the text semantic vector; Traversing all pixels of all scales of the third image embedding vector to obtain a correlation set between all pixels of all scales and the text embedding vector; The dependency graph is determined based on the dependency set.

5. The semantic alignment method according to claim 1, characterized in that After determining the area in the remote sensing image corresponding to the text description data according to the correlation output graph, the semantic alignment method further includes: Determine a multi-scale semantic alignment image according to the correlation output map and the remote sensing image; According to the multi-scale semantically aligned image, a region in the remote sensing image corresponding to the text description data is determined.

6. A semantic alignment device for remote sensing images, characterized in that: The semantic alignment device comprises: A feature extraction and embedding module is responsive to a remote sensing image input by a user to determine a first image embedding vector based on the remote sensing image; the feature extraction and embedding module is responsive to text description data input by a user to determine a first text embedding vector based on the text description data; a cross-modal feature alignment module, performing cross-modal feature alignment processing on the first text embedding vector and the first image embedding vector to obtain a second text embedding vector and a second image embedding vector; a multi-scale semantic alignment module, which determines a correlation output graph between the text description data and the remote sensing image based on the second text embedding vector and the second image embedding vector; An output module determines, based on the correlation output graph, an area in the remote sensing image corresponding to the text description data.

7. The semantic alignment device according to claim 6, characterized in that: The multi-scale semantic alignment module includes: a preprocessing unit, configured to determine a text semantic vector based on the second text embedding vector, and to perform a reshape process on the second image embedding vector to obtain a third image embedding vector; a correlation assignment unit, determining a correlation graph between the text semantic vector and the third image embedding vector; The multi-scale integration unit determines the correlation output graph according to the correlation graph.

8. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for semantic alignment of remote sensing images according to any one of claims 1 to 5 is implemented.

9. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the semantic alignment method of remote sensing images as described in any one of claims 1-5.

10. A computer program product, characterized in that The method comprises a computer program stored on a computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the semantic alignment method for remote sensing images according to any one of claims 1 to 5.