Multi-modal Retrieval Method and Device for Flood Events Based on a Semantic Enhancer

By adopting a multimodal search method based on semantic enhancer in the retrieval of remote sensing data of flood event-related remote sensing data, the problems of insufficient semantic information of events and difficult to obtain keywords are solved, and more efficient data retrieval matching and efficiency are achieved.

CN119597908BActive Publication Date: 2025-06-17AEROSPACE INFORMATION RES INST CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143766.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-17
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

In the retrieval of flood event-related remote sensing data, the prior art has problems such as insufficient semantic information of event and difficult to obtain keywords such as event-related spatiotemporal range, resulting in low matching degree of data retrieval and low retrieval efficiency.

Method used

A multimodal search method based on semantic enhancer is used to perform multi-level semantic enhancement of the keywords entered by users through the global flood event knowledge base and Internet news information to generate semantic enhancement text. Then, the satellite remote sensing images are quality-first screened and preliminary searched, and image features and text semantic features are extracted in combination with ResNet and BERT models, and multi-source remote sensing fusion and multi-modal feature fusion are performed, and the search results are finally obtained through similarity matching.

Benefits of technology

It improves the matching degree and efficiency of remote sensing data retrieval related to flood events, can match relevant data more accurately, and enhances the search rate and efficiency of the retrieval system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119597908B_ABST
    Figure CN119597908B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-modal retrieval method and device for flood events based on a semantic enhancer, which relates to the technical field of data retrieval. The method includes: performing multi-level semantic enhancement on keywords through a global flood event knowledge base and Internet news information included in the semantic enhancer, and then performing quality-first retrieval and preliminary screening to obtain multi-source remote sensing images to be matched. Next, corresponding optical and radar image features and context semantic features of the semantically enhanced text are extracted from the image modality and the text modality respectively. Then, the features of the image modality and the text modality are fused to obtain multi-modal fusion features, and the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched is mapped. Finally, the retrieval result of the keywords is determined. Through this application, the problems of low retrieval matching degree and low retrieval efficiency caused by lack of semantics and difficulty in obtaining event-related spatio-temporal ranges in the retrieval of flood event remote sensing data are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data retrieval, and in particular, to a multimodal retrieval method and device for flood events based on a semantic enhancer. Background Art

[0002] With the continuous development of earth observation technology, the types of satellite sensors continue to increase, and the amount of remote sensing data shows exponential growth, featuring multi-source, multi-scale, multi-temporal, and high-dimensional characteristics, marking the arrival of the "Earth Big Data" era. In addition, as governments and research institutions around the world gradually open up data resources, PB-level remote sensing data can be freely obtained through cloud platforms such as Amazon Web Services (AWS) and Google Earth Engine (GEE), providing rich data support for large-scale resource and environment research at the global scale.

[0003] However, the existing remote sensing data retrieval technology lags far behind the data growth rate. Especially in complex scenarios such as flood events, there are still significant difficulties in quickly screening out the data that meets the requirements from the massive and multi-source remote sensing data. Although high-resolution remote sensing images contain rich potential semantic information, traditional retrieval methods mainly rely on shallow features such as spatio-temporal range, sensor parameters, or label categories, lacking deep semantic support, resulting in unsatisfactory matching degree and efficiency in processing complex events (such as floods).

[0004] In addition, during the remote sensing data retrieval process for flood events, the problem of insufficient event semantic information is often faced. Flood events involve key details such as specific locations, time ranges, causes, and affected areas, but this information is difficult to directly present or clearly express in traditional remote sensing data. At the same time, the retrieval usually relies on specific spatio-temporal ranges and keywords describing the event, and these information are often difficult to accurately obtain in complex remote sensing scenarios, resulting in the retrieval system being difficult to accurately match relevant data. The problems of insufficient information and missing keywords directly affect the matching degree of the retrieval results, leading to a low recall rate of the system and a decline in retrieval efficiency, such that a large amount of relevant data is not effectively retrieved. Summary of the Invention

[0005] The present invention provides a multimodal retrieval method and device for flood events based on a semantic enhancer to solve the problems in the prior art that in the retrieval of remote sensing data related to flood events, the event semantic information is insufficient, and keywords such as event-related spatio-temporal ranges are difficult to obtain, resulting in low data retrieval matching degree and low retrieval efficiency.

[0006] The present invention provides a multimodal retrieval method for flood events based on a semantic enhancer. The method includes the following steps:

[0007] For the keywords input by the user, perform multi-level semantic enhancement on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text;

[0008] Perform quality-first screening on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and perform preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images;

[0009] Extract the optical image features and radar image features of the multi-source remote sensing images to be matched based on the ResNet model, and extract the context semantic features of the semantically enhanced text based on the BERT model;

[0010] Stitch the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features;

[0011] Perform mapping processing on the multi-modal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched, and select the multi-source remote sensing image with the highest similarity and the corresponding semantically enhanced text as the retrieval result of the keywords.

[0012] In some embodiments, the performing multi-level semantic enhancement on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text includes:

[0013] At the first level, calculate the overall similarity between the keyword and the flood event records in the global flood event knowledge base included in the semantic enhancer, where the overall similarity is calculated by weighted calculation of edit distance, text cosine similarity, fuzzy matching similarity, and sequence matching similarity;

[0014] At the second level, when the overall similarity is higher than the set threshold, retrieve the flood event record with the highest overall similarity from the global flood event knowledge base as the semantically enhanced text;

[0015] At the third level, when the overall similarity is lower than or equal to the set threshold, perform semantic enhancement on the keyword through the Internet news information included in the semantic enhancer to obtain a semantically enhanced text.

[0016] In some embodiments, the performing quality-first screening on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images includes:

[0017] Set corresponding retrieval parameters according to the event location and duration of the flood event to retrieve satellite remote sensing images;

[0018] In the satellite remote sensing images, obtain optical images with cloud cover lower than the threshold, radar images in the online state, and radar images in the offline state and already activated as candidate multi-source remote sensing images.

[0019] In some embodiments, the preliminary retrieval of the candidate multi-source remote sensing images based on the semantic enhanced text to obtain the multi-source remote sensing images to be matched includes:

[0020] Extract spatio-temporal range information from the semantic enhanced text and convert the spatio-temporal range information into longitude and latitude coordinates;

[0021] Use the longitude and latitude coordinates as pre-screening conditions and screen the candidate multi-source remote sensing images according to the pre-screening conditions to obtain the multi-source remote sensing images to be matched.

[0022] In some embodiments, the multi-modal fusion of the multi-source remote sensing fusion feature and the context semantic feature based on the BERT model to obtain the multi-modal fusion feature includes:

[0023] Call the BERT model to embed the multi-source remote sensing fusion feature and the context semantic feature to obtain the corresponding embedded multi-modal feature;

[0024] Perform multi-modal feature fusion on the embedded multi-modal feature through the BERT model to obtain the multi-modal fusion feature.

[0025] In some embodiments, the mapping process of the multi-modal fusion feature to obtain the similarity between the semantic enhanced text and the multi-source remote sensing image to be matched includes:

[0026] Select the prediction label of the flood event from the global flood event knowledge base;

[0027] Perform a mapping process on the multi-modal fusion feature to obtain the matching similarity for the prediction label, and use the matching similarity as the label matching similarity between the semantic enhanced text and the multi-source remote sensing image to be matched.

[0028] The present invention also provides a multi-modal retrieval device for flood events based on a semantic enhancer, and the device includes the following modules:

[0029] A semantic enhancement module for performing multi-level semantic enhancement on the keywords input by the user through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain semantic enhanced text;

[0030] A feature extraction module for preferentially screening satellite remote sensing images of flood events by quality to obtain candidate multi-source remote sensing images, and preliminarily retrieving the candidate multi-source remote sensing images based on the semantic-enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images;

[0031] Extract the optical image features and radar image features of the multi-source remote sensing images to be matched based on the ResNet model, and extract the context semantic features of the semantic-enhanced text based on the BERT model;

[0032] A feature fusion module for splicing the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and performing multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features;

[0033] A multi-modal retrieval module for performing mapping processing on the multi-modal fusion features to obtain the similarity between the semantic-enhanced text and the multi-source remote sensing images to be matched, and selecting the multi-source remote sensing image with the highest similarity and the corresponding semantic-enhanced text as the retrieval result of the keyword.

[0034] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-modal retrieval method for flood events based on a semantic enhancer as described in any one of the above.

[0035] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-modal retrieval method for flood events based on a semantic enhancer as described in any one of the above.

[0036] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the multi-modal retrieval method for flood events based on a semantic enhancer as described in any one of the above.

[0037] The multi-modal retrieval method and device for flood events based on a semantic enhancer provided by the present invention first perform multi-level semantic enhancement related to flood events on the retrieval keywords based on the semantic enhancer of the global flood event knowledge base and Internet news information. Secondly, the quality-preferred screening and preliminary retrieval of multi-source remote sensing images are carried out using the semantically enhanced retrieval text. Thirdly, based on the ResNet model and the BERT model, a multi-modal feature extraction method that combines remote sensing image features and flood event text semantics is constructed, which can obtain deep-level image visual semantic features while obtaining context semantic features related to flood events. Subsequently, a multi-modal feature fusion method for multi-source remote sensing images and text information is constructed based on the BERT model to fuse the text modal features and the image modal features. Finally, combining the multi-modal feature similarity matching and the multi-modal feature mapping method, multi-modal remote sensing image retrieval combining the text semantics related to flood events is carried out. This method can solve problems such as the lack of event semantic information in the retrieval of remote sensing data related to flood events, and the low data retrieval matching degree and low retrieval efficiency caused by the difficulty in obtaining keywords such as event-related spatio-temporal ranges. Brief Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art one by one. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 It is a schematic flowchart of the multi-modal retrieval method for flood events based on a semantic enhancer provided by the present invention.

[0040] Figure 2 It is a schematic diagram of the implementation principle of the multi-modal retrieval method for flood events based on a semantic enhancer provided by the present invention.

[0041] Figure 3 It is a schematic diagram of the processing of the semantic enhancer provided by the present invention.

[0042] Figure 4 It is a schematic diagram of the framework principle of the multi-modal retrieval provided by the present invention.

[0043] Figure 5 It is a schematic diagram of the structure of the multi-modal retrieval device for flood events based on a semantic enhancer provided by the present invention.

[0044] Figure 6 It is a schematic diagram of the physical structure of an electronic device provided by the present invention. Detailed Embodiments

[0045] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] The following describes the multi-modal retrieval method for flood events based on a semantic enhancer provided by the present invention with reference to the accompanying drawings.

[0047] Figure 1 is a schematic flowchart of the multi-modal retrieval method for flood events based on a semantic enhancer provided by the present invention. As Figure 1 shown, the method includes the following steps 101 to 104.

[0048] Step 101: For the keywords input by the user, perform multi-level semantic enhancement on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text.

[0049] First, it is necessary to obtain retrieval keywords related to flood events. Here, reference can be made to Figure 2 , Figure 2 is a schematic implementation diagram of the multi-modal retrieval method for flood events based on a semantic enhancer provided by the present invention. As Figure 2 shown, when the user needs to query news report records related to flood events, the retrieval keyword key (hereinafter referred to as the keyword) can be directly input into the retrieval engine or search engine of the terminal. After obtaining the keyword key related to flood events, for the keyword key input by the user, perform multi-level semantic enhancement on the keyword through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text.

[0050] Specifically, the keyword key input by the user is input into the semantic enhancer. The semantic enhancer is constructed based on the global flood event knowledge base. After receiving the input keyword key, the semantic enhancer first performs a first-level retrieval on the global flood event knowledge base through the keyword key. During the retrieval process, a similarity calculation is performed, specifically, calculating the overall similarity between the keyword and the flood event records in the global flood event knowledge base included in the semantic enhancer. The overall similarity is obtained by weighted calculation of the edit distance, text cosine similarity, fuzzy matching similarity, and sequence matching similarity.

[0051] Here, for the keyword and the flood event records in the global flood event knowledge base, four similarity calculations of edit distance, text cosine similarity, fuzzy string matching, and sequence matching are respectively performed. The following combinesFigure 3 Specifically introduce the process of semantic enhancement by the semantic enhancer. Figure 3 It is a processing schematic diagram of the semantic enhancer provided by the present invention. As Figure 3 shown, after the user retrieves and inputs the keyword key, it enters the semantic enhancer. First, determine the similarity between the keyword key and the global flood event knowledge base. Here, the global flood event knowledge base is pre-constructed and specifically composed of multiple flood event records query. The flood event record query uses the flood event occurrence location and occurrence time as retrieval parameters and is retrieved from Internet flood event news. These flood event records query describe rich semantic information related to flood events. Therefore, the process of similarity calculation is also the process of semantic information matching. Calculate the overall similarity between the keyword key and the global flood event knowledge base, that is, calculate the semantic similarity between the keyword key and each flood event record query in the library. When performing similarity calculation, four similarity calculations of edit distance, text cosine similarity, fuzzy string matching, and sequence matching are performed respectively.

[0052] For each flood event record query in the global flood event knowledge base, determine the similarity between the keyword key and the flood event record query. Among them, the similarity includes edit distance, text cosine similarity, fuzzy string matching degree, and sequence matching degree. When calculating the similarity, simultaneously calculate the edit distance LD(key, query), text cosine similarity COS(key, query), fuzzy string matching degree Fuzzy(key, query), and sequence matching degree Sequence(key, query) between the retrieval keyword key and the flood event record query. Then obtain the preset similarity measurement weights to assign weights to the calculation results of each semantic similarity, expressed as 、 、 、 . Next, perform weighted accumulation on the edit distance, text cosine similarity, fuzzy string matching, and sequence matching to obtain the overall similarity Similar(key, query) between the keyword and the flood event record in the global flood event knowledge base. The calculation method is expressed as the following formula (1):

[0053] (1)

[0054] In the above formula (1), is the measurement weight of the edit distance 、 is the measurement weight of the text cosine similarity 、 is the fuzzy string matching degree The measurement weight, is the sequence matching degree The measurement weight, where i represents the serial number of the keyword key, indicating the i-th keyword, j represents the serial number of the flood event record query, indicating the j-th flood event record.

[0055] In actual weighted accumulation operations, the assigned measurement weights 、 、 、 can be adjusted according to the results obtained from the operations, and then the weighted accumulation operation is performed again. After multiple layers of weight adjustment, the optimal matching result, that is, the optimal overall similarity Similar(key, query), is finally obtained.

[0056] Based on the overall similarity, the first-level retrieval of keywords is realized. Next, the second-level retrieval can be entered, and the flood event records with higher overall similarity (greater than the set threshold) are screened out from the global flood event knowledge base as semantic enhancement texts. If the calculated overall similarity (less than or equal to the set threshold) is low, it indicates that no flood event records similar to the user-input keyword key can be matched from the global flood event knowledge base. Therefore, the third-level retrieval is performed. The keyword key is passed into the Internet, and the keyword key is semantically enhanced through the Internet news information included in the semantic enhancer to obtain the semantic enhancement text of the keyword. The Internet news retrieval is used as the third-level retrieval, and finally the retrieval event statement (news report) can be retrieved as the semantic enhancement text.

[0057] As Figure 3 shown, during the second-level retrieval, when the overall similarity Similar(key, query) between the retrieval keyword and the flood event record in the global flood event knowledge base is greater than the threshold T, it indicates that there are flood event records semantically similar to the retrieval keyword in the global flood event knowledge base at this time. Then, the flood event record query with the highest overall similarity is retrieved from the global flood event knowledge base as the semantic enhancement text of the keyword key, that is, the flood event record with the highest similarity is screened out from the global flood event knowledge base as the semantic enhancement text.

[0058] When the overall similarity Similar(key, query) between the keyword and the global flood event knowledge base is less than or equal to the threshold T, it indicates that there is no flood event record semantically similar to the keyword in the global flood event knowledge base at this time, and then the third-level retrieval is performed. Internet news retrieval is executed for the keyword whose overall similarity Similar(key, query) is less than or equal to the threshold T. Here, when the overall similarity Similar(key, query) is less than or equal to the set threshold, the keyword is transmitted to the Internet, and the semantic information of the keyword is enhanced through the Internet news information included in the semantic enhancer, and the corresponding news record information (i.e., flood event record) is retrieved as the semantic enhancement text of the keyword key.

[0059] In the embodiment of the present invention, through the semantic enhancer, multi-level retrieval is respectively performed using the global flood event knowledge base and Internet news information to achieve multi-level semantic enhancement of the user's retrieval keyword, and finally a semantic enhancement text semantically matching the retrieval keyword is obtained, enriching the semantic information in the retrieval keyword and achieving a closer match between the user's needs and the retrieval query.

[0060] Step 102: Perform quality-priority screening on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and perform preliminary retrieval on the candidate multi-source remote sensing images based on the semantic enhancement text to obtain multi-source remote sensing images to be matched.

[0061] Continue to refer to Figure 2 , after determining the semantic enhancement text, multi-source remote sensing images to be matched are screened based on the semantic enhancement text. Here, first, quality-priority screening is performed on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images.

[0062] Specifically, first, the corresponding satellite remote sensing images can be retrieved from the cloud platform according to the retrieval parameters such as the occurrence time, location, influence range, and specific details of the flood event of the flood event. Considering that the quality of these satellite remote sensing images is different and there may be factors such as cloudy and rainy weather in the images, which cannot accurately describe the corresponding flood image features, corresponding quality screening conditions need to be set to perform quality-priority screening on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images and eliminate the images with poor quality. Among them, the candidate multi-source remote sensing images include optical images and radar images.

[0063] Next, according to the semantic-enhanced text, the obtained radar images and optical images are preliminarily screened. First, the time, location, affected area, etc. of the flood event are determined from the semantic-enhanced text as screening conditions for preliminary retrieval. During the preliminary retrieval process, multi-source remote sensing images that meet the screening conditions are retrieved from the candidate multi-source remote sensing images as the multi-source remote sensing images to be matched. Of course, the multi-source remote sensing images to be matched include optical images and radar images for subsequent retrieval.

[0064] In some embodiments, during the process of preferentially screening the satellite remote sensing images of the flood event to obtain candidate multi-source remote sensing images, first, the satellite remote sensing images related to the flood event need to be obtained. However, these satellite remote sensing images are cross-cloud platform, belonging to cross-platform remote sensing images. There are many types of remote sensing image data, including Landsat series satellite images (belonging to optical images) and Sentinel-1 satellite images (belonging to radar images), and there are also quality differences (such as more clouds or fewer clouds). In addition, due to the different states of data provision, the radar images may be offline or online in the cloud platform. To download and obtain these remote sensing image data on the cloud platform, a multi-threaded method needs to be used to call relevant interfaces for parallel downloading.

[0065] In the embodiment of the present invention, retrieval parameters are set based on the event location and duration of the flood event to retrieve satellite remote sensing images. Among the satellite remote sensing images, optical images with cloud cover below the threshold, radar images in the online state, and radar images in the offline state and already activated are obtained as candidate multi-source remote sensing images.

[0066] Specifically, for the satellite remote sensing images across cloud platforms, they are obtained in parallel using the remote sensing event data acquisition tool. Whether it is for Landsat series satellite images or Sentinel-1 satellite images, first, retrieval parameters are set based on the event location and duration of the flood event, and the remote sensing event data acquisition tool is used for retrieval to obtain satellite remote sensing images related to the flood event.

[0067] Among them, the position expansion of the central coordinates of the flood event occurrence location is used to describe the central position of the flood event occurrence and the diffusion range from the central position, representing the spatial information of the flood event. The duration is used to represent the duration of the flood event occurrence process, representing the time information of the flood event. Therefore, using the spatial information and time information of the flood event as retrieval parameters can effectively and accurately screen the satellite remote sensing images of the flood event for easy download and acquisition.

[0068] Considering that the quality of these satellite remote sensing images varies, and there may be factors such as cloudy and rainy weather in the images, which cannot accurately describe the characteristics of the corresponding flood images. Therefore, corresponding quality screening conditions need to be set. Here, the cloud cover threshold and the online status are set to screen and obtain candidate multi-source remote sensing images.

[0069] The satellite remote sensing images retrieved according to the retrieval parameters are divided into optical images and radar images. When the satellite remote sensing image is an optical image, the satellite remote sensing image is screened according to the cloud cover threshold. Download the optical images with cloud cover not exceeding the preset cloud cover threshold as candidate multi-source remote sensing images. Here, for the optical images of the Landsat series satellite images, that is, when the satellite remote sensing image is an optical image, the extended and duration of the central point coordinates of the flood event are used as the spatial and temporal parameters for retrieval. Using the Pystac tool that supports spatio-temporal asset catalog processing and data management as the remote sensing event data acquisition tool, the Landsat satellite image data in the "landsat-c2-l2" dataset on the Microsoft Planetary Computing Platform is retrieved to obtain the satellite remote sensing image. And a cloud cover threshold of 20% is preset as the pre-screening condition, and the optical images with cloud cover not exceeding 20% of the cloud cover threshold are selected and downloaded as candidate multi-source remote sensing images.

[0070] The download method can be to set the chunk_size to 512 for a single cross-cloud platform remote sensing image and perform chunked download to achieve breakpoint resume, thereby reducing the consumption of IO resources. And multiple cross-cloud platform remote sensing images are obtained concurrently through multi-threading to complete the efficient download of remote sensing image data.

[0071] When the satellite remote sensing image is a radar image, it is necessary to judge the online status of the radar image and obtain the radar image according to the online status. Here, for the radar images of Sentinel-1 satellite images, the position extension and duration of the central point coordinates of the flood event are also used as the spatial and temporal parameters for retrieval. Using the Sentinelsat API that supports the access and download of European Space Agency (ESA) Sentinel satellite data as the remote sensing event data acquisition tool, the Sentinel-1 data retrieval provided by the Copernicus Open Access Hub is realized to obtain the satellite remote sensing image. Next, it is necessary to judge the online status of the satellite remote sensing image. Here, by parsing the Sentinel-1 satellite remote sensing data record in the retrieval result, the Online parameter in the record is used to judge whether the radar image is online.

[0072] If the online status is the online state, then the online radar images are downloaded and obtained in parallel. Here, a data download list is constructed for the online radar images, and the online satellite remote sensing images are downloaded in parallel through the Sentinelsat API in a multi-threaded manner as candidate multi-source remote sensing images.

[0073] When the online state is the offline state, the radar images in the offline state in the satellite remote sensing images are pre-activated, and a preset time interval (such as 40 minutes) is set to determine the activation state in an asynchronous polling manner, that is, to judge whether it is activated. If it has been activated, the activated radar images are added to the data download list, and the radar images with the activation state of activated are downloaded in parallel as candidate multi-source remote sensing images. If it has not been activated, it will be activated again after the time interval and wait for the next poll.

[0074] In the embodiment of the present invention, after obtaining the semantic enhanced text, corresponding retrieval parameters are set to retrieve satellite remote sensing images, and the satellite remote sensing images are preliminarily screened by cloud cover and online state, which can improve the data quality of remote sensing images and facilitate subsequent retrieval operations based on the semantic enhanced text.

[0075] In some embodiments, after obtaining candidate multi-source remote sensing images through quality priority screening, next, the candidate multi-source remote sensing images are preliminarily retrieved based on the semantic enhanced text to obtain to-be-matched multi-source remote sensing images.

[0076] First, the spatio-temporal range information of the flood event (that is, the range of diffusion of the center coordinate position of the flood event occurrence location and the duration) is extracted from the semantic enhanced text, and the spatio-temporal range information is converted into longitude and latitude coordinates as pre-screening conditions.

[0077] Here, the spatial information and time information of the flood event are described, and using the interface provided by map tools or map software, they are converted into longitude and latitude coordinates as pre-screening conditions. Then, the obtained radar images and optical images are screened according to the pre-screening conditions to obtain to-be-matched multi-source remote sensing images for subsequent multi-modal retrieval in combination with the semantic enhanced text.

[0078] As Figure 2 shown, based on the retrieval event statement (that is, the semantic enhanced text), the satellite remote sensing images obtained in parallel are screened to narrow the retrieval range to obtain to-be-screened image data (that is, to-be-matched multi-source remote sensing images), and the to-be-matched multi-source remote sensing images also include optical images and radar images.

[0079] In the embodiment of the present invention, after obtaining candidate multi-source remote sensing images, the data is first preliminarily retrieved through the spatial information and time information of the flood event, thereby narrowing the retrieval range of the semantic enhanced text, reducing the retrieval difficulty, and being beneficial to the accuracy of subsequent retrieval results.

[0080] In step 103, the optical image features and radar image features of the to-be-matched multi-source remote sensing images are extracted based on the ResNet model, and the context semantic features of the semantic enhanced text are extracted based on the BERT model. The to-be-matched multi-source remote sensing images include optical images and radar images.

[0081] After obtaining the multi-source remote sensing images to be matched through the preliminary retrieval in step 102, next, multi-modal retrieval is performed in combination with the semantic-enhanced text. However, before multi-modal retrieval, the features of the multi-source remote sensing images to be matched and the semantic-enhanced text need to be extracted separately.

[0082] Here, different models are respectively called to extract the corresponding features. Based on the ResNet model, the optical image features and radar image features of the multi-source remote sensing images to be matched are extracted, and based on the BERT model, the context semantic features of the semantic-enhanced text are extracted.

[0083] In some embodiments, for the optical images and radar images in the multi-source remote sensing images to be matched, the ResNet model is called to extract the corresponding optical image features and radar image features respectively. As Figure 2 shown, the multi-source remote sensing images to be matched are used as the image modality, and the retrieval event statement is used as the text modality, and input into the encoder and decoder of the multi-modal retrieval model for feature extraction. Subsequently, feature fusion is performed through the BERT encoder to obtain the fused features.

[0084] The following combines Figure 4 to specifically illustrate the processing process of the multi-modal retrieval model. Figure 4 is the framework schematic diagram of the multi-modal retrieval provided by the present invention. As Figure 4 shown, in the multi-modal retrieval model, the semantic-enhanced text is used as the text data of the text modality, and the multi-source remote sensing images to be matched are used as the image data of the image modality. Here, the multi-source remote sensing images are subjected to feature extraction through the Residual Network (ResNet) model to obtain the multi-source remote sensing fused features. In the actual application environment, the multi-source remote sensing images to be matched are divided into optical images and radar images. For the optical images in the multi-source remote sensing images to be matched, the optical images are subjected to feature encoding through the ResNet model to obtain the optical image features. For the radar images in the multi-source remote sensing images to be matched, the radar images are subjected to feature encoding through the ResNet model to obtain the radar image features.

[0085] At the same time, for the text data of the semantic-enhanced text, the Bidirectional Encoder Representations from Transformers (BERT) model is used to extract the context semantic features of the semantic-enhanced text as the text modality features. Thus, the feature acquisition of the multi-source remote sensing images and the text modality under the same model framework is realized, and the multi-modal features based on images and texts are obtained.

[0086] In the embodiments of the present invention, the features of the image modality and the text modality are independently extracted through the ResNet model and the BERT model, realizing the accurate extraction of the semantic information and visual information of the flood time, facilitating the subsequent further matching of the semantic information and the visual information, thereby effectively improving the accuracy of the graphic-text matching and ensuring the retrieval effectiveness of the flood event keywords.

[0087] Step 104: splice the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features.

[0088] Through step 103, after extracting multi-modal features (i.e., optical image features, radar image features, and context semantic features) in the same model framework, first splice the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and then perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features, realizing the fusion of multi-modal features.

[0089] As Figure 4 shown, first, fuse the features of the radar and optical images through the method of vector splicing to obtain multi-source remote sensing fusion features. Then perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain and output multi-modal fusion features. Thus, the fusion of multi-modal features is realized.

[0090] In some embodiments, the BERT model is used to perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features to obtain multi-modal fusion features. Specifically, the BERT model is used to embed (Embedding) the multi-source remote sensing fusion features and the context semantic features, and then the embedded features are merged to obtain the corresponding embedded multi-modal features.

[0091] Next, the BERT model is used to perform multi-modal feature fusion on the embedded multi-modal features to obtain multi-modal fusion features. Here, the embedded multi-modal features are input into the BERT model, and the embedded multi-modal features are encoded by the encoder of the BERT model to obtain multi-modal fusion features, thereby realizing the multi-modal feature fusion of the image and the text.

[0092] In the embodiments of the present invention, the BERT encoder encodes the embedded text modality and image modality features and maps them to the same feature space, fully mining the potential semantic association between the image modality and the text modality, thereby realizing multi-modal feature fusion to obtain consistent semantic features, which not only improves the retrieval accuracy but also provides rich semantic support for flood event retrieval.

[0093] Step 105: Perform mapping processing on the multi-modal fusion features to obtain the similarity between the semantic-enhanced text and the multi-source remote sensing images to be matched, and select the multi-source remote sensing image with the highest similarity and the corresponding semantic-enhanced text as the retrieval result of the keyword.

[0094] As Figure 4 shown, the BERT model outputs the multi-modal fusion features, and then performs mapping processing on the multi-modal fusion features to obtain the label matching similarity between the semantic-enhanced text and the multi-source remote sensing images to be matched.

[0095] Specifically, to achieve accurate retrieval matching, it is necessary to calculate the matching similarity. Since the multi-modal fusion features have been calculated through Step 104, the multi-modal fusion features can be classified through a mapping module (such as a fully connected layer or a multi-layer perceptron) and mapped to the probability of a certain flood event category label. The flood event category can be the type of flood, such as various flood types caused by waterlogging, levee breach, or other natural or human factors in a certain area. The higher the probability corresponding to the flood event category label, the higher the label matching similarity between the semantic-enhanced text and the multi-source remote sensing images to be matched, the higher the matching similarity between the two, and the more accurate the retrieval effect.

[0096] Therefore, here, the multi-source remote sensing image with the highest similarity and the corresponding semantic-enhanced text are selected as the retrieval result of the keyword key. Here, under the corresponding flood event category label, the multi-source remote sensing images to be matched are sorted according to the size of the label matching similarity, and the k multi-source remote sensing images with the largest label matching similarity (top-k) can be screened out as the return value of the multi-modal retrieval model. This return value is the flood event semantic retrieval result of the keyword key, that is, the k multi-source remote sensing images with the highest label matching similarity and the semantic-enhanced text are used as the flood event semantic retrieval result of the keyword key.

[0097] As Figure 2 shown, after feature fusion of the text modal features and the image modal features to obtain the multi-modal fusion features, mapping processing is performed on the fusion features to output the flood event semantic retrieval result corresponding to the keyword key. In the embodiment of the present invention, the flood event semantic retrieval result includes: news reports (i.e., retrieval event statements), optical images, and radar images, as multi-modal retrieval results. Among them, the news report (i.e., the retrieval event statement) can be the semantic-enhanced text corresponding to the keyword key, and the optical image and the radar image are the multi-source remote sensing images retrieved from the multi-source remote sensing images to be matched and most matched with the semantic-enhanced text.

[0098] In the embodiments of the present invention, first, based on the semantic enhancer of the global flood event knowledge base and Internet news information, semantic enhancement related to flood events is performed on the retrieval keywords. Secondly, the semantically enhanced retrieval text is used for multi-source remote sensing image quality prioritized screening and preliminary retrieval. Thirdly, based on the ResNet model and the BERT model, a multi-modal feature extraction method combining remote sensing image features and flood event text semantics is constructed, which can obtain deep-level image visual semantic features while obtaining context semantic features related to flood events. Subsequently, a multi-modal feature fusion method for multi-source remote sensing images and text information is constructed based on the BERT model to perform feature fusion on text modal features and image modal features. Finally, combining the multi-modal feature similarity matching and multi-modal feature mapping methods, multi-modal remote sensing image retrieval combining the text semantics related to flood events is performed. This method can solve problems such as low data retrieval matching degree and low retrieval efficiency caused by the lack of event semantic information in the retrieval of remote sensing data related to flood events and the difficulty in obtaining keywords such as event-related spatio-temporal ranges.

[0099] In some embodiments, determining the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched can be achieved by mapping to a certain prediction label in the global flood event knowledge base, which is specifically described below.

[0100] First, select the prediction labels of flood events from the global flood event knowledge base. Here, a prediction label can be set according to the labels matched in the global flood event knowledge base in the semantic enhancer or common flood event labels. For example, for flood event types such as waterlogging, dike breaches, or other natural or man-made factors occurring in a certain area, the corresponding flood event type is used as the prediction label.

[0101] Then, perform mapping processing on the multi-modal fusion features to obtain the matching similarity for the prediction label. Here, the mapping processing can perform mapping calculations on the multi-modal fusion features through a linear layer (such as a fully connected layer) or a multi-layer perceptron, and output the matching similarity for the prediction label as the label matching similarity between the semantically enhanced text and the multi-source remote sensing images to be matched.

[0102] Of course, there is more than one type of flood event, and there is more than one selected prediction label. During the mapping calculation, the matching similarity corresponding to each prediction label will be output respectively.

[0103] In the embodiments of the present invention, based on the prediction labels of flood events, the retrieval results of keywords are determined according to the multi-modal fusion features, realizing the retrieval decision-making for flood events using the multi-modal joint judgment results, associating the semantic description information of complex remote sensing events with the retrieval conditions, enhancing the semantic information density of the retrieval statement, and realizing fast and accurate semantic retrieval for facing a large amount of remote sensing data.

[0104] The following describes the multi-modal retrieval device for flood events based on a semantic enhancer provided by the present invention. The multi-modal retrieval device for flood events based on a semantic enhancer described below can be correspondingly referred to the multi-modal retrieval method for flood events based on a semantic enhancer described above.

[0105] See Figure 5 , Figure 5 which is a schematic structural diagram of the multi-modal retrieval device for flood events based on a semantic enhancer provided by the present invention. As Figure 5 shown, the multi-modal retrieval device for flood events based on a semantic enhancer includes: a semantic enhancement module 501, a feature extraction module 502, a feature fusion module 503, and a multi-modal retrieval module 504. Among them, the semantic enhancement module 501 is used to perform multi-level semantic enhancement on the keywords input by the user through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text; the feature extraction module 502 is used to perform quality-first screening on the satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and perform preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images; extract the optical image features and radar image features of the multi-source remote sensing images to be matched based on the ResNet model, and extract the context semantic features of the semantically enhanced text based on the BERT model; the feature fusion module 503 is used to splice the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and perform multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features; the multi-modal retrieval module 504 is used to perform mapping processing on the multi-modal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched, and select the multi-source remote sensing image with the highest similarity and the corresponding semantically enhanced text as the retrieval result of the keywords.

[0106] It should be noted that the beneficial effects of the multi-modal retrieval device for flood events based on a semantic enhancer here can correspond to those of the multi-modal retrieval method for flood events based on a semantic enhancer in the above text. Therefore, the beneficial effects of the multi-modal retrieval device for flood events based on a semantic enhancer will not be elaborated here.

[0107] Figure 6 which is a schematic physical structure diagram of an electronic device provided by the present invention. As Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute a multimodal retrieval method for flood events based on a semantic enhancer. The method includes: for the keywords input by the user, performing multi-level semantic enhancement on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text; performing quality-priority screening on satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and performing preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images; extracting optical image features and radar image features of the multi-source remote sensing images to be matched based on the ResNet model, and extracting context semantic features of the semantically enhanced text based on the BERT model; splicing the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and performing multimodal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multimodal fusion features; performing mapping processing on the multimodal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched, and selecting the multi-source remote sensing images to be matched with the highest similarity and the corresponding semantically enhanced text as the retrieval results of the keywords.

[0108] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-modal retrieval method for flood events based on a semantic enhancer provided by the above-mentioned various methods. The method includes: for the keywords input by the user, performing multi-level semantic enhancement on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text; performing quality-priority screening on satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and performing preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images; extracting optical image features and radar image features of the multi-source remote sensing images to be matched based on the ResNet model, and extracting context semantic features of the semantically enhanced text based on the BERT model; splicing the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and performing multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features; performing mapping processing on the multi-modal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched, and selecting the multi-source remote sensing image with the highest similarity and the corresponding semantically enhanced text as the retrieval result of the keywords.

[0110] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-modal retrieval method for flood events based on a semantic enhancer provided by the above-mentioned various methods. The method includes: for the keywords input by a user, performing multi-level semantic enhancement on the keywords through a global flood event knowledge base and Internet news information included in the semantic enhancer to obtain a semantically enhanced text; preferentially screening satellite remote sensing images of flood events to obtain candidate multi-source remote sensing images, and performing a preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, where the multi-source remote sensing images to be matched include optical images and radar images; extracting optical image features and radar image features of the multi-source remote sensing images to be matched based on a ResNet model, and extracting context semantic features of the semantically enhanced text based on a BERT model; splicing the radar image features and the optical image features to obtain multi-source remote sensing fusion features, and performing multi-modal fusion on the multi-source remote sensing fusion features and the context semantic features based on the BERT model to obtain multi-modal fusion features; performing a mapping process on the multi-modal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing images to be matched, and selecting the multi-source remote sensing image with the highest similarity and the corresponding semantically enhanced text as the retrieval result of the keywords.

[0111] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal retrieval method for flood events based on semantic enhancer, characterized in that: The method comprises: For the keywords input by the user, multi-level semantic enhancement is performed on the keywords through the global flood event knowledge base and Internet news information included in the semantic enhancer to obtain semantically enhanced text; The satellite remote sensing images of the flood event are screened by quality priority to obtain candidate multi-source remote sensing images, and the candidate multi-source remote sensing images are preliminarily retrieved based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, wherein the multi-source remote sensing images to be matched include optical images and radar images; Extracting optical image features and radar image features of the multi-source remote sensing image to be matched based on the ResNet model, and extracting contextual semantic features of the semantically enhanced text based on the BERT model; The radar image feature and the optical image feature are spliced ​​to obtain a multi-source remote sensing fusion feature, and the multi-source remote sensing fusion feature and the context semantic feature are multimodally fused based on the BERT model to obtain a multimodal fusion feature; Mapping the multimodal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing image to be matched, and selecting the multi-source remote sensing image to be matched with the highest similarity and the corresponding semantically enhanced text as the search result of the keyword; The preliminary retrieval of the candidate multi-source remote sensing images based on the semantically enhanced text to obtain the multi-source remote sensing images to be matched includes: Extracting spatiotemporal range information from the semantically enhanced text, and converting the spatiotemporal range information into longitude and latitude coordinates; The latitude and longitude coordinates are used as pre-screening conditions, and the candidate multi-source remote sensing images are screened according to the pre-screening conditions to obtain a multi-source remote sensing image to be matched; The mapping process is performed on the multimodal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing image to be matched, including: Selecting a prediction label of a flood event from the global flood event knowledge base; The multimodal fusion features are mapped to obtain a matching similarity for the predicted label, and the matching similarity is used as a label matching similarity between the semantically enhanced text and the multi-source remote sensing image to be matched.

2. The multimodal retrieval method for flood events based on semantic enhancer according to claim 1 is characterized in that: The semantic enhancer includes a global flood event knowledge base and Internet news information, and performs multi-level semantic enhancement on the keywords to obtain semantically enhanced text, including: In the first stage, the overall similarity between the keyword and the flood event records in the global flood event knowledge base included in the semantic enhancer is calculated, and the overall similarity is obtained by weighted calculation of edit distance, text cosine similarity, fuzzy matching similarity and sequence matching similarity; In the second stage, when the overall similarity is higher than a set threshold, the flood event record with the highest overall similarity is retrieved from the global flood event knowledge base as the semantically enhanced text; In the third stage, when the overall similarity is lower than or equal to the set threshold, the keywords are semantically enhanced by the Internet news information included in the semantic enhancer to obtain a semantically enhanced text.

3. The multimodal retrieval method for flood events based on semantic enhancer according to claim 1 is characterized in that: The satellite remote sensing images of flood events are screened by quality priority to obtain candidate multi-source remote sensing images, including: According to the location and duration of the flood event, corresponding search parameters are set to retrieve satellite remote sensing images; Among the satellite remote sensing images, optical images with cloud cover below a threshold, radar images in an online state, and radar images in an offline state and activated are obtained as candidate multi-source remote sensing images.

4. The multimodal retrieval method for flood events based on semantic enhancer according to claim 1 is characterized in that: The multi-modal fusion of the multi-source remote sensing fusion features and the contextual semantic features based on the BERT model to obtain the multi-modal fusion features includes: Calling the BERT model to embed the multi-source remote sensing fusion features and the contextual semantic features to obtain corresponding embedded multimodal features; The embedded multimodal features are subjected to multimodal feature fusion through the BERT model to obtain multimodal fusion features.

5. A multimodal retrieval device for flood events based on semantic enhancer, characterized in that: The device comprises: A semantic enhancement module, for performing multi-level semantic enhancement on keywords input by a user through a global flood event knowledge base and Internet news information included in a semantic enhancer, to obtain semantically enhanced text; A feature extraction module is used to perform quality-first screening on the satellite remote sensing images of the flood event to obtain candidate multi-source remote sensing images, and to perform preliminary retrieval on the candidate multi-source remote sensing images based on the semantically enhanced text to obtain multi-source remote sensing images to be matched, wherein the multi-source remote sensing images to be matched include optical images and radar images; Extracting optical image features and radar image features of the multi-source remote sensing image to be matched based on the ResNet model, and extracting contextual semantic features of the semantically enhanced text based on the BERT model; A feature fusion module, used for splicing the radar image feature with the optical image feature to obtain a multi-source remote sensing fusion feature, and performing multi-modal fusion on the multi-source remote sensing fusion feature and the contextual semantic feature based on a BERT model to obtain a multi-modal fusion feature; A multimodal retrieval module is used to map the multimodal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing image to be matched, and select the multi-source remote sensing image to be matched with the highest similarity and the corresponding semantically enhanced text as the retrieval result of the keyword; The preliminary retrieval of the candidate multi-source remote sensing images based on the semantically enhanced text to obtain the multi-source remote sensing images to be matched includes: Extracting spatiotemporal range information from the semantically enhanced text, and converting the spatiotemporal range information into longitude and latitude coordinates; The latitude and longitude coordinates are used as pre-screening conditions, and the candidate multi-source remote sensing images are screened according to the pre-screening conditions to obtain a multi-source remote sensing image to be matched; The mapping process is performed on the multimodal fusion features to obtain the similarity between the semantically enhanced text and the multi-source remote sensing image to be matched, including: Selecting a prediction label of a flood event from the global flood event knowledge base; The multimodal fusion features are mapped to obtain a matching similarity for the predicted label, and the matching similarity is used as a label matching similarity between the semantically enhanced text and the multi-source remote sensing image to be matched.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the multimodal retrieval method for flood events based on a semantic enhancer as described in any one of claims 1 to 4 is implemented.

7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multimodal retrieval method for flood events based on a semantic enhancer as described in any one of claims 1 to 4 is implemented.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the multimodal retrieval method for flood events based on a semantic enhancer as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Multi-source heterogeneous non-point source pollution big data association and retrieval method based on spatial and temporal characteristics and supervision platform

    CN110334090A

  • Multi-source heterogeneous remote sensing data association construction and multi-user data matching method

    CN111666313A