Remote sensing image target identification data set standardized construction method based on image-text fusion

By constructing a remote sensing image target recognition data set based on image-text fusion, the problem of lack of image-text fusion and irregular labeling of existing data sets is solved, and the multimodal characteristics of the data are enhanced and the accuracy of target recognition is improved.

CN120047775APending Publication Date: 2025-05-27BEIHANG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510213332.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing remote sensing image target recognition data set lacks image-text fusion, single data, language description information, irregular labeling, and inconsistent processing methods, resulting in limited application of artificial intelligence in the field of remote sensing.

Method used

Build a remote sensing image target recognition data set based on image-text fusion, collect remote sensing images through tools such as Google Earth, use open source annotation software for semantic annotation, call large language models to generate diversified language descriptions, and perform standardized annotation and classification management.

Benefits of technology

The multimodal characteristic enhancement of remote sensing image data is achieved, providing a solid foundation for subsequent intelligent analysis and deep learning model training, improving the accuracy and reliability of target recognition, and filling the gaps in this field at home and abroad.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047775A_ABST
    Figure CN120047775A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image target recognition data set standardized construction method based on image-text fusion, and the method comprises the steps: obtaining and adjusting a remote sensing image through an open source tool, carrying out the semantic tag and segmentation point labeling of a target object, analyzing and summarizing the universal and specific features of the target, and obtaining a target recognition data set; dividing points are converted and marked as binary mask images and target detection rectangular frames, diversified language descriptions are generated in combination with a large language model, and the generated language descriptions, the remote sensing images, the mask images and the target detection rectangular frames are correspondingly stored and managed in a classified mode. According to the data set, remote sensing images (images) and target description (texts) are combined, and a complete process of'data acquisition, data analysis, data standardization labeling and data classification management 'of the multi-modal image-text remote sensing data set is provided. According to the method, the blank in the aspect of remote sensing image target recognition data sets based on vision and language fusion at home and abroad is filled up, and application of remote sensing image intelligent analysis in multiple fields is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides an artificial intelligence data set and a construction method thereof, and particularly relates to a remote sensing image target recognition data set based on image-text fusion and a standardized construction method thereof, belonging to the field of remote sensing image intelligent big data. Background Art

[0002] Remote sensing images refer to electromagnetic wave signals reflected from the Earth's surface obtained by satellites or aircraft, including various data types such as images and videos, and are commonly used for analyzing and detecting ground object targets. With the rapid development of remote sensing technology, remote sensing images have been widely applied in fields such as resource exploration, environmental monitoring, urban planning, and agriculture. However, due to the large amount of complex ground object information contained in remote sensing images, how to accurately identify the targets therein has become one of the core issues in remote sensing image processing.

[0003] In recent years, artificial intelligence, especially deep learning technology, has achieved remarkable results in remote sensing image target recognition. Driven by large-scale remote sensing data, deep learning models can effectively extract complex features, thereby improving the accuracy and efficiency of target recognition. However, the performance of artificial intelligence models highly depends on high-quality large-scale training data sets, which is also applicable to remote sensing image target recognition. Currently, the construction of remote sensing image target recognition data sets based on image-text fusion is a relatively new field, aiming to fuse multi-modal information (such as visual and language information) to accurately sense and identify specific targets described by language, and improve the accuracy and reliability of target recognition.

[0004] Most of the existing remote sensing image target recognition data sets only involve single-type remote sensing data, lack language descriptions corresponding to target objects, and the standardized processing process of the data sets is also imperfect. Remote sensing images have complex and diverse characteristics, and it is necessary to standardize the data and perform refined annotation to ensure its applicability to different target recognition tasks. However, there is no high-quality remote sensing image target recognition data set for visual and language fusion at home and abroad, which has become one of the main bottlenecks in the application of artificial intelligence in the remote sensing field.

[0005] The remote sensing image target recognition data set based on image-text fusion and a construction method thereof provided by the present invention aim to meet the needs of remote sensing image intelligent analysis and target recognition, and fill the gap in this field at home and abroad. By constructing this data set, the problems existing in the prior art such as single remote sensing data, lack of language description information, non-standard annotation, and unsystematic processing methods can be effectively solved, and the further application and development of remote sensing images in the field of intelligent big data can be promoted. Summary of the Invention

[0006] Object of the Invention

[0007] The object of the present invention is to provide a remote sensing image target recognition dataset based on image-text fusion, which aims to fill the gap at home and abroad in the standardized and process-based construction of remote sensing image target recognition datasets for the need to accurately sense and identify specific targets (determined by language description) in remote sensing scenarios. In addition, the invention provides a complete design process and construction method for the dataset covering "data acquisition - data analysis - data normalization annotation - data classification management".

[0008] Technical solution

[0009] The present invention relates to a method for standardizing the construction of a remote sensing image target recognition dataset based on image-text fusion, and the steps are as follows:

[0010] Step 1: Search using the Google Earth open-source tool according to the target categories to be collected, and determine the target area. According to the actual sizes of different category objects, reasonably adjust the ground resolution and the viewing angle tilt angle to ensure that the display ratio of the target in the remote sensing image is appropriate for effective analysis. After adjustment, export the image and save it to a local storage device.

[0011] Step 2: Import the remote sensing image saved to the local storage device into an open-source annotation software, perform semantic label and segmentation point annotation on the target object, and convert the segmentation point annotation into a binary mask image and a target detection rectangle through a data processing algorithm.

[0012] Step 3: Conduct a comprehensive analysis of the remote sensing images of each target category, summarize the general features of each category to form corresponding descriptive phrases. Subsequently, based on the segmentation annotation results, further analyze the specific features of the target, including the absolute position, relative position, and geographical environment of the target in the image, and describe them using corresponding phrases.

[0013] Step 4: Through the application programming interface, call a large language model, input the descriptive phrase corresponding to each specific target and a pre-compiled prompt template together, and automatically, efficiently, and diversely generate language descriptions that meet the requirements.

[0014] Step 5: Correlate the generated language descriptions with the annotated remote sensing images one by one, and classify and store them according to the semantic labels and geomorphic features of each target.

[0015] Through the above steps, a remote sensing image target recognition dataset based on the fusion of vision and language is constructed. This dataset includes two-dimensional remote sensing images, binary segmentation mask images, target detection bounding boxes, and specific target descriptions generated by large language models. For this type of dataset, a full-process standardized design scheme of "data collection - data analysis - data normalization annotation - data classification management" is provided, and specific construction methods are proposed to ensure the standardization of all aspects of the dataset, including collection, annotation, analysis, and classification management.

[0016] Among them, in "Step 1", the statement "According to the actual sizes of different types of objects, reasonably adjust the ground resolution and the perspective tilt angle to ensure that the display ratio of the target in the remote sensing image is appropriate" is implemented as follows: On the one hand, the main purpose of adjusting the ground resolution is to ensure that the target in the remote sensing image is within an appropriate proportion range. For smaller targets, such as cars driving on the road, a higher resolution can display their shapes and positions. For larger targets, such as bridges or railway stations, a lower resolution is sufficient to show their overall structures, avoiding the target occupying too large a proportion in the remote sensing image and taking up too many pixel spaces, which may affect the overall analysis. On the other hand, considering that the same target may present completely different appearance features from different perspectives. Through multi-perspective coverage, the dataset can provide richer and more challenging samples for the training of subsequent algorithms, improving the performance of the model in real applications.

[0017] Among them, in "Step 2", the statement "Perform semantic label and segmentation point annotation on the target object, and convert the segmentation point annotation into a binary mask image and a target detection bounding box through a data processing algorithm" is implemented as follows: Since the segmentation point annotation format output by the annotation software cannot directly meet the training requirements of mainstream deep learning models, it is necessary to use library functions to recreate a binary image. Through a conversion algorithm, the area enclosed by the annotated segmentation points is filled. Among them, the background area is represented by a value of 0 (black), and the target area is represented by a value of 1 (white) to generate a binary mask image suitable for the training of deep learning segmentation models. In addition, through the analysis of the area enclosed by the segmentation points, the minimum bounding rectangle of this area is generated. This rectangle can be used as the annotation box for the target detection task and is used for the training of the target detection model.

[0018] Among them, for the operations described in "Step 3", namely, "conduct comprehensive analysis on the remote sensing images of each target category, summarize the general features of each category, and form corresponding descriptive phrases. Subsequently, based on the segmentation annotation results, further analyze the specific features of the targets", the specific implementation is as follows: Summarize the general features of the targets, mainly including the inherent features of such objects, including color, shape, etc., such as "red", "rectangle". Traverse and analyze the specific features of all target objects, including the absolute position of the target in the image (such as the target is located on the left side of the image), relative position (such as the positional relationship relative to other objects), and the geographical environment features of the remote sensing scene where the target is located (such as desert, city, countryside, etc.).

[0019] Among them, for the operations described in "Step 4", namely, "through the application programming interface, call the large language model, and input the feature descriptive phrases corresponding to each specific target together with the pre-compiled prompt template to automatically, efficiently, and diversely generate language descriptions that meet the requirements", the specific implementation is as follows: First, compile the prompt template for the large language model to ensure that no descriptive phrase information of the target is added or omitted during the input process, and at the same time guide the large language model to diversely use vocabulary and grammar structures when generating language descriptions to improve the diversity and richness of language generation. Second, set the number of times the large language model generates language descriptions for a single target. In the case of less target description information, the recommended reference value is 2, that is, generate two different language description versions for each target. Finally, input the descriptive phrase information of the target and the prompt template into the large prediction model together to achieve automatic language annotation. After all annotations are completed, conduct manual sampling inspection.

[0020] Among them, for the operations described in "Step 5", namely, "correspond the generated language descriptions with the annotated remote sensing images one by one, and classify and store them according to the semantic labels and geomorphic features of each target", the specific implementation is as follows: Since each target in this dataset contains two or more language descriptions, to ensure the one-to-one correspondence between the original remote sensing image, binary segmentation mask, target detection box, and language description, use the image-language description pair as the unique identifier (ID) of an object for numbering management. At the same time, record the semantic label and geomorphic feature to which each target belongs. Generally speaking, each object ID in the dataset should include the following content: original remote sensing image, binary segmentation mask annotation, target detection box, corresponding multiple language descriptions, semantic label, and geomorphic feature.

[0021] Advantages of the Invention

[0022] The advantages of the present invention are as follows: It provides a remote sensing image target recognition dataset based on image-text fusion, promotes the development of artificial intelligence technology in the fields of intelligent analysis of remote sensing images and target sensing and recognition, and fills the gap in high-quality remote sensing image datasets that integrate visual and language information at home and abroad. This dataset combines two-dimensional remote sensing images, binary segmentation masks, target detection bounding boxes, and diverse target descriptions, and proposes a dataset construction method of "data collection - data analysis - data normalization annotation - data classification management", providing a standardized process and technical solution, and setting an industry standard for the development and application of remote sensing image analysis and target recognition datasets. At the same time, the design method of this dataset not only enhances the multi-modal characteristics of remote sensing image data, but also lays a solid foundation for subsequent intelligent analysis and deep learning model training, and is expected to play an important role in multiple remote sensing application fields such as resource exploration and environmental monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Flowchart for constructing a remote sensing image target recognition dataset based on image-text fusion

[0024] Figure 2 Schematic diagram of standardized annotation of remote sensing images

[0025] Figure 3 Schematic diagram for generating remote sensing target language descriptions based on large language models DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention relates to a remote sensing image target recognition dataset based on image-text fusion and a construction method thereof, as shown in Figure 1 , Figure 2 , Figure 3 , and the steps are as follows:

[0027] Step 1: According to the target category, use open-source tools such as Google Earth to search for and determine the target area, adjust the resolution and perspective to ensure that the target display ratio is appropriate, and export the image and save it locally.

[0028] Step 2: Import the image into open-source annotation software, perform semantic label and segmentation annotation, and use an algorithm to convert the segmentation points into a binary mask image and a target detection bounding box.

[0029] Step 3: Analyze images of various categories, summarize common features, form descriptive phrases, and further describe the specific features of the target based on the segmentation annotation results, including relative position and geographical environment features.

[0030] Step 4: Call a large language model through an application programming interface, input the target description phrase and template, and automatically generate diverse language descriptions.

[0031] Step 5: Correlate the generated language descriptions with the images and classify and store them according to semantic tags and geomorphic features.

[0032] Among them, the method of Step 1 is as follows:

[0033] According to the pre-selected acquisition target category, search for the target area in Google Earth, reasonably adjust the ground resolution and tilt angle to ensure that the display ratio of the same category of targets in the remote sensing image is appropriate and the viewing coverage is wide. After the adjustment is completed, use the export function of Google Earth to save the remote sensing image to the local storage device at a resolution of 3840x2160. For the convenience of data analysis and collation, a single remote sensing image only contains targets of the same category. Each remote sensing image is named in the format of "category + serial number", and a separate folder is created for each target category, and the serial numbers are stored incrementally starting from 0. For example, the first remote sensing image of the bridge category is named "bridge_00001.jpg".

[0034] Among them, the method of Step 2 is as follows:

[0035] Use open-source annotation software to perform segmentation area annotation on the remote sensing image. This process includes three main steps: segmentation point area annotation, segmentation mask image generation, and target detection rectangle conversion, as shown in Figure 2 Schematic diagram of standardized annotation of remote sensing images. First, for the segmentation point area annotation, manually and finely annotate the targets in the remote sensing image, draw segmentation points point by point along the target contour, and determine a closed area that only contains the target from the segmentation point set. Then, for the segmentation mask image generation, based on the segmentation point set, call the library function to generate a binary image, where the background area is set to 0 (black) and the target area is set to 1 (white). After generating the binary image, call the library function to save the segmentation mask image to the local. Finally, for the target detection rectangle conversion, based on the segmentation point area annotation result, generate the target detection rectangle by calculating the minimum bounding rectangle of the target area. Each rectangle is represented by a quadruple (x1, y1, x2, y2), where x1 and y1 are the coordinates of the upper left corner of the rectangle, and x2 and y2 are the coordinates of the lower right corner, accurately positioning the target in the image.

[0036] Considering that there may be multiple targets of the same category in a remote sensing image, when saving the segmentation mask image, it is named in the format of "category name + image serial number + segmentation mask serial number". For example, the first target of the first remote sensing image of the bridge category is named "bridge_00001_01.png". Similarly, the annotation information of the target detection rectangle is stored in a text file format, and the file name format is "bridge_00001_01.txt". This method ensures the accuracy and standardization of remote sensing image segmentation annotation and target detection, facilitating subsequent data management and analysis.

[0037] Among them, the method for step three is as follows:

[0038] For different category targets in the dataset, their general features such as color and shape are manually summarized. Subsequently, by visualizing each target in the remote sensing image one by one (framed using the target detection rectangular boxes generated in step two), its specific features are manually analyzed, including the absolute position of the target (such as the target is located on the left side of the image), relative position (such as the positional relationship relative to other objects), and the geographical environment features of the remote sensing scene where the target is located (such as desert, city, countryside, etc.). For all descriptive information of each target, it is stored in the data format of dictionary key-value pairs. For target features that cannot be described, the value of the corresponding key is set to a null value. After obtaining all the descriptive information of a target, this data is stored in a dictionary form in a.json file. The naming rule follows the format of "category name + image serial number + target serial number". For example, the feature description file of the first target in the first remote sensing image of the bridge category is named "bridge_00001_01.json". This method ensures that the feature information of each target is stored in a structured manner, facilitating subsequent retrieval and analysis.

[0039] Among them, the method for step four is as follows:

[0040] A prompt template for the large language model is compiled to ensure that all input information is accurately reflected in the generated language description. The brief prompt template is as follows: "Hello, you will play the role of a language generator and you have strong prompt language generation capabilities. I need you to generate a logically coherent, grammatically correct, and accurately worded sentence from a series of descriptive phrases or words. All information in the descriptive phrases or words must be reflected in the generated sentence without addition or reduction." Subsequently, the large language model is called through the application programming interface, and the prompt template corresponding to the target and the descriptive information are input together to generate the language description of the target. When a target requires multiple language descriptions, the prompt template can be fine-tuned by adding diverse prompt words (such as "please use different expression methods") to enrich the expression form without changing the actual meaning of the language description. For a single target, its two corresponding language descriptions are stored as.txt files in string format, with the naming format of "category name + image serial number + target serial number + language description serial number". For example, the first language description file of the first target in the first remote sensing image of the bridge category is named "bridge_00001_01_01.txt". This method ensures that the description of each target is accurate and compliant, and provides diverse language expressions for further analysis and application.

[0041] Among them, the method for step five is as follows:

[0042] The data generated in the aforementioned steps are corresponded one by one and sorted out, and finally a remote sensing image target recognition dataset based on image-text fusion is constructed. In this dataset, multiple language descriptions of each target have a unique identifier (ID), and each ID corresponds to the following information: the original remote sensing image, the binary segmentation mask annotation, the target detection box, the language description, the inherent features (such as color, shape, etc.), and the specific features (such as the absolute position, relative position, geographical environment, etc. of the target in the image). All information related to this target will be stored in the format of dictionary key-value pairs, and the values corresponding to the information that cannot be obtained are set to empty to ensure the integrity and consistency of the data structure. Then, all the generated IDs and their corresponding dictionary structures are added to a list one by one for unified management. Finally, the entire list is exported in json format to form a complete remote sensing image target recognition dataset based on image-text fusion. This dataset structurally integrates remote sensing images and language descriptions, has high flexibility and scalability, and can provide comprehensive data support for subsequent intelligent analysis, model training, and multi-modal information fusion.

[0043] Currently, the remote sensing image target recognition dataset based on image-text fusion constructed using the aforementioned steps includes 14,560 IDs, a total of 5,281 two-dimensional remote sensing images, and 15 categories including "ships", "bridges", "trains", etc. The construction of this dataset ensures the traceability and integrity of each target's information, and provides a standardized and normalized data solution for the field of remote sensing image target recognition.

Claims

1. A method for normalizing and constructing a remote sensing image target recognition dataset based on image-text fusion, characterized in that: Step 1: Use open source remote sensing tools such as Google Earth to search and determine the target area according to the preset target category. During the acquisition process, reasonably adjust the ground resolution and viewing angle to ensure that the target in the remote sensing image is displayed in the image at an appropriate scale. Step 2: Import the saved high-resolution remote sensing images into the open source annotation software to annotate the target objects with detailed semantic labels and segmentation points. During the annotation process, the target outline is clearly defined through manual annotation, and the segmentation algorithm is used to convert these annotation points into a binary mask image and a target detection rectangle. Step 3: Conduct a systematic analysis of each type of remote sensing image and summarize the general and specific features of the target in that category. General features include physical features such as shape and color, while specific features involve the relative position, absolute position and environmental characteristics of the target in the image. Step 4: Use the application programming interface to call the large language model and input the description phrases and prompt templates corresponding to each target to automatically generate a variety of language descriptions, ensuring that the diversity and richness of language expression are improved without changing the description information. Step 5: Classify and manage the data. Each target data object in the dataset is numbered and managed by a unique identifier (ID) to ensure that it contains the following content: two-dimensional remote sensing image, segmentation mask image, target detection rectangle, and multiple language descriptions and their semantic labels. All data is stored by classification to ensure that the dataset can still be efficiently retrieved and applied in the case of multiple categories and multiple annotations.

2. The remote sensing image target recognition dataset based on image-text fusion and the construction method thereof according to claim 1, characterized in that: In step 1, remote sensing images are collected based on open source platforms such as Google Earth. The ground resolution and viewing angle are adjusted to ensure that the proportion of the target in the image is appropriate. After the image is collected, the image is classified and managed using a standardized naming system. The naming rule is "category + serial number", such as "bridge_0001.jpg", to facilitate batch processing and management of data of different categories.

3. The remote sensing image target recognition dataset based on image-text fusion and the construction method thereof according to claim 1, characterized in that: In step 2, the specific method of segmentation point annotation is to use open source annotation software to manually select the contour point set of the target object, and then use the library function to fill the internal area of ​​the target object with white (value 1) and the background area with black (value 0) to generate a binary mask image. At the same time, the minimum bounding rectangle of the target is generated based on the segmentation points for subsequent target detection task annotation.

4. The remote sensing image target recognition dataset based on image-text fusion and the construction method thereof according to claim 1, characterized in that: In step 3, a series of description phrases are formed by manually summarizing the characteristics of the target category. The description phrases of each category include the general physical characteristics of the target object (such as shape and color), as well as the location characteristics of the specific target in the category in the image and the surrounding environment characteristics. All these description information are stored in a dictionary format to ensure the correlation between different features for subsequent retrieval and analysis.

5. The remote sensing image target recognition dataset based on image-text fusion and the construction method thereof according to claim 1, characterized in that: In step 4, the large language model is called through the application programming interface to generate language descriptions. First, a prompt template is preset to ensure that the language model covers all key features when generating descriptions. For each target, the model generates multiple versions of descriptions to ensure the diversity and accuracy of the description information, and saves the generated description files in the format of "category + image number + description number" for easy management and retrieval.

6. The remote sensing image target recognition dataset based on image-text fusion and the construction method thereof according to claim 1, characterized in that: In step 5, each target data in the dataset is uniquely identified and managed by an ID, and each ID includes multiple pieces of information such as the original remote sensing image, the binary segmentation mask image, the target detection rectangle and its language description. This data storage structure ensures that all relevant information of each target can be quickly retrieved and supports subsequent intelligent analysis and model training.

Citation Information

Cited By

  • Video understanding description generation method, device and equipment based on remote sensing time sequence data

    CN121214055A