Extensible remote sensing deep learning sample library construction method based on target region planning

Through target area planning and multimodal image acquisition, combined with large language models and knowledge graph construction, the problems of high construction cost and poor scalability of remote sensing sample libraries have been solved, and the scalability and automated labeling capabilities of remote sensing sample libraries have been improved.

CN120612563APending Publication Date: 2025-09-09AEROSPACE INFORMATION RES INST CAS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510591064.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The existing remote sensing sample library construction method has the problems of high construction cost and poor scalability. It is difficult to flexibly expand to cover new categories, resulting in low data utilization.

Method used

A target area planning method is adopted to determine the target area range based on image resolution and multiple target area center points, and multi-modal and multi-scale remote sensing image acquisition is carried out. Vector boundary clipping and coding naming are combined, and corpus data is generated using a large language model. A knowledge graph is constructed and potential samples are mined through an open vocabulary target detection model.

Benefits of technology

It achieves the scalability and semantic relevance of the remote sensing sample library, improves the automatic labeling capability of the sample library, reduces the sample production cost and improves data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612563A_ABST
    Figure CN120612563A_ABST
Patent Text Reader

Abstract

The invention provides an extensible remote sensing deep learning sample library construction method based on target region planning, which is applied to the technical field of remote sensing image processing, and comprises the following steps: determining target region ranges of a plurality of target regions based on image resolution and a plurality of target region center points; performing cutting and coding naming on the target region remote sensing image based on the vector boundary of the target region range and a preset coding rule to obtain coarse samples corresponding to the plurality of target regions respectively; inputting the coarse sample subjected to feature enhancement processing into a large language model to obtain corpus data of the coarse sample output by the large language model; performing feature alignment with the coarse sample based on the semantic relationship to obtain a feature semantic mapping result of the coarse sample; constructing a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping result; and inputting the coarse sample knowledge graph into the open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model. According to the invention, dynamic extensible labeling of large-scale remote sensing samples can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a method for constructing an expandable remote sensing deep learning sample library based on target area planning. Background Art

[0002] The Remote Sensing Sample Library is a database that systematically stores and manages remote sensing images and their annotated information. It integrates multi-source remote sensing images acquired from platforms such as satellites and drones, performs pre-processing such as radiometric and geometric correction, and combines it with manual or semi-automatic annotation (such as feature classification and target location marking) to form a standardized, labeled dataset.

[0003] The existing sample library construction plan is task-driven. The preset sample types are difficult to fully cover the diversity of regional / global surface environments. When encountering new categories, the sample library cannot be flexibly expanded. It is necessary to re-expand the sample classification system, collect remote sensing images, and refine the sample library construction process, resulting in high sample production costs and low data utilization. It can be seen from this that the remote sensing sample library construction method in related technologies has technical problems such as high construction cost and poor scalability. Summary of the Invention

[0004] The present invention provides a method for constructing an extensible remote sensing deep learning sample library based on target area planning, which is used to solve the defects of the existing remote sensing sample library construction method in the prior art, such as high construction cost and poor scalability, and realize dynamic and extensible labeling of large-scale remote sensing samples.

[0005] The present invention provides a method for constructing an expandable remote sensing deep learning sample library based on target area planning, comprising the following steps.

[0006] Based on the image resolution and the center points of multiple target areas, the target area ranges of multiple target areas are determined, wherein the target area ranges match the image resolution; multimodal and multi-scale remote sensing images are collected based on the target area ranges to obtain target area remote sensing images; the target area remote sensing images are cropped and coded and named based on the vector boundaries of the target area ranges and preset coding rules to obtain coarse samples corresponding to the multiple target areas respectively; the coarse samples after feature enhancement processing are input into a large language model to obtain corpus data of the coarse samples output by the large language model; entities and relationships are extracted from the corpus data based on natural language processing to obtain semantic relationship information of the coarse samples; and features are aligned with the coarse samples based on the semantic relationships to obtain feature semantic mapping results of the coarse samples; a coarse sample knowledge graph is constructed based on the semantic relationship information and the feature semantic mapping results; the coarse sample knowledge graph is input into an open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model.

[0007] According to a method for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention, before determining the target area ranges of multiple target areas based on image resolution and multiple target target area center points, the method also includes: generating initial target area center points based on the earth's spherical surface based on a spherical uniform distribution algorithm or grid division technology to obtain multiple initial target area center points; and adjusting the multiple initial target area center points according to a preset regional image acquisition difficulty and preset natural factors to obtain multiple target target area center points.

[0008] According to a method for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention, the target area range of multiple target areas is determined based on image resolution and multiple target area center points, including: performing the following steps for each of the multiple target area center points to obtain the target area range of each target area center point: obtaining a remote sensing image corresponding to the target area center point; when the resolution of the remote sensing image is greater than a first spatial resolution threshold, taking the remote sensing image as a low-resolution image, and setting the target area range of the target area center point to a first range area; when the resolution of the remote sensing image is less than the first spatial resolution threshold and greater than a second spatial resolution threshold, taking the remote sensing image as a medium-resolution image, and setting the target area range of the target area center point to a second range area; when the resolution of the remote sensing image is less than the second spatial resolution threshold, taking the remote sensing image as a high-resolution image, and setting the target area range of the target area center point to a third range area.

[0009] According to a method for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention, the coarse sample includes: a coarse sample image and a coarse sample code. The target area remote sensing image is cropped and coded and named based on the vector boundary of the target area range and a preset coding rule to obtain coarse samples corresponding to the multiple target areas, including: using the vector boundary of the target area range as a cropping basis, segmenting the target area remote sensing image to obtain a coarse sample image; coding and naming the coarse sample image to obtain a coarse sample code, wherein the coarse sample code includes: target area number, image acquisition time, image modality and image resolution.

[0010] According to a method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, the method includes inputting a coarse sample after feature enhancement processing into a large language model to obtain corpus data of the coarse sample output by the large language model, including: performing image enhancement according to the modal attributes of the coarse sample to obtain the coarse sample after feature enhancement processing, wherein the modal attributes include: optical image and radar image; using annotation information in an open map database as auxiliary prompt words, inputting the coarse sample after feature enhancement processing into a large language model, and obtaining corpus data of the coarse sample output by the large language model, wherein the corpus data includes at least one of the following: image attribute description, image content summary, and description of geographic entity and spatial relationship.

[0011] According to a method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, after inputting the coarse sample knowledge graph into an open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model, the method further includes: slicing the potential samples according to preset slicing requirements to obtain sample slices, wherein the preset slicing requirements include at least one of the following: slice size, slice shape, and degree of ground feature integrity; inputting the sample slices and the coarse sample knowledge graph into a pre-trained annotation model to obtain the annotation content of the sample slices.

[0012] The present invention also provides an expandable remote sensing deep learning sample library construction device based on target area planning, comprising the following modules: a target area module, for determining the target area range of multiple target areas based on image resolution and multiple target area center points, wherein the target area range matches the image resolution; an image module, for performing multi-modal and multi-scale remote sensing image acquisition based on the target area range to obtain a target area remote sensing image; a coarse sample module, for cropping and encoding the target area remote sensing image based on the vector boundary of the target area range and a preset encoding rule to obtain coarse samples corresponding to the multiple target areas; a corpus module, for collecting the target area remote sensing image after feature enhancement processing; and a corpus module, for performing multi-modal and multi-scale remote sensing image acquisition based on the target area range to obtain a target area remote sensing image. The processed coarse sample is input into the large language model to obtain the corpus data of the coarse sample output by the large language model; a mapping module is used to extract entities and relationships from the corpus data based on natural language processing to obtain semantic relationship information of the coarse sample; and feature alignment is performed with the coarse sample based on the semantic relationship to obtain the feature semantic mapping result of the coarse sample; a construction module is used to construct a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping result; an expansion module is used to input the coarse sample knowledge graph into an open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements any of the above-described methods for constructing a scalable remote sensing deep learning sample library based on target area planning.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for constructing a scalable remote sensing deep learning sample library based on target area planning as described in any of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned methods for constructing a scalable remote sensing deep learning sample library based on target area planning.

[0016] The present invention provides a method for constructing a scalable remote sensing deep learning sample library based on target area planning. This method achieves systematic data coverage through target area planning and multimodal image acquisition, improves sample standardization efficiency by combining coding cropping and feature enhancement, utilizes a large language model to generate corpus and extract semantic relationships, constructs a knowledge graph that integrates multimodal features and semantic information, and finally mines potential samples through an open vocabulary detection model, thereby enhancing the scalability, semantic relevance, and automated labeling capabilities of the remote sensing sample library. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 It is a flow chart of the method for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention.

[0019] Figure 2 It is the technical roadmap provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the rough sample annotation and raw data corpus provided by the present invention.

[0021] Figure 4 This is a schematic diagram of the knowledge graph construction provided by the present invention.

[0022] Figure 5 It is a schematic diagram of the process of sample labeling production provided by the present invention.

[0023] Figure 6It is a structural schematic diagram of the device for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention.

[0024] Figure 7 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0025] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0026] By building a high-quality remote sensing sample library, we can provide rich and accurate training data for deep learning models, significantly improving the effectiveness of remote sensing image feature extraction. In the remote sensing field, deep learning technology has been widely applied to important tasks such as scene understanding, object detection, and land cover classification. A high-quality sample library can support models to better complete these tasks, improving the accuracy and efficiency of interpretation. Furthermore, building a remote sensing sample library will help enhance the practicality of intelligent remote sensing interpretation systems, enabling their application in more fields.

[0027] While deep learning techniques have been used in remote sensing to support various tasks and have achieved some success by building large amounts of sample data to train deep learning networks, existing sample libraries suffer from numerous issues. First, different sample libraries employ different classification systems, making it difficult for trained deep learning models to share sample sets and prone to classification bias when processing data outside the sample library's coverage. Second, existing sample library construction schemes are task-driven, and the pre-set sample types fail to fully capture the diversity of regional and global surface environments. This inflexible expansion of sample libraries to accommodate new categories requires re-expanding the sample classification system, acquiring remote sensing imagery, and refining the sample library construction process. This leads to high sample production costs and low data utilization. Furthermore, most existing sample libraries are constructed based on the ImageNet model, which inadequately reflects the multi-scale, multi-sensor, and multi-temporal characteristics of remote sensing imagery. Furthermore, most lack geographic and temporal attributes, which weakens the robustness of the models.

[0028] Currently, sample construction work is generally carried out when specific needs arise. First, a remote sensing interpretation classification system that meets the requirements is constructed according to the existing classification system; secondly, sample annotation areas are selected according to geographical characteristics, land feature types and sample quantity, and remote sensing image data of the corresponding annotation areas are obtained; then, the remote sensing images containing the target land features are identified and sliced, and then the slices are finely annotated to generate samples; then, quality inspection is carried out using design indicators to ensure the accuracy and reliability of the data; finally, intelligent interpretation and change discovery are carried out through land feature classification and change detection models, forming a "sample-model-knowledge" sample library construction method.

[0029] Traditional dataset construction processes generally lack unified specifications and standardized processes, resulting in unstable data sources and difficulty ensuring dataset quality. Furthermore, because traditional construction methods establish a direct, rigid connection between raw images and samples, raw image data is difficult to reuse, limiting dataset scalability and maintaining high construction costs.

[0030] The purpose of the present invention is to address the shortcomings of existing methods and provide a method for constructing a scalable remote sensing deep learning sample library based on target area planning and image corpusization, so as to realize dynamic and scalable annotation of large-scale remote sensing samples.

[0031] Optionally, the method for constructing a scalable remote sensing deep learning sample library based on target area planning in an embodiment of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device. Taking the example of the method for constructing a scalable remote sensing deep learning sample library based on target area planning in this embodiment being executed by a server.

[0032] Figure 1 This is a flow chart of the method for constructing an extensible remote sensing deep learning sample library based on target area planning provided by the present invention. Figure 1 As shown, the method includes the following steps.

[0033] Step 101 : determining target area ranges of multiple target areas based on image resolution and multiple target area center points, wherein the target area ranges match the image resolution.

[0034] In an embodiment of the present invention, based on the spherical distribution characteristics of the earth, a spherical uniform distribution algorithm or grid division technology is used to generate initial target area center points on a global scale, and the generated initial center points are adjusted to optimize the distribution of target area center points by taking into account natural factors such as geographical terrain (such as mountains, plains, oceans, etc.), climate zones (such as tropical, temperate, and cold zones), and the feasibility of image acquisition in the region.

[0035] For example, a spherical uniform distribution algorithm (such as the Fibonacci grid algorithm) is used to generate initial target area center points globally to ensure uniform geographic coverage. GIS tools are used to screen initial points based on topography (mountains, plains, water bodies) and climate zones (tropical, temperate).

[0036] For example, reduce redundant areas: sparse target areas in low-data-value areas such as oceans and polar regions; densify key areas: increase target area density in urban areas and ecologically sensitive areas.

[0037] Specific target areas will be supplemented according to industry needs (agriculture, urban planning), such as adding target area center points in concentrated farmland areas.

[0038] In some embodiments, based on the adjusted target area center point, GIS software (such as ArcGIS) is used to generate a vector boundary polygon. For example, the target area side length is automatically calculated according to the resolution level (such as 50km corresponds to a low-resolution target area); the boundary shape preferably uses a regular quadrilateral to ensure the integrity of the image cropping.

[0039] Based on the spatial resolution characteristics of remote sensing images, the target area is carefully delineated to ensure that the target area matches the resolution of the image data, thereby improving the effectiveness of the sample data.

[0040] Step 102: collect multi-modal and multi-scale remote sensing images based on the target area to obtain a remote sensing image of the target area.

[0041] In an embodiment of the present invention, based on the acquisition of multimodal and multiscale remote sensing images of the target area, multimodal (such as optical, radar, hyperspectral, etc.) and multi-scale (with different spatial and temporal resolutions) remote sensing images covering the target area are collected to provide multi-source data support for the sample library.

[0042] Step 103 : cropping and coding the remote sensing image of the target area based on the vector boundary of the target area and a preset coding rule to obtain coarse samples corresponding to the plurality of target areas.

[0043] In an embodiment of the present invention, the remote sensing image is cut into image segments (coarse samples) within the target area according to the target area vector boundary, and a unique code is generated for each coarse sample. The coding rules include key information such as the target area number, image acquisition time, and image type. The coarse samples are named in a standardized manner to ensure a unified file name format for easy subsequent management and retrieval.

[0044] The coarse samples are classified and compiled according to multi-dimensional attributes such as target area, image type, and acquisition time. A preliminary quality check is performed on the samples to be stored to ensure data integrity and validity, and the classified coarse samples are stored in the sample database.

[0045] Step 104: Input the coarse sample after feature enhancement processing into the large language model to obtain corpus data of the coarse sample output by the large language model.

[0046] In an embodiment of the present invention, after extracting a coarse sample, histogram equalization and contrast stretching are applied to the optical image to enhance the visual effect. Spectral curves are extracted and texture features are enhanced using gray-level co-occurrence matrices and wavelet transforms. For radar images, surface elevation changes are acquired using InSAR technology, and polarimetric radar decomposition is used to analyze scattering characteristics.

[0047] Multi-dimensional corpus generation based on large language models uses coarse sample data after feature enhancement as the core input, and uses pre-trained large-scale multi-domain language models to generate multi-dimensional, semantically rich corpus data.

[0048] For example, first, for the geographic entity information in the coarse sample, such as mountains, rivers, cities, forests, etc., a detailed geographic description corpus is generated, and the spatial position of the geographic entity in the image and the spatial relationship between it and other surrounding geographic entities are also explained. Secondly, for the image feature information in the coarse sample, including spectral features, texture features, and shape features, the large language model can convert the visual features in the image into textual expressions.

[0049] Step 105 , extracting entities and relationships from the corpus data based on natural language processing to obtain semantic relationship information of the coarse sample; and aligning features with the coarse sample based on the semantic relationship to obtain feature semantic mapping results of the coarse sample.

[0050] In an embodiment of the present invention, natural language processing technology is used to extract core entities (such as place names, feature names) and semantic relationships in the corpus, and the semantic relationships are aligned with the sample features based on the attribute information of the sample to form a semantic description of the sample.

[0051] Step 106: construct a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping results.

[0052] Based on the rich semantic relationship information and feature semantic mapping results obtained through semantic relationship extraction and association alignment, a knowledge graph for the remote sensing sample library is constructed. Entity types include geographic entities, remote sensing image entities, industry application entities, and semantic concept entities. Relationship types include geospatial relationships, semantic logical relationships, image-entity mapping relationships, and industry application relationships. Attribute types define the various attributes of each entity, such as the name, location, area, and altitude of geographic entities, and the resolution, capture time, and sensor type of image entities.

[0053] Step 107: Input the coarse sample knowledge graph into the open vocabulary object detection model to obtain potential samples output by the open vocabulary object detection model.

[0054] In an embodiment of the present invention, based on the knowledge of geographic entities, land features, semantic relationships, etc. in the above-mentioned knowledge graph, target detection and image recognition technology are used to automatically search for potential coarse sample areas (i.e., potential samples) in the image.

[0055] Through the above steps of the embodiment of the present invention, systematic data coverage is achieved through target area planning and multimodal image acquisition, the efficiency of sample standardization is improved by combining coding cropping and feature enhancement, a large language model is used to generate corpus and extract semantic relationships, a knowledge graph is constructed to integrate multimodal features and semantic information, and finally potential samples are mined through an open vocabulary detection model, thereby enhancing the scalability, semantic relevance and automatic labeling capabilities of the remote sensing sample library.

[0056] According to a method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, before determining the target area ranges of multiple target areas based on image resolution and multiple target area center points, the method further includes: Based on the spherical uniform distribution algorithm or grid division technology, the initial target area center point is generated based on the earth's spherical surface to obtain multiple initial target area center points; According to the difficulty of obtaining the preset regional image and the preset natural factors, multiple initial target area center points are adjusted to obtain multiple target area center points.

[0057] First, a random algorithm is used to uniformly generate longitude and latitude pairs around the globe, which serve as the basis for the initial distribution of target area centers. This process is designed to ensure the breadth and spatial uniformity of the initial target area selection and avoid bias in regional coverage.

[0058] Next, the Earth's surface is divided into regular grid cells using discrete grid technology, and the initial target center points within each grid are further adjusted and optimized. During this process, the richness of the grid area's content is evaluated and classified: for areas with less content (such as vast oceans and polar regions), the grid density is appropriately reduced or redundant grids are deleted; for areas with richer content (such as densely populated urban areas), the density and distribution of target center points are appropriately increased. This stage of screening and optimization results in a more scientific second-round target center point distribution.

[0059] Finally, the feasibility of data acquisition for each target area is assessed, and additional data is selected for areas with limited data. In particular, priority is given to adding target areas to areas with limited data coverage, such as underdeveloped regions and extreme geographical locations (such as remote islands and high-altitude areas). This ensures sufficient extreme samples are obtained and improves the comprehensiveness of the sample library. After comprehensively considering all screening factors, the final target area distribution is determined to form a target point set with high coverage, high representativeness, and diversity.

[0060] According to the present invention, a method for constructing a scalable remote sensing deep learning sample library based on target area planning is provided, which determines the target area ranges of multiple target areas based on image resolution and multiple target area center points, including: For each of the multiple target area center points, perform the following steps to obtain the target area range of each target area center point: Obtain the remote sensing image corresponding to the center point of the target area; When the resolution of the remote sensing image is greater than the first spatial resolution threshold, the remote sensing image is used as a low-resolution image, and the target area range of the target area center point is set as the first range area; When the resolution of the remote sensing image is less than the first spatial resolution threshold and greater than the second spatial resolution threshold, the remote sensing image is used as a medium-resolution image, and the target area range of the target area center point is set as the second range area; When the resolution of the remote sensing image is less than the second spatial resolution threshold, the remote sensing image is used as a high-resolution image, and the target area range of the target area center point is set as the third range area.

[0061] In this embodiment of the present invention, the image source is further determined based on the center point of each target area. The spatial resolution of different imaging products is then combined to rationally design the corresponding target area coverage, taking into account data acquisition costs, to ensure the completeness of image coverage and adaptability to the target area. In specific implementation, the target area range is adjusted in a graded manner based on the spatial resolution of the image, forming a system that matches resolution and coverage.

[0062] Ultra-low-resolution images (spatial resolution > 100m) are usually suitable for large-scale regional feature observation and scene recognition and classification tasks. Therefore, the target area is set to 500km*500km to fully cover the geographical and environmental features in a wide area.

[0063] For low-resolution images (spatial resolution of 10m-100m), their coverage is more suitable for observation and classification analysis of regional geographic features. Therefore, the corresponding target area is set to 50km × 50km. This range can preserve local area details while taking into account large-scale environmental characteristics. For medium-resolution images (spatial resolution of 5m - 10m), this resolution is mainly aimed at the observation needs of more specific landform types or detailed scenes. Therefore, the corresponding target area is set to 10km × 10km, focusing more on the spatial distribution characteristics of the medium scale.

[0064] For high-resolution imagery (spatial resolution of 1m-5m), its application scenarios typically include building identification and traffic network analysis, which require more refined spatial resolution capabilities. Therefore, the target area is further reduced to 5km × 5km to emphasize the capture of local detailed information.

[0065] Finally, for ultra-high-resolution imagery (spatial resolution ≤ 1m), which is widely used for fine-grained feature extraction and high-precision scene analysis, to maximize the image's detail capabilities, the corresponding target area is set to 2km × 2km, so that every part covered by the image can accurately describe the detailed features of the target area.

[0066] Through the embodiment of the present invention, the design of adjusting the target area coverage based on spatial resolution ensures that image acquisition achieves a balance between wide-area coverage and local refinement and overall data acquisition cost, providing adaptive support for multi-scale and multi-resolution analysis of remote sensing data.

[0067] According to a method for constructing an extensible remote sensing deep learning sample library based on target area planning provided by the present invention, the coarse sample includes: a coarse sample image and a coarse sample code. The target area remote sensing image is cropped and coded and named based on the vector boundary of the target area range and a preset coding rule to obtain coarse samples corresponding to multiple target areas, including: The vector boundary of the target area is used as the basis for cutting, and the remote sensing image of the target area is segmented to obtain a coarse sample image; The coarse sample image is coded and named to obtain a coarse sample code, wherein the coarse sample code includes: target area number, image acquisition time, image modality, and image resolution.

[0068] Based on a globally defined target area, the remote sensing imagery covering the target area is segmented using its vector boundaries as a cropping basis to generate a coarse sample image. Specifically, the target area's vector boundaries are spatially overlaid with the image data to extract the image content within the target area. The cropping operation considers the pixel integrity of the image data to ensure that important features are not lost in the boundary area due to cropping. Furthermore, larger image data is cropped in blocks based on the size of the target area.

[0069] After cropping, a unique code is generated for each coarse sample for identification and management. This code consists of several key information fields, including the target area number, image acquisition time, imaging modality (such as optical, radar, hyperspectral, etc.), and other important attributes. In this example, the specific encoding rules can be designed as follows: [Target Area Number]-[Image Acquisition Time]-[Image Modality]-[Image Resolution]. For example, T123-20250102-OPT-0.5m indicates a coarse sample with target area number T123, acquisition time January 2, 2025, image type optical image, and spatial resolution of 0.5 meters.

[0070] Furthermore, to improve management efficiency, standardized naming is used for raw sample files. This naming convention is typically consistent with encoding conventions, ensuring that all key information is included in the file name, facilitating file retrieval and classification. For example, a cropped image file might be named T123_20250102_OPT_0.5m.tif. This consistent file name format not only improves systematic data management but also facilitates automated processing processes (such as data migration and batch annotation).

[0071] After cropping and naming are complete, the resulting coarse samples undergo a preliminary inspection. This inspection includes checking whether the image fully covers the target area, whether there are any missing pixels or distortions, and whether the naming rules are correct. If any problems are found, they are promptly corrected or re-cropped to ensure that each coarse sample meets the requirements of subsequent processing.

[0072] Through the embodiments of the present invention, remote sensing images are precisely cropped by vector boundaries to ensure that the spatial positioning of the coarse sample is strictly aligned with the target area planning, eliminating geometric misalignment errors; the coding naming rules embed multidimensional metadata such as target area number, time, modality, resolution, etc. into the file name, realizing unique identification of sample identity and rapid retrieval of attributes.

[0073] According to a method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, a coarse sample after feature enhancement processing is input into a large language model to obtain corpus data of the coarse sample output by the large language model, including: Performing image enhancement according to the modal attributes of the coarse sample to obtain the coarse sample after feature enhancement processing, wherein the modal attributes include: optical image and radar image; The annotation information in the open map database is used as auxiliary prompt words, and the coarse samples after feature enhancement processing are input into the large language model to obtain the corpus data of the coarse samples output by the large language model. The corpus data includes at least one of the following: image attribute description, image content summary, and description of geographic entities and spatial relationships.

[0074] In an embodiment of the present invention, for optical images, methods such as histogram equalization and contrast stretching are used to improve the visual effects and information representation capabilities of the images, and certain spectral bands are individually adjusted according to the characteristic requirements of the specific images. For example, the vegetation bands are normalized to highlight the vegetation coverage information.

[0075] For radar imagery, feature enhancement is achieved based on the different physical properties of radar imaging. For example, InSAR (Interferometric Synthetic Aperture Radar) technology is used to extract surface elevation changes from images, generating high-precision digital elevation models (DEMs) to support terrain analysis. Polarimetric radar decomposition technology is also used to analyze the scattering characteristics of images. For example, surface reflectance characteristics (such as the different scattering patterns of vegetation, buildings, and bare ground) are extracted to accurately identify ground object types.

[0076] In addition, for all modal image data, texture feature extraction techniques are combined to further enrich sample information. For example, gray-level co-occurrence matrices (GLCMs) are used to analyze image texture characteristics, wavelet transforms are applied to decompose images into different frequency components, and texture details at different scales are analyzed.

[0077] After feature extraction and enhancement, the newly generated feature data is stored in a standardized format and a mapping relationship is established with the original coarse sample.

[0078] In this embodiment of the present invention, to address the shortcomings of existing techniques in labeling remote sensing sample libraries, a multidimensional corpus generation method based on transfer learning and a large language model is used. This method combines a large visual language model with open map databases (such as OpenStreetMap) to achieve multidimensional semantic description and label generation for remote sensing images. The specific steps are as follows: Feature enhancement of sample data: Using feature-enhanced raw sample data as the core input, and using annotations from the Open Map Database as auxiliary prompts, this is fed into the visual language model. The model generates descriptions containing multi-dimensional semantic information through the visual encoding layer and the text decoding layer.

[0079] Multidimensional semantic description generation: For remote sensing images, the generated multidimensional semantic information includes the following aspects: Image attribute description: First, basic image attribute information is generated, including shooting source, shooting time, shooting angle, image type, etc. These descriptions are obtained by fine-tuning the visual language model. Image content summary: This further summarizes the general content of the image, including the type, distribution, spectral characteristics, and texture characteristics of the main content. Geographic entities and spatial relationships are described.

[0080] Using data from open map databases, we generate the spatial locations of geographic entities in an image and their relationships. The specific implementation steps include obtaining latitude and longitude information, geographic information tags, and integrating semantic information.

[0081] Through the embodiments of the present invention, by implementing differentiated enhancements for the modal characteristics of optical and radar images (such as color correction of optical images and texture enhancement of radar images), the complementarity of multimodal data is fully exploited; combined with the semantic priors of open map annotations to guide the large language model to generate structured corpus, cross-modal alignment of the visual features of remote sensing images and geographic semantics is achieved.

[0082] According to a method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, after inputting the coarse sample knowledge graph into an open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model, the method further includes: Slicing the potential sample according to preset slicing requirements to obtain sample slices, wherein the preset slicing requirements include at least one of the following: slice size, slice shape, and feature integrity; The sample slices and the coarse sample knowledge graph are input into the pre-trained annotation model to obtain the annotation content of the sample slices.

[0083] In this embodiment of the present invention, automatically discovered coarse samples (i.e., potential samples) are sliced ​​and divided into multiple subsample slices using an image segmentation algorithm based on preset slice size, shape, or feature integrity requirements. Based on semantic information from the knowledge graph and a pre-trained annotation model, the subsample slices are automatically annotated with information such as feature type (e.g., specific tree species, crop type), geographic attributes (e.g., area, altitude range), and semantic relationships (e.g., interactions with surrounding features), generating standardized, fully annotated sample slices.

[0084] Perform a comprehensive quality check on sample slices and annotated samples, checking image clarity, color accuracy, and accuracy of annotation information. Samples with quality issues are processed through image restoration and annotation correction techniques. If they still fail, they are discarded. Qualified samples are stored in a sample database using a specific data format and storage structure, and the database index and metadata information are updated.

[0085] refer to Figure 2 , Figure 2 This is the technical roadmap provided by the present invention, which includes: global target area planning, coarse sample collection, sample knowledge construction and sample annotation generation.

[0086] Global target area planning includes generating a globally uniform target area center, adjusting target area centers, adaptively resizing target areas, and supplementing industry target areas. Coarse sample collection includes acquiring multimodal and multiscale remote sensing images, extracting metadata, cropping and coding coarse samples, and cleaning and compiling coarse samples. Sample knowledge construction includes coarse sample feature enhancement, LLM-based multidimensional corpus generation, semantic relationship extraction and association alignment, and knowledge graph construction. Sample annotation production includes automated coarse sample discovery, sample slicing, automated sample annotation, and sample quality inspection.

[0087] The fine samples are stored in the fine sample database, the corpus samples are stored in the coarse sample database, the coarse samples are stored in the basic database, and the global uniform target area is stored in the basic database.

[0088] The following describes an example of the practical application of the method for constructing a scalable remote sensing deep learning sample library based on target area planning provided by the present invention, which specifically includes the following steps.

[0089] Step 1: Generate and adjust the center point of the global uniform target area.

[0090] First, a random algorithm is used to uniformly generate longitude and latitude pairs around the globe, which serve as the basis for the initial distribution of target area centers. This process is designed to ensure the breadth and spatial uniformity of the initial target area selection and avoid bias in regional coverage.

[0091] Secondly, the Earth's surface is divided into regular grid cells using discrete grid technology, and the initial target center points within each grid are further adjusted and optimized. During this process, the richness of the grid area's content is assessed and classified. For regions with less content (such as vast oceans and polar regions), the grid density is appropriately reduced or redundant grids are deleted. For regions with richer content (such as densely populated urban areas), the density and distribution of target center points are appropriately increased. In selecting and adjusting target areas, both natural geographic and human factors are fully considered to ensure diversity and representativeness of target content. Regarding natural factors, terrain features (such as mountains, plains, hills, and oceans) and climate zones (such as tropical, temperate, and frigid zones) are integrated to cover the Earth's major natural landscape types. Regarding human factors, transportation infrastructure (such as highways, railways, and overpasses) and the built environment (such as different types of residential and commercial areas and landmark buildings) are considered to reflect the diversity of regional characteristics and human activities.

[0092] Finally, the feasibility of data acquisition in each target area will be assessed, and additional data will be selected for areas with limited data. In particular, additional target areas will be prioritized for areas with limited data coverage, such as underdeveloped regions and extreme geographical locations (such as remote islands and high-altitude areas). This will ensure sufficient extreme samples are obtained and improve the comprehensiveness of the sample library. After comprehensively considering all screening factors, the final target area distribution will be determined to form a target point set with high coverage, high representativeness, and diversity, laying a solid foundation for subsequent data acquisition and analysis.

[0093] Step 2: Determine the target area based on image resolution.

[0094] Based on the center point of each target area, the image source is further determined. Taking into account the spatial resolution of different image products, the corresponding target area coverage is rationally designed while taking into account data acquisition costs to ensure the completeness of image coverage and adaptability to the target area. In specific implementation, the target area range is graded and adjusted according to the spatial resolution of the image, forming a system that matches resolution and coverage. For ultra-low resolution images (spatial resolution > 100m), they are usually suitable for large-scale regional feature observation and scene recognition and classification tasks. Therefore, the target area is set to 500km*500km to fully cover the geographical and environmental features in a wide area. For low-resolution images (spatial resolution of 10m-100m), their coverage is more suitable for the observation and classification analysis of regional geographical elements. Therefore, the corresponding target area is set to 50km×50km. This range can preserve local regional details while taking into account large-scale environmental characteristics. For medium-resolution images (spatial resolution of 5m-10m), this resolution is mainly aimed at the observation needs of more specific landform types or detailed scenes. Therefore, the corresponding target area is set to 10km×10km, focusing more on medium-scale spatial distribution characteristics. For high-resolution images (spatial resolution of 1m-5m), their application scenarios usually include building recognition, traffic network analysis, etc., which require more refined spatial resolution capabilities. Therefore, the target area is further reduced to 5km×5km to highlight the capture of local detailed information. Finally, for ultra-high resolution images (spatial resolution ≤ 1m), widely used for fine-grained feature extraction and high-precision scene analysis. To maximize the image's detail, the corresponding target area is set to 2km × 2km, ensuring that every portion of the image coverage accurately depicts the detailed features of the target area. By adjusting the target area coverage based on spatial resolution, image acquisition achieves a balance between wide-area coverage, local refinement, and overall data acquisition cost, providing adaptive support for multi-scale and multi-resolution analysis of remote sensing data.

[0095] Step 3: Supplement industry target areas.

[0096] In order to ensure that the sample library can meet the actual application needs of multiple industries and scenarios, the initially determined target area will be further optimized and improved.

[0097] Based on specific industry needs and actual application scenarios, a new round of refined adjustments and supplementation of the target area will be conducted. Through the collaboration and refined adjustments of experts from multiple industries, the target area of ​​the sample library will not only be broadly representative, but also accurately meet the actual needs of multiple fields such as agriculture, forestry, urban planning, and ecological environment. Ultimately, a multi-level and multi-angle optimization of the target area will be achieved.

[0098] Step 4: Acquire multi-modal and multi-scale remote sensing images based on the target area.

[0099] Based on the identified data sources, a scientific and stable target area data acquisition strategy is developed to ensure the consistency and efficiency of data collection. This strategy covers several key aspects, including the periodic arrangement of acquisition time, data type, and transmission protocol.

[0100] In terms of scheduling, flexible collection cycles can be set based on the specific characteristics of the target area and application requirements. For example, in areas with significant seasonal changes, collection can be carried out on a seasonal or monthly basis; while in areas with rapidly changing environmental dynamics, the collection interval can be shortened to capture key events or changing trends.

[0101] In terms of data types, multimodal remote sensing imagery covering the target area is collected, including data from different sensors such as optical, radar, and hyperspectral sensors to meet diverse research needs. Optical imagery provides clear visible light information of ground objects, radar imagery offers all-weather and all-day imaging capabilities, and hyperspectral imagery supports more detailed ground object analysis through its rich spectral dimension information. Furthermore, by collecting data at multiple scales, taking into account differences in spatial and temporal resolution, comprehensive coverage is achieved, from macroscopic observation to microscopic analysis.

[0102] Step 5: Rough sample cutting and coding naming.

[0103] Based on a globally defined target area, the remote sensing imagery covering the target area is segmented using its vector boundaries as a cropping basis to generate a coarse sample image. Specifically, the target area's vector boundaries are spatially overlaid with the image data to extract the image content within the target area. The cropping operation considers the pixel integrity of the image data to ensure that important features are not lost in the boundary area due to cropping. Furthermore, larger image data is cropped in blocks based on the size of the target area.

[0104] After cropping, a unique code is generated for each coarse sample for identification and management. This code consists of several key information fields, including the target area number, image acquisition time, imaging modality (such as optical, radar, hyperspectral, etc.), and other important attributes. In this example, the specific encoding rules can be designed as follows: [Target Area Number]-[Image Acquisition Time]-[Image Modality]-[Image Resolution]. For example, T123-20250102-OPT-0.5m indicates a coarse sample with target area number T123, acquisition time January 2, 2025, image type optical image, and spatial resolution of 0.5 meters.

[0105] Furthermore, to improve management efficiency, standardized naming is used for raw sample files. This naming convention is typically consistent with encoding conventions, ensuring that all key information is included in the file name, facilitating file retrieval and classification. For example, a cropped image file might be named T123_20250102_OPT_0.5m.tif. This consistent file name format not only improves systematic data management but also facilitates automated processing processes (such as data migration and batch annotation).

[0106] After cropping and naming are complete, the resulting coarse samples undergo a preliminary inspection. This inspection includes checking whether the image fully covers the target area, whether there are any missing pixels or distortions, and whether the naming rules are correct. If any problems are found, they are promptly corrected or re-cropped to ensure that each coarse sample meets the requirements of subsequent processing.

[0107] Step 6: Compile and store rough samples.

[0108] After the coarse samples are cropped and coded, they are systematically organized to ensure efficient data management and utilization. First, the coarse samples are categorized and organized based on multi-dimensional attributes such as target area number, image type, acquisition time, and resolution. Samples within each target area are grouped by modality, such as optical imagery and radar imagery, and further refined into sub-directories covering different time periods. Layered management is also implemented based on spatial resolution.

[0109] Next, the organized rough samples undergo a comprehensive data integrity and quality check. This includes checking for file corruption, ensuring the cropped area fully covers the target area, and whether the image contains blur or noise. The sample naming and metadata attributes are verified to ensure consistency with the actual image content, ensuring that the encoding rules are correct. If quality issues are identified, such as damaged or duplicated image files, unqualified samples are repaired or removed to ensure high-quality data storage.

[0110] After data inspection is complete, qualified rough samples are stored in the remote sensing sample database. Using a hierarchical directory structure that separates metadata from the data itself, samples are organized and stored according to target area and image attributes. Metadata is stored in a separate database, while data is stored locally. Database index information, including target area range, image type, and temporal distribution, is also updated to support efficient retrieval and dynamic expansion.

[0111] Finally, establish a comprehensive backup mechanism to ensure data security, regularly storing sample data and database indexes off-site or in the cloud. Also, set up access control to ensure only authorized users can view or operate the database, minimizing losses caused by data loss or misoperation.

[0112] Step 7: Coarse sample feature enhancement.

[0113] After the rough sample is cropped and stored, the image data needs to be enhanced to improve its information expression and the accuracy of subsequent processing. Feature enhancement mainly includes image enhancement processing and feature extraction for different modal data such as optical images and radar images.

[0114] First, for optical images, methods such as histogram equalization and contrast stretching are used to improve the visual effects and information presentation capabilities of the images. Based on the specific image feature requirements, certain spectral bands are adjusted individually. For example, the vegetation band is normalized to highlight the vegetation coverage information.

[0115] For radar imagery, feature enhancement is achieved based on the different physical properties of radar imaging. For example, InSAR (Interferometric Synthetic Aperture Radar) technology is used to extract surface elevation changes from images, generating high-precision digital elevation models (DEMs) to support terrain analysis. Polarimetric radar decomposition technology is also used to analyze the scattering characteristics of images. For example, surface reflectance characteristics (such as the different scattering patterns of vegetation, buildings, and bare ground) are extracted to accurately identify ground object types.

[0116] In addition, for all modal image data, texture feature extraction techniques are combined to further enrich sample information. For example, gray-level co-occurrence matrices (GLCMs) are used to analyze image texture characteristics, wavelet transforms are applied to decompose images into different frequency components, and texture details at different scales are analyzed.

[0117] After feature extraction and enhancement, the newly generated feature data is stored in a standardized format and a mapping relationship is established with the original coarse sample.

[0118] Step 8: Multi-dimensional corpus generation based on large language model.

[0119] To address the shortcomings of existing techniques for labeling remote sensing imagery, we employed a multidimensional corpus generation method based on transfer learning and a large language model. This method combines a large visual language model with open map databases (such as OpenStreetMap) to achieve multidimensional semantic description and label generation for remote sensing images. The specific steps are as follows: Feature enhancement of sample data: Using feature-enhanced raw sample data as the core input, and using annotations from the Open Map Database as auxiliary prompts, this is fed into the visual language model. The model generates descriptions containing multi-dimensional semantic information through the visual encoding layer and the text decoding layer.

[0120] Multi-dimensional semantic description generation: For remote sensing images, the generated multi-dimensional semantic information includes the following aspects: Image attribute description: First, the basic attribute information of the image is generated, including shooting source, shooting time, shooting angle, image type, etc. These descriptions are obtained by fine-tuning the visual language model. Image content summary: Further summarize the general content of the image, including the type, distribution, spectral characteristics and texture characteristics of the main content. Description of geographic entities and spatial relationships Using data from an open map database, we generate the spatial locations of geographic entities in an image and their relationships. Specific implementation steps include obtaining latitude and longitude information, geographic information tags, and integrating semantic information.

[0121] refer to Figure 3 , Figure 3 This is a schematic diagram of the rough sample annotation and raw data corpus provided by the present invention, which includes: Description 1: This image is a satellite view of an urban area with a distinct grid pattern. The upper portion shows a dense concentration of buildings, roads, and green areas such as parks or sports fields. Several large rectangular areas appear to be residential campuses or institutional buildings. The lower portion features more open space and agricultural land, indicating a transition from urban to rural environments.

[0122] Description 2: This aerial photograph captures a cityscape. A major road cuts through the middle, dividing the landscape into two main sections. The upper section consists of numerous blocks filled with closely spaced buildings, likely apartments or offices, interspersed with patches of vegetation. The lower section shows large plots of farmland surrounded by smaller buildings.

[0123] Description 3: This image presents a bird's-eye view of an organized urban layout. To the north, there are several square areas containing high-density buildings, possibly representing commercial or residential areas. To the south, the landscape shifts towards agricultural use, marked by neat plots of farmland and scattered houses.

[0124] Caption 4: A detailed satellite map shows the transition from a structured urban environment to a suburban and rural landscape. The central area shows tightly clustered buildings separated by straight streets, while the southern area highlights vast farmlands divided into regular plots.

[0125] Step 9: Semantic relationship extraction and association alignment.

[0126] In this step, the image description generated in step 8 is used as input. A natural language processing model (such as PromptCLUE-base) is used to extract nouns from the corpus. Keyword information, such as place names, geographic tag names, and feature names, is further extracted from the nouns. The relationships between keywords are also identified and used as attribute information for the sample. Based on the extracted semantic relationships, they are aligned with the sample features to generate corresponding semantic descriptions. Finally, the multidimensional corpus, keywords, attribute relationship information, and remote sensing image metadata (such as image source, storage address, latitude and longitude, unique ID, shooting time, sensor information, etc.) generated in steps 8 and 9 are integrated and packaged into a standardized JSON file to provide data support for the subsequent knowledge graph construction process.

[0127] Step 10: Constructing the coarse sample knowledge graph.

[0128] refer to Figure 4 , Figure 4 This is a schematic diagram of the knowledge graph construction provided by the present invention.

[0129] like Figure 4 As shown, the satellite image (red center node) is used as the core and is divided into three main parts through the blue main branches: the urban part, which includes the street network, commercial areas, high-density buildings, residential areas, and green spaces and vegetation; the urban-rural transition part, which is centered on the urban-rural fringe and includes suburbs and is classified as an urban-rural fringe; and the rural part, which includes farmland, cultivated land, and scattered houses.

[0130] Based on the semantic relationship information extracted in step nine, the feature-semantic mapping results, and the generated JSON files for each image, a knowledge graph for the remote sensing sample library is constructed. This knowledge graph design covers geographic entities, remote sensing image entities, semantic concept entities, and industry application entities, clarifying the spatial, logical, mapping, and application relationships between entities and incorporating detailed attribute information such as geographic location, resolution, and domain classification. The construction process utilizes a lightweight knowledge graph generation tool to automate the process from JSON data to a complete knowledge graph, providing support for the semantic organization of remote sensing data and knowledge-driven applications.

[0131] Step 11: Automatic discovery of coarse samples.

[0132] Step 11 uses target detection and image recognition technology, combined with geographic entities, land features, and semantic relationships in the knowledge graph, to automatically search for potential coarse sample areas in the image, improving discovery efficiency and accuracy. This method uses an open vocabulary target detection model to take keywords and geographic tags extracted from the knowledge graph as input labels. Unlike traditional closed-set detection models, the open vocabulary target detection model can detect targets that have not appeared during training through similarity matching in the semantic space, achieving open vocabulary detection. After obtaining the detection box and category information, the segmentation model is used to segment the image and generate a mask image for each category. At the same time, the category label, bounding box, unique ID, and other annotations corresponding to the mask are generated to provide data support for the subsequent organization and expansion of sample library materials.

[0133] Step 12: Automatically slice and label the sample.

[0134] In this step, the automatically discovered coarse samples are sliced ​​according to preset slice size, shape, or feature integrity requirements. Based on the semantic information in the knowledge graph and a pre-trained annotation model, the subsample slices are automatically annotated. This annotation includes information such as feature type (e.g., specific tree species, crop type), geographic attributes (e.g., area, altitude range), and semantic relationships (e.g., interactions with surrounding features), thereby generating standardized, fully annotated sample slices.

[0135] refer to Figure 5 , Figure 5 This is a flow chart of the sample annotation production process provided by the present invention, which includes: primary sample library, LLMs, corpus sample library, knowledge extraction, remote sensing knowledge graph, primary sample automatic discovery, SAM model, semi-automatic sample generation, sports field samples and fine sample library.

[0136] Step 13: Sample quality inspection and storage.

[0137] This step performs a comprehensive quality check on the sliced ​​and annotated samples to ensure image clarity, color accuracy, and the accuracy of the annotation information. Samples with quality issues are processed using techniques such as image restoration and annotation correction. Those that still fail are discarded. Finally, qualified samples are stored in a sample database using a specific data format and storage structure, and the database's index and metadata are updated.

[0138] In the above-mentioned embodiment of the present invention, this method collects coarse samples based on global random target area planning and industry target area supplementation, efficiently slices and automatically inspects multimodal and multi-resolution remote sensing satellite images, and transmits, organizes, and stores the produced coarse samples in a basic database. Instead of producing samples directly from original images, the information extraction process from image-corpus-knowledge is realized based on a large language model and human-in-the-loop. The produced corpus samples are stored in a corpus sample library, a sample knowledge base is constructed based on a knowledge graph, and fine samples are annotated and produced based on the coarse sample knowledge base. The knowledge base is used to achieve dynamic scalability of the sample set and has a certain open annotation capability.

[0139] Existing approaches to building remote sensing sample libraries rely on direct image annotation, but this approach presents technical bottlenecks such as sample target classification, fixed spatial scale, high production costs, low data utilization, and an inability to integrate existing sample library resources. Compared to traditional direct image annotation methods, this invention, through image corpusization, makes sample library construction more dynamic, scalable, and reconfigurable, addressing existing issues such as high remote sensing data annotation costs, slow updates, and a single service model. This improves the utilization efficiency of remote sensing data resources and the automation rate of sample annotation, reducing the costs of sample library construction and maintenance.

[0140] The following describes the scalable remote sensing deep learning sample library construction device based on target area planning provided by the present invention. The scalable remote sensing deep learning sample library construction device based on target area planning described below and the scalable remote sensing deep learning sample library construction method based on target area planning described above can be referenced to each other.

[0141] refer to Figure 6 , Figure 6 It is a structural schematic diagram of the device for constructing an expandable remote sensing deep learning sample library based on target area planning provided by the present invention.

[0142] A target area module 601 is configured to determine target area ranges of multiple target areas based on image resolution and multiple target area center points, wherein the target area ranges match the image resolution; An imaging module 602 is configured to acquire multi-modal and multi-scale remote sensing images based on the target area to obtain a remote sensing image of the target area; The coarse sample module 603 is used to crop and encode the remote sensing image of the target area based on the vector boundary of the target area and the preset encoding rules to obtain coarse samples corresponding to multiple target areas; Corpus module 604, used to input the coarse sample after feature enhancement processing into the large language model to obtain corpus data of the coarse sample output by the large language model; Mapping module 605 is used to extract entities and relationships from corpus data based on natural language processing to obtain semantic relationship information of the coarse sample; and to align features with the coarse sample based on the semantic relationship to obtain feature semantic mapping results of the coarse sample; A construction module 606 is used to construct a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping result; The expansion module 607 is used to input the coarse sample knowledge graph into the open vocabulary object detection model to obtain potential samples output by the open vocabulary object detection model.

[0143] Specifically, the above-mentioned scalable remote sensing deep learning sample library construction device based on target area planning provided by the present invention can implement all the method steps implemented in the above-mentioned scalable remote sensing deep learning sample library construction method embodiment based on target area planning, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those of the method embodiment will not be described in detail here.

[0144] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 7 As shown, the electronic device may include: a processor (processor) 710 , a communication interface (Communications Interface) 720 , a memory (memory) 730 and a communication bus 740 , wherein the processor 710 , the communication interface 720 and the memory 730 communicate with each other via the communication bus 740 . The processor 710 can call the logic instructions in the memory 730 to execute a scalable remote sensing deep learning sample library construction method based on target area planning, the method including: determining the target area range of multiple target areas based on image resolution and multiple target area center points, wherein the target area range matches the image resolution; performing multimodal and multi-scale remote sensing image acquisition based on the target area range to obtain the target area remote sensing image; cropping and encoding and naming the target area remote sensing image based on the vector boundary of the target area range and a preset encoding rule to obtain coarse samples corresponding to multiple target areas; inputting the coarse samples after feature enhancement processing into the large language model to obtain corpus data of the coarse samples output by the large language model; performing entity and relationship extraction on the corpus data based on natural language processing to obtain semantic relationship information of the coarse samples; and performing feature alignment on the coarse samples based on the semantic relationship to obtain feature semantic mapping results of the coarse samples; constructing a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping results; inputting the coarse sample knowledge graph into the open vocabulary target detection model to obtain potential samples output by the open vocabulary target detection model.

[0145] Furthermore, the logic instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0146] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the scalable remote sensing deep learning sample library construction method based on target area planning provided by the above methods. The method includes: determining the target area range of multiple target areas based on image resolution and multiple target area center points, wherein the target area range matches the image resolution; performing multi-modal and multi-scale remote sensing image acquisition based on the target area range to obtain a target area remote sensing image; based on the vector boundary of the target area range and the preset coding rules The remote sensing images of the target area are cropped and coded to obtain coarse samples corresponding to multiple target areas; the coarse samples after feature enhancement are input into the large language model to obtain the corpus data of the coarse samples output by the large language model; entities and relationships are extracted from the corpus data based on natural language processing to obtain semantic relationship information of the coarse samples; and features are aligned with the coarse samples based on the semantic relationships to obtain the feature semantic mapping results of the coarse samples; a coarse sample knowledge graph is constructed based on the semantic relationship information and the feature semantic mapping results; the coarse sample knowledge graph is input into the open vocabulary target detection model to obtain the potential samples output by the open vocabulary target detection model.

[0147] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by the processor to execute the scalable remote sensing deep learning sample library construction method based on target area planning provided by the above methods, the method comprising: determining the target area range of multiple target areas based on image resolution and multiple target area center points, wherein the target area range matches the image resolution; performing multi-modal and multi-scale remote sensing image acquisition based on the target area range to obtain the target area remote sensing image; cropping and encoding the target area remote sensing image based on the vector boundary of the target area range and the preset coding rules. Encode and name to obtain coarse samples corresponding to multiple target areas; input the coarse samples after feature enhancement processing into the large language model to obtain the corpus data of the coarse samples output by the large language model; extract entities and relationships from the corpus data based on natural language processing to obtain semantic relationship information of the coarse samples; and align features with the coarse samples based on the semantic relationship to obtain the feature semantic mapping results of the coarse samples; construct a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping results; input the coarse sample knowledge graph into the open vocabulary target detection model to obtain the potential samples output by the open vocabulary target detection model.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. That is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0149] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for constructing a scalable remote sensing deep learning sample library based on target area planning, characterized in that: include: Determining target area ranges of the plurality of target areas based on the image resolution and the center points of the plurality of target areas, wherein the target area ranges match the image resolution; Perform multi-modal and multi-scale remote sensing image acquisition based on the target area range to obtain a remote sensing image of the target area; Based on the vector boundary of the target area range and a preset coding rule, the remote sensing image of the target area is cropped and coded and named to obtain rough samples corresponding to the multiple target areas respectively; Inputting the coarse sample after feature enhancement processing into the large language model to obtain corpus data of the coarse sample output by the large language model; Performing entity and relationship extraction on the corpus data based on natural language processing to obtain semantic relationship information of the coarse sample; and performing feature alignment on the coarse sample based on the semantic relationship to obtain a feature semantic mapping result of the coarse sample; Constructing a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping result; The coarse sample knowledge graph is input into an open vocabulary object detection model to obtain potential samples output by the open vocabulary object detection model.

2. The method for constructing a scalable remote sensing deep learning sample library based on target area planning according to claim 1 is characterized in that: Before determining the target ranges of the plurality of target areas based on the image resolution and the plurality of target area center points, the method further includes: Based on the spherical uniform distribution algorithm or grid division technology, the initial target area center point is generated based on the earth's spherical surface to obtain multiple initial target area center points; The multiple initial target area center points are adjusted according to the preset regional image acquisition difficulty and preset natural factors to obtain multiple target area center points.

3. The method for constructing a scalable remote sensing deep learning sample library based on target area planning according to claim 1 is characterized in that: Determining target ranges of multiple target areas based on image resolution and multiple target area center points includes: The following steps are performed for each of the plurality of target area center points to obtain a target area range for each target area center point: Acquire a remote sensing image corresponding to the center point of the target area; When the resolution of the remote sensing image is greater than a first spatial resolution threshold, the remote sensing image is used as a low-resolution image, and the target area range of the target area center point is set as a first range area; When the resolution of the remote sensing image is less than the first spatial resolution threshold and greater than the second spatial resolution threshold, the remote sensing image is used as a medium-resolution image, and the target area range of the target area center point is set as a second range area; When the resolution of the remote sensing image is less than the second spatial resolution threshold, the remote sensing image is used as a high-resolution image, and the target area range of the target area center point is set as a third range area.

4. The method for constructing a scalable remote sensing deep learning sample library based on target area planning according to claim 1, characterized in that: The coarse sample includes: a coarse sample image and a coarse sample code. The target area remote sensing image is cropped and coded based on the vector boundary of the target area range and the preset coding rules to obtain the coarse samples corresponding to the multiple target areas, including: Using the vector boundary of the target area as a basis for cropping, the remote sensing image of the target area is segmented to obtain a coarse sample image; The coarse sample image is coded and named to obtain a coarse sample code, wherein the coarse sample code includes: target area number, image acquisition time, image modality and image resolution.

5. The method for constructing an extensible remote sensing deep learning sample library based on target area planning according to claim 1, characterized in that: Inputting the coarse sample after feature enhancement processing into the large language model to obtain corpus data of the coarse sample output by the large language model includes: Performing image enhancement according to the modal attributes of the coarse sample to obtain a coarse sample after feature enhancement processing, wherein the modal attributes include: optical image and radar image; The annotation information in the open map database is used as auxiliary prompt words, and the coarse sample after feature enhancement processing is input into a large language model to obtain corpus data of the coarse sample output by the large language model. The corpus data includes at least one of the following: image attribute description, image content summary, and description of geographic entity and spatial relationship.

6. The method for constructing an extensible remote sensing deep learning sample library based on target area planning according to claim 1, characterized in that: After inputting the coarse sample knowledge graph into an open vocabulary object detection model to obtain potential samples output by the open vocabulary object detection model, the method further includes: Slicing the potential sample according to preset slicing requirements to obtain sample slices, wherein the preset slicing requirements include at least one of the following: slice size, slice shape, and feature integrity; The sample slice and the coarse sample knowledge graph are input into a pre-trained annotation model to obtain the annotation content of the sample slice.

7. A device for constructing an extensible remote sensing deep learning sample library based on target area planning, characterized in that: include: A target area module, configured to determine target area ranges of a plurality of target areas based on image resolution and a plurality of target area center points, wherein the target area ranges match the image resolution; An imaging module is used to collect multi-modal and multi-scale remote sensing images based on the target area to obtain a remote sensing image of the target area; A coarse sample module is used to crop and encode the remote sensing image of the target area based on the vector boundary of the target area range and a preset encoding rule to obtain coarse samples corresponding to the multiple target areas respectively; A corpus module, configured to input the coarse sample after feature enhancement processing into a large language model, and obtain corpus data of the coarse sample output by the large language model; A mapping module is configured to extract entities and relationships from the corpus data based on natural language processing to obtain semantic relationship information of the coarse sample; and perform feature alignment with the coarse sample based on the semantic relationship to obtain a feature semantic mapping result of the coarse sample; A construction module, configured to construct a coarse sample knowledge graph based on the semantic relationship information and the feature semantic mapping result; An expansion module is used to input the coarse sample knowledge graph into an open vocabulary object detection model to obtain potential samples output by the open vocabulary object detection model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, it implements the method for constructing a scalable remote sensing deep learning sample library based on target area planning as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for constructing a scalable remote sensing deep learning sample library based on target area planning as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for constructing a scalable remote sensing deep learning sample library based on target area planning as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Remote sensing open vocabulary target detection method based on multi-modal large language model

    CN121640482A

  • Remote sensing open vocabulary object detection method based on multi-modal large language model

    CN121640482B

  • A dynamic sample database construction method, device and medium for a remote sensing cross-scale interpretation large model

    CN122364480A