Remote sensing open vocabulary object detection method based on multi-modal large language model

The remote sensing open vocabulary target detection method based on a multimodal large language model solves the problem of poor adaptability of traditional remote sensing target recognition to non-preset ground features, generates high-precision open vocabulary detection results, and improves the adaptability and accuracy of remote sensing observation.

CN121640482BActive Publication Date: 2026-05-05SHAANXI TIRAIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHAANXI TIRAIN TECH CO LTD
Filing Date
2026-02-04
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional remote sensing target recognition methods rely heavily on manually pre-set fixed vocabulary databases, which cannot effectively identify non-pre-set ground features, resulting in poor adaptability in complex scenarios.

Method used

A remote sensing open vocabulary target detection method based on a multimodal large language model is adopted. By acquiring remote sensing images of the target area selected by the user, the target vocabulary set and matching degree set are identified. Combined with environmental remote sensing images, language large model recognition resources are configured to generate open vocabulary and calculate confidence scores, and the optimal detection results are output.

Benefits of technology

It achieves high-precision detection of non-preset ground features, reduces the false recognition rate in complex scenarios, improves the practicality and scene adaptability of remote sensing observation, and the dynamic resource allocation balances recognition accuracy and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640482B_ABST
    Figure CN121640482B_ABST
Patent Text Reader

Abstract

This application provides a remote sensing open vocabulary target detection method based on a multimodal large language model, belonging to the field of remote sensing image detection technology. The method includes: acquiring remote sensing images of a user-selected target area; performing remote sensing target recognition to obtain a target vocabulary set and a target vocabulary matching degree set; configuring the environmental remote sensing image division range to divide and acquire environmental remote sensing images within the environment of the target area, obtaining an environmental vocabulary set and an environmental vocabulary matching degree set; configuring language model recognition resources, randomly combining the environmental vocabulary set and the target vocabulary set, inputting them into the configured remote sensing large language model, and outputting a target open vocabulary set and a frequency set; calculating the open vocabulary confidence score, selecting the optimal open vocabulary as the target detection result for the target area. This solves the technical problem that existing remote sensing target recognition methods can only detect predefined categories and cannot address the recognition needs of non-predefined ground features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image detection, and more particularly to a remote sensing open vocabulary target detection method based on a multimodal large language model. Background Technology

[0002] Remote sensing target identification of target areas can quickly extract spatial distribution and category information of ground objects, providing core data support for key scenarios such as ecological monitoring, urban planning, and disaster emergency response.

[0003] However, traditional remote sensing target recognition methods have significant limitations: they rely heavily on manually pre-set fixed vocabulary databases and can only detect predefined categories such as stadiums and forests, failing to meet the recognition needs of a large number of non-pre-set features in real-world scenarios.

[0004] Therefore, there is an urgent need for an open vocabulary target detection method that integrates remote sensing image features to meet diverse remote sensing observation needs. Summary of the Invention

[0005] This invention addresses the technical problem that existing remote sensing target recognition methods can only detect predefined categories and cannot meet the recognition needs of non-predefined ground features. It provides a remote sensing open vocabulary target detection method based on a multimodal large language model.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0007] This invention provides a remote sensing open vocabulary target detection method based on a multimodal large language model, including:

[0008] Acquire remote sensing images of the target area selected by the user, perform remote sensing target recognition, and obtain a target vocabulary set and a target vocabulary matching degree set;

[0009] Based on the number of target words in the target vocabulary set, configure the environmental remote sensing image division range, divide and acquire environmental remote sensing images within the environment of the target area, perform remote sensing environment identification, and obtain the environmental vocabulary set and the environmental vocabulary matching degree set.

[0010] Based on the number of words in the target vocabulary set and the environment vocabulary set, configure language large model recognition resources, randomly combine the environment vocabulary set and the target vocabulary set, input them into the configured remote sensing large language model, and output the target open vocabulary set and the frequency set of occurrence.

[0011] Based on the target word matching set, the environmental word matching set, and the frequency of occurrence set, the confidence score of open words is calculated, and the optimal open words are selected as the target detection result for the target region.

[0012] The beneficial effects of this invention are:

[0013] Compared to existing technologies, this application first acquires remote sensing images of the target area selected by the user, performs remote sensing target identification, and obtains a target vocabulary set and a target vocabulary matching degree set, providing stable and quantifiable basic target information for subsequent environmental analysis and open vocabulary generation. Second, based on the number of target words in the target vocabulary set, it configures the environmental remote sensing image division range, divides and acquires environmental remote sensing images within the environment of the target area, performs remote sensing environment identification, and obtains an environmental vocabulary set and an environmental vocabulary matching degree set. This ensures sufficient environmental context information is obtained and acquires a more macroscopic environmental vocabulary, providing reliable data support for subsequent open vocabulary generation. Third, based on the number of words in the target vocabulary set and the environmental vocabulary set, it configures language model recognition resources, randomly combines the environmental vocabulary set and the target vocabulary set, inputs them into the configured remote sensing language model, and outputs a target open vocabulary set and a frequency set. This considers the total complexity of the target and environmental vocabulary and dynamically allocates the language model resources accordingly, generating a target open vocabulary that breaks through the preset library. Finally, based on the target word matching set, the environmental word matching set, and the frequency of occurrence set, the confidence of open words is calculated, and the optimal open words are selected as the target detection result of the target region. Through multi-dimensional fusion and quantitative evaluation, the target open words that are both consistent with the image facts and semantic logic are output.

[0014] Through the above technical solution, this application first identifies the target area image selected by the user to obtain a target vocabulary set, then dynamically adjusts the environmental image range according to the number of target words to obtain an environmental vocabulary set. Subsequently, the two types of words are randomly combined and input into a remote sensing large language model adapted to the resources, which can generate non-preset open words such as alpine meadows and plain farmland. Moreover, by calculating confidence in multiple dimensions to select the optimal results, it ensures that the open words fit the actual features of the image and have semantic and logical consensus, reducing the misidentification rate in complex scenes. At the same time, dynamic resource configuration takes into account both recognition accuracy and computational efficiency, avoiding resource waste or insufficient computing power. In this way, it solves the pain point of traditional remote sensing target recognition relying on a preset vocabulary library and having poor adaptability to non-preset categories, and can output scene-specific, high-precision open vocabulary detection results, improving the practicality and scene adaptability of remote sensing observation. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the remote sensing open vocabulary target detection method based on a multimodal large language model provided by the present invention.

[0016] Figure 2 This is a schematic diagram illustrating the process of inputting a remote sensing image of a target into a remote sensing target recognition network group and obtaining a target vocabulary set and a target vocabulary matching degree set in the remote sensing open vocabulary target detection method based on a multimodal large language model provided by the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0019] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0020] Examples, such as Figure 1 As shown, this embodiment of the invention provides a remote sensing open vocabulary target detection method based on a multimodal large language model, including:

[0021] S10: Acquire remote sensing images of the target area selected by the user, perform remote sensing target recognition, and obtain a target vocabulary set and a target vocabulary matching degree set.

[0022] Traditional remote sensing target recognition methods rely on a pre-set vocabulary and can only detect fixed categories such as stadiums and forests. The limited vocabulary severely restricts the observation experience in complex scenes, and its adaptability is insufficient, especially in scenes without pre-set categories.

[0023] To address the aforementioned issues, this application acquires remote sensing images of the target area selected by the user, performs remote sensing target identification, and obtains a target vocabulary set and a target vocabulary matching degree set.

[0024] Specifically, step S10 in the method includes:

[0025] Acquire remote sensing images of the target area selected by the user;

[0026] The target remote sensing image is input into a remote sensing target recognition network group, and the recognition output obtains a target vocabulary set and a target vocabulary matching degree set, wherein the matching degree of each target vocabulary is the proportion of each type of target vocabulary.

[0027] In this embodiment, the target remote sensing image of the target area selected by the user is first acquired. Specifically, the user determines the target area to be detected through interactive operations such as selecting a box on the remote sensing map. Then, a high-resolution image of the target area is extracted from the remote sensing image database, or a high-resolution image is captured by a drone as the target remote sensing image. This ensures that the object to be identified is accurately matched with the user's needs, avoids invalid analysis of irrelevant areas, and improves the targeting of the detection.

[0028] Secondly, the target remote sensing image is input into a remote sensing target recognition network group, and the recognition outputs a target vocabulary set and a target vocabulary matching degree set. The remote sensing target recognition network group is a collection of multiple independently trained remote sensing target recognition networks. Each network can independently recognize the target remote sensing image and output one target vocabulary. The parallel recognition of the remote sensing target recognition network group can reduce the accidental misjudgments caused by interference from image noise, shadows, cloud cover, etc., by a single model, thus improving the stability and accuracy of the recognition results.

[0029] Among them, the target vocabulary set is a collection of possible target categories identified based on the features of remote sensing images of the target, such as farmland and orchard, which fully reflects the potential target types within the target area.

[0030] The matching degree of each target word is the percentage of occurrences of each type of target word in the output results of the remote sensing target recognition network group. For example, if there are 10 remote sensing target recognition networks in the remote sensing target recognition network group, and 7 of them identify farmland, then the matching degree of farmland is 70%. Generally speaking, the higher the matching degree of the target word, the larger the proportion of the target word in the target remote sensing image and the clearer the visual features. The higher the credibility of its recognition results, the more quantifiable the basis for subsequent environmental analysis and open vocabulary generation.

[0031] Specifically, such as Figure 2 As shown, the step of "inputting the target remote sensing image into a remote sensing target recognition network group, and obtaining a target vocabulary set and a target vocabulary matching degree set by recognition output" includes:

[0032] Based on the remote sensing target identification and recording data, a set of sample target remote sensing images is collected, and the target words in each sample target remote sensing image are labeled to obtain a set of sample target words.

[0033] The sample target remote sensing image set and the sample target vocabulary set are randomly partitioned N times with replacement to obtain N target vocabulary recognition training datasets;

[0034] Based on machine learning, N remote sensing target recognition networks are constructed. The N target vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing target recognition network group is obtained.

[0035] The target remote sensing image is input into the remote sensing target recognition network group, and multiple target words are output. Target words that are in the same category are clustered, and the occurrence frequency ratio of each type of target word is calculated to obtain the target word set and the target word matching degree set.

[0036] In this embodiment of the application, a large number of remote sensing images covering different scenes and different land features are first collected based on remote sensing target identification and recording data, which serve as a sample target remote sensing image set. Then, the target words in each sample target remote sensing image are manually labeled to obtain a sample target word set.

[0037] Secondly, the sample target remote sensing image set and sample target vocabulary set are randomly partitioned N times with replacement to obtain N target vocabulary recognition training datasets. Each target vocabulary recognition training dataset will be used to independently train a remote sensing target recognition network. Since each target vocabulary recognition training dataset is generated through random partitioning with replacement, different target vocabulary recognition training datasets will have certain differences, thus ensuring differentiated training data for the N remote sensing target recognition networks and avoiding homogenization due to identical training data. N can be dynamically determined based on the actual sample size, computing resources, and recognition accuracy requirements. For example, if N=10, then 10 independent random partitioning with replacement will generate 10 distinct target vocabulary recognition training datasets from the sample target remote sensing image set and sample target vocabulary set.

[0038] Furthermore, based on machine learning, N remote sensing target recognition networks are constructed, and N target vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing target recognition network group is obtained.

[0039] For example, considering the characteristics of remote sensing images such as high resolution, multispectral features, and large differences in the scale of ground objects, an improved convolutional neural network (CNN) architecture can be used to construct N remote sensing target recognition networks. These networks mainly consist of an input layer, a multi-scale convolutional feature extraction layer, a pooling layer, a global feature fusion layer, and an output layer.

[0040] The input layer receives remote sensing images and performs preprocessing such as standardization, random cropping, and data augmentation on the images to improve the robustness of the remote sensing target recognition network to changes in lighting and shooting angle deviations in the remote sensing images.

[0041] The multi-scale convolutional feature extraction layer is used to capture ground feature details and global features. It adopts a structure of three convolutional blocks and multi-scale convolution in parallel. Based on the size differences of remote sensing ground features, multi-dimensional features are extracted by convolutional kernels of different sizes. The first convolutional block includes two consecutive convolutional layers (convolutional kernel size 3×3, number of kernels 64 and 128 respectively), stride 1, padding=1, used to extract the low-level texture features of the image. Each convolutional layer is followed by batch normalization and ReLU activation function. After the first convolutional block, three sets of convolutional kernels (1×1, 3×3, 5×5), each with 256 kernels, are set in parallel to capture small-scale details (such as leaf texture), medium-scale contours (such as single fruit trees), and large-scale regional features (such as the boundary of an entire orchard). Then, multi-scale features are fused by channel concatenation to avoid the omission of different-sized targets by a single convolutional kernel.

[0042] Among them, the pooling layer is used for feature dimensionality reduction and key information preservation. After each round of convolutional block / multi-scale convolutional layer, a max pooling layer is set with a pooling kernel size of 2×2 and a stride of 2. By preserving the maximum feature value in the local region, the feature map size is halved.

[0043] The global feature fusion layer is used to integrate deep semantic information. It consists of two fully connected layers: the first layer has 1024 neurons and the second layer has 512 neurons. The first layer flattens the high-dimensional feature map (e.g., 16×16×512) output by the pooling layer into a one-dimensional vector before global feature fusion. A dropout layer (dropout rate=0.5) is added to the fully connected layer to randomly shield some neurons and prevent the model from over-relying on a certain type of feature, which could lead to overfitting.

[0044] The output layer is used to map the target word category. It maps the 512-dimensional feature vector output by the fully connected layer to the preset target word category space (such as K target words including farmland, orchard, farmland, building, etc.). The Softmax activation function is used to output the probability value of each category (the sum of the probabilities of all categories is 1). The category with the highest probability is the target word recognition result of the network for the input image.

[0045] For example, during training, N target word recognition training datasets are used to independently train N remote sensing target recognition networks. For instance, the target word recognition training datasets are divided into training and validation sets in an 8:2 ratio. The sample target remote sensing images in the training set are used as input features, and the corresponding sample target words are used as supervision labels. The cross-entropy loss function is used to quantify the difference between the model output and the true labels. The Adam optimizer is used, with an initial learning rate of 1e-4. Cosine annealing scheduling is employed, with a training batch size of 32 (adapting to GPU memory) and 50 iterations. After each iteration, the model accuracy is evaluated using an independent validation set. If the validation set accuracy does not improve for five consecutive iterations, an early stopping strategy is triggered to terminate training, preventing overfitting. The model parameters at this point are saved, resulting in a successfully trained and performing remote sensing target recognition network. This training process is repeated, and after independently training N remote sensing target recognition networks using N target word recognition training datasets, they are combined to form a remote sensing target recognition network group.

[0046] Finally, the target remote sensing image is input into the remote sensing target recognition network group. Each remote sensing target recognition network independently outputs a target word. In this way, multiple target words are obtained. Then, the target words with the same type are clustered, and the occurrence rate of each type of target word is calculated to obtain the target word set and the target word matching degree set.

[0047] For example, the target remote sensing image is input into a remote sensing target recognition network group, and 10 target words are output. The same target words are grouped into one category: 7 farmlands, 1 orchard, 1 vegetable garden, and 1 wasteland. The occurrence rate of farmlands is 7 / 10 = 0.7, that of orchards is 1 / 10 = 0.1, that of vegetable gardens is 1 / 10 = 0.1, and that of wasteland is 1 / 10 = 0.1. The occurrence rate of each category of target words is the target word matching degree. Finally, the target word set {farmland, orchard, vegetable garden, wasteland} and the target word matching degree set {70%, 10%, 10%, 10%} are obtained.

[0048] In summary, compared to existing technologies, this application acquires remote sensing images of a user-selected target area, performs remote sensing target identification, and obtains a target vocabulary set and a target vocabulary matching degree set. Thus, by using a remote sensing target identification network group, it solves the problem of susceptibility to image noise interference in traditional single-model identification, providing stable and quantifiable basic target information for subsequent environmental analysis and open vocabulary generation.

[0049] S20: Based on the number of target words in the target vocabulary set, configure the environmental remote sensing image division range, divide and acquire the environmental remote sensing images within the environment where the target area is located, perform remote sensing environment identification, and obtain the environmental vocabulary set and the environmental vocabulary matching degree set.

[0050] The more complex the target area, the more susceptible its geographic semantics are to the influence of the surrounding macro-environment. Relying solely on the image features of the target area itself makes it difficult to accurately define the scene attributes of geographic features. Therefore, it is necessary to combine a larger range of environmental remote sensing imagery to capture sufficient macro-environmental features. By associating environmental semantics with target semantics, contextual support can be provided for the accurate generation of subsequent open vocabulary, avoiding semantic ambiguity caused by insufficient environmental information.

[0051] To address the aforementioned issues, this application configures the environmental remote sensing image segmentation range based on the number of target words within the target vocabulary set, segments and acquires environmental remote sensing images within the environment of the target area, performs remote sensing environment identification, and obtains an environmental vocabulary set and an environmental vocabulary matching degree set.

[0052] Specifically, step S20 in the method includes:

[0053] Obtain the number of target words in the target vocabulary set, calculate the ratio with the preset number of target words, and obtain the target vocabulary coefficient;

[0054] Obtain the preset environmental remote sensing imagery for delineation;

[0055] Based on the target vocabulary coefficients, the preset environmental remote sensing image division range is adjusted and calculated to obtain the environmental remote sensing image division range.

[0056] According to the defined range of the environmental remote sensing images, the environmental remote sensing images within the environment of the target area are divided and acquired.

[0057] The environmental remote sensing image is input into a remote sensing environment recognition network group, and the output is processed to obtain an environmental vocabulary set and an environmental vocabulary matching degree set.

[0058] In this embodiment, the number of target words in the target vocabulary set is first obtained, and the ratio of this number to the preset number of target words is calculated to obtain the target vocabulary coefficient. The target vocabulary coefficient is calculated as: Target Vocabulary Coefficient = Number of Target Words / Preset Number of Target Words. The preset number of target words is a baseline value set based on historical remote sensing identification data (such as the average number of target words in similar areas) or actual needs. For example, setting the preset number of target words to 3 represents the target vocabulary size for a medium-complexity area. For instance, if the number of target words in the target vocabulary set is 4 and the preset number of target words is 3, then the target vocabulary coefficient = 4 / 3 ≈ 1.33. Thus, the complexity of the target area is quantified through the target vocabulary coefficient. If the target vocabulary coefficient > 1, it indicates that the target area has many types of land cover, requiring more environmental information to assist in judgment; if the target vocabulary coefficient < 1, it indicates that the target area has a simple land cover, requiring no excessively large environmental scope to avoid information redundancy.

[0059] Secondly, the preset environmental remote sensing image delineation range is obtained. This range is a predefined area covered by the environmental image around the target region, typically defined using a radius plus shape, with the center point of the target region as the reference. The setting of the preset environmental remote sensing image delineation range needs to balance information sufficiency and computational efficiency. Setting it too small may result in insufficient environmental information, while setting it too large will lead to a surge in image data, increasing subsequent recognition time. Therefore, those skilled in the art can dynamically set the preset environmental remote sensing image delineation range according to the type of the target region and actual needs. For example, if the target region is a suburban mixed-feature area, typically including farmland, low-rise buildings, artificial green spaces, etc., belonging to a medium-complexity scene, then a circular area with a radius of 1000 meters and the center point of the target region is used as the preset environmental remote sensing image delineation range.

[0060] Next, based on the target vocabulary coefficient, the preset environmental remote sensing image division range is adjusted to obtain the environmental remote sensing image division range. Specifically, the preset environmental remote sensing image division range is scaled by the target vocabulary coefficient. For example, if the target vocabulary coefficient is 1.33, the preset environmental remote sensing image division range is a circular area with the center point of the target area as the center and a radius of 1000 meters. Multiplying the target vocabulary coefficient by the radius of the preset environmental remote sensing image division range, we get the radius of the environmental remote sensing image division range = 1000 × 1.33 = 1330 meters. Thus, depending on the complexity of the target area, the preset environmental remote sensing image division range is appropriately scaled by the target vocabulary coefficient. The more complex the target area, the larger the environmental remote sensing image division range, and the more surrounding information is obtained, ensuring that complex areas have sufficient context to improve recognition accuracy. The simpler the target area, the smaller the environmental remote sensing image division range, to reduce invalid data processing and improve data processing efficiency.

[0061] Furthermore, based on the defined range of the environmental remote sensing imagery, environmental remote sensing images within the environment of the target area are obtained. For example, according to the defined range of the environmental remote sensing imagery (such as a circular range with a radius of 1330 meters), the corresponding area of ​​the image is cropped from the remote sensing image database. For instance, based on the geographic coordinates of the remote sensing imagery, the center point coordinates of the target area are first located, and then the boundary coordinates are calculated according to the defined range of the environmental remote sensing imagery. Finally, the pixel area within the boundary coordinates is extracted from the large-scale remote sensing imagery as the environmental remote sensing imagery.

[0062] Finally, the environmental remote sensing imagery is input into a remote sensing environmental recognition network, and the output is processed to obtain an environmental vocabulary set and an environmental vocabulary matching score set. The environmental vocabulary consists of macroscopic terrain and landform categories such as mountains, plains, rivers, and hills. The environmental vocabulary matching score is the percentage of occurrences of each type of environmental vocabulary, reflecting the reliability of the environmental vocabulary.

[0063] Specifically, the step of "inputting the environmental remote sensing image into a remote sensing environment recognition network group and outputting the processed environmental vocabulary set and environmental vocabulary matching degree set" includes:

[0064] Based on remote sensing environment identification and recording data, a set of sample environment remote sensing images is collected, and environmental terms within each sample environment remote sensing image are labeled to obtain a set of sample environment terms.

[0065] The sample environment remote sensing image set and the sample environment vocabulary set are randomly partitioned N times with replacement to obtain N environmental vocabulary recognition training datasets.

[0066] Based on machine learning, N remote sensing environment recognition networks are constructed. The N environmental vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing environment recognition network group is obtained.

[0067] The environmental remote sensing image is input into the remote sensing environment recognition network group, and multiple environmental terms are output. Environmental terms that are in the same category are clustered, and the occurrence rate of each type of environmental term is calculated to obtain the environmental term set and the environmental term matching degree set.

[0068] In this embodiment of the application, remote sensing image data covering different environmental types (such as mountains, plains, rivers, hills, etc.) are first collected based on remote sensing environmental identification and recording data to form a sample environmental remote sensing image set. Environmental terms in each sample environmental remote sensing image are macroscopically labeled, such as plains, mountains, etc., to obtain a sample environmental terminology set.

[0069] Secondly, the sample environment remote sensing image set and the sample environment vocabulary set are randomly partitioned N times with replacement to obtain N environmental vocabulary recognition training datasets. N can be dynamically determined based on the actual sample size, computing resources, and recognition accuracy requirements. For example, if N=10, then 10 distinct environmental vocabulary recognition training datasets are generated from the sample environment remote sensing image set and the sample environment vocabulary set through 10 independent random sampling with replacement.

[0070] Furthermore, based on machine learning, N remote sensing environment recognition networks are constructed, and N environmental vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing environment recognition network group is obtained.

[0071] For example, a remote sensing environment recognition network group can be constructed using the same model architecture as the remote sensing target recognition network group in step S10. For instance, N remote sensing environment recognition networks can be constructed using a CNN architecture. The remote sensing environment recognition network mainly consists of an input layer, a multi-scale convolutional feature extraction layer, a pooling layer, a global feature fusion layer, and an output layer. The structure of the input layer, pooling layer, global feature fusion layer, and output layer is similar to that of the remote sensing target recognition network in step S10. The multi-scale convolutional feature extraction layer can be adjusted to adapt to the macroscopic features of the environmental image, for example, by adding large-size convolutional kernels (such as 7×7) or dilated convolutions to improve the ability to capture large-scale terrain contours.

[0072] For example, during training, N environmental vocabulary recognition training datasets are used to independently train N remote sensing environmental recognition networks. For instance, the environmental vocabulary recognition training datasets are divided into training and validation sets in an 8:2 ratio. The sample environmental remote sensing images in the training set are used as input features, and the corresponding sample environmental words are used as supervision labels. The cross-entropy loss function is used to quantify the difference between the model output and the true labels. The Adam optimizer is used, with an initial learning rate of 1e-4. Cosine annealing scheduling is used, the training batch size is set to 32 (adapting to GPU memory), and the number of iterations is set to 50. After each iteration, the model accuracy is evaluated using an independent validation set. If the validation set accuracy does not improve for five consecutive iterations, an early stopping strategy is triggered to terminate training, avoiding overfitting. The model parameters at this point are saved, resulting in a successfully trained and performing remote sensing environmental recognition network. This training process is repeated, and after independently training N remote sensing environmental recognition networks using N environmental vocabulary recognition training datasets, they are combined to form a remote sensing environmental recognition network group.

[0073] Finally, the environmental remote sensing image is input into a remote sensing environment recognition network group. Each remote sensing environment recognition network independently outputs an environmental term, resulting in multiple environmental terms. Terms with the same characteristics are clustered, and the proportion of occurrence of each type of environmental term is calculated to obtain an environmental term set and an environmental term matching degree set. For example, inputting the environmental remote sensing image into the remote sensing environment recognition network group outputs 10 environmental terms, and terms with the same characteristics are grouped into one category: 8 plains, 1 mountain range, and 1 hill. The proportion of occurrence of plains is 8 / 10 = 0.8, the proportion of occurrence of mountains is 1 / 10 = 0.1, and the proportion of occurrence of hills is 1 / 10 = 0.1. The proportion of occurrence of each type of environmental term is the environmental term matching degree, ultimately yielding an environmental term set {plains, mountains, hills} and an environmental term matching degree set {80%, 10%, 10%}.

[0074] In summary, compared to existing technologies, this application configures the environmental remote sensing image segmentation range based on the number of target words within the target vocabulary set, divides and acquires environmental remote sensing images within the environment of the target area, performs remote sensing environment identification, and obtains an environmental vocabulary set and an environmental vocabulary matching degree set. In this way, the environmental image range is dynamically adjusted according to the complexity of the target area, ensuring sufficient environmental context information is obtained and acquiring a more macroscopic environmental vocabulary, providing reliable data support for subsequent open vocabulary generation.

[0075] S30: Based on the number of words in the target vocabulary set and the environment vocabulary set, configure the language large model recognition resources, randomly combine the environment vocabulary set and the target vocabulary set, input them into the configured remote sensing large language model, and output the target open vocabulary set and the frequency set of occurrence.

[0076] The target vocabulary can only reflect the specific land cover categories within the target area, such as farmland and orchards, and lacks semantic association with the macro environment in which the target area is located. This isolated land cover recognition makes it unable to distinguish geographical scenes and is difficult to meet the needs of scene-based and refined land cover description in remote sensing observation.

[0077] To address the aforementioned issues, this application configures language large-scale model recognition resources based on the number of words in the target vocabulary set and the environment vocabulary set, randomly combines the environment vocabulary set and the target vocabulary set, inputs them into the configured remote sensing large-scale language model, and outputs the target open vocabulary set and the frequency set of occurrence.

[0078] Specifically, step S30 in the method includes:

[0079] Obtain the maximum number of target words and the maximum number of environment words for identifying remote sensing targets and remote sensing environments within a historical period;

[0080] Extract the number of target words and the number of environment words in the target vocabulary set and the environment vocabulary set, calculate the ratios of these values ​​to the maximum number of target words and the maximum number of environment words, and calculate the average values ​​to obtain the language large model recognition resource coefficient.

[0081] Obtain a remote sensing large language model group, calculate the number of large models based on the resource identification coefficients of the large language models, and randomly select the remote sensing large language models from the large model group.

[0082] The environmental vocabulary set and the target vocabulary set are randomly combined to obtain multiple vocabulary groups, which are then input into the large-scale remote sensing language model. The output is the target open vocabulary set and the frequency set, where each frequency includes the proportion of each type of target open vocabulary.

[0083] In this embodiment, the maximum number of target words and the maximum number of environmental words identified within a historical period are first obtained. The maximum number of target words is the highest number of target words output in a single target identification operation within a historical period. For example, in historical data, a complex area was identified with 5 target words: farmland, orchard, building, water body, and road; therefore, the maximum number of target words is 5. The maximum number of environmental words is the highest number of environmental words output in a single environmental identification operation within a historical period. For example, in historical data, a certain area was identified with 4 environmental types: mountains, hills, rivers, and grassland; therefore, the maximum number of environmental words is 4. The maximum number of target words and the maximum number of environmental words reflect the vocabulary scale of the most complex task previously processed, providing a comparable benchmark for the resource requirements of the current task. The closer the current number of target words and environmental words is to the historical maximum, the more complex the task, requiring more resources.

[0084] Secondly, the number of target words and the number of environment words in the target vocabulary set and the environment vocabulary set are extracted. The ratios to the maximum number of target words and the maximum number of environment words are calculated, and the average is calculated to obtain the language large-scale model recognition resource coefficient. For example, if the number of target words in the target vocabulary set is 4, the number of environment words in the environment vocabulary set is 3, the maximum number of target words is 5, and the maximum number of environment words is 4, then the ratio of the number of target words to the maximum number of target words = 4 / 5 = 0.8, and the ratio of the number of environment words to the maximum number of environment words = 3 / 4 = 0.75. The average is then calculated, and the language large-scale model recognition resource coefficient is obtained as (0.8 + 0.75) / 2 = 0.775. The language large-scale model recognition resource coefficient integrates the total complexity of the target and the environment. The closer the language large-scale model recognition resource coefficient is to 1, the closer the complexity of the current task is to the historical high level. The closer the language large-scale model recognition resource coefficient is to 0, the simpler the current task is.

[0085] Next, a large group of remote sensing language models is obtained. Based on the resource coefficients of the language models, the number of large models is calculated, and a random selection of these large remote sensing language models is made. Specifically, the number of remote sensing language models participating in inference is determined by the resource coefficients of the language models. This ensures that more remote sensing language models are used for high-complexity tasks (with high resource coefficients for language model recognition) to improve recognition accuracy and precision, while fewer remote sensing language models are used for low-complexity tasks (with low resource coefficients for language model recognition) to avoid resource waste.

[0086] Finally, the environmental vocabulary set and the target vocabulary set are randomly combined to obtain multiple vocabulary groups. These groups are then input into a large remote sensing language model with a large number of models. The output includes the target open vocabulary set and the frequency set, where each frequency includes the proportion of each type of target open vocabulary.

[0087] For example, environmental vocabulary from the environmental vocabulary set and target vocabulary from the target vocabulary set are randomly combined. For instance, the target vocabulary set {farmland, orchard, vegetable garden, wasteland} and the environmental vocabulary set {plain, mountain, hill} are combined in pairs to obtain 12 vocabulary groups covering all possible combinations: farmland + plain, farmland + mountain, farmland + hill, orchard + plain, orchard + mountain, orchard + hill, ... Then, all vocabulary groups are input into a large number of remote sensing language models. Each remote sensing language model performs semantic reasoning on each vocabulary group and outputs one target open vocabulary that breaks through the preset library. For example, a remote sensing language model outputs "plain" instead of "farmland + plain". Farmland is output as alpine meadows, etc., by combining farmland and mountains. Then, all target open vocabulary output by the remote sensing large-scale language model is aggregated, deduplicated, and a target open vocabulary set {plain farmland, alpine meadow, ...} is obtained. All open vocabulary is then clustered, and the frequency of each open vocabulary among all open vocabulary is calculated. For example, if 12 vocabulary groups are input into 8 remote sensing large-scale language models, 96 target open vocabulary are output. Plain farmland appears 80 times, alpine meadow appears 5 times, so the frequency of plain farmland = 80 / 96 = 0.83, the frequency of alpine meadow = 5 / 96 = 0.05, etc., thus obtaining the frequency set {0.83, 0.05, ...}. In this way, by randomly combining environmental and target vocabulary sets and using multiple remote sensing large-scale language models in parallel, the semantic relationships between environmental and target vocabulary are fully explored, generating open vocabulary that is more relevant to the actual scene.

[0088] Specifically, the step of "obtaining a large remote sensing language model group, calculating the number of large models based on the resource identification coefficients of the large language models, and randomly selecting remote sensing language models from the large number of large models" includes:

[0089] Based on remote sensing vocabulary data, a set of sample vocabulary groups is collected, and the target open vocabulary of each sample vocabulary group is labeled to obtain a set of sample target open vocabulary.

[0090] Based on natural language models, multiple remote sensing large language models are constructed. The sample vocabulary set and the sample target open vocabulary set are randomly divided. The multiple remote sensing large language models are trained in a supervised manner. After verification and convergence, a remote sensing large language model group is obtained.

[0091] The number of large language models is obtained by multiplying the resource coefficient of the large language model by the number of large remote sensing language models and rounding it down. Then, the number of large remote sensing language models is randomly selected from the large language models.

[0092] In this embodiment, firstly, based on historical remote sensing vocabulary data, a large number of target vocabulary + environmental vocabulary combinations covering different land cover and environment scenarios are collected, such as farmland + mountain range, orchard + plain, building + river, etc., as a sample vocabulary set. The target open vocabulary of each sample vocabulary set is manually labeled. For example, farmland + mountain range is labeled as alpine meadow, orchard + plain is labeled as plain orchard, building + river is labeled as riverside building, etc., as a sample target open vocabulary set. The labeling process should reflect the scene-specific semantics of the remote sensing field as much as possible and avoid general vocabulary to ensure that the model learns professional remote sensing terminology.

[0093] Secondly, based on natural language models, multiple remote sensing large language models are constructed. The sample vocabulary set and the sample target open vocabulary set are randomly divided, and supervised training is performed on multiple remote sensing large language models. After verification and convergence, a remote sensing large language model group is obtained. The specific number of remote sensing large language models can be dynamically determined according to actual computing power conditions, recognition accuracy, etc.

[0094] For example, in order to adapt to remote sensing scenarios, remote sensing big language models can use general natural language big models (such as LLaMA-7B and GPT-2) as the basic framework, retain the core semantic reasoning capabilities of general natural language big models, and then construct multiple (e.g., 10) remote sensing big language models with the same architecture by injecting remote sensing term vectors (such as distributed representations of professional terms like meadow and hills) and adjusting the attention mechanism (enhancing the weight of phrases related to the target and the environment). During training, the sample vocabulary set and the sample target open vocabulary set are randomly split into multiple training data sets, the same number as the number of remote sensing large language models (e.g., 10 sets). Each training data set independently trains one remote sensing large language model. Taking a single training data set as an example, the sample vocabulary set in the training data is used as the input feature, and the corresponding sample target open vocabulary is used as the supervision label. The cross-entropy loss function (adapted to the probability distribution optimization of discrete vocabulary) is used to quantify the difference between the model output and the supervision label. The AdamW optimizer (initial learning rate 5e-5, weight decay 0.01) is used to iteratively optimize the model parameters through gradient descent. When the validation loss no longer decreases for 5 consecutive rounds, the model is considered to have converged, and training is terminated, resulting in a trained remote sensing large language model. This process is repeated to obtain multiple remote sensing large language models. Finally, the multiple remote sensing large language models are integrated to obtain a remote sensing large language model group.

[0095] Finally, the number of large-scale language models is obtained by multiplying the language model recognition resource coefficient by the number of remote sensing large-scale language models and rounding it down. A random selection of these large-scale language models is then performed. Rounding can be done using methods such as rounding to the nearest integer, rounding up, or rounding down, which can be dynamically selected by those skilled in the art based on actual needs. For example, if the language model recognition resource coefficient is 0.775 and the number of remote sensing large-scale language models is 10, then the number of large-scale models = 10 × 0.775 ≈ 8 (rounded down). This ensures that the number of large-scale models retrieved is proportional to the task complexity; more models are used to cover semantic possibilities for complex tasks, while fewer models are used to control computational costs for simple tasks. Then, a random selection of a number of large-scale language models (e.g., 8) is performed from the remote sensing large-scale language model group to participate in subsequent open vocabulary generation. Random selection ensures the diversity and objectivity of the generated open vocabulary.

[0096] In summary, compared to existing technologies, this application configures language model recognition resources based on the number of words in the target vocabulary set and the environment vocabulary set. The environment vocabulary set and the target vocabulary set are randomly combined and input into the configured remote sensing language model, outputting a target open vocabulary set and a frequency set. Thus, the total complexity of the target and environment vocabulary is considered, and the language model resources are dynamically allocated accordingly, generating a target open vocabulary that surpasses the preset library.

[0097] S40: Calculate the open word confidence score based on the target word matching score set, the environmental word matching score set, and the frequency of occurrence set, and select the optimal open word as the target detection result for the target region.

[0098] The aforementioned steps involve obtaining a target vocabulary matching set (reflecting the recognition credibility of target objects) through remote sensing image recognition of the target area, obtaining an environmental vocabulary matching set (reflecting the recognition credibility of the macro environment) through environmental remote sensing image recognition, and obtaining a target open vocabulary frequency set (reflecting the consensus of multiple models on semantic associations) through remote sensing large language model inference. Therefore, the optimal open vocabulary can be selected as the target detection result for the target area.

[0099] To address the aforementioned issues, this application calculates the open word confidence score based on the target word matching score set, the environmental word matching score set, and the frequency of occurrence set, and selects the optimal open word as the target detection result for the target region.

[0100] Specifically, step S40 in the method includes:

[0101] Based on the target vocabulary matching set and the environment vocabulary matching set, calculate the mean of the target vocabulary matching degree and the environment vocabulary matching degree corresponding to each target open vocabulary to obtain the basic matching degree set.

[0102] Based on the basic matching degree set and the occurrence frequency set, calculate the open word confidence score for each target open word to obtain the open word confidence score set;

[0103] Select the target open vocabulary with the highest confidence level to obtain the optimal open vocabulary, which is then used as the target detection result for the target region.

[0104] In this embodiment of the application, since the target open vocabulary is generated by the combination of target vocabulary and environment vocabulary input into the remote sensing big language model, each target open vocabulary can be traced back to a unique target vocabulary and environment vocabulary. Then, based on the matching target vocabulary matching degree set and environment vocabulary matching degree set, the average of the target vocabulary matching degree and environment vocabulary matching degree of each target open vocabulary is calculated to obtain the basic matching degree set.

[0105] For example, if the target vocabulary set is {farmland, orchard, vegetable garden, wasteland} and the target vocabulary matching degree set is {70%, 10%, 10%, 10%}, and the environmental vocabulary set is {plain, mountain, hill} and the environmental vocabulary matching degree set is {80%, 10%, 10%}, and the target open vocabulary for plain farmland is farmland and the environmental vocabulary is plain, then the basic matching degree for plain farmland = (70% + 80%) / 2 = 75%. If the target open vocabulary for alpine meadow is farmland and the environmental vocabulary is mountain, then the basic matching degree for alpine meadow = (70% + 10%) / 2 = 40%. In this way, the basic matching degree of all target open vocabulary is calculated and summarized to obtain the basic matching degree set, for example, {75% (plain farmland), 45% (alpine meadow), ...}. The basic matching degree can reflect the correlation strength between the target open vocabulary and the actual features of the target remote sensing image and the environmental remote sensing image. The higher the basic matching degree, the stronger the correlation, and the higher the credibility of the target open vocabulary.

[0106] Secondly, based on the basic matching degree set and the occurrence frequency set, the open-word confidence score for each target open-word is calculated to obtain the open-word confidence score set. Here, the open-word confidence score = basic matching degree × occurrence frequency. The open-word confidence score combines the basic matching degree and occurrence frequency, and can more comprehensively reflect the reliability of the target open-word. The higher the open-word confidence score, the stronger the reliability and accuracy of the open-word in matching the actual geographical scene of the target area. For example, if the target open-word set is {plain farmland, alpine meadow, ...}, the occurrence frequency set is {0.83, 0.05, ...}, and the basic matching degree set is {75% (plain farmland), 45% (alpine meadow), ...}, then the open-word confidence score for plain farmland = 75% × 0.83 = 0.62, and the open-word confidence score for alpine meadow = 45% × 0.05 = 0.02. Thus, by traversing the target open-word sets, the open-word confidence score set is obtained.

[0107] Finally, the target open word with the highest open word confidence is selected to obtain the optimal open word, which is used as the target detection result for the target region. For example, the open word confidence set can be sorted in descending order, the target open word with the highest open word confidence can be selected as the optimal open word, and finally the optimal open word and its open word confidence are output as the target detection result for the target region.

[0108] In summary, compared to existing technologies, this application calculates the open-word confidence score based on the target word matching set, the environmental word matching set, and the frequency of occurrence set, and selects the optimal open-word as the target detection result for the target region. Thus, through multi-dimensional fusion and quantitative evaluation, it outputs target open-words that are both consistent with image facts and semantic logic.

[0109] In summary, the embodiments of this application have at least the following technical effects:

[0110] Compared to existing technologies, this application first acquires remote sensing images of the target area selected by the user, performs remote sensing target identification, and obtains a target vocabulary set and a target vocabulary matching degree set. Thus, by using a remote sensing target identification network group, it solves the problem of susceptibility to image noise interference in traditional single-model identification, providing stable and quantifiable basic target information for subsequent environmental analysis and open vocabulary generation.

[0111] Secondly, this application configures the environmental remote sensing image segmentation range based on the number of target words in the target vocabulary set, divides and acquires environmental remote sensing images within the environment of the target area, performs remote sensing environment identification, and obtains an environmental vocabulary set and an environmental vocabulary matching degree set. In this way, the environmental image range is dynamically adjusted according to the complexity of the target area, ensuring sufficient environmental context information is obtained and a more macroscopic environmental vocabulary is acquired, providing reliable data support for subsequent open vocabulary generation.

[0112] Furthermore, this application configures language model recognition resources based on the number of words in the target vocabulary set and the environment vocabulary set. The environment vocabulary set and the target vocabulary set are randomly combined and input into the configured remote sensing language model, outputting a target open vocabulary set and a frequency set. In this way, the total complexity of the target and environment vocabulary is considered, and the language model resources are dynamically allocated accordingly, generating a target open vocabulary that breaks through the preset library.

[0113] Finally, based on the target vocabulary matching set, the environmental vocabulary matching set, and the frequency of occurrence set, this application calculates the open vocabulary confidence score and selects the optimal open vocabulary as the target detection result for the target region. Thus, through multi-dimensional fusion and quantitative evaluation, the output targets open vocabulary that conforms to both image facts and semantic logic.

[0114] Through the above technical solution, this application first identifies the target area image selected by the user to obtain a target vocabulary set, then dynamically adjusts the environmental image range according to the number of target words to obtain an environmental vocabulary set. Subsequently, the two types of words are randomly combined and input into a remote sensing large language model adapted to the resources, which can generate non-preset open words such as alpine meadows and plain farmland. Moreover, by calculating confidence in multiple dimensions to select the optimal results, it ensures that the open words fit the actual features of the image and have semantic and logical consensus, reducing the misidentification rate in complex scenes. At the same time, dynamic resource configuration takes into account both recognition accuracy and computational efficiency, avoiding resource waste or insufficient computing power. In this way, it solves the pain point of traditional remote sensing target recognition relying on a preset vocabulary library and having poor adaptability to non-preset categories, and can output scene-specific, high-precision open vocabulary detection results, improving the practicality and scene adaptability of remote sensing observation.

[0115] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0116] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0120] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0121] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.

Claims

1. A remote sensing open vocabulary target detection method based on a multimodal large language model, characterized in that, The method includes: Acquire remote sensing images of the target area selected by the user, perform remote sensing target recognition, and obtain a target vocabulary set and a target vocabulary matching degree set; Based on the number of target words in the target vocabulary set, configure the environmental remote sensing image division range, divide and acquire the environmental remote sensing images within the environment where the target area is located, perform remote sensing environment identification, and obtain the environmental vocabulary set and the environmental vocabulary matching degree set, wherein each environmental vocabulary matching degree is the proportion of the occurrence of each type of environmental vocabulary. Based on the number of words in the target vocabulary set and the environment vocabulary set, configure language large model recognition resources, randomly combine the environment vocabulary set and the target vocabulary set, input them into the configured remote sensing large language model, and output the target open vocabulary set and the frequency set of occurrence. Based on the target word matching set, the environmental word matching set, and the frequency of occurrence set, the confidence of open words is calculated, and the optimal open words are selected as the target detection result for the target region. This includes acquiring remote sensing images of the target area selected by the user, performing remote sensing target recognition, and obtaining a target vocabulary set and a target vocabulary matching degree set, including: Acquire remote sensing images of the target area selected by the user; The target remote sensing image is input into a remote sensing target recognition network group, and the recognition output obtains a target vocabulary set and a target vocabulary matching degree set, wherein the matching degree of each target vocabulary is the proportion of each type of target vocabulary.

2. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 1, characterized in that, The target remote sensing image is input into a remote sensing target recognition network group, and the recognition output obtains a target vocabulary set and a target vocabulary matching degree set, including: Based on the remote sensing target identification and recording data, a set of sample target remote sensing images is collected, and the target words in each sample target remote sensing image are labeled to obtain a set of sample target words. The sample target remote sensing image set and the sample target vocabulary set are randomly partitioned N times with replacement to obtain N target vocabulary recognition training datasets; Based on machine learning, N remote sensing target recognition networks are constructed. The N target vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing target recognition network group is obtained. The target remote sensing image is input into the remote sensing target recognition network group, and multiple target words are output. Target words that are in the same category are clustered, and the occurrence frequency ratio of each type of target word is calculated to obtain the target word set and the target word matching degree set.

3. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 1, characterized in that, Based on the number of target words in the target vocabulary set, the environmental remote sensing image segmentation range is configured, and environmental remote sensing images within the environment of the target area are acquired. Remote sensing environment identification is performed to obtain an environmental vocabulary set and an environmental vocabulary matching degree set, including: Obtain the number of target words in the target vocabulary set, calculate the ratio with the preset number of target words, and obtain the target vocabulary coefficient; Obtain the pre-defined environmental remote sensing image range; Based on the target vocabulary coefficients, the preset environmental remote sensing image division range is adjusted and calculated to obtain the environmental remote sensing image division range. According to the defined range of the environmental remote sensing images, the environmental remote sensing images within the environment of the target area are divided and acquired. The environmental remote sensing image is input into a remote sensing environment recognition network group, and the output is processed to obtain an environmental vocabulary set and an environmental vocabulary matching degree set.

4. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 3, characterized in that, The environmental remote sensing image is input into a remote sensing environmental recognition network group, and the output is processed to obtain an environmental vocabulary set and an environmental vocabulary matching degree set, including: Based on remote sensing environment identification and recording data, a set of sample environment remote sensing images is collected, and environmental terms within each sample environment remote sensing image are labeled to obtain a set of sample environment terms. The sample environment remote sensing image set and the sample environment vocabulary set are randomly partitioned N times with replacement to obtain N environmental vocabulary recognition training datasets. Based on machine learning, N remote sensing environment recognition networks are constructed. The N environmental vocabulary recognition training datasets are used to conduct supervised training. After the training is completed, a remote sensing environment recognition network group is obtained. The environmental remote sensing image is input into the remote sensing environment recognition network group, and multiple environmental terms are output. Environmental terms that are in the same category are clustered, and the occurrence rate of each type of environmental term is calculated to obtain the environmental term set and the environmental term matching degree set.

5. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 1, characterized in that, Based on the number of words in the target vocabulary set and the environment vocabulary set, configure language large-scale model recognition resources, randomly combine the environment vocabulary set and the target vocabulary set, input them into the configured remote sensing large-scale language model, and output the target open vocabulary set and the frequency set, including: Obtain the maximum number of target words and the maximum number of environment words for identifying remote sensing targets and remote sensing environments within a historical period; Extract the number of target words and the number of environment words in the target vocabulary set and the environment vocabulary set, calculate the ratios of these values ​​to the maximum number of target words and the maximum number of environment words, and calculate the average values ​​to obtain the language large model recognition resource coefficient. Obtain a remote sensing large language model group, calculate the number of large models based on the resource identification coefficients of the large language models, and randomly select the remote sensing large language models from the large model group. The environmental vocabulary set and the target vocabulary set are randomly combined to obtain multiple vocabulary groups, which are then input into the large-scale remote sensing language model. The output is the target open vocabulary set and the frequency set, where each frequency includes the proportion of each type of target open vocabulary.

6. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 1, characterized in that, Obtain a large remote sensing language model group, calculate the number of large models based on the resource identification coefficients of the large language models, and randomly select remote sensing language models from the group, including: Based on remote sensing vocabulary data, a set of sample vocabulary groups is collected, and the target open vocabulary of each sample vocabulary group is labeled to obtain a set of sample target open vocabulary. Based on natural language models, multiple remote sensing large language models are constructed. The sample vocabulary set and the sample target open vocabulary set are randomly divided. The multiple remote sensing large language models are trained in a supervised manner. After verification and convergence, a remote sensing large language model group is obtained. The number of large language models is obtained by multiplying the resource coefficient of the large language model by the number of large remote sensing language models and rounding it down. Then, the number of large remote sensing language models is randomly selected from the large language models.

7. The remote sensing open vocabulary target detection method based on a multimodal large language model according to claim 1, characterized in that, Based on the target vocabulary matching set, the environmental vocabulary matching set, and the frequency of occurrence set, the confidence score of open words is calculated, and the optimal open words are selected as the target detection result for the target region, including: Based on the target vocabulary matching set and the environment vocabulary matching set, calculate the mean of the target vocabulary matching degree and the environment vocabulary matching degree corresponding to each target open vocabulary to obtain the basic matching degree set. Based on the basic matching degree set and the occurrence frequency set, calculate the open word confidence score for each target open word to obtain the open word confidence score set; Select the target open vocabulary with the highest confidence level to obtain the optimal open vocabulary, which is then used as the target detection result for the target region.

Citation Information

Patent Citations

  • Open vocabulary object detection method based on SAM candidate box generation and candidate region-word clustering

    CN120451679A

  • Extensible remote sensing deep learning sample library construction method based on target region planning

    CN120612563A