Remote sensing image KNN retrieval method of visual language model
By applying KNN search method in the field of remote sensing, remote sensing images are combined with visual language, and using deep learning models and multi-head attention mechanisms and other technologies, the problem of insufficient transfer learning and language comprehension capabilities in the existing technology is solved, and efficient remote sensing image retrieval is achieved.
Patent Information
- Application Number
- CN202510058337.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art is difficult to effectively transfer deep learning models to specific language scenarios in the field of remote sensing, and lacks language comprehension capabilities and cannot be applied to retrieval applications.
The KNN search method is used to combine remote sensing images with visual language, and combine deep learning models and multi-head attention mechanisms, contrast learning, and CLIP models to search remote sensing images in visual language.
In the KNN search task in the remote sensing field, it improves the accuracy and efficiency of the search, reduces the cost of manual labeling, and enhances the universality and adaptability of the model.
Smart Images

Figure CN119942343A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision, remote sensing science and natural language processing, and in particular to remote sensing image analysis and remote sensing feature extraction technology in remote sensing science and technology. More specifically, the present invention relates to a KNN retrieval method that combines remote sensing images with visual language. The method combines a deep learning model with a multi-head attention mechanism, contrastive learning, image feature extraction and a CLIP (Contrastive Language-Image Pretraining) model for remote sensing image retrieval tasks in visual language and is widely used in agriculture, forestry, urban planning, environmental monitoring, geological survey, disaster warning and other fields. Background Art
[0002] With the surge in the amount of image-text pair data and the diversity of visual and language tasks, scholars have introduced a large number of deep learning models in this research field. In addition, in recent years, transfer learning has also achieved great success in the field of remote sensing, such as image classification, object detection and other tasks, as well as question answering, machine translation and so on in natural language processing. The great success of the basic vision-language model has promoted the research and application of remote sensing and multimodal representation learning. However, how to effectively migrate these basic models to specific language scenarios remains a challenge. The CLIP (ContrastiveLanguage-Image Pretraining) study proposed the idea of cross-modal contrastive learning, that is, training images and natural language text as joint inputs, thereby improving the versatility and adaptability of the model.
[0003] In the field of remote sensing, self-supervised learning and masked image modeling have been used to build basic models. However, these models mainly learn low-level features and require labeled data for fine-tuning. In addition, they are not suitable for retrieval applications due to the lack of language understanding. Remote sensing data is usually taken by satellites or drones, involving a large amount of complex visual information, which often needs to be interpreted through manual annotation or complex image processing techniques.
[0004] Chinese patent document CN118736625A describes a method for pedestrian retrieval in images and texts based on fusion of key point attention guidance. It introduces human key point information into the CLIP model through a cross-attention mechanism, adds additional supervision signals, reduces interference from occlusion and background information, and improves the accuracy of pedestrian image and text retrieval while reducing the cost of manual labeling.
[0005] Chinese patent document CN119153120A describes a customized microscopy case classification retrieval system based on multimodal learning. The CLIP model is used to process retrieval requirements input in the form of images, and after fusing them with retrieval requirements input in the form of text, its feature representation information is extracted, and then compared with the feature representation of the image in the historical data after CLIP encoding to obtain the historical image that is closest to the fused retrieval requirement.
[0006] Although there have been studies on retrieval methods of visual language, none of them have been implemented in the field of remote sensing images. The present invention collects visual language remote sensing data sets, performs base64 processing and jsonl processing on them, and then serializes the data. The processed data is then sent to the CLIP large model for training, and then a KNN retrieval model is obtained and feature extraction is performed. Finally, KNN retrieval is implemented and evaluated, and finally a KNN retrieval method is obtained. Summary of the invention
[0007] The present invention aims to provide a remote sensing image KNN retrieval method of a visual language model, which trains a remote sensing image dataset through downstream task KNN retrieval, and then evaluates the remote sensing image test data through evaluations such as recall rate.
[0008] In order to achieve the above purpose, the technical solution adopted by the present invention includes the following steps:
[0009] Step S1: Data preprocessing: perform base64 processing on the remote sensing image dataset in the test dataset, perform jsonl processing on the text dataset, and upload them to the database after processing;
[0010] Step S2: Image text data serialization: Serialize the preprocessed data set to facilitate random reading during training;
[0011] Step S3: CLIP large model training: the serialized data is sent to the visual language model for training to obtain a training model;
[0012] Step S4: Image and text feature extraction: extract features from the test data through the training model;
[0013] Step S5: KNN retrieval evaluation: Perform KNN retrieval evaluation on the files with extracted features.
[0014] The image data base64 processing in step S1 is as follows:
[0015] Step S11: read the remote sensing image data file and convert it into base64 encoding;
[0016] Step S12: traverse all files in the remote sensing image folder, convert the base64 encoding of each remote sensing image file, and write it into the specified .tsv file;
[0017] The image text data serialization process in step S2 is as follows:
[0018] Step S21: Serializing the image file and the text file processed by base64;
[0019] Step S22: converting the serialized data into a memory-indexed LMDB database file;
[0020] The step S3 model pre-training is as follows:
[0021] Step S31: input a remote sensing image, divide the image into blocks of fixed size and convert the image blocks into an embedding space;
[0022] Step S32: adding position coding to explicitly introduce the position information of the tile;
[0023] Step S33: Process the layer normalized result using a multi-head self-attention mechanism and a feedforward neural network and introduce dropout to avoid overfitting or underfitting;
[0024] Step S34: Activate the MLP with the improved QuickGELU, and output the probability distribution result through the fully connected layer and the softmax layer; the formula of the improved QuickGELU activation function is:
[0025] (1)
[0026] Where y represents the output result, x is the input variable, and σ represents the Sigmoid function, that is,
[0027] (2)
[0028] (3)
[0029] (4)
[0030] represents the dynamically adjusted parameters, where α is the parameter set during initialization.
[0031] Step S35: Train using standard supervised learning methods and adjust and optimize in Adam.
[0032] The step S4 image and text feature extraction process is as follows:
[0033] Step S41: After the input image passes through three convolution layers and one pooling layer, a lower resolution feature map is output, where the convolution layer is a variable convolution, and its formula is:
[0034] (5)
[0035] Among them, p is the padding parameter in the standard convolution, and Represent the output and input channel indices respectively, i and j represent the positions of the output feature map in the height and width directions respectively, is the convolution kernel weight.
[0036] Step S42: Further extract more complex high-level features through multiple Bottleneck residual blocks.
[0037] Step S43: Use the self-attention mechanism to aggregate spatial information and transform the feature map into a vector of fixed length.
[0038] Step S44: Output a vector containing high-dimensional features for further retrieval tasks.
[0039] The step S5KNN retrieval evaluation is as follows:
[0040] Step S51: Calculate the distance between the feature-extracted vector and the vector in the data set;
[0041] Step S52: determine whether the first K neighbors are relevant based on ground truth;
[0042] Step S53: Evaluate its precision, recall rate and F1-Score. The specific definition of evaluating its recall rate is as follows:
[0043] (6)
[0044] (7)
[0045] (8)
[0046] Where R is the recall rate, TP is the number of positive classes predicted as positive, FN is the number of negative classes predicted as negative, R@k is the recall point, r(k) is the number of relevant images among the first k images returned, r is the total number of images related to the query, and MR is the average recall rate. The higher the MR score, the better the retrieval effect.
[0047] The present invention has the following excellent effects and advantages:
[0048] Compared with the traditional activation function algorithm, the improved QuickGELU algorithm of the present invention introduces a dynamically adjusted 'alpha' value, which allows the activation function to change dynamically according to the characteristics of the input data; adds processing for negative inputs, which may help retain some features so that the model can perform better in certain input situations; combines linear and nonlinear elements so that the activation function can better capture complex patterns in the data.
[0049] Compared with the traditional multi-head self-attention mechanism, the present invention can handle large-scale, multi-dimensional complex data by adding two dropouts between the attention output and the MLP. The model has strong generalization ability, is not prone to overfitting problems, and has high classification accuracy. Compared with the standard convolution layer, the present invention uses a deformable convolution to enhance the performance of the network, making it more flexible to handle the spatial deformation of the feature map, which may improve the performance of some tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 It is a structural schematic diagram of the present invention;
[0052] Figure 2 It is a conceptual schematic diagram of the CLIP model of the present invention;
[0053] Figure 3 It is the VIT framework diagram of the present invention;
[0054] Figure 4 is a training flow chart of the present invention;
[0055] Figure 5 It is the UCM_captions dataset;
[0056] Figure 6 This is the result graph of CLIP model training indicators. DETAILED DESCRIPTION
[0057] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0059] like Figure 1 As shown, the present invention discloses a remote sensing image KNN retrieval method of a visual language model, comprising the following steps: S1 data preprocessing: performing base64 processing on the remote sensing image data set in the visual language test data set, performing jsonl processing on the text data set, and uploading the processed data to the database; S2 image text data serialization: serializing the preprocessed data set to facilitate random reading during training; S3 CN-CLIP large model training: sending the serialized data to the visual language model for training to obtain a training model; S4 image and text feature extraction: extracting features from the test data through the training model; S5 KNN retrieval evaluation: performing KNN retrieval evaluation on the feature-extracted files.
[0060] like Figure 2 As shown, the CLIP model training process of the present invention is divided into the following steps:
[0061] First, input the remote sensing image dataset and text dataset files.
[0062] Secondly, the remote sensing image and text data are encoded in the encoder respectively, and their features are obtained.
[0063] Finally, the image features are compared with the text features for learning or other cross-modal operations to learn the semantic association between images and texts, thus enabling the model to understand and process cross-modal information.
[0064] like Figure 3 As shown, the model pre-training in the present invention requires the basic framework of the VIT model, and its operation process is divided into the following steps:
[0065] First, the image is divided into small blocks of equal size and flattened.
[0066] Secondly, it is encoded by the transformer encoder and its tensor features are output.
[0067] Finally, the output from the last layer of the transformer encoder is passed through a multi-layer perceptron (MLP) head, usually a simple fully connected layer, to perform the final task.
[0068] Example:
[0069] Combination Figure 1The present invention describes the remote sensing image KNN retrieval steps based on the visual language model, combined with Figure 5 The UCM_captions data used in the present invention is described. The present invention selects the UCM_captions data set for experimental processing, which contains 21 types of remote sensing image data sets of geographic / land use types, with an image resolution of 256x256 pixels, covering a variety of geographic areas, such as urban areas, agricultural areas, forests, grasslands, etc. Each image represents a geographic category, and each image is equipped with a text description. These descriptions are manually written to accurately reflect the content and scene features of the image. Each text description usually includes information such as the main objects, environment, and scene contained in the image.
[0070] Recombination Figure 4 The training specific flowchart implements the following steps:
[0071] Step S1: Data preprocessing
[0072] Step S11: Base64 processing of image data
[0073] First, read the remote sensing image data file, convert it into base64 encoding, traverse the folder storing the remote sensing images, and perform base64 encoding conversion on each remote sensing image file.
[0074] Secondly, store the converted base64 code in a .tsv file in the specified format: 1680 / 9j / 4AAQSkZJ...YQj7314oA / / 2Q==
[0075] Finally, make sure that the .tsv file can fully save the encoding information of all image files.
[0076] Step S12: Text data jsonl processing
[0077] Extract the text description data corresponding to the remote sensing image and convert each record into JSONL format. The format of each record in the JSONL file is: {"text_id": 8410, "text": "It is a river with some plants onone bank and sands on the other side.", "image_ids":
[1682] }. Upload the processed text data to the database for subsequent training.
[0078] Step S2: Image text data serialization
[0079] The .tsv and .jsonl files generated in step S1 are serialized. The serialized files can randomly read data according to training requirements, and the serialized data is used to generate a database file in LMDB format. The LMDB file can be used as an efficient memory index to facilitate subsequent fast access operations.
[0080] Step S3: CLIP large model training
[0081] Fine-tune the training parameters of the model, set the number of training GPUs, the LMDB data path for training, specify the ViT-B-16 visual training scale and the text training scale, the number of training steps is 50 rounds, the training learning rate is set to 5e-5, and set the weight result storage path to start training.
[0082] like Figure 6 Shown are the training indicator results of the CLIP model of the present invention.
[0083] Step S4: Image and text feature extraction and processing
[0084] Write the weight file path obtained by training in step 3 into the feature extraction code to start image and text feature extraction and obtain the image feature file. Each line stores the features of an image in json format. The format is as follows: {"image_id": 1680,"feature": [0.0198, ..., -0.017, 0.0248]}, and the text feature file has the following format: {"text_id": 8410, "feature": [0.1314, ..., 0.0018, -0.0002]}
[0085] Step S5 KNN retrieval evaluation
[0086] Step S51: KNN search
[0087] Write the feature file obtained in step 4 into KNN retrieval, calculate the top-k recall results of text-to-image and image-to-text retrieval, and save the output results in the specified jsonl file. Each line represents the top-k image id of a text recall, in the following format: {"text_id": 8410, "image_ids": [1358, 1645, 1004, 1585, 1531, 1674,1245, 1678, 765, 1165]} Each line represents the top-k text id of an image recall, in the following format: {"image_id": 1680, "text_ids": [6884, 7454, 8754, 5916, 6188,8305, 7179, 4599, 4956,4743]}
[0088] Step S52: Recall calculation evaluation
[0089] Write the jsonl file obtained by KNN retrieval and implement Recall calculation evaluation according to Recall@1 / 5 / 10. The results are as follows:
[0090] {"success":true,"score":40.47619047619048,"scoreJson":{"score":40.47619047619048 ,"mean_recall":40.47619047619048,"r1":14.285714285714285,"r5":35.714285714285715, "r10": 71.42857142857143}}.
[0091] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.
Claims
1. A remote sensing image KNN retrieval method based on a visual language model, characterized in that: include: S1 data preprocessing, S2 image text data serialization, S3 CLIP large model training, S4 image and text feature extraction, S5 KNN retrieval evaluation.
2. The remote sensing image KNN retrieval method of a visual language model according to claim 1, characterized in that: The step S1 data preprocessing includes: Step S11: read the remote sensing image data file and convert it into base64 encoding; Step S12: traverse all files in the remote sensing image folder, convert the base64 encoding of each remote sensing image file, and write it into the specified .tsv file.
3. The remote sensing image KNN retrieval method of a visual language model according to claim 1, characterized in that: The step S2 of serializing the image text data includes: Step S21: Serializing the image file and the text file processed by base64; Step S22: Convert the serialized data into a memory-indexed LMDB database file.
4. The remote sensing image KNN retrieval method of a visual language model according to claim 1, characterized in that: The steps of S3CLIP large model training include: Step S31: input a remote sensing image, divide the image into blocks of fixed size and convert the image blocks into an embedding space; Step S32: adding position coding to explicitly introduce the position information of the tile; Step S33: Process the layer normalized result using a multi-head self-attention mechanism and a feedforward neural network and introduce dropout to avoid overfitting or underfitting; Step S34: Activate the MLP with the improved QuickGELU, and output the probability distribution result through the fully connected layer and the softmax layer; the formula of the improved QuickGELU activation function is: (1) Where y represents the output result, x is the input variable, and σ represents the Sigmoid function, that is, (2) (3) (4) represents the dynamically adjusted parameters, where α is the parameter set during initialization; Step S35: Train using standard supervised learning methods and adjust and optimize in Adam.
5. The remote sensing image KNN retrieval method of a visual language model as claimed in claim 1, characterized in that: The step S4 of image and text feature extraction processing includes: Step S41: After the input image passes through three convolution layers and one pooling layer, a lower resolution feature map is output, where the convolution layer is a variable convolution, and its formula is: (5) Among them, p is the padding parameter in the standard convolution, and Represent the output and input channel indices respectively, i and j represent the positions of the output feature map in the height and width directions respectively, is the convolution kernel weight; Step S42: further extracting more complex high-level features through multiple Bottleneck residual blocks; Step S43: Use the self-attention mechanism to aggregate spatial information and transform the feature map into a vector of fixed length; Step S44: Output a vector containing high-dimensional features for further retrieval tasks.
6. The remote sensing image KNN retrieval method of a visual language model according to claim 1, characterized in that: The step S5KNN retrieval evaluation comprises: Step S51: Calculate the distance between the feature-extracted vector and the vector in the data set; Step S52: determine whether the first K neighbors are relevant based on ground truth; Step S53: evaluate the recall rate and mean_recall; the specific definition of evaluating the recall rate is as follows: (6) (7) (8) Where R is the recall rate, TP represents the number of positive classes predicted as positive classes, FN represents the number of negative classes predicted as negative classes, R@k represents the recall point, r(k) represents the number of relevant images among the first k images returned, r represents the total number of images related to the query, and MR represents the average recall rate; the higher the MR score, the better the retrieval effect.
Citation Information
Patent Citations
Image-text pedestrian retrieval method based on fusion key point attention guidance
CN118736625A
Customized microscopic examination case grading retrieval system based on multi-modal learning
CN119153120A
Landmark retrieval recognition and positioning method based on space self-attention
CN115761492A
Generating texture mesh using one or more neural networks
CN117635871A
Remote sensing question answer generation method and device, medium and equipment
CN118709766A
Cited By
Satellite image quality evaluation method and system based on image search engine
CN120612619A