Model Training Method, High-Quality Satellite Image Retrieval Method and Device

By training the satellite image model, object category vectors and classification prediction values are generated, and model parameters are optimized, the problem of low accuracy in satellite image retrieval is solved, and higher matching degree and retrieval accuracy are achieved.

CN119478720BActive Publication Date: 2025-07-29BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411646361.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-07-29
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

In the prior art, the accuracy of satellite image retrieval is not high, and the matching degree between the query image and the target image is insufficient, mainly because the content characteristics of the query image and the image to be retrieved are not high.

Method used

A model training method is adopted to generate object category vectors and classification prediction values by inputting sample satellite images into the first neural network model, and calculate model loss values based on these values, and adjust model parameters to generate target neural network models, which are used to generate object category and classification prediction values for query images. The model consists of a convolutional neural network, a bidirectional encoder representation model based on Transformer, and an encoder, and combines multiple loss functions to optimize the model performance.

Benefits of technology

It improves the accuracy of satellite image retrieval, improves the degree of matching between the target satellite image and the query image, and enhances the accuracy of image retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478720B_ABST
    Figure CN119478720B_ABST
Patent Text Reader

Abstract

The present application provides a model training method, a high-quality satellite image retrieval method and device. The first object category of each object in the sample satellite image can be obtained through the first neural network model, as well as the first classification prediction value corresponding to each first object category. Further, the first model loss value of the first neural network model can be calculated according to the first object category, the first classification prediction value and the sample object category of the sample satellite image itself. Thus, the model parameters of the first neural network model can be adjusted through the first model loss value to obtain the target neural network model. Since the processing of the sample satellite image is accurate to each object in the image, the prediction accuracy of the target neural network model can be improved. When using the target neural network model to process the query image to retrieve satellite images, the matching degree between the finally obtained target satellite image and the query image can be improved, and the accuracy of image retrieval query can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular, to a model training method, a high-quality satellite image retrieval method, and an apparatus. Background Art

[0002] In recent years, the rapid development of satellite technology has driven the growth of various satellite applications. As one of the main fields of satellite applications, satellite remote sensing can obtain earth information all-weather, all-day, at high altitudes, and on a large scale, playing an unparalleled role in information acquisition and providing decision-making support for major issues such as disaster prevention. Thanks to the development of satellite software and hardware technologies, satellites have certain computing and storage capabilities and can store a large amount of observation data. However, due to the uniqueness of the space environment and the communication capacity limitations of satellites themselves, how to successfully transmit target data back to the ground station during the satellite's overpass time has become an urgent problem to be solved.

[0003] In the related art, according to the annotation information of satellite remote sensing image data, content features of a query image and an image to be retrieved can be extracted through SIFT or deep learning methods, and the similarity can be calculated to determine whether they belong to matching samples. Then, the target image data obtained by the query is sent to the ground station through a communication system to complete the image data transmission.

[0004] However, in the above method, the accuracy of the content features of the query image and the image to be retrieved is not high, the target image data obtained by the query often does not match the query image, and the accuracy of the image retrieval query is not high. Summary of the Invention

[0005] In view of the above problems, embodiments of this application provide a model training method, a high-quality satellite image retrieval method, an apparatus, an electronic device, and a readable storage medium to overcome or at least partially solve the above problems.

[0006] In a first aspect, an embodiment of this application provides a model training method, and the method includes:

[0007] Input a sample satellite image into a first neural network model to obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object;

[0008] Based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value, determine a first model loss value of the first neural network model;

[0009] Adjust the model parameters of the first neural network model based on the first model loss value to obtain a target neural network model; wherein, the target neural network model is used to generate the target object categories of the query image and the target classification prediction values respectively corresponding to each of the target object categories.

[0010] Optionally, the first neural network model is composed of a convolutional neural network model, a bidirectional encoder representation model based on Transformer, and an encoder model based on Transformer. The step of inputting the sample satellite image into the first neural network model to obtain the first object category vector output by the first neural network model and the first classification prediction value corresponding to the first object category vector includes:

[0011] Input the sample satellite image into the convolutional neural network to obtain the sample satellite image feature vector output by the convolutional neural network model;

[0012] Input each object category name into the bidirectional encoder representation model based on Transformer to obtain the sample object category vector output by the bidirectional encoder representation model based on Transformer;

[0013] Concatenate the sample satellite image feature vector and the sample object category vector to obtain a first sample concatenated vector;

[0014] Input the first sample concatenated vector into the encoder based on Transformer to obtain the second sample concatenated vector output by the encoder based on Transformer;

[0015] Determine the first object category vector from the second sample concatenated vector, and based on the first object category vector, determine the first object category corresponding to the sample satellite image and the first classification prediction values respectively corresponding to each of the first object categories.

[0016] Optionally, the step of determining the first model loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image includes:

[0017] Determine the binary cross-entropy loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image;

[0018] Determine the second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector;

[0019] Based on the first classification prediction value, determine the maximum likelihood estimation loss value of the first neural network model;

[0020] Based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value, determine the first model loss value of the first neural network model.

[0021] Optionally, the determining the second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector includes:

[0022] Perform semantic transformation on the first object category vector to obtain a second object category vector; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector;

[0023] Calculate the first cosine similarity between the second object category vector and the sample object category vector;

[0024] Based on the first cosine similarity and the first classification prediction value, determine the second model loss value of the first neural network model.

[0025] Optionally, the determining the second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value includes:

[0026] Determine the second cosine similarity corresponding to the pure sample satellite image and the third cosine similarity corresponding to the noisy sample satellite image from the first cosine similarity; wherein, the pure sample satellite image and the noisy sample satellite image are both included in the sample satellite image, and the noisy sample satellite image is generated based on the pure sample satellite image and a noise vector;

[0027] Based on the second cosine similarity, the third cosine similarity, and the noise vector, calculate the third model loss value of the first neural network model;

[0028] Based on the third model loss value and the first classification prediction value, determine the second model loss value of the first neural network model.

[0029] Optionally, the determining the maximum likelihood estimation loss value of the first neural network model based on the first classification prediction value includes:

[0030] Determine the first Gumbel Softmax values corresponding to the first classification prediction values respectively;

[0031] When the first Gumbel Softmax value is greater than or equal to the first threshold, input the first object category vector corresponding to the first Gumbel Softmax value into a multi-layer perceptron to obtain a second classification probability value output by the multi-layer perceptron;

[0032] Based on the second classification probability value, determine a first category identifier corresponding to the first object category vector;

[0033] Input the sample satellite image into the multi-layer perceptron to obtain a third classification probability value output by the multi-layer perceptron;

[0034] Based on the third classification probability value, determine the maximum likelihood estimation loss value of the first neural network model.

[0035] In a second aspect, an embodiment of the present application provides a high-quality satellite image retrieval method, and the method includes:

[0036] Input a query image into a target neural network model to obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein, the target neural network model is obtained based on any one of the above model training methods;

[0037] Determine a query category identifier of the query image as the first target object category corresponding to the first target classification prediction value greater than or equal to the first threshold;

[0038] Based on the cosine similarity between the first target object category vector and a standard category vector corresponding to a standard object category name, determine a query similarity identifier of the query image;

[0039] Based on the query category identifier and the query similarity identifier, perform a retrieval in an image database of a target satellite to obtain a target satellite image.

[0040] Optionally, the performing a retrieval in an image database of a target satellite based on the query category identifier and the query similarity identifier to obtain a target satellite image includes:

[0041] Determine a first similarity between an image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and a second similarity between an image similarity identifier of each first satellite image and the query similarity identifier;

[0042] Determine a first satellite image with the first similarity greater than or equal to a second threshold and the second similarity greater than or equal to a third threshold as the target satellite image.

[0043] In a third aspect, an embodiment of the present application provides a model training device, which includes:

[0044] An input-output module, configured to input a sample satellite image into a first neural network model, and obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object;

[0045] A determination module, configured to determine a first model loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image;

[0046] An adjustment module, configured to adjust model parameters of the first neural network model based on the first model loss value to obtain a target neural network model; wherein, the target neural network model is used to generate a target object category of a query image and target classification prediction values respectively corresponding to each of the target object categories.

[0047] Optionally, the first neural network model is composed of a convolutional neural network model, a Transformer-based bidirectional encoder representation model, and a Transformer-based encoder model. The input-output module includes:

[0048] A first input-output sub-module, configured to input a sample satellite image into the convolutional neural network, and obtain a sample satellite image feature vector output by the convolutional neural network model;

[0049] A second input-output sub-module, configured to input each object category name into the Transformer-based bidirectional encoder representation model, and obtain a sample object category vector output by the Transformer-based bidirectional encoder representation model;

[0050] A splicing sub-module, configured to splice the sample satellite image feature vector and the sample object category vector to obtain a first sample splicing vector;

[0051] A third input-output sub-module, configured to input the first sample splicing vector into the Transformer-based encoder, and obtain a second sample splicing vector output by the Transformer-based encoder;

[0052] A first determination sub-module, configured to determine a first object category vector from the second sample splicing vector, and based on the first object category vector, determine a first object category corresponding to the sample satellite image and first classification prediction values respectively corresponding to each of the first object categories.

[0053] Optionally, the determining module includes:

[0054] A second determining sub-module, configured to determine the binary cross-entropy loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value;

[0055] A third determining sub-module, configured to determine the second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector;

[0056] A fourth determining sub-module, configured to determine the maximum likelihood estimation loss value of the first neural network model based on the first classification prediction value;

[0057] A fifth determining sub-module, configured to determine the first model loss value of the first neural network model based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value.

[0058] Optionally, the third determining sub-module includes:

[0059] A semantic transformation unit, configured to perform semantic transformation on the first object category vector to obtain a second object category vector; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector;

[0060] A calculation unit, configured to calculate a first cosine similarity between the second object category vector and the sample object category vector;

[0061] A first determining unit, configured to determine the second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value.

[0062] Optionally, the first determining unit includes:

[0063] A first determining sub-unit, configured to determine a second cosine similarity corresponding to a pure sample satellite image and a third cosine similarity corresponding to a noisy sample satellite image from the first cosine similarity; wherein, the pure sample satellite image and the noisy sample satellite image are both included in the sample satellite image, and the noisy sample satellite image is generated based on the pure sample satellite image and a noise vector;

[0064] A calculation sub-unit, configured to calculate a third model loss value of the first neural network model based on the second cosine similarity, the third cosine similarity, and the noise vector;

[0065] A second determination subunit, configured to determine a second model loss value of the first neural network model based on the third model loss value and the first classification prediction value.

[0066] Optionally, the fourth determination sub-module includes:

[0067] A second determination unit, configured to determine first Gumbel Softmax values corresponding to the first classification prediction values respectively;

[0068] A first input / output unit, configured to, when the first Gumbel Softmax value is greater than or equal to a first threshold, input a first object category vector corresponding to the first Gumbel Softmax value into a multi-layer perceptron, and obtain a second classification probability value output by the multi-layer perceptron;

[0069] A third determination unit, configured to determine a first category identifier corresponding to the first object category vector based on the second classification probability value;

[0070] A second input / output unit, configured to input the sample satellite image into the multi-layer perceptron, and obtain a third classification probability value output by the multi-layer perceptron;

[0071] A fourth determination unit, configured to determine a maximum likelihood estimation loss value of the first neural network model based on the third classification probability value.

[0072] In a fourth aspect, an embodiment of the present application provides a high-quality satellite image retrieval device, where the device includes:

[0073] An input / output module, configured to input a query image into a target neural network model, and obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein, the target neural network model is obtained based on the model training method according to any one of claims 1 to 6;

[0074] A first determination module, configured to determine a second target object category corresponding to a first target classification prediction value greater than or equal to a first threshold as a query category identifier of the query image;

[0075] A second determination module, configured to determine a query similarity identifier of the query image based on a cosine similarity between the first target object category vector and a standard category vector corresponding to a standard object category name;

[0076] A retrieval module, configured to perform retrieval in an image database of a target satellite based on the query category identifier and the query similarity identifier, and obtain a target satellite image.

[0077] Optionally, the retrieval module includes:

[0078] A first determination sub-module, configured to determine a first similarity between the image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and a second similarity between the image similarity identifier of each first satellite image and the query similarity identifier;

[0079] A second determination sub-module, configured to determine a first satellite image whose first similarity is greater than or equal to a second threshold and whose second similarity is greater than or equal to a third threshold as the target satellite image.

[0080] In a fifth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the model training method or the high-quality satellite image retrieval method described in any one of the above.

[0081] In a sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the model training method or the high-quality satellite image retrieval method described in any one of the above is implemented.

[0082] Specific beneficial effects are as follows:

[0083] In an embodiment of the present application, by inputting a sample satellite image into a first neural network model, a first object category output by the first neural network model and first classification prediction values respectively corresponding to each first object category are obtained. Based on the sample object category, the first object category, and the first classification prediction values of the sample satellite image, a first model loss value of the first neural network model is determined. Based on the first model loss value, the model parameters of the first neural network model are adjusted to obtain a target neural network model. Wherein, the target neural network model is used to generate a target object category of a target query image and target classification prediction values respectively corresponding to each target object category based on the target query image. The first object category of each object in the sample satellite image and the first classification prediction values respectively corresponding to each first object category can be obtained through the first neural network model, and the first model loss value of the first neural network model can be further calculated according to the first object category, the first classification prediction values, and the sample object category of the sample satellite image itself. Thus, the model parameters of the first neural network model can be adjusted through the first model loss value to obtain the target neural network model. Since the processing of the sample satellite image is accurate to each object in the image, the prediction accuracy of the finally obtained target neural network model can be improved. When using the target neural network model to process a query image to retrieve satellite images, the matching degree between the finally obtained target satellite image and the query image can be improved, and the accuracy of image retrieval query can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the description of the embodiments of the present application will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0085] Figure 1 is a flowchart of another model training method provided by an embodiment of the present application;

[0086] Figure 2 is a flowchart of another model training method provided by an embodiment of the present application;

[0087] Figure 3 is a flowchart of a high-quality satellite image retrieval method provided by an embodiment of the present application;

[0088] Figure 4 is a logic block diagram of a model training device provided by an embodiment of the present application;

[0089] Figure 5 is a logic block diagram of a high-quality satellite image retrieval device provided by an embodiment of the present application;

[0090] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present application. Specific implementation manners

[0091] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings in the embodiments of the present application. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0092] Refer to Figure 1 , Figure 1 It is a schematic flowchart of a model training method provided by an embodiment of the present application. The method includes:

[0093] Step 101, input a sample satellite image into a first neural network model to obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object.

[0094] In the embodiments of the present application, the objects in the sample satellite image may refer to some specific objects. For example, the sample satellite image may include roads, railways, houses, vehicles, fields, etc., and these are all objects in the sample satellite image. Each object has a corresponding object category. The sample satellite image can be input into the first neural network model, so that a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector can be obtained. The first neural network model can be a convolutional neural network model, a recurrent neural network model, or a combined model of a convolutional neural network model and a recurrent neural network model. The first object category vector may include the categories of each object in the sample satellite image, and the first classification prediction value corresponding to the first object category vector may refer to the probabilities of the sample satellite image including various object categories.

[0095] Step 102, determine a first model loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value.

[0096] In the embodiments of the present application, the sample object category vector of the sample satellite image can be used as the annotation information of the sample satellite image. Based on the sample object category vector of the sample satellite image, the first object category vector output by the first neural network model, and the first classification prediction value, the first model loss value of the first neural network model can be calculated. For any object category, there are two cases: the sample satellite image contains this object category, or the sample satellite image does not contain this object category. Therefore, the first model loss value can be a binary cross-entropy loss value. In addition, considering that the first object category and the sample object category may not completely overlap, the loss value of the first neural network model in object classification can also be designed based on the sample object category and the first object category, which is called the second model loss value. Thus, the first model loss value can also be a combination of the binary cross-entropy loss value and the second model loss value. The combination method can be direct addition or addition with a regularization parameter.

[0097] Step 103: Based on the first model loss value, adjust the model parameters of the first neural network model to obtain a target neural network model; wherein, the target neural network model is used to generate the target object category of the query image and the target classification prediction values respectively corresponding to each of the target object categories.

[0098] In the embodiments of the present application, the model parameters of the first neural network model can be adjusted according to the first model loss value, so as to obtain a target neural network model. Among them, the adjustment direction of the above model parameters can be the direction that makes the first model loss value decrease. In addition, the adjustment process of the above model parameters can be repeated multiple times, that is, the process of steps 101 to 103 can be repeatedly executed to repeatedly train the first neural network model. When the training process of the first neural network model meets the preset convergence condition, the training of the first neural network model can be stopped, and the first neural network model obtained after the last model parameter adjustment is used as the target neural network model. The target neural network model can be used to process the query image, so as to generate the target object category of the query image and the target classification prediction values respectively corresponding to the target object category.

[0099] In an embodiment of the present application, by inputting a sample satellite image into a first neural network model, a first object category output by the first neural network model and first classification prediction values respectively corresponding to each first object category are obtained. Based on the sample object category, the first object category, and the first classification prediction values of the sample satellite image, a first model loss value of the first neural network model is determined. Based on the first model loss value, the model parameters of the first neural network model are adjusted to obtain a target neural network model. Wherein, the target neural network model is used to generate a target object category of a target query image and target classification prediction values respectively corresponding to each target object category. The first object category of each object in the sample satellite image and the first classification prediction values respectively corresponding to each first object category can be obtained through the first neural network model, and the first model loss value of the first neural network model can be further calculated according to the first object category and the first classification prediction values in combination with the sample object category of the sample satellite image itself. Thus, the model parameters of the first neural network model can be adjusted through the first model loss value to obtain the target neural network model. Since the processing of the sample satellite image is accurate to each object in the image, the prediction accuracy of the finally obtained target neural network model can be improved. When using the target neural network model to process a query image to retrieve satellite images, the matching degree between the finally obtained target satellite image and the query image can be improved, and the accuracy of image retrieval queries can be improved.

[0100] Referring to Figure 2 , Figure 2 FIG. is a schematic flowchart of another model training method provided by an embodiment of the present application. The method may include:

[0101] Step 201: Input the sample satellite image into the convolutional neural network to obtain a sample satellite image feature vector output by the convolutional neural network model.

[0102] In an embodiment of the present application, the first neural network model may be composed of a convolutional neural network model, a bidirectional encoder representation model based on Transformer, and an encoder based on Transformer. The convolutional neural network model may be used to extract feature vectors of an image. The sample satellite image may be input into the convolutional neural network model in the first neural network model, so that a sample satellite image feature vector output by the convolutional neural network model can be obtained. The dimension of the sample satellite image feature vector may be determined based on the dimension of the sample satellite image and the number of image features. For example, if the dimension of the sample image is , and the number of image features is d, the dimension of the sample satellite image feature vector may be .

[0103] Step 202: Input each object category name into the Transformer-based bidirectional encoder representation model to obtain the sample object category vector output by the Transformer-based bidirectional encoder representation model.

[0104] In an embodiment of the present application, the Transformer-based bidirectional encoder representation model is the BERT (Bidirectional Encoder Representations from Transformers) model, which can convert natural language text into numerical representations (which can be in the form of pure numbers, vectors, or matrices). Each object category name can be input into the BERT model in the first neural network model, so that the sample object category vector output by the BERT model can be obtained.

[0105] Step 203: Concatenate the sample satellite image feature vector and the sample object category vector to obtain a first sample concatenated vector.

[0106] In an embodiment of the present application, the sample satellite image feature vector and the object category vector can be concatenated to obtain a first sample concatenated vector. Among them, the elements in the object category vector can be located in the latter half of the first sample concatenated vector, and the elements in the sample satellite image feature vector can be located in the first half of the first sample concatenated vector.

[0107] For example, if the sample satellite image feature vector is , and the object category vector is , then the first sample concatenated vector can be .

[0108] Step 204: Input the first sample concatenated vector into the Transformer-based encoder to obtain the second sample concatenated vector output by the Transformer-based encoder.

[0109] In an embodiment of the present application, the self-attention mechanism of the Transformer-based encoder can be used to learn the correlation between the object category and the image features. On this basis, the first concatenated vector can be input into the Transformer-based encoder in the first neural network model, so that the second sample concatenated vector output by the Transformer-based encoder can be obtained.

[0110] Continuing with the above example, the second sample concatenated vector can be .

[0111] Step 205: Determine the first object category vector from the second sample splicing vector, and based on the first object category vector, determine the first object category corresponding to the sample satellite image and the first classification prediction value corresponding to each of the first object categories.

[0112] In an embodiment of the present application, the first object category vector can be obtained from the second sample splicing vector. After obtaining the first object category vector, the first object category corresponding to the sample satellite image and the first classification prediction value corresponding to each of the first object categories can be determined according to the first category vector.

[0113] Continuing with the above example, the first object category vector can be , which contains a total of C first object categories. The first classification prediction value can be obtained by linearizing the first object category vector and then calculating it through the sigmoid function, expressed as , where represents a linear function, represents the sigmoid function, . If , it means that the object category i is not included in the first object category. If , it means that the object category i may be included in the first object category.

[0114] In an embodiment of the present application, by inputting the sample satellite image into a convolutional neural network, the sample satellite image feature vector output by the convolutional neural network model is obtained. Each object category name is input into the bidirectional encoder representation model based on Transformer, and the sample object category vector output by the bidirectional encoder representation model based on Transformer is obtained. The sample satellite image feature vector and the sample object category vector are spliced to obtain the first sample splicing vector. The first sample splicing vector is input into the encoder based on Transformer, and the second sample splicing vector output by the encoder based on Transformer is obtained. The first object category vector is determined from the second sample splicing vector, and based on the first object category vector, the first object category corresponding to the sample satellite image and the first classification prediction value corresponding to each of the first object categories are determined. The first neural network model combined with a convolutional neural network, a bidirectional encoder representation model based on Transformer, and an encoder based on Transformer can be used to process the sample satellite image step by step to obtain the first object category vector, the first object category, and the first classification prediction value, making full use of the performance advantages of each mature model and improving the accuracy of the first object category vector, the first object category, and the first classification prediction value to a certain extent.

[0115] Step 206: Determine the binary cross-entropy loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value.

[0116] In an embodiment of the present application, the binary cross-entropy loss value of the first neural network model can be calculated according to the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value. The calculation method is shown in the following formulas (1) and (2):

[0117] (Formula 1)

[0118] (Formula 2)

[0119] In the above formulas (1) and (2), represents the binary cross-entropy loss value, represents the binary cross-entropy loss function, is the number of sample satellite images, C is the number of all object categories that the sample satellite image may contain, represents the first classification prediction value that the sample satellite image i contains the category j, represents whether the sample satellite image i contains the object category j, The value of can be obtained according to and the first threshold . If , then , indicating that the sample satellite image i contains the object category j. If , then , indicating that the sample satellite image i does not contain the object category j. When calculating, the corresponding relationship between the first classification prediction value and each object category included in the sample satellite image can be determined according to the corresponding relationship of the elements in the sample object category vector and the first object category vector. Furthermore, according to the above corresponding relationship, the first classification prediction value can be substituted into formulas (1) and (2) to calculate the binary cross-entropy loss value of the first neural network model.

[0120] Step 207: Determine the second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector.

[0121] In an embodiment of the present application, the second model loss value of the first neural network model can be calculated according to the first classification prediction value, the sample object category vector of the sample satellite image, and the first object category vector. The calculation method is shown in the following formulas (3) and (4):

[0122] (Formula 3)

[0123] (Equation 4)

[0124] In the above Equations 3 and 4, represents the second model loss value, represents the number of sample satellite images without noise vectors, represents the first classification prediction value that the sample satellite image i contains the object category j, represents the sample feature representation of the object category j in the sample satellite image i, which can be obtained by transforming the name of the object category j through the BERT model. represents the first feature representation of the object category j in the sample satellite image i, which can be obtained after processing the sample satellite image by the first neural network model. represents the calculation process of Equation 4. When calculating, the corresponding relationship between the first classification prediction value and each object category included in the sample satellite image can be determined according to the corresponding relationship of each element in the sample object category vector and the first object category vector. Furthermore, according to the above corresponding relationship, the first classification prediction value can be substituted into Equations 3 and 4 to calculate the binary cross-entropy loss value of the first neural network model.

[0125] Optionally, in step 207, the following sub-steps may be included:

[0126] Sub-step 2071: Semantically transform the first object category vector to obtain a second object category vector; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector.

[0127] In the embodiments of the present application, the first object category vector can be semantically transformed to obtain a second object category vector. When transforming, the semantic space where the sample object category vector is located can be used as a reference space, and the transformation is performed through a linear function between the semantic space of the first object category vector and the semantic space of the sample object category vector, so that the semantic space of the obtained second object category vector is the same as the semantic space of the sample object category vector.

[0128] Sub-step 2072: Calculate the first cosine similarity between the second object category vector and the sample object category vector.

[0129] In the embodiments of the present application, the first cosine similarity between the second object category vector and the sample object category vector can be calculated, such as the part in Equation 4.

[0130] Sub-step 2073: Determine the second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value.

[0131] In an embodiment of the present application, the second model loss value of the first neural network model can be calculated according to the first cosine similarity and the first classification prediction value, and the calculation method is as shown in the above formulas 3 and 4, where, is the first cosine similarity, is the first classification prediction value. The meanings of the other parameter symbols can be referred to the content of the embodiment under step 207, and will not be elaborated here.

[0132] Optionally, in sub-step 2073, the following sub-steps may be included:

[0133] Sub-step A1, determining the second cosine similarity corresponding to the pure sample satellite image and the third cosine similarity corresponding to the noise sample satellite image from the first cosine similarity; wherein, the pure sample satellite image and the noise sample satellite image are both included in the sample satellite image, and the noise sample satellite image is generated based on the pure sample satellite image and the noise vector.

[0134] In an embodiment of the present application, the sample satellite image may include a pure sample satellite image and a noise sample satellite image. The pure sample satellite image does not contain a noise vector, and the noise sample satellite image can be obtained by adding a noise vector to the pure sample satellite image. Therefore, the first cosine similarity also includes the second cosine similarity corresponding to the pure sample satellite image and the third cosine similarity corresponding to the noise sample satellite image. Thus, the second cosine similarity and the third cosine similarity can be determined from the first cosine similarity. The second cosine similarity can be expressed as , representing the cosine similarity of the object category j included in the sample satellite image i, and the third cosine similarity can be expressed as , representing the cosine similarity of the noise object category j included in the sample satellite image i.

[0135] Sub-step A2, calculating the third model loss value of the first neural network model based on the second cosine similarity, the third cosine similarity, and the noise vector.

[0136] In an embodiment of the present application, the third model loss value of the first neural network model can be calculated according to the second cosine similarity, the third cosine similarity, and the noise vector, and the calculation method is as shown in the following formula 5:

[0137] (Formula 5)

[0138] In the above formula 5, represents the third model loss value, represents the feature representation corresponding to the noise vector in the object category j, which can be obtained by transforming the noise vector through the BERT model. Indicates taking a positive number. Among them, the noise vector can be manually set and added to the sample satellite image, so that a sample satellite image containing the noise vector can be obtained. The interpretations of the remaining parameter symbols can refer to the interpretations of the parameter symbols in Equations 3 and 4, which will not be elaborated here.

[0139] Sub-step A3: Determine the second model loss value of the first neural network model based on the third model loss value and the first classification prediction value.

[0140] In the embodiments of the present application, the second loss value of the first neural network model can be calculated according to the third model loss value and the first classification prediction value, and its calculation method is shown in Equation 6 below:

[0141] (Equation 6)

[0142] In the above Equation 6, represents the third model loss value. The interpretations of the remaining parameter symbols can refer to the interpretations of the parameter symbols in Equations 3 and 4, which will not be elaborated here.

[0143] In the embodiments of the present application, the second cosine similarity corresponding to the pure sample satellite image and the third cosine similarity corresponding to the noisy sample satellite image are determined from the first cosine similarity; among them, the pure sample satellite image and the noisy sample satellite image are both included in the sample satellite image, and the noisy sample satellite image is generated based on the pure sample satellite image and the noise vector. Based on the second cosine similarity, the third cosine similarity, and the noise vector, the third model loss value of the first neural network model is calculated. Based on the third model loss value and the first classification prediction value, the second model loss value of the first neural network model is determined. The second model loss value of the first neural network model can be calculated using the different image features of the pure sample satellite image and the noisy sample satellite image, which improves the accuracy and reliability of the second model loss value to a certain extent and can also improve the prediction performance of the finally obtained target neural network model.

[0144] In an embodiment of the present application, by performing semantic transformation on the first object category vector, a second object category vector is obtained; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector, the first cosine similarity between the second object category vector and the sample object category vector is calculated, and based on the first cosine similarity and the first classification prediction value, the second model loss value of the first neural network model is determined. The first object category vector can be transformed by means of semantic transformation, and then the first cosine similarity between the obtained second object category vector after transformation and the sample object category vector can be calculated. Finally, the second model loss value of the first neural network model can be calculated according to the first cosine similarity, avoiding the calculation error caused by different semantic spaces and improving the accuracy of the second model loss value to a certain extent.

[0145] Step 208, based on the first classification prediction value, determine the maximum likelihood estimation loss value of the first neural network model.

[0146] In an embodiment of the present application, the maximum likelihood estimation loss value of the first neural network can be calculated according to the first classification prediction value, and its calculation method is shown in the following formulas 5 to 7:

[0147] (Formula 5)

[0148] (Formula 6)

[0149] (Formula 7)

[0150] In the above formulas 5 to 7, represents the maximum likelihood estimation loss value, represents the number of sample satellite images, represents the sample satellite image represents the number of object categories in the sample satellite image represents the sample satellite image represents the maximum Gumbel Softmax value for object category j in the sample satellite image. is the similarity identifier for object category j. is the Gumbel Softmax function, represents the first classification prediction value indicating that the sample satellite image contains object category j. represents the temperature coefficient, and k is an inherent parameter in the GumbelSoftmax function, .

[0151] Optionally, step 208 may include the following sub-steps:

[0152] Sub-step 2081: Determine the first Gumbel Softmax values corresponding to the first classification prediction values.

[0153] In the embodiments of the present application, first, the first Gumbel Softmax values corresponding to the first classification prediction values can be calculated. The calculation method is as shown in Equation 7 above.

[0154] Sub-step 2082: When the first Gumbel Softmax value is greater than or equal to the first threshold, input the first object category vector corresponding to the first Gumbel Softmax value into a multi-layer perceptron to obtain the second classification probability value output by the multi-layer perceptron.

[0155] In the embodiments of the present application, if the first Gumbel Softmax value is greater than or equal to the first threshold, the first object category vector corresponding to the first Gumbel Softmax value can be input into a multi-layer perceptron, so as to obtain the second classification probability value output by the multi-layer perceptron. The above process can be expressed in the form of Equation 8 as follows:

[0156] (Equation 8)

[0157] In Equation 8 above, represents the second classification probability value for object category j, , which serves as an identifier. is the propagation function of the multi-layer perceptron, is the feature representation of object category j in the second object category vector.

[0158] Sub-step 2083: Based on the second classification probability value, determine the first category identifier corresponding to the first object category vector.

[0159] In the embodiments of the present application, the first category representation corresponding to the first object category vector can be obtained according to the second classification probability value, as shown in Equation 9 below:

[0160] = (Equation 9)

[0161] In Equation 9 above, represents the maximum first Gumbel Softmax value. The meanings of the other parameter symbols can be referred to the interpretations of the parameter symbols in Equations 6 and 8, and will not be elaborated here.

[0162] Sub-step 2084: Input the sample satellite image into the multi-layer perceptron to obtain the third classification probability value output by the multi-layer perceptron.

[0163] In an embodiment of the present application, a sample satellite image can be input into a multi-layer perceptron to obtain a third classification probability value output by the multi-layer perceptron, as shown in the following formula 10:

[0164] (Formula 10)

[0165] In the above formula 10, can be used as a reference quantity during the propagation of the multi-layer perceptron, represents the sample satellite image i, represents the third classification probability value.

[0166] Sub-step 2085, based on the third classification probability value, determine the maximum likelihood estimation loss value of the first neural network model.

[0167] In an embodiment of the present application, the maximum likelihood estimation loss value of the first neural network model can be calculated according to the third classification probability value, and its calculation method is as shown in the above formula 5. The meanings of each parameter symbol can refer to the embodiment content under step 208 and will not be elaborated here.

[0168] In an embodiment of the present application, by determining the first Gumbel Softmax value corresponding to the first classification prediction value, when the first Gumbel Softmax value is greater than or equal to the first threshold, the first object category vector corresponding to the first Gumbel Softmax value is input into the multi-layer perceptron to obtain a second classification probability value output by the multi-layer perceptron. Based on the second classification probability value, determine the first category identifier corresponding to the first object category vector. Input the sample satellite image into the multi-layer perceptron to obtain a third classification probability value output by the multi-layer perceptron. Based on the third classification probability value, determine the maximum likelihood estimation loss value of the first neural network model. The Gumbel Softmax value can be used to replace the first classification prediction value and the first threshold for numerical judgment. Furthermore, the third classification probability value corresponding to the sample satellite image can be obtained through the multi-layer perceptron, and the maximum likelihood estimation loss value of the first neural network model can be calculated according to the third classification probability value. Since the Gumbel Softmax function is differentiable, using the Gumbel Softmax value to replace the discrete and discontinuous first classification prediction value can improve the accuracy and reliability of the maximum likelihood estimation loss value to a certain extent, and can improve the prediction performance of the finally obtained target neural network model.

[0169] Step 209, based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value, determine the first model loss value of the first neural network model.

[0170] In an embodiment of the present application, the first model loss value of the first neural network model can be calculated based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value, and its calculation method is shown in Equation 11 below:

[0171] (Equation 11)

[0172] In Equation 11 above, is the first model loss value, is the binary cross-entropy loss value, is the second model loss value, is the maximum likelihood estimation loss value, is the regularization parameter of the second model loss value, which is used to balance the proportion of the second model loss value in the first model loss value.

[0173] In an embodiment of the present application, based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image, the binary cross-entropy loss value of the first neural network model is determined, based on the first classification prediction value, the sample object category vector, and the first object category vector, the second model loss value of the first neural network model is determined, based on the first classification prediction value, the maximum likelihood estimation loss value of the first neural network model is determined, and based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value, the first model loss value of the first neural network model is determined. The binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value of the first neural network model can be calculated, and the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value can be combined to obtain the first model loss value of the first neural network model, which improves the accuracy and usability of the first model loss value to a certain extent, and can improve the prediction performance of the finally obtained target neural network model.

[0174] Step 210, based on the first model loss value, adjust the model parameters of the first neural network model to obtain a target neural network model.

[0175] In an embodiment of the present application, the implementation content of this step can refer to the embodiment content of step 103, which will not be elaborated here.

[0176] Refer to Figure 3 , Figure 3 is a schematic flowchart of a high-quality satellite image retrieval method provided by an embodiment of the present application. The method may include:

[0177] Step 301: Input the query image into the target neural network model to obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein, the target neural network model is obtained based on any one of the model training methods described above.

[0178] In an embodiment of the present application, the query image can be an image with a certain number of object categories. The query image can be input into the target neural network model, so that a first target object category vector output by the target neural network model and a target classification prediction value corresponding to the first target object category vector can be obtained.

[0179] Step 302: Determine the query category identifier of the query image as the first target object category corresponding to the first target classification prediction value greater than or equal to the first threshold.

[0180] In an embodiment of the present application, the number of first target classification prediction values can be multiple, corresponding to each object category respectively. The first target classification prediction values can be compared with the first threshold, and the first target object category corresponding to the first target classification prediction value greater than or equal to the first threshold is determined as the query category identifier of the query image, which is used to query the corresponding target satellite image in the target satellite.

[0181] Step 303: Determine the query similarity identifier of the query image based on the cosine similarity between the first target object category vector and the standard category vector corresponding to the standard object category name.

[0182] In an embodiment of the present application, the cosine similarity between the first target object category vector and the standard category vector corresponding to the standard object category name can be calculated, and then the query similarity identifier of the query image can be generated according to this cosine similarity.

[0183] For example, the generation method of the query similarity identifier is shown in Equation 12 below:

[0184] (Equation 12)

[0185] In Equation 12, is the cosine similarity between the feature of the first target object category vector and the feature representation of the standard category vector, where i represents the position of the element in the vector, represents the feature representation at position i in the standard category vector, represents the feature representation at position i in the first target object category vector. represents the floor function, is a preset proportional constant. For example, if when, .

[0186] Step 304: Retrieve in the image database of the target satellite based on the query category identifier and the query similarity identifier to obtain the target satellite image.

[0187] In the embodiment of the present application, in the image database of the target satellite, the image category identifier and the image similarity identifier of the first satellite images included in its database can also be pre-generated according to the methods of steps 302 and 303. Thus, through the query category identifier and the query similarity identifier, a match can be made in the image database of the target satellite. If the match is successful, the first satellite image that matches can be used as the target satellite image.

[0188] Optionally, step 304 may include the following sub-steps:

[0189] Sub-step 3041: Determine the first similarity between the image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and the second similarity between the image similarity identifier of each first satellite image and the query similarity identifier.

[0190] In the embodiment of the present application, the first similarity between the image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and the second similarity between the image similarity identifier of each first satellite image and the query similarity identifier can be calculated. Among them, the categories of the first similarity and the second similarity can be cosine similarity.

[0191] Sub-step 3042: Determine the first satellite image whose first similarity is greater than or equal to the second threshold and whose second similarity is greater than or equal to the third threshold as the target satellite image.

[0192] In the embodiment of the present application, screening can be performed among the first satellite images, and the first satellite image whose first similarity is greater than or equal to the second threshold and whose second similarity is greater than or equal to the third threshold is determined as the target satellite image.

[0193] In an embodiment of the present application, by determining the first similarity between the image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and the second similarity between the image similarity identifier of each first satellite image and the query similarity identifier, the first satellite image with the first similarity greater than or equal to the second threshold and the second similarity greater than or equal to the third threshold is determined as the target satellite image. The target satellite image can be screened from the first satellite images according to the first similarity between the query category identifier and the image category identifier of the first satellite image, and the second similarity between the query similarity identifier and the image similarity identifier of the first satellite image, which can improve the matching degree between the target image and the query image and the accuracy of target satellite image retrieval.

[0194] In an embodiment of the present application, by inputting the query image into the target neural network model, the first target object category vector output by the target neural network model and the first target classification prediction value corresponding to the first target object category vector are obtained; wherein, the target neural network model is obtained based on any one of the above model training methods. The first target object category corresponding to the first target classification prediction value greater than or equal to the first threshold is determined as the query category identifier of the query image. Based on the cosine similarity between the first target object category vector and the standard category vector corresponding to the standard object category name, the query similarity identifier of the query image is determined. Based on the query category identifier and the query similarity identifier, a search is performed in the image database of the target satellite to obtain the target satellite image. The query category identifier and the query similarity identifier corresponding to the query image can be obtained through the target neural network model, and a search can be performed in the image database of the target satellite according to the query category identifier and the query similarity identifier to obtain the target satellite image. Since two types of query identifiers are used, the matching degree between the target satellite image and the query image can be improved to a certain extent, and the accuracy of satellite image retrieval query can be improved.

[0195] Refer to Figure 4 , Figure 4 FIG. is a logical block diagram of a model training device provided in an embodiment of the present application. The device 400 may include:

[0196] An input / output module 401, configured to input a sample satellite image into the first neural network model to obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object;

[0197] A determination module 402, configured to determine a first model loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value;

[0198] An adjustment module 403, configured to adjust model parameters of the first neural network model based on the first model loss value to obtain a target neural network model, where the target neural network model is used to generate a target object category of a query image and target classification prediction values respectively corresponding to each of the target object categories.

[0199] Optionally, the first neural network model is composed of a convolutional neural network model, a bidirectional encoder representation model based on Transformer, and an encoder model based on Transformer. The input-output module 401 includes:

[0200] A first input-output sub-module, configured to input a sample satellite image into the convolutional neural network to obtain a sample satellite image feature vector output by the convolutional neural network model;

[0201] A second input-output sub-module, configured to input each object category name into the bidirectional encoder representation model based on Transformer to obtain a sample object category vector output by the bidirectional encoder representation model based on Transformer;

[0202] A splicing sub-module, configured to splice the sample satellite image feature vector and the sample object category vector to obtain a first sample splicing vector;

[0203] A third input-output sub-module, configured to input the first sample splicing vector into the encoder based on Transformer to obtain a second sample splicing vector output by the encoder based on Transformer;

[0204] A first determination sub-module, configured to determine a first object category vector from the second sample splicing vector, and based on the first object category vector, determine a first object category corresponding to the sample satellite image and first classification prediction values respectively corresponding to each of the first object categories.

[0205] Optionally, the determination module 402 includes:

[0206] A second determination sub-module, configured to determine a binary cross-entropy loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction values of the sample satellite image;

[0207] A third determination sub-module, configured to determine a second model loss value of the first neural network model based on the first classification prediction values, the sample object category vector, and the first object category vector;

[0208] A fourth determination sub-module, configured to determine a maximum likelihood estimation loss value of the first neural network model based on the first classification prediction value;

[0209] A fifth determination sub-module, configured to determine a first model loss value of the first neural network model based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value.

[0210] Optionally, the third determination sub-module includes:

[0211] A semantic transformation unit, configured to perform semantic transformation on the first object category vector to obtain a second object category vector; wherein, a semantic space of the second object category vector is the same as a semantic space of the sample object category vector;

[0212] A calculation unit, configured to calculate a first cosine similarity between the second object category vector and the sample object category vector;

[0213] A first determination unit, configured to determine a second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value.

[0214] Optionally, the first determination unit includes:

[0215] A first determination sub-unit, configured to determine a second cosine similarity corresponding to a pure sample satellite image and a third cosine similarity corresponding to a noise sample satellite image from the first cosine similarity; wherein, the pure sample satellite image and the noise sample satellite image are both included in the sample satellite image, and the noise sample satellite image is generated based on the pure sample satellite image and a noise vector;

[0216] A calculation sub-unit, configured to calculate a third model loss value of the first neural network model based on the second cosine similarity, the third cosine similarity, and the noise vector;

[0217] A second determination sub-unit, configured to determine a second model loss value of the first neural network model based on the third model loss value and the first classification prediction value.

[0218] Optionally, the fourth determination sub-module includes:

[0219] A second determination unit, configured to determine first Gumbel Softmax values corresponding to the first classification prediction values respectively;

[0220] The first input / output unit is configured to, when the first Gumbel Softmax value is greater than or equal to the first threshold, input the first object category vector corresponding to the first Gumbel Softmax value into a multi-layer perceptron to obtain a second classification probability value output by the multi-layer perceptron;

[0221] The third determination unit is configured to determine a first category identifier corresponding to the first object category vector based on the second classification probability value;

[0222] The second input / output unit is configured to input the sample satellite image into the multi-layer perceptron to obtain a third classification probability value output by the multi-layer perceptron;

[0223] The fourth determination unit is configured to determine a maximum likelihood estimation loss value of the first neural network model based on the third classification probability value.

[0224] The model training device provided by the embodiments of the present application can implement Figures 1 to 2 each process implemented by the method embodiments, and for the sake of avoiding repetition, it will not be elaborated here.

[0225] Referring to Figure 5 , Figure 5 FIG. is a logic block diagram of a model training device provided by an embodiment of the present application. The device 500 may include:

[0226] An input / output module 501 is configured to input a query image into a target neural network model to obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein, the target neural network model is obtained based on the model training method according to any one of claims 1 to 6;

[0227] The first determination module 502 is configured to determine a second target object category corresponding to a first target classification prediction value greater than or equal to the first threshold as a query category identifier of the query image;

[0228] The second determination module 503 is configured to determine a query similarity identifier of the query image based on a cosine similarity between the first target object category vector and a standard category vector corresponding to a standard object category name;

[0229] A retrieval module 504 is configured to perform a retrieval in an image database of a target satellite based on the query category identifier and the query similarity identifier to obtain a target satellite image.

[0230] Optionally, the retrieval module 504 includes:

[0231] The first determination sub-module is configured to determine a first similarity between an image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and a second similarity between an image similarity identifier of each first satellite image and the query similarity identifier;

[0232] The second determination sub-module is configured to determine, as target satellite images, the first satellite images for which the first similarity is greater than or equal to a second threshold and the second similarity is greater than or equal to a third threshold.

[0233] The model training device provided by the embodiments of the present application can implement Figure 3 each process implemented by the method embodiments. To avoid repetition, details are not described herein again.

[0234] The model training device and the high-quality satellite image retrieval device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than a terminal. Exemplarily, the electronic device may be a GPU BOX, a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0235] The model training device and the high-quality satellite image retrieval device in the embodiments of the present application may be devices with an operating system. The operating system may be an Android operating system, a Linux, a Windows operating system, etc., and may also be other possible operating systems. The embodiments of the present application do not make specific limitations.

[0236] The embodiments of the present application provide an electronic device. Refer to Figure 6, the electronic device 60 includes: a processor 601, a memory 602, and a computer program 6021 stored on the memory 602 and executable on the processor 601. When the processor 601 executes the program, it implements the model training method or the high-quality satellite image retrieval method of the foregoing embodiments.

[0237] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps in the model training method or the high-quality satellite image retrieval method disclosed in the embodiments of the present application are implemented.

[0238] The embodiments of the present application further provide a computer program product. When the computer program product runs on an electronic device, it causes the processor to implement the steps in the model training method or the high-quality satellite image retrieval method disclosed in the embodiments of the present application when executed.

[0239] The various embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0240] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, devices, electronic devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0241] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0242] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process or multiple processes and / or blocks. Figure 1 One process or multiple processes and / or blocks Figure 1 Steps for implementing the functions specified in one block or multiple blocks.

[0243] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0244] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.

[0245] The above has introduced in detail a model training method, a high-quality satellite image retrieval method and a device provided by the present application. Specific examples are used in this document to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A model training method, characterized in that, The method includes: Inputting a sample satellite image into a first neural network model to obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object; Determining a first model loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image; Adjusting the model parameters of the first neural network model based on the first model loss value to obtain a target neural network model; wherein, the target neural network model is used to generate a target object category of a query image and target classification prediction values respectively corresponding to each of the target object categories; The determining the first model loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image includes: Determining a binary cross-entropy loss value of the first neural network model based on the sample object category vector, the first object category vector, and the first classification prediction value of the sample satellite image; Determining a second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector; Determining a maximum likelihood estimation loss value of the first neural network model based on the first classification prediction value; Determining the first model loss value of the first neural network model based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value; The determining the second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector includes: Performing semantic transformation on the first object category vector to obtain a second object category vector; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector; Calculating a first cosine similarity between the second object category vector and the sample object category vector; Determining the second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value.

2. The method according to claim 1, wherein The first neural network model is composed of a convolutional neural network model, a bidirectional encoder representation model based on Transformer, and an encoder model based on Transformer. The inputting a sample satellite image into the first neural network model to obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector includes: Inputting the sample satellite image into the convolutional neural network to obtain a sample satellite image feature vector output by the convolutional neural network model; Inputting each object category name into the bidirectional encoder representation model based on Transformer to obtain a sample object category vector output by the bidirectional encoder representation model based on Transformer; Splicing the sample satellite image feature vector and the sample object category vector to obtain a first sample splicing vector; Inputting the first sample splicing vector into the Transformer-based encoder to obtain a second sample splicing vector output by the Transformer-based encoder; A first object category vector is determined from the second sample splicing vector, and based on the first object category vector, a first object category corresponding to the sample satellite image and a first classification prediction value corresponding to each of the first object categories are determined.

3. The method according to claim 1, wherein The determining, based on the first cosine similarity and the first classification prediction value, a second model loss value of the first neural network model includes: Determining a second cosine similarity corresponding to a clean sample satellite image and a third cosine similarity corresponding to a noisy sample satellite image from the first cosine similarity; wherein the clean sample satellite image and the noisy sample satellite image are both included in the sample satellite image, and the noisy sample satellite image is generated based on the clean sample satellite image and a noise vector; Calculating a third model loss value of the first neural network model based on the second cosine similarity, the third cosine similarity, and the noise vector; Based on the third model loss value and the first classification prediction value, a second model loss value of the first neural network model is determined.

4. The method according to claim 1, wherein The determining, based on the first classification prediction value, a maximum likelihood estimation loss value of the first neural network model includes: Determining first Gumbel Softmax values corresponding to the first classification prediction values respectively; When the first Gumbel Softmax value is greater than or equal to a first threshold, inputting a first object category vector corresponding to the first Gumbel Softmax value into a multilayer perceptron to obtain a second classification probability value output by the multilayer perceptron; determining a first category identifier corresponding to the first object category vector based on the second classification probability value; Inputting the sample satellite image into the multi-layer perceptron to obtain a third classification probability value output by the multi-layer perceptron; Based on the third classification probability value, a maximum likelihood estimation loss value of the first neural network model is determined.

5. A high-quality satellite image retrieval method, characterized in that: The method comprises: Inputting a query image into a target neural network model to obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein the target neural network model is obtained based on the model training method according to any one of claims 1 to 4; Determining a first target object category corresponding to a first target classification prediction value that is greater than or equal to a first threshold as a query category identifier of the query image; determining a query similarity identifier of the query image based on a cosine similarity between the first target object category vector and a standard category vector corresponding to a standard object category name; Based on the query category identifier and the query similarity identifier, a search is performed in an image database of the target satellite to obtain an image of the target satellite.

6. The method according to claim 5, wherein Retrieving in the image database of the target satellite based on the query category identifier and the query similarity identifier to obtain a target satellite image, including: Determining a first similarity between the image category identifier of each first satellite image included in the image database of the target satellite and the query category identifier, and a second similarity between the image similarity identifier of each first satellite image and the query similarity identifier; Determining the first satellite images with the first similarity greater than or equal to a second threshold and the second similarity greater than or equal to a third threshold as target satellite images.

7. A model training device, characterized in that, The device includes: An input / output module, configured to input a sample satellite image into a first neural network model, and obtain a first object category vector output by the first neural network model and a first classification prediction value corresponding to the first object category vector; wherein, the sample satellite image includes at least one object; A determination module, configured to determine a first model loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value; An adjustment module, configured to adjust model parameters of the first neural network model based on the first model loss value to obtain a target neural network model; wherein, the target neural network model is used to generate a target object category of a query image and target classification prediction values respectively corresponding to each of the target object categories; The determination module includes: A second determination sub-module, configured to determine a binary cross-entropy loss value of the first neural network model based on the sample object category vector of the sample satellite image, the first object category vector, and the first classification prediction value; A third determination sub-module, configured to determine a second model loss value of the first neural network model based on the first classification prediction value, the sample object category vector, and the first object category vector; A fourth determination sub-module, configured to determine a maximum likelihood estimation loss value of the first neural network model based on the first classification prediction value; A fifth determination sub-module, configured to determine a first model loss value of the first neural network model based on the binary cross-entropy loss value, the second model loss value, and the maximum likelihood estimation loss value; The third determination sub-module includes: A semantic transformation unit, configured to perform semantic transformation on the first object category vector to obtain a second object category vector; wherein, the semantic space of the second object category vector is the same as the semantic space of the sample object category vector; A calculation unit, configured to calculate a first cosine similarity between the second object category vector and the sample object category vector; A first determination unit, configured to determine a second model loss value of the first neural network model based on the first cosine similarity and the first classification prediction value.

8. A high-quality satellite image retrieval device, characterized in that, The device includes: An input / output module, configured to input a query image into a target neural network model, and obtain a first target object category vector output by the target neural network model and a first target classification prediction value corresponding to the first target object category vector; wherein the target neural network model is obtained based on the model training method according to any one of claims 1 to 4; a first determining module, configured to determine a second target object category corresponding to a first target classification prediction value that is greater than or equal to a first threshold as a query category identifier of the query image; a second determining module, configured to determine a query similarity identifier of the query image based on a cosine similarity between the first target object category vector and a standard category vector corresponding to a standard object category name; The retrieval module is used to search the image database of the target satellite based on the query category identifier and the query similarity identifier to obtain the target satellite image.