Cross-modal remote sensing image-text retrieval method based on expert-guided trusted learning

Through expert-guided trusted learning methods, the Dilicre distribution is constructed for cross-modal retrieval of remote sensing images and text descriptions, solving the performance attenuation problem of existing methods under noise interference, and achieving higher retrieval accuracy and adaptability.

CN120336449APending Publication Date: 2025-07-18TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510471145.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing cross-modal remote sensing graphic and text retrieval method based on deep learning is affected by noise in practical applications, and its performance attenuation is severe, and it is unable to effectively handle text descriptions of diversity and noise, resulting in a decrease in retrieval accuracy.

Method used

Using a method based on expert-guided trusted learning, the features of remote sensing images and text descriptions are extracted through feature extractors and expert feature extractors, and the Dirichlet distribution is constructed, and the relationship learning in-modal and intermodal trusted learning is performed, which increases uncertain consistency learning, and balances prediction accuracy and uncertainty.

Benefits of technology

The cross-modal retrieval accuracy of remote sensing images and text descriptions is improved, the retrieval performance of the model under noisy data is enhanced, and the model is adapted to diverse application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336449A_ABST
    Figure CN120336449A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal remote sensing image-text retrieval method based on expert-guided trusted learning, and relates to the technical field of information processing and mining of remote sensing big data, the method comprises the following steps: using a feature extractor and an expert feature extractor to respectively extract features described by a remote sensing image and a text; aiming at the extracted features, constructing corresponding Dirichlet distribution based on internal similarity of the features; aiming at the remote sensing image features and the text description features extracted by the feature extractor, respectively constructing Dirichlet distribution based on the similarity of the image to the text and the similarity of the text to the image; and on the basis of the constructed Dirichlet distribution, performing intra-modal relationship learning, inter-modal credible learning and uncertainty consistency learning guided by experts. By the adoption of the scheme, cross-modal retrieval between the remote sensing image and the text description is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of information processing and mining of remote sensing big data, and particularly to a cross-modal remote sensing image-text retrieval method and device based on expert-guided trustworthy learning. Background Art

[0002] With the exponential growth of the global remote sensing data volume, these data have brought great opportunities for ecological monitoring, climate analysis, disaster prevention, etc. At the same time, how to obtain key information from the massive remote sensing data is a crucial challenge. Content-based remote sensing image retrieval is precisely to solve this problem. Its goal is to find the most relevant samples in a huge dataset, and it is a key technology for remote sensing big data management. In recent years, single-modal remote sensing retrieval methods have made great progress. However, single-modal remote sensing retrieval requires the retrieval query samples to be consistent with the modality of the data archive, which limits the flexibility of the retrieval model and reduces its adaptability in diverse application scenarios.

[0003] Similar to the fact that humans usually need multiple senses to obtain a comprehensive understanding of the environment, different modalities of remote sensing data have different characteristics. For example, optical images are easy to interpret and rich in texture information, but are greatly affected by weather; synthetic aperture radar (SAR) images can be acquired all-weather, but have low resolution and are difficult to interpret; text descriptions can be directly understood by people, but are not intuitive and highly subjective. Comprehensive utilization of different modalities of data helps researchers to comprehensively understand the target of interest. Cross-modal remote sensing retrieval is precisely to achieve in-depth mining and efficient application of different modalities of data through accurate matching of different modalities. The definition of cross-modal remote sensing retrieval is to given a query sample of one modality (such as an optical image), find the most relevant sample in the database of another modality (such as an SAR image). This cross-modal retrieval technology can match between different modalities of data such as optical images, SAR images, and text descriptions, and has been widely applied in fields such as disaster rescue, military surveillance, and urban planning. The cross-modal remote sensing image-text retrieval task is a hot spot in cross-modal retrieval tasks. Remote sensing images can provide rich spatial information, including detailed features such as the distribution, shape, and structure of ground objects, while text can describe complex semantics, such as objects, events, and relationships in a scene. By querying text with an image, rich scene descriptions can be obtained; by querying an image with text, an intuitive impression of the scene can be obtained. Cross-modal remote sensing image-text retrieval is the key to realizing intelligent remote sensing data interpretation.

[0004] The problem of cross-modal remote sensing image-text retrieval is a very challenging problem because there are significant modal differences between images and texts. A key issue in the design of cross-modal remote sensing image-text retrieval methods is how to eliminate the heterogeneity gap caused by modal differences. Existing deep learning-based cross-modal remote sensing image-text retrieval methods are usually trained on large-scale screened and processed clean training data. Although these methods have achieved significant improvements on benchmark datasets, they often fail to consider that the data in real applications may contain noise. Due to sensor problems or environmental changes, images may be disturbed by noise, and in actual retrieval tasks, text descriptions often have great diversity. Therefore, the quality of the image and text data obtained in actual retrieval cannot be guaranteed. During the inference process, the retrieval performance of existing methods decays severely when facing noisy test data. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the related art to some extent.

[0006] To this end, the first object of this application is to propose a cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning, which solves the technical problem that existing methods are severely affected by noise interference and performance decay, and realizes cross-modal retrieval between remote sensing images and text descriptions.

[0007] The second object of this application is to propose a computer device.

[0008] To achieve the above object, the first aspect embodiment of this application proposes a cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning, including: using a feature extractor and an expert feature extractor to extract the features of remote sensing images and text descriptions respectively; for each extracted feature, constructing a corresponding Dirichlet distribution based on its internal similarity; for the remote sensing image features and text description features extracted by the feature extractor, constructing Dirichlet distributions based on the similarity of image to text and the similarity of text to image respectively; performing expert-guided intra-modal relationship learning, trustworthy learning between modalities, and uncertainty consistency learning based on the constructed Dirichlet distributions.

[0009] To achieve the above object, the second aspect embodiment of this invention proposes a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning.

[0010] The cross-modal remote sensing image-text retrieval method and device based on expert-guided trustworthy learning in the embodiments of the present application introduce trustworthy learning to address the uncertainty in retrieval data, and propose an expert-guided trustworthy retrieval method. In the training stage, cross-modal trustworthy learning, uncertainty consistency learning, and expert-guided intra-modal relationship learning are designed to learn the sample pairing relationship and uncertainty measurement. Specifically, in the model training part, deep neural networks are used to extract high-order features of remote sensing images and text descriptions respectively, and a model pre-trained on a large-scale dataset is selected as the backbone neural network to speed up the training; for the image V i and the text data T i input into the feature extraction network, the image and text feature extractors of CLIP are selected as the backbone network to extract features, and the extracted features are denoted as v i and t i ; perform cross-modal image-text trustworthy retrieval modeling on the remote sensing image features and text description features, calculate the similarity between the image and text features and the self-similarity within each modality through inner product, and construct the corresponding Dirichlet distribution according to the obtained similarity matrix; design cross-modal trustworthy learning (EDL), uncertainty consistency learning (UC), and expert-guided intra-modal relationship learning (RL). In this embodiment, not only can the prediction result of the model be obtained according to the expectation of the Dirichlet distribution, but also the prediction uncertainty index is added, so as to balance the prediction accuracy and uncertainty.

[0011] Additional aspects and advantages of the present application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, in which:

[0013] Figure 1 is a schematic flowchart of a cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning provided in Embodiment 1 of the present application;

[0014] Figure 2 is a schematic diagram of a cross-modal remote sensing image-text retrieval model based on expert-guided trustworthy learning in the embodiments of the present application; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0016] The cross-modal remote sensing image-text retrieval method and device based on expert-guided trustworthy learning according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0017] Figure 1 It is a schematic flowchart of a cross-modal remote sensing image-text retrieval method provided by Embodiment 1 of the present application.

[0018] As Figure 1 shown, the cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning includes the following steps:

[0019] Step 101, use a feature extractor and an expert feature extractor to extract the features of the remote sensing image and the text description respectively;

[0020] In the embodiments of the present application, using a feature extractor and an expert feature extractor to extract the features of the remote sensing image and the text description respectively includes:

[0021] Input the remote sensing image into an image feature extractor and an expert image feature extractor respectively, and output the image feature and the expert image feature;

[0022] Input the text description into a text feature extractor and an expert text feature extractor respectively, and output the text feature and the expert text feature.

[0023] Step 102, for each extracted feature, construct a corresponding Dirichlet distribution based on its internal similarity;

[0024] In the embodiments of the present application, for the extracted image feature, expert image feature, text feature, and expert text feature, calculate the self-similarity through inner product, and construct a corresponding Dirichlet distribution according to the obtained similarity matrix.

[0025] Step 103, for the remote sensing image feature and text description feature extracted by the feature extractor, construct Dirichlet distributions based on the similarity of image to text and the similarity of text to image respectively;

[0026] In the embodiments of the present application, perform cross-modal image-text trustworthy retrieval modeling on the remote sensing image feature and text description feature. Calculate the similarity between the image and text features through inner product, and construct a corresponding Dirichlet distribution according to the obtained similarity matrix. Taking the trustworthy retrieval modeling from image to text as an example (the modeling process from text to image is similar), specifically includes:

[0027] Calculate the similarity of image to text through inner product:

[0028]

[0029] Among them, K represents K texts, i represents the i-th image, v iis the image feature, t i is the text feature;

[0030] Obtain the evidence of each text feature through the evidence activation function:

[0031]

[0032] Generate a Dirichlet distribution based on subjective logic and calculate the uncertainty:

[0033]

[0034] Among them,

[0035] The corresponding Dirichlet distribution is expressed as:

[0036]

[0037] Among them, B(·) represents the K-dimensional beta function, and SP K represents the K-dimensional simplex.

[0038] Step 104, perform expert-guided intra-modal relationship learning, cross-modal credible learning, and uncertainty consistency learning based on the constructed Dirichlet distribution.

[0039] In the embodiment of the present application, the training is a cross-modal and intra-modal credible retrieval learning process for images and texts within one batch, and the batch size is K.

[0040] In the embodiment of the present application, in the intra-modal relationship learning, a neural network pre-trained for unimodal similarity is introduced as an expert model. These expert models have been pre-trained on a large scale for unimodal similarity and are reliable guides for evaluating unimodal similarity. The expert model and the corresponding modal feature extractor respectively construct a Dirichlet distribution according to the internal similarity of the extracted features. The expert-guided intra-modal relationship learning includes intra-modal relationship learning for the image modality and intra-modal relationship learning for the text modality, and the processes of the two are similar.

[0041] Taking the intra-modal relationship learning of the image modality as an example, calculate the intra-modal similarity between the output of the image feature extractor and the output of the expert image model, and construct a Dirichlet distribution: and Constrain the intra-modal distribution of the image features to approach the expert-guided intra-modal distribution, which is expressed as:

[0042]

[0043] Among them, K represents K images, is the Dirichlet distribution constructed based on the self-similarity of the image features, It is a Dirichlet distribution constructed based on the self-similarity of expert image features.

[0044] In the embodiments of the present application, the cross-modal reliable learning includes image-to-text reliable learning and text-to-image reliable learning, and the processes of the two are similar. The cross-modal uncertainty consistency learning includes image-to-text uncertainty consistency learning and text-to-image uncertainty consistency learning, and the processes of the two are similar.

[0045] In the embodiments of the present application, taking the image-to-text reliable learning as an example, it includes:

[0046] Define the pairing label between samples as indicating that the j-th text sample is paired with the i-th image sample. The log-likelihood function of the i-th image sample is expressed as:

[0047]

[0048] According to the maximum likelihood estimation, the loss function within the batch is obtained:

[0049]

[0050] where

[0051] Add the Kullback–Leibler constraint to unpaired samples so that unpaired sample pairs do not generate evidence. Among them, this constraint is expressed as:

[0052]

[0053] where

[0054] Based on Construct the reliable learning loss EDL of image-text as:

[0055]

[0056] where b1 is a hyperparameter, and n e is the current training epoch.

[0057] In the embodiments of the present application, taking the image-to-text uncertainty consistency learning as an example, it includes:

[0058] Based on the reliable learning framework, the decision result and uncertainty estimation are expressed as:

[0059]

[0060] where

[0061] To ensure that the uncertainty estimated by the method can be used as a stable indicator for measuring retrieval reliability, the uncertainty should be inversely proportional to the retrieval accuracy. That is, samples with high uncertainty have low retrieval accuracy, and samples with low uncertainty have high retrieval accuracy. Assume that in a batch of size K, the c-th sample is the correct retrieved sample. The uncertainty consistency loss function UC for image-to-text is designed as follows:

[0062]

[0063] where is the indicator function.

[0064] Specifically, in the embodiments of this application, the overall loss function during learning is constructed as:

[0065] L = L EDL + αL UC + βL RL

[0066] where α and β are hyperparameters, and L EDL is the cross-modal reliable learning loss function, expressed as:

[0067]

[0068] is the reliable learning loss from text to image,

[0069] L UC is the uncertainty consistency learning loss between modalities, expressed as:

[0070]

[0071] is the uncertainty consistency loss for text-to-image,

[0072] L RL is the expert-guided intra-modal relationship learning, expressed as:

[0073]

[0074] is the intra-modal relationship learning loss for the text modality.

[0075] The cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning in the embodiments of this application introduces trustworthy learning to address the uncertainty in retrieval data, and proposes an expert-guided trustworthy retrieval method. In the training stage, inter-modal trustworthy learning, uncertainty consistency learning, and expert-guided intra-modal relationship learning are designed to learn the sample pairing relationship and uncertainty measurement. Specifically, in the model training part, deep neural networks are used to extract high-order features of remote sensing images and text descriptions respectively. A model pre-trained on a large-scale dataset is selected as the backbone neural network to accelerate the training speed; the image V i and the text data T i are input into the feature extraction network. The image and text feature extractors of CLIP are selected as the backbone network to extract features. The extracted features are denoted as v i and t i ; cross-modal image-text trustworthy retrieval modeling is performed on the remote sensing image features and text description features. The similarity between the image and text features and the self-similarity within each modality are calculated through inner product, and the corresponding Dirichlet distribution is constructed according to the obtained similarity matrix; inter-modal trustworthy learning (EDL), uncertainty consistency learning (UC), and expert-guided intra-modal relationship learning (RL) are designed. In this embodiment, not only the prediction result of the model can be obtained according to the expectation of the Dirichlet distribution, but also the prediction uncertainty index is added, so as to balance the prediction accuracy and uncertainty.

[0076] The cross-modal remote sensing image-text retrieval model based on expert-guided trustworthy learning constructed in the embodiments of this application is as Figure 2 shown. The training process of this model includes:

[0077] Step 1, first, according to the overall loss function, the AdamW optimizer is used on the training set, and appropriate step sizes and hyperparameters are selected to train the model. In this embodiment, the step size is set to 10 -6 , and the hyperparameters ɑ and β in the training stage are both set to 1. For the training set used in this embodiment, a total of 40 epochs are trained.

[0078] Step 2, after the training is completed, taking the text retrieval image task as an example, the images in the retrieval dataset are input into the corresponding feature extraction network. In this embodiment, the structure of the image and text feature extractors of CLIP is adopted as the remote sensing image and text description feature extraction network, and the final dimension of the extracted features is 512 for both. An image retrieval feature set is obtained, and the r-th feature in the image retrieval feature set is denoted as v r . The query text is input into the corresponding text feature extractor to obtain the query text feature t q .

[0079] Step 3, calculate the cosine similarity between the query text description feature and each feature in the image retrieval feature set. The calculation expression is as follows:

[0080]

[0081] Where ‖·‖ represents the 2-norm. The similarities between the query sample features and all sample features in the target modality dataset are sorted from large to small to obtain the final retrieval results. The higher the similarity, the more relevant the retrieval results.

[0082] After the training is completed, the query image samples can be input into the trained retrieval model to obtain features, calculate the similarity with the features in the target modality dataset, and sort them from high to low according to the similarity to obtain the final retrieval results.

[0083] This embodiment conducts simulation experiments on the remote sensing image and text retrieval public dataset RSICD, which contains 10921 optical remote sensing images, each of which corresponds to 5 text descriptions. The simulation experiment results are shown in Appendix 1.

[0084] Table 1 Simulation results

[0085]

[0086] In Table 1, the retrieval performance of this embodiment is compared with that of six mainstream cross-modal remote sensing image and text retrieval methods. The recall rate of different numbers of returned samples is used as an evaluation indicator to compare the remote sensing image and text retrieval accuracy of different methods. It can be seen that the cross-modal remote sensing image and text retrieval method based on expert-guided trusted learning in this embodiment significantly improves the retrieval accuracy, thereby strongly proving that the present invention has high practical value in practice.

[0087] In order to implement the above embodiments, the present invention further proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the above embodiments is implemented.

[0088] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0089] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0090] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code that includes one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0091] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.

[0092] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0093] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0094] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0095] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A cross-modal remote sensing image-text retrieval method based on expert-guided trustworthy learning, characterized in that, Including: Using a feature extractor and an expert feature extractor to extract the features of remote sensing images and text descriptions respectively; For each of the extracted features, constructing a corresponding Dirichlet distribution based on its internal similarity; For the remote sensing image features and text description features extracted by the feature extractor, constructing Dirichlet distributions based on the similarity of image to text and text to image respectively; Based on the constructed Dirichlet distributions, performing expert-guided intra-modal relationship learning and inter-modal credibility learning and uncertainty consistency learning.

2. The method according to claim 1, characterized in that, Using a feature extractor and an expert feature extractor to extract the features of remote sensing images and text descriptions respectively, including: Inputting the remote sensing images into an image feature extractor and an expert image feature extractor respectively, and outputting image features and expert image features; Inputting the text descriptions into a text feature extractor and an expert text feature extractor respectively, and outputting text features and expert text features.

3. The method according to claim 2, wherein The step of constructing a Dirichlet distribution for each of the extracted features based on its internal similarity includes: For the extracted image features, expert image features, text features and expert text features, calculating the self-similarity through inner product, and constructing the corresponding Dirichlet distribution according to the obtained similarity matrix.

4. The method according to claim 1, wherein Determining the similarity of image to text for the remote sensing image features and text description features extracted by the feature extractor, and constructing a Dirichlet distribution based on the similarity of image to text, including: Calculating the similarity of image to text through inner product: Among them, K represents K texts, u represents the u-th image, and v i is the image feature, and t i is the text feature; Obtaining the evidence of each text feature through an evidence activation function: Generating a Dirichlet distribution based on subjective logic and calculating the uncertainty: Among them, The corresponding Dirichlet distribution is expressed as: where B(·) represents the K-dimensional beta function, and SP K represents the K-dimensional simplex.

5. The method according to claim 3, characterized in that, The expert-guided intra-modal relationship learning includes intra-modal relationship learning of images and intra-modal relationship learning of text; Performing the intra-modal relationship learning of images includes: Constraining the intra-modal distribution of images to approach the intra-modal distribution of images guided by experts, expressed as: where K represents K images, is a Dirichlet distribution constructed based on the self-similarity of image features, is a Dirichlet distribution constructed based on the self-similarity of expert image features.

6. The method according to claim 4, wherein The inter-modal credibility learning includes image-to-text credibility learning and text-to-image credibility learning, and the inter-modal uncertainty consistency learning includes image-to-text uncertainty consistency learning and text-to-image uncertainty consistency learning.

7. The method according to claim 6, wherein Performing image-to-text credibility learning includes: Define the paired label between samples as indicating that the j-th text sample is paired with the i-th image sample, and the log-likelihood function of the i-th image sample is expressed as: According to the maximum likelihood estimation, obtaining the loss function within the batch: Among them, Increasing the Kullback–Leibler constraint for unpaired samples so that the unpaired sample pairs do not generate evidence, where this constraint Among them, Based on The construction of the trustworthy learning loss for image-text is as follows: Among them, b1 is a hyperparameter, and n e is the current training round.

8. The method according to claim 7, wherein Performing image-to-text uncertainty consistency learning includes: Based on the credibility learning framework, the decision result and uncertainty estimation are expressed as: Among them, Assuming that in a batch of size K, the c-th sample is the correct retrieval sample, designing an uncertainty consistency loss function as: wherein, is an indicator function.

9. The method according to any one of claims 5 and 8, characterized in that The method further includes: Constructing the overall loss function during learning as: L = L EDL + αL UC + βL RL where α and β are hyperparameters, and L EDL is the trust learning loss function between modalities, expressed as: is a trusted learning loss for text-image L UC is the uncertainty consistency learning loss between modalities, expressed as: For the uncertainty consistency loss of text-to-image, L RL For expert-guided in-modal relationship learning, denoted as: It is the loss for learning the intra-modal relationship for the text modality.

10. A computer device, characterized in that, Including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in any one of claims 1-9 is implemented.

Citation Information

Cited By

  • Image generation quality evaluation method for semantic evidence learning

    CN121937462A