Zero-shot task enhancement method and system based on Sinkhorn algorithm

By applying the Sinkhorn algorithm in the zero-sample task of the big model and calculating and subtracting the target hub value vector, the hub problem that the big model has in the zero-sample task is solved, the recall rate and diagnostic accuracy are improved, and more comprehensive data coverage and more relevant search results are achieved.

CN119474423BActive Publication Date: 2025-05-09HANGZHOU MIND MEDICAL ALLIANCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510077540.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-09
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Large models have pivot problems in zero-sample tasks, resulting in low recall rates, especially in the field of medical graphic search, where certain types of disease may be ignored, affecting diagnostic and treatment decisions.

Method used

The zero-sample task enhancement method based on the Sinkhorn algorithm is used to calculate the similarity matrix between the query set and the target set and normalize it to obtain the converged Gibbs matrix, thereby accurately estimating the target hub value vector and reducing the impact of the hub point in the similarity matrix.

Benefits of technology

Effectively alleviates the pivot problems, improves the recall rate and diagnostic accuracy of the zero-sample task, enables the search results to cover the data distribution more comprehensively, provide more relevant medical images and information, and helps discover more disease patterns or treatment options.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474423B_ABST
    Figure CN119474423B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and discloses a zero-sample task enhancement method and system based on the Sinkhorn algorithm, the method comprising: obtaining a target medical image and text data set; the target medical image and text data set comprises a query set and a target set, the query data in the query set is text data, and the target data in the target set is image data; calculating the similarity matrix between the query set and the target set; performing Sinkhorn normalization on the similarity matrix to obtain a converged Gibbs matrix; obtaining a target hub value vector according to the converged Gibbs matrix; obtaining a hub-less similarity matrix according to the similarity matrix minus the target hub value vector; obtaining the corresponding target data according to the input query data and the hub-less similarity matrix. The present invention can effectively alleviate the hub problem, and at the same time provides a method for accurately estimating the hub values ​​of all target data in the target set when the query set distribution is unknown during testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a zero-sample task enhancement method and system based on a Sinkhorn algorithm. Background Art

[0002] Zero-shot tasks refer to tasks that allow machine learning models to directly reason about unknown data without retraining. Using large models, such as large language models (LLM) or vision-language models (VLM), to directly reason about user data without fitting the large model to the user data is a typical zero-shot task. Currently, large models including LLM and VLM have a hubness problem when representing data, that is, in actual retrieval, some targets in the target set will be frequently recalled. These targets that are retrieved multiple times are called hub targets, and the appearance of these hub targets will result in a low overall recall rate for large models, a phenomenon called the hub problem. This hub problem will greatly affect the performance of large models in zero-shot tasks. For example, when using large models to perform zero-shot classification tasks on data, a lot of data will be incorrectly classified into very few classes, resulting in low accuracy of the classification task; when using large models for retrieval tasks, large models will allow individual targets to be retrieved frequently, while the probability of most targets being retrieved is 0, resulting in a low retrieval recall rate. The present invention is dedicated to solving this pivot problem and enhancing the performance of large models in zero-shot tasks, including zero-shot classification tasks and zero-shot retrieval tasks. Especially in the field of medical image and text retrieval, some disease types or features may be ignored in the retrieval because they are not pivot points, resulting in differences in retrieval performance for different disease types, and incorrect correlations may mislead doctors to make incorrect diagnosis or treatment decisions.

[0003] The current methods for solving the pivot problem can be roughly divided into two categories: uniform representation methods and balanced probability methods. The uniform representation method aims to make the representation distribution as uniform as possible in the representation space, that is, to make the variance of the representation distribution as large as possible, so as to fundamentally alleviate the pivot problem. This type of method starts with the model architecture and analyzes that the reason why the current deep learning model will have convergence and pivot problems is that the model's normalization technology, such as batch normalization and layer normalization, will reduce the variance of the representation distribution, while the nonlinear layers in the large model, such as the linear rectified unit or the softmax function, will make the data representation further converge. In order to alleviate the distribution convergence, the uniform representation method designs special normalization technology or uses less normalization technology, and designs targeted loss functions and training methods, so that the representation output by the model is as uniformly distributed as possible in the representation space during training. The uniform representation method requires changes to the model structure or parameters, so it is not suitable for zero-shot tasks. The balanced probability method takes a different approach. When testing the model, the target data is compensated for its probability of selection according to its probability of selection, so that the selection probability of all test data is as consistent as possible, thereby alleviating the hub problem. However, since the balanced probability method only considers the hub of the target side but not the query side, it can only obtain a suboptimal solution instead of an optimal solution. In addition, due to the large difference in the distribution of the pseudo query set and the target set, the single set hub estimation has a large deviation, which makes the hub problem not completely solved.

[0004] Therefore, there is an urgent need for a zero-shot task enhancement method and system based on the Sinkhorn algorithm, which can effectively alleviate the hub problem and provide a method for accurately estimating the hub value of all target data in the target set when the query set distribution is unknown during testing. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a zero-sample task enhancement method and system based on the Sinkhorn algorithm, which can effectively alleviate the hub problem. At the same time, a method for accurately estimating the hub value of all target data in the target set when the query set distribution is unknown during testing is provided.

[0006] The present invention provides a zero-sample task enhancement method based on the Sinkhorn algorithm, which specifically includes the following steps:

[0007] S1. Obtain a target medical image and text dataset; wherein the target medical image and text dataset includes a query set and a target set, the query data in the query set is text data, and the target in the target set is image data;

[0008] S2, calculate the similarity matrix between the query set and the target set;

[0009] S3, performing Sinkhorn normalization on the similarity matrix, and calculating the converged Gibbs matrix;

[0010] S4, obtaining the target pivot value vector according to the converged Gibbs matrix;

[0011] S5, subtracting the target hub value vector from the similarity matrix to obtain a hub-less similarity matrix;

[0012] S6. Obtain corresponding target data according to the input query data and the similarity matrix without hubs.

[0013] Furthermore, when the number of query data in the query set is equal to 1, in S2, calculating the similarity matrix between the query set and the target set includes:

[0014] S2a, randomly generate a number of texts as a query supplement library;

[0015] S2b, generate images corresponding to the text in the query supplement library and add them to the target set;

[0016] S2c, generate any image as the target supplement library, and satisfy that every query in the query supplement library can find the corresponding target in the target supplement library;

[0017] S2d, taking the query supplement library and the original query set as a new query set, and taking the target supplement library and the target set obtained in S2b as a new target set;

[0018] S2e, calculate the similarity matrix between the new query set and the new target set.

[0019] Furthermore, when the number of query data in the query set is equal to 1, in S4, the target hub value vector is obtained according to the converged Gibbs matrix, including:

[0020] S4a, obtaining the first target pivot value vector according to the converged Gibbs matrix;

[0021] S4b, according to the number of target data in the target set obtained in S2b, the first target pivot value vector is trimmed, and the first n data of the first target pivot value vector are retained; wherein n represents the number of target data in the target set obtained in S2b;

[0022] S4c, taking the first n data of the first target pivot value vector as the final target pivot value vector.

[0023] Furthermore, in S3, the similarity matrix is ​​Sinkhorn normalized, and the converged Gibbs matrix is ​​calculated to include:

[0024] S31, setting query side probability constraints, target side probability constraints, maximum number of iterations and entropy hyperparameters;

[0025] S32, calculating the Gibbs matrix according to the similarity matrix and the entropy hyperparameter;

[0026] The calculation formula is as follows:

[0027] K = exp(S / λ);

[0028] Among them, K represents the Gibbs matrix, S represents the similarity matrix, and λ represents the entropy hyperparameter;

[0029] S33, determine whether the current number of iterations reaches the maximum number of iterations; if so, use the current Gibbs matrix as the converged Gibbs matrix; if not, perform row normalization and column normalization on the current Gibbs matrix, and determine whether the row-normalized and column-normalized Gibbs matrix converges, if so, use the current Gibbs matrix as the converged Gibbs matrix, if not, re-execute S33.

[0030] Furthermore, in S33, the Gibbs matrix is ​​row-normalized and column-normalized, and the calculation formula is as follows:

[0031] The calculation formula for row normalization is as follows:

[0032] ;

[0033] The calculation formula for column normalization is as follows:

[0034] ;

[0035] Where b represents the target side probability constraint, K represents the Gibbs matrix, r represents the number of iterations, and β represents the target side probability in the iteration. represents the target side probability after the rth iteration, α represents the query side probability in the iteration, represents the query side probability after the rth iteration.

[0036] Furthermore, when the number of query data in the query set is equal to 1, in S5, the target hub value vector is subtracted from the similarity matrix to obtain a hub-less similarity matrix, including:

[0037] S5a, calculating a first similarity matrix according to the original query set and the target set included in the target medical image and text dataset;

[0038] S5b, subtracting the target hub value vector from the first similarity matrix to obtain a hub-less similarity matrix.

[0039] Further, in S5, the target hub value vector is subtracted from the similarity matrix to obtain the hub-less similarity matrix, where the hub-less similarity matrix is ​​expressed as:

[0040] ;

[0041] in, represents the similarity matrix without hubs, Y ij represents the similarity between the i-th query and the j-th target after removing the hub.

[0042] Furthermore, in S6, the corresponding target data obtained according to the input query data and the similarity matrix without the hub includes:

[0043] According to the input query data q i , sort the elements in the i-th row of the similarity matrix without the hub from large to small, and use the target data corresponding to the largest element as the target data corresponding to the input query data.

[0044] The present invention also provides a zero-sample task enhancement system based on the Sinkhorn algorithm, which is used to implement any of the zero-sample task enhancement methods based on the Sinkhorn algorithm described above, and the system includes:

[0045] A data acquisition module is used to acquire a target medical image and text data set; wherein the target medical image and text data set includes a query set and a target set, the query data in the query set is text data, and the target data in the target set is image data;

[0046] A similarity matrix calculation module, connected to the data acquisition module, is used to calculate the similarity matrix between the query set and the target set;

[0047] A Gibbs matrix calculation module is connected to the similarity matrix calculation module and is used to perform Sinkhorn normalization on the similarity matrix to calculate a converged Gibbs matrix;

[0048] A pivot value vector calculation module is connected to the Gibbs matrix calculation module and is used to obtain a target pivot value vector according to the converged Gibbs matrix;

[0049] A hub removal module, connected to the similarity matrix calculation module and the target hub value vector calculation module, is used to subtract the target hub value vector from the similarity matrix to obtain a hub-removed similarity matrix;

[0050] The query module is connected to the hub removal module and is used to obtain corresponding target data according to the input query data and the hub removal similarity matrix.

[0051] The embodiments of the present invention have the following technical effects:

[0052] The present invention obtains a Gibbs matrix by applying Sinkhorn normalization to a similarity matrix, and obtains an accurate target hub value vector based on the Gibbs matrix. By subtracting the target hub value vector from the similarity matrix, the influence of hub points on nearest neighbor search in high-dimensional space is reduced, so that text queries and image targets can be matched more accurately, and the hub problem can be effectively alleviated, thereby improving the accuracy of diagnosis. Alleviating the hub problem can make the retrieval results no longer overly biased towards certain frequently occurring data points, but can better cover the entire data distribution, provide more diverse and highly query-related medical images and information, and help doctors discover more possible disease patterns or treatment options; at the same time, a method is provided for accurately estimating the hub values ​​of all target data in a target set when the query set distribution is unknown during testing. When the number of query data is small, the query supplement mechanism can still effectively construct a similarity matrix, thereby supporting task execution in small sample conditions, and can also respond quickly to rare diseases or newly emerging disease types to handle zero-sample learning scenarios for rare diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0054] Figure 1 It is a schematic diagram of the zero-shot retrieval process of a large model in the prior art;

[0055] Figure 2 is a flow chart of a zero-sample task enhancement method based on the Sinkhorn algorithm provided by an embodiment of the present invention;

[0056] Figure 3 It is a logic diagram of a zero-sample task enhancement method based on the Sinkhorn algorithm provided in an embodiment of the present invention;

[0057] Figure 4 It is a structural diagram of a zero-sample task enhancement system based on the Sinkhorn algorithm provided by an embodiment of the present invention;

[0058] Figure 5 This is a first PLIP zero-sample retrieval enhanced visualization comparison schematic diagram provided by an embodiment of the present invention;

[0059] Figure 6 This is a second PLIP zero-sample retrieval enhanced visualization comparison schematic diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0061] Figure 1 This is a schematic diagram of the large model zero-shot retrieval process in the prior art, see Figure 1 , Taking the zero-shot enhancement task as an example, the multimodal large model encodes the query and all the targets in the target set into representations respectively, so as to accurately calculate the similarity between the query and the target, and then completes the retrieval by basic similarity sorting, where, Represents the query q i With target t j However, many studies have found that in actual retrieval, some targets in the target set will be frequently recalled. These targets that are retrieved multiple times are called hub targets, and the appearance of these hub targets will lead to a low overall recall rate of large models. This phenomenon is called the hub problem.

[0062] This paper proposes a zero-sample task enhancement method based on the Sinkhorn algorithm. Figure 2 is a flowchart of a zero-sample task enhancement method based on the Sinkhorn algorithm provided by an embodiment of the present invention. Figure 3 is a logic diagram of a zero-sample task enhancement method based on the Sinkhorn algorithm provided by an embodiment of the present invention, see Figure 2 and Figure 3 , specifically including the following steps:

[0063] S1. Obtain the target medical image and text dataset.

[0064] The target medical graphic data set includes a query set and a target set. The query data in the query set is text data, such as a pathology report, which is the conclusion drawn after observing tissue samples under a microscope, usually involving the determination of the nature of the tumor (benign or malignant). The target data in the target set is the image data corresponding to the query text data.

[0065] S2. Calculate the similarity matrix between the query set and the target set.

[0066] In some embodiments, when the number of query data in the query set is greater than 1, any large model can be used to directly calculate the similarity matrix based on the query set and the target set in the target medical graphic data set.

[0067] Similarity Matrix , where S represents the similarity matrix between the query set Q and the target set G, Q represents the representation of the query set, and G represents the representation of the target set.

[0068] In some embodiments, when the number of query data in the query set is equal to 1, calculating the similarity matrix between the query set and the target set includes:

[0069] S2a, randomly generate some texts as the query supplementary library Q h .

[0070] S2b, generate images corresponding to the text in the query supplementary library and add them to the target set.

[0071] S2c, generate any image as the target to supplement the library G h , and every query in the query supplement library can find the corresponding target in the target supplement library.

[0072] S2d: The query supplement library and the original query set are used as a new query set, and the target supplement library and the target set obtained in S2b are used as a new target set.

[0073] S2e, calculate the similarity matrix between the new query set and the new target set.

[0074] Among them, any large model can be used to calculate the similarity matrix based on the new query set and the new target set, wherein the similarity matrix between the new query set and the new target set includes: the similarity matrix between the query set and the target set, the similarity matrix between the query supplement library and the target set, and the similarity matrix between the query supplement library and the target supplement library.

[0075] Similarity Matrix , , ;

[0076] Among them, S represents the similarity matrix between the query set Q and the target set G, Q represents the representation of the query set, G represents the representation of the target set, and S bt Indicates the query supplementary database Q h The similarity matrix between the target set G, S bb Indicates the query supplementary database Q h Complement library G with target h The similarity matrix between .

[0077] In order to obtain the target joint distribution , and make In the similarity matrix When the difference is as small as possible, it can satisfy both the query side probability balance and the target side probability balance. The problem to be optimized can be described as:

[0078] ;

[0079] ;

[0080] in, represents the probability that the i-th query is semantically similar to the j-th target, express The information entropy of the query sample representation set , M represents the number of query data, and the target sample representation set , N represents the number of target data, q i represents the i-th query, t j represents the jth target, Represents a matrix whose elements are all 1.

[0081] S3. Perform Sinkhorn normalization on the similarity matrix and calculate the converged Gibbs matrix.

[0082] For the problem to be optimized, it can be transformed into a classic matrix scaling problem and solved using the classic Sinkhorn algorithm. When the number of query data in the query set is greater than 1, the Sinkhorn-Knopp algorithm is directly used for Sinkhorn normalization (SN); when the number of query data in the query set is equal to 1, since it is necessary to construct a query supplement library and a target supplement library, that is, to forge a query set and a forged target set, and to combine the forged dual sets for Sinkhorn normalization, the method that requires forging dual sets for Sinkhorn normalization is called DualBank Sinkhorn Normlization (DBSN).

[0083] When the number of query data in the query set is greater than 1, the Sinkhorn normalization of the similarity matrix specifically includes:

[0084] S31. Set query-side probability constraints, target-side probability constraints, maximum number of iterations, and entropy hyperparameters.

[0085] For example, the query side probability constraint a= , target side probability constraint b = , maximum number of iterations , entropy hyperparameter .

[0086] S32. Calculate the Gibbs matrix based on the similarity matrix and the entropy hyperparameter.

[0087] First, initialize the process variable β (0) = .

[0088] The calculation formula of the Gibbs matrix is ​​as follows:

[0089] ;

[0090] Among them, K represents the Gibbs matrix, S represents the similarity matrix, and λ represents the entropy hyperparameter.

[0091] S33. Determine whether the current number of iterations has reached the maximum number of iterations.

[0092] If yes, then the Gibbs matrix at this time is taken as the converged Gibbs matrix;

[0093] The converged Gibbs matrix ;

[0094] in, is the diagonalization operation.

[0095] If not, the Gibbs matrix at this time is row-normalized and column-normalized, and it is determined whether the Gibbs matrix after row-normalization and column-normalization converges. If so, the Gibbs matrix at this time is used as the converged Gibbs matrix. If not, S33 is executed again.

[0096] In some embodiments, the calculation formula for row normalization is as follows:

[0097] ;

[0098] The calculation formula for column normalization is as follows:

[0099] ;

[0100] Where b represents the target side probability constraint, K represents the Gibbs matrix, r represents the number of iterations, and β represents the target side probability in the iteration. represents the target side probability after the rth iteration, α represents the query side probability in the iteration, represents the query side probability after the rth iteration.

[0101] When the number of query data in the query set is equal to 1, the Sinkhorn normalization of the similarity matrix is ​​similar to that when the number of query data in the query set is greater than 1. ]Estimate all target pivot values[ ]; where [;] indicates matrix connection, [ ] indicates S bt With S bb The connection matrix, The pivot vector representing all target data in the estimated target complement library, which has no effect on subsequent steps.

[0102] S4. Obtain the target hub value vector according to the converged Gibbs matrix.

[0103] In some embodiments, when the number of query data in the query set is greater than 1, the target pivot value vector is obtained according to the converged Gibbs matrix, and the calculation formula is as follows:

[0104] ;

[0105] in, represents the target pivot value vector, and R represents the maximum number of iterations.

[0106] In some embodiments, when the number of query data in the query set is equal to 1, obtaining the target pivot value vector according to the converged Gibbs matrix includes:

[0107] S4a. Obtain the first target hub value vector according to the converged Gibbs matrix.

[0108] S4b. According to the number of target data in the target set obtained in S2b, the first target pivot value vector is trimmed to retain the first n data of the first target pivot value vector.

[0109] Wherein, n represents the number of target data in the target set obtained in S2b.

[0110] S4c, taking the first n data of the first target pivot value vector as the final target pivot value vector.

[0111] S5. Subtract the target hub value vector from the similarity matrix between the query set and the target set to obtain a hub-less similarity matrix.

[0112] In some embodiments, the hub-less similarity matrix ;

[0113] The similarity matrix without hubs is expressed as:

[0114] ;

[0115] in, represents the similarity matrix without hubs, Y ij represents the similarity between the i-th query and the j-th target after removing the hub.

[0116] S6. Obtain corresponding target data according to the input query data and the similarity matrix without hubs.

[0117] In some embodiments, according to the input query data q i , sort the elements in the i-th row of the similarity matrix without the hub from large to small, and use the target data corresponding to the largest element as the target data corresponding to the input query data.

[0118] For example, this embodiment takes the zero-sample retrieval capability of the pre-trained multimodal pathology model PLIP on two clinical pathology datasets, Pubmed and BookSet, as an example. By comparing the zero-sample retrieval results of PLIP before and after the SN algorithm enhancement, the significant advantage of the SN algorithm is verified.

[0119] First, let's briefly introduce two open source image-text retrieval datasets. Pubmed contains a total of 3,309 images, each with a corresponding English text description. The Bookset dataset contains 4,270,000 images, and each image also has a pathology description given by a clinician.

[0120] We use the PLIP-VIT / B-32 model to evaluate text-to-image search and image-to-text search on two test sets. PLIP is a large multimodal pathology model proposed by Stanford University. It can understand and generate the connection between descriptions and pathology texts through joint training of large-scale pathology images and pathology text descriptions. PLIP aims to improve the ability of pathology image understanding and generation models through cross-modal learning, especially the mapping between text and pathology images. Therefore, it is suitable as a basic model to evaluate zero-shot retrieval tasks.

[0121] The basic process of retrieval and the flowchart of zero-sample retrieval enhancement using SN algorithm or DBSN algorithm are as follows: Figure 3 As shown. Combined Figure 3 We first use PLIP's image encoder to obtain the encodings of all images in the Pubmed dataset as the target set, and the text encoder to encode all text descriptions as the query set. The similarity between the query set and the target set is calculated, and the retrieval of text image search is completed according to the similarity sorting. The SN method normalizes the similarity between the query set and the target set before retrieval, estimates the hub value of each target data, and then subtracts the hub value from the original similarity matrix and sorts it to obtain a more accurate sorting result. DBSN does not directly use the query set (that is, all text encodings in the Pubmed dataset) to estimate the target hub value, but borrows other query supplementary sets (such as all text encodings of the Bookset dataset) and target supplementary sets (corresponding to all image encodings of the Bookset dataset) to estimate the target hub value, thereby solving the hub problem and improving retrieval performance.

[0122] The results of the zero-shot retrieval test are shown in Table 1 below. It can be seen that DBSN and SN have significant performance improvements compared to the results before enhancement. In particular, compared to the current best method for solving the hub problem, IS (Inverted Softmax), the effect of SN is more obvious; and because DBSN does not directly use the query set estimation, the improvement is limited compared to the method using the query set estimation, but it still has significant performance improvement compared to the basic retrieval.

[0123] Table 1 PLIP zero-shot retrieval enhancement comparison results;

[0124]

[0125] Figure 5 is a first PLIP zero-sample retrieval enhanced visualization comparison schematic diagram provided by an embodiment of the present invention, Figure 6 is a second PLIP zero-shot retrieval enhancement visualization comparison diagram provided by an embodiment of the present invention, wherein Baseline represents a naive retrieval, that is, an existing retrieval method, IS represents a retrieval enhanced based on an inverted softmax method, and SN represents a retrieval enhanced based on the Sinkhorn normalization method of this solution. Figure 5 For the query data "regardless of clear, eosinophilic, epithelioid or spindle cells, clear cell sarcoma always has large eosinophilic nucleoli (AD)", the image data obtained by naive retrieval, retrieval enhanced by the inverted softmax method, and retrieval enhanced by the Sinkhorn normalization method of this scheme; Figure 6 For the query data "low-grade Mullerian gland sarcoma. The tumor has a biphasic growth of malignant mesenchymal components and benign glandular components. Polypoid intracavitary protrusions of malignant mesenchymal components and condensation of mesenchymal components around the glands can be seen", the image data obtained by naive retrieval, retrieval enhanced by the inverted softmax method, and retrieval enhanced by the Sinkhorn normalization method of this scheme, the visualization results of text-to-image search on Pubmed using CLIP are as follows Figure 5-6 As shown, it can be seen that because the hub problem is solved, semantically related images that are not retrieved by basic retrieval and the IS enhancement method can be well retrieved, thereby improving the retrieval performance.

[0126] The present invention obtains a Gibbs matrix by applying Sinkhorn normalization to a similarity matrix, and obtains an accurate target hub value vector based on the Gibbs matrix. By subtracting the target hub value vector from the similarity matrix, the influence of hub points on nearest neighbor search in high-dimensional space is reduced, so that text queries and image targets can be matched more accurately, and the hub problem can be effectively alleviated, thereby improving the accuracy of diagnosis. Alleviating the hub problem can make the retrieval results no longer overly biased towards certain frequently occurring data points, but can better cover the entire data distribution, provide more diverse and highly query-relevant medical images and information, and help doctors discover more possible disease patterns or treatment options; at the same time, a method is provided for accurately estimating the hub values ​​of all target data in a target set when the query set distribution is unknown during testing. When the number of query data is small, the query supplement mechanism can still effectively construct a similarity matrix, thereby supporting task execution in small sample conditions, and can also respond quickly to rare diseases or newly emerging disease types to handle zero-sample learning scenarios for rare diseases.

[0127] The embodiment of the present invention further provides a zero-sample task enhancement system based on the Sinkhorn algorithm, which is used to implement the zero-sample task enhancement method based on the Sinkhorn algorithm described in the above embodiment. Figure 4 is a structural diagram of a zero-sample task enhancement system based on the Sinkhorn algorithm provided by an embodiment of the present invention, see Figure 4 , the system includes:

[0128] A data acquisition module is used to acquire a target medical image and text data set; wherein the target medical image and text data set includes a query set and a target set, the query in the query set is text data, and the target in the target set is image data;

[0129] A similarity matrix calculation module, connected to the data acquisition module, is used to calculate the similarity matrix between the query set and the target set;

[0130] A Gibbs matrix calculation module is connected to the similarity matrix calculation module and is used to perform Sinkhorn normalization on the similarity matrix to calculate a converged Gibbs matrix;

[0131] A pivot value vector calculation module is connected to the Gibbs matrix calculation module and is used to obtain a target pivot value vector according to the converged Gibbs matrix;

[0132] A hub removal module, connected to the similarity matrix calculation module and the target hub value vector calculation module, is used to obtain a hub-removed similarity matrix by subtracting the target hub value vector from the similarity matrix between the query set and the target set;

[0133] The query module is connected to the hub removal module and is used to obtain the corresponding target according to the target query and the hub removal similarity matrix.

[0134] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0135] It should also be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", etc. should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be a connection between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A zero-shot task enhancement method based on the Sinkhorn algorithm, characterized in that: The specific steps include: S1. Obtain a target medical image-text dataset; wherein the target medical image-text dataset includes a query set and a target set, the query data in the query set is text data, and the target data in the target set is image data; S2, calculate the similarity matrix between the query set and the target set; When the number of query data in the query set is equal to 1, it is also necessary to generate a number of texts to construct a query supplement library and images corresponding to the texts and add them to the target set; S3, performing Sinkhorn normalization on the similarity matrix to calculate a converged Gibbs matrix; S4, obtaining a target hub value vector according to the converged Gibbs matrix; Specifically include: When the number of query data in the query set is greater than 1, the target pivot value vector is obtained according to the converged Gibbs matrix. The calculation formula is as follows: ; in, represents the target pivot value vector, R represents the maximum number of iterations, and λ represents the entropy hyperparameter; When the number of query data in the query set is equal to 1, the following sub-steps are included: S4a, obtaining a first target hub value vector according to the converged Gibbs matrix; S4b, according to the number of target data in the target set after supplementation obtained in S2, the first target pivot value vector is trimmed, and the first n data of the first target pivot value vector are retained; wherein n represents the number of target data in the target set after supplementation obtained in S2; S4c, taking the first n data of the first target pivot value vector as the final target pivot value vector; S5, subtracting the target hub value vector from the similarity matrix between the query set and the target set to obtain a hub-less similarity matrix; S6. Obtain corresponding target data according to the input query data and the hub-less similarity matrix.

2. The zero-shot task enhancement method based on the Sinkhorn algorithm according to claim 1, characterized in that: When the number of query data in the query set is equal to 1, in S2, calculating the similarity matrix between the query set and the target set includes: S2a, randomly generate a number of texts as a query supplement library; S2b, generating an image corresponding to the text in the query supplement library and adding it to the target set; S2c, generate any image as the target supplement library, and satisfy that every query in the query supplement library can find the corresponding target in the target supplement library; S2d, taking the query supplement library and the original query set as a new query set, and taking the target supplement library and the target set obtained in S2b as a new target set; S2e, calculating the similarity matrix between the new query set and the new target set; wherein the similarity matrix between the new query set and the new target set includes: the similarity matrix between the query set and the target set, the similarity matrix between the query supplement library and the target set, and the similarity matrix between the query supplement library and the target supplement library.

3. The zero-shot task enhancement method based on the Sinkhorn algorithm according to claim 1, characterized in that: In S3, the similarity matrix is ​​subjected to Sinkhorn normalization, and the converged Gibbs matrix is ​​calculated to include: S31, setting query side probability constraints, target side probability constraints, maximum number of iterations and entropy hyperparameters; S32, calculating the Gibbs matrix according to the similarity matrix and the entropy hyperparameter; The calculation formula is as follows: K = exp(S / λ); Among them, K represents the Gibbs matrix, S represents the similarity matrix, and λ represents the entropy hyperparameter; S33, determine whether the current number of iterations reaches the maximum number of iterations; if so, use the current Gibbs matrix as the converged Gibbs matrix; if not, perform row normalization and column normalization on the current Gibbs matrix, and determine whether the row-normalized and column-normalized Gibbs matrix converges, if so, use the current Gibbs matrix as the converged Gibbs matrix, if not, re-execute S33.

4. The zero-shot task enhancement method based on the Sinkhorn algorithm according to claim 3, characterized in that: In S33, the Gibbs matrix is ​​row normalized and column normalized, and the calculation formula is as follows: The calculation formula for row normalization is as follows: ; The calculation formula for column normalization is as follows: ; Where b represents the target side probability constraint, K represents the Gibbs matrix, r represents the number of iterations, and β represents the target side probability in the iteration. represents the target side probability after the rth iteration, α represents the query side probability in the iteration, represents the query side probability after the rth iteration.

5. The zero-shot task enhancement method based on the Sinkhorn algorithm according to claim 1, characterized in that: In S5, the target hub value vector is subtracted from the similarity matrix between the query set and the target set to obtain a hub-less similarity matrix, wherein the hub-less similarity matrix is ​​expressed as: ; in, represents the similarity matrix without hubs, Y ij represents the similarity between the i-th query and the j-th target after removing the hub.

6. The zero-shot task enhancement method based on the Sinkhorn algorithm according to claim 5, characterized in that: In S6, obtaining corresponding target data according to the input query data and the hub-less similarity matrix includes: According to the query data q i , sort the elements in the i-th row of the hub-less similarity matrix from large to small, and use the target data corresponding to the largest-ranked element as the target data corresponding to the input query data.

7. A zero-sample task enhancement system based on the Sinkhorn algorithm, used to implement the zero-sample task enhancement method based on the Sinkhorn algorithm as described in any one of claims 1 to 6, characterized in that: The system includes: A data acquisition module, used to acquire a target medical image and text data set; wherein the target medical image and text data set includes a query set and a target set, the query data in the query set is text data, and the target data in the target set is image data; A similarity matrix calculation module is connected to the data acquisition module and is used to calculate the similarity matrix between the query set and the target set; when the number of query data in the query set is equal to 1, it is also necessary to generate a number of texts to construct a query supplement library and images corresponding to the texts and add them to the target set; A Gibbs matrix calculation module, connected to the similarity matrix calculation module, for performing Sinkhorn normalization on the similarity matrix to calculate a converged Gibbs matrix; A pivot value vector calculation module is connected to the Gibbs matrix calculation module and is used to obtain a target pivot value vector according to the converged Gibbs matrix; specifically comprising: When the number of query data in the query set is greater than 1, the target pivot value vector is obtained according to the converged Gibbs matrix. The calculation formula is as follows: ; in, represents the target pivot value vector, R represents the maximum number of iterations, and λ represents the entropy hyperparameter; When the number of query data in the query set is equal to 1, it includes: Obtaining a first target pivot value vector according to the converged Gibbs matrix; According to the number of target data in the target set after supplementation obtained in the similarity matrix calculation module, the first target pivot value vector is trimmed to retain the first n data of the first target pivot value vector; wherein n represents the number of target data in the target set after supplementation obtained in the similarity matrix calculation module; The first n data of the first target pivot value vector are used as the final target pivot value vector; A hub removal module, connected to the similarity matrix calculation module and the target hub value vector calculation module, for subtracting the target hub value vector from the similarity matrix between the query set and the target set to obtain a hub-free similarity matrix; The query module is connected to the hub removal module and is used to obtain corresponding target data according to the input query data and the hub removal similarity matrix.

Citation Information

Patent Citations

  • Training model and small sample classification method and device

    CN112598091A

  • Tea disease and pest classification method, system and equipment based on small sample learning and medium

    CN116129196A