Text-guided ultrasonic image prototype learning system and method based on partial optimal transmission

By constructing the prototype space of images and text and using some optimal transmission modules for alignment, the problems of representation differences and high computational complexity in multimodal data alignment are solved, efficient image and text matching and transformation are achieved, and the generalization ability and alignment accuracy of the model are improved.

CN120217003APending Publication Date: 2025-06-27NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216584.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art has problems such as representation differences, unsatisfactory alignment and high computational complexity in multimodal data alignment, which limits its effect in practical applications.

Method used

By constructing the prototype space of images and text, and using some optimal transmission modules to achieve probability distribution alignment of the two, eliminating semantic differences between different modal data, and achieving effective matching and conversion of images and text.

Benefits of technology

It significantly improves the generalization ability of the model and multi-scene adaptability, improves the semantic alignment accuracy between images and text, reduces computational overhead, and provides a new solution for multimodal data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217003A_ABST
    Figure CN120217003A_ABST
Patent Text Reader

Abstract

The invention discloses a text-guided ultrasonic image prototype learning system and method based on partial optimal transmission, and the system comprises a visual prototype module which is used for extracting features from image data and constructing an image prototype space; the semantic prototype module is used for extracting features from the text data and constructing a text prototype space; and the partial optimal transmission module is used for calculating a partial optimal transmission scheme between the image prototype space and the text prototype space so as to realize alignment of probability distribution of the image prototype space and the text prototype space. According to the system, the probability distribution of a text prototype is adjusted to be approximate to the probability distribution of an image prototype through an optimal transmission module, so that effective matching and conversion between an image and a text are realized. The method has the beneficial effects that the generalization ability of the model in different modal data processing tasks is improved, visual features can be migrated to a text processing task, or semantic knowledge is applied to image recognition and understanding tasks, and the performance of the model in various application scenes is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing, natural language processing, and optimal transport theory, and belongs to the fields of computer science and artificial intelligence technology. Specifically, it relates to a text-guided ultrasound image prototype learning system and method based on partial optimal transport. Background Art

[0002] With the development of artificial intelligence and big data technologies, multi-modal data processing has gradually become a research hotspot. Multi-modal data includes various forms such as text, images, audio, etc., and has rich semantic information and a wide range of application scenarios. However, due to the heterogeneity of data in different modalities, such as the spatial features in images and the temporal semantic features in text, there are huge representational differences between them. How to achieve efficient alignment and conversion of multi-modal data has always been an important research problem.

[0003] In the prior art, specific models designed separately are usually used to process single-modal data. For example, convolutional neural networks (CNNs) are used to process image data, or recurrent neural networks (RNNs) and Transformer models are used to process text data. Although these methods perform excellently in single-modal tasks, they are often limited in multi-modal tasks. Existing multi-modal data alignment methods usually rely on directly matching image and text features, without fully considering the distribution differences of different modal features, resulting in unsatisfactory alignment effects. In addition, traditional alignment methods have a high computational complexity when dealing with large-scale data, making it difficult to meet the actual application requirements.

[0004] Optimal Transport (OT), as a mathematical framework, can achieve efficient alignment of different modal data by minimizing the transport cost between probability distributions. However, traditional optimal transport methods often face problems of low computational efficiency and excessive resource consumption when dealing with high-dimensional complex data, which limits their application in actual multi-modal tasks.

[0005] To solve the above problems, the present invention proposes a text-guided ultrasound image prototype learning system and method based on partial optimal transport. By separately constructing the prototype spaces of images and text, and using a partial optimal transport module to achieve the alignment of their probability distributions, this method can efficiently eliminate the semantic differences between different modal data, thereby achieving effective matching and conversion of images and text. This method not only improves the generalization ability of the model but also significantly reduces the computational overhead, providing a new solution for multi-modal data processing. Summary of the Invention

[0006] The present invention provides a text-guided ultrasound image prototype learning system and method based on partial optimal transport to solve the problems of difficult multi-modal data alignment, high computational complexity, and insufficient model generalization ability in the prior art.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A text-guided ultrasound image prototype learning system based on partial optimal transport, comprising:

[0009] An image input module 100, configured to receive and preprocess image data;

[0010] A text generation module 110, configured to generate text data;

[0011] A visual prototype module 120, configured to extract features from the image data and construct an image prototype space;

[0012] A semantic prototype module 130, configured to extract features from the text data and construct a text prototype space;

[0013] A partial optimal transport module 140, configured to calculate a partial optimal transport scheme between the image prototype space and the text prototype space and align the probability distributions of the two;

[0014] A semantic alignment module 150, configured to store and output the aligned feature distribution;

[0015] A cross-modal transfer application module 160, configured to perform feature transfer and application in multi-modal tasks.

[0016] Further, the image input module 100 is configured to receive a pathological image input and transfer it to the visual prototype module for processing;

[0017] The text generation module 110 is configured to generate a text description according to the input prompt statement and transfer it to the semantic prototype module.

[0018] Further, the visual prototype module 120 converts the input pathological image into an image token through VQGAN;

[0019] The semantic prototype module 130 converts the generated text description into a text token by using BERT.

[0020] Further, the semantic prototype module 130 is configured to guide GPT to generate a pathology-related feature description through a prompt statement, and use the BERT model to tokenize and embed the generated text to generate corresponding text tokens, and the text tokens represent the semantic information in the image description;

[0021] The GPT-generated pathological feature descriptions include the different manifestations of benign and malignant pathological ultrasound images, and the text descriptions generated by GPT include detailed descriptions of pathological features.

[0022] Further, the partial optimal transport module 140 semantically aligns image tokens and text tokens using CLIP by calculating the optimal transport relationship between image tokens and text tokens.

[0023] The partial optimal transport module includes:

[0024] An optimization algorithm for calculating the partial optimal transport relationship between image tokens and text tokens;

[0025] A calculation module for executing the CLIP model and outputting the similarity between image tokens and text tokens.

[0026] A text-guided ultrasound image prototype learning method based on partial optimal transport completes the alignment and migration of multimodal data through the following steps:

[0027] Step 1, Input data:

[0028] Receive image data and text data as input: Preprocess the input image, usually including denoising, normalization, and cropping steps; the text data is generated by GPT through prompts and then passed to the text prototype learning module for processing.

[0029] Step 2, Image feature extraction and prototype construction and text feature extraction and prototype construction:

[0030] Among them: Image feature extraction and prototype construction include the following steps:

[0031] Step 211, Input pathological ultrasound image:

[0032] The input pathological ultrasound image is denoted as I, with a size of H×W×C,

[0033] where: H is the image height; W is the image width; C is the number of image channels;

[0034] Step 212, VQGAN encoder:

[0035] The input pathological ultrasound image is processed by the VQGAN encoder to convert it into discrete image tokens; the encoder of VQGAN maps the image to a high-dimensional feature space and represents these features as discrete vectors in a quantized manner, and each vector represents the local features of the image.

[0036] Step 213, Encoder operation:

[0037] The input image I undergoes several layers of convolution, non-linear activation, and pooling operations to obtain a feature representation F enc ;

[0038] F enc = CNN(I)

[0039] where CNN represents the convolution operation, and F enc is the feature representation of the image;

[0040] The feature F enc is mapped to a discrete codeword space, and through quantization operations, the image features are discretized into an image token z i :

[0041] z i = Quantize(F enc )

[0042] where z i is the quantized image token and belongs to a vector in a predefined discrete dictionary;

[0043] Step 214, Generation of image tokens:

[0044] After being processed by the VQGAN encoder, the input image I is mapped to multiple image tokens, denoted as: z1, z2,..., z n , and each image token represents a certain local feature of the image; in this process, the image is decomposed into a discrete prototype space, and these image tokens capture the global and local information of the image to a certain extent;

[0045] Step 215, Generation of the image prototype space:

[0046] The quantized image tokens constitute the image prototype space, and each image token is a discrete vector z i :

[0047] Z = {z1, z2,..., z n}

[0048] where Z is the set of image tokens and represents the discrete representation of the input pathological image I in the prototype space;

[0049] Among them: Text feature extraction and prototype construction include the following steps:

[0050] Step 221, Input of pathological description prompt statements:

[0051] Input a set of predefined pathological description prompt statements, aiming to guide the generation of pathological feature descriptions;

[0052] Step 222, Text Generation: Generate a pathological description through GPT:

[0053] Input the pathological description prompt statement into the GPT model, and GPT generates a detailed pathological feature description based on the input prompt;

[0054] Step 223, Text to Token:

[0055] Encode through the BERT model, and the generated text description will be input into the BERT model; the BERT model will perform word segmentation on the text and convert each word segment into a text token, and these tokens will be the low-dimensional representation of the pathological feature description and can be processed by subsequent modules;

[0056] The process of the BERT model in Step 223:

[0057] Accept the input text, perform word segmentation on the text, and then convert the segmented text into the corresponding token vector representation. Each word or sub-word will be mapped to a vector T of a fixed dimension i :

[0058] T i = BERTEmbedding(T i )

[0059] where: T i is the i-th token, and T i is the word embedding representation of this token;

[0060] Finally, the text prototype learning module outputs a set of text tokens T, and these tokens are the encoding results of BERT for the pathological feature description and will be used for semantic alignment with the image tokens;

[0061] Step 3: Partial Optimal Transport, including the following steps:

[0062] Step 31, Input Image Tokens and Text Tokens

[0063] Image Tokens: From the visual prototype module, representing a set of discrete feature vectors of the pathological ultrasound image

[0064]

[0065] where N is the number of image tokens, is the first image token, is the second image token, is the i-th image token, is the Nth image token, representing the local features of the image;

[0066] Text token: from the text prototype learning module, representing a set of discrete feature vectors of the generated pathological feature description

[0067]

[0068] where M is the number of text tokens, the 1st text token, is the 2nd text token, is the jth text token, the Mth text token, representing the semantics of the text description;

[0069] Step 32, calculate the distance between the image token and the text token:

[0070] Evaluate their matching degree by calculating the semantic distance between the image token and the text token. For this purpose, use the CLIP model to map the image token and the text token into a shared embedding space, so as to calculate the similarity between them;

[0071] The CLIP model: The CLIP model processes images and texts simultaneously, maps images and texts into a common semantic space, and given an image token and a text token The CLIP model calculates the similarity between them:

[0072]

[0073] where, and are the embedding vectors of the image token and the text token respectively, · represents the dot product of vectors, and ‖·‖ is the norm of vectors;

[0074] Step 33, partial optimal transport calculation:

[0075] Since the number of image tokens and text tokens is usually not equal, partial optimal transport is needed to find the optimal alignment. The goal of partial optimal transport is to match image tokens with text tokens by minimizing the transport cost;

[0076] Cost matrix construction: Calculate a cost matrix C according to the similarity between the image token and the text token, where: each element C ijDenote the image tokens and the text tokens The transmission cost between them is usually constructed using the similarity between the image and the text as follows:

[0077]

[0078] The cost matrix here reflects the matching degree between the image tokens and the text tokens;

[0079] Partial optimal transport algorithm: Use the partial optimal transport algorithm to optimize the matching between the image tokens and the text tokens; The goal of the optimization process is to solve an alignment matrix P, where: P ij denotes the matching weights of the image tokens and the text tokens ;

[0080]

[0081] where P ij denotes the matching probabilities of the image tokens and the text tokens , and satisfies the normalization constraint;

[0082]

[0083] By optimizing this goal, the best matching relationship between the image tokens and the text tokens is obtained;

[0084] Step 34. Semantic alignment:

[0085] Based on the alignment matrix P optimized by partial optimal transport, align the image tokens and the text tokens according to this matrix, so as to achieve their semantic matching, that is, the matching degree between the image tokens and the text tokens is determined by the weights in the matrix P;

[0086] For each pair of matching image tokens and text tokens, the model further analyzes their semantic relationships;

[0087] Step 35. Output the aligned image tokens and text tokens;

[0088] Finally, the partial optimal transport module will output the image tokens and text tokens after semantic alignment:

[0089]

[0090] where, T image_aligned and T text_alignedrespectively represent the aligned visual prototype space and semantic prototype space; and respectively represent the aligned image tokens and text tokens;

[0091] Step 4: Feature Transfer and Application

[0092] The aligned feature distribution output by the semantic alignment module 150 is input into the cross-modal transfer application module 160 for the feature transfer task: for the input pathological image, its pathological type and pathological features can be obtained.

[0093] Compared with the prior art, the present invention has the following beneficial effects:

[0094] The present invention provides a system and method capable of efficiently aligning multi-modal data distributions and improving cross-modal task performance, significantly enhancing the generalization ability and multi-scenario adaptability of the model.

[0095] The present invention is used to improve the semantic alignment accuracy between images and texts, especially in the field of medical image analysis, and solves the problem that existing images and texts cannot be effectively aligned. Brief Description of the Drawings

[0096] Figure 1 is a schematic structural diagram of the present invention;

[0097] Figure 2 is a schematic structural diagram of the image input module in the present invention;

[0098] Figure 3 is a schematic structural diagram of the text generation module in the present invention;

[0099] Figure 4 is a schematic structural diagram of the visual prototype module in the present invention;

[0100] Figure 5 is a schematic structural diagram of the semantic prototype module in the present invention;

[0101] Figure 6 is a schematic structural diagram of a partial optimal transport module in the present invention;

[0102] Figure 7 is a schematic structural diagram of the cross-modal transfer application module in the present invention;

[0103] Wherein: 100 - image input module, 110 - text generation module, 120 - visual prototype module, 130 - semantic prototype module, 140 - partial optimal transport module, 150 - semantic alignment module, 160 - cross-modal transfer application module. Detailed Embodiments

[0104] The present invention will be further described in conjunction with the embodiments.

[0105] Embodiment 1

[0106] As Figure 1-7 shown, a text-guided ultrasound image prototype learning system based on partial optimal transport includes:

[0107] An image input module 100 for receiving and preprocessing image data, where the image data is a pathological image; specifically, the image input module 100 is used to receive the input of the pathological image and transfer it to the visual prototype module for processing;

[0108] A text generation module 110 for generating text data; specifically, the text generation module 110 is used to generate a text description according to the input prompt statement and transfer it to the semantic prototype module, and the text description is a pathological feature description;

[0109] A visual prototype module 120 for extracting features from the image data and constructing an image prototype space; specifically, the visual prototype module 120 converts the input pathological image into an image token through VQGAN;

[0110] A semantic prototype module 130 for extracting features from the text data and constructing a text prototype space; specifically, the semantic prototype module 130 uses BERT to convert the generated text description into a text token; the semantic prototype module 130 is used to guide GPT to generate pathological-related feature descriptions through prompt statements, and use the BERT model to tokenize and embed the generated text to generate corresponding text tokens, and the text tokens represent the semantic information in the image description; the pathological-related feature descriptions generated by GPT include the different manifestations of benign and malignant pathological ultrasound images, and the text descriptions generated by GPT include the detailed descriptions of pathological features.

[0111] A partial optimal transport module 140 for calculating a partial optimal transport scheme between the image prototype space and the text prototype space and realizing the alignment of their probability distributions;

[0112] Specifically, the partial optimal transport module 140 calculates the optimal transport relationship between the image token and the text token, and uses CLIP to perform semantic alignment on the image token and the text token; the partial optimal transport module includes: an optimization algorithm for calculating the partial optimal transport relationship between the image token and the text token; a calculation module for executing the CLIP model and outputting the similarity between the image token and the text token.

[0113] The semantic alignment module 150 is used to store and output the aligned feature distribution;

[0114] The cross-modal transfer application module 160 is used to perform feature transfer and application in multi-modal tasks.

[0115] In the present invention, the image input module can receive different types of medical images and transfer them to the visual prototype module for processing; the text generation module can automatically generate a detailed pathological description according to the prompt statements provided by doctors or medical experts, convert it into text tokens through the BERT model, and finally achieve efficient alignment of image tokens and text tokens through the semantic alignment module.

[0116] The above-mentioned trained text-guided prototype learning system based on partial optimal transport can calculate the distance between a new image input and the trained prototype to judge the pathological condition of the new image.

[0117] The visual prototype module is used to convert the input image (pathological ultrasound image) into image tokens through VQGAN. These image tokens represent the feature information of the image in the discrete vector space; specifically: the visual prototype module processes the input pathological ultrasound image through the encoder of VQGAN to obtain a set of discrete image tokens. These image tokens represent the features of the pathological image in a discrete vector space; the VQGAN model of the visual prototype module generates a generator and a discriminator to convert the input pathological image into discrete image tokens. These tokens contain the high-level semantic information of the image; the visual prototype module extracts features from the input pathological ultrasound image, maps the image to a discrete vector space using the encoder of VQGAN, and generates image tokens. These image tokens can effectively capture the key features in the image and are then used for subsequent learning and reasoning; the visual prototype module extracts features from the input pathological ultrasound image, maps the image to a discrete vector space using the encoder of VQGAN, and generates image tokens. These image tokens can effectively capture the key features in the image and are then used for subsequent learning and reasoning;

[0118] The semantic prototype module is used to guide GPT to generate pathology-related feature descriptions through prompt statements, and use the BERT model to tokenize and embed the generated text to generate corresponding text tokens, where the text tokens represent the semantic information in the image description; the text generated by the text generation module through the GPT model includes the specific features of the pathology image, and the text tokens represent the potential semantic information of the image; the text prototype learning module guides GPT to generate natural language descriptions about pathology features through given prompt statements; the text description can elaborate on the possible pathology features in the image and convert it into text tokens based on the BERT model.

[0119] The partial optimal transport module calculates the optimal transport relationship between image tokens and text tokens, and uses CLIP to perform semantic alignment on image tokens and text tokens; specifically: the partial optimal transport module uses the optimal transport algorithm to solve the problem of inconsistent numbers of image tokens and text tokens, and realizes semantic alignment between image tokens and text tokens by calculating their similarity. The CLIP model is used to calculate the similarity between image tokens and text tokens in the shared embedding space, thereby realizing semantic alignment between images and text. The partial optimal transport module solves the problem of inconsistent numbers of image tokens and text tokens by calculating the "transport cost" between image tokens and text tokens, and uses the CLIP model for semantic alignment, so that there is higher semantic consistency between images and text in the shared embedding space, thereby improving the accuracy and efficiency of image-text matching.

[0120] The feature description includes the different manifestations of benign and malignant pathology ultrasound images, and the text description generated by GPT includes a detailed description of the pathology features.

[0121] Example 2

[0122] A text-guided ultrasound image prototype learning method based on partial optimal transport completes the alignment and migration of multimodal data through the following steps:

[0123] Step 1, Input data:

[0124] Receive image data and text data as input: Preprocess the input image, usually including denoising, normalization, and cropping steps; the text data is generated by GPT through prompts and then passed to the text prototype learning module for processing;

[0125] Step 2, Image feature extraction and prototype construction and text feature extraction and prototype construction:

[0126] Among them: Image feature extraction and prototype construction include the following steps:

[0127] Step 211: Input the pathological ultrasound image:

[0128] The input pathological ultrasound image is denoted as I, with a size of H×W×C,

[0129] where: H is the image height; W is the image width; C is the number of channels of the image (for example, an RGB image has 3 channels);

[0130] Step 212: VQGAN encoder:

[0131] The input pathological ultrasound image is processed by the VQGAN encoder and converted into discrete image tokens; the encoder of VQGAN maps the image to a high-dimensional feature space and uses quantization to represent these features as discrete vectors (codebook), and each vector represents the local features of the image;

[0132] Step 213: Encoder operation:

[0133] The input image I undergoes several layers of convolution, non-linear activation, and pooling operations to obtain a feature representation F enc ;

[0134] F enc = CNN(I)

[0135] where CNN represents the convolution operation, and F enc is the feature representation of the image;

[0136] Map the feature F enc to a discrete codeword space, and through quantization operations, discretize the image features into an image token z i :

[0137] z i = Quantize(F enc )

[0138] where z i is the quantized image token and belongs to a certain vector in the predefined discrete dictionary (codebook);

[0139] Step 214: Generation of image tokens:

[0140] After being processed by the VQGAN encoder, the input image I is mapped to multiple image tokens, denoted as: z1, z2,..., z n, each image token represents a certain local feature of the image; in this process, the image is decomposed into a discrete prototype space, and these image tokens capture the global and local information of the image to a certain extent;

[0141] Step 215, generate the image prototype space:

[0142] The quantized image tokens constitute the image prototype space, and each image token is a discrete vector z i :

[0143] Z = {z1, z2,..., z n}

[0144] where Z is the set of image tokens, representing the discrete representation of the input pathological image I in the prototype space;

[0145] Among them: text feature extraction and prototype construction include the following steps:

[0146] Step 221, input the pathological description prompt statement:

[0147] Input a set of predefined pathological description prompt statements, aiming to guide the generation of pathological feature descriptions;

[0148] Step 222, text generation: generate the pathological description through GPT:

[0149] Input the pathological description prompt statement into the GPT model, and GPT generates a detailed pathological feature description based on the input prompt;

[0150] Step 223, convert text to tokens:

[0151] Encode through the BERT model, and the generated text description (such as the text generated by GPT) will be input into the BERT model; the BERT model will perform word segmentation on the text and convert each word segment into a text token, and these tokens will be the low-dimensional representation of the pathological feature description and can be processed by subsequent modules;

[0152] The process of the BERT model in step 223:

[0153] Accept the input text, perform word segmentation on the text, and then convert the segmented text into the corresponding token vector representation. Each word or sub-word will be mapped to a vector T of a fixed dimension i :

[0154] T i = BERTEmbedding(T i )

[0155] where: Ti is the i-th token, T i is the word embedding representation of this token;

[0156] Finally, the text prototype learning module outputs a set of text tokens T, which are the encoding results of the BERT for the pathological feature description and will be used for semantic alignment with the image tokens;

[0157] Step 3: Partial optimal transport, including the following steps:

[0158] Step 31. Input the image tokens and text tokens

[0159] Image tokens: from the visual prototype module, representing a set of discrete feature vectors of the pathological ultrasound image

[0160]

[0161] where N is the number of image tokens, is the first image token, is the second image token, is the i-th image token, is the N-th image token, representing the local features of the image;

[0162] Text tokens: from the text prototype learning module, representing a set of discrete feature vectors of the generated pathological feature description

[0163]

[0164] where M is the number of text tokens, the first text token, is the second text token, is the j-th text token, the M-th text token, representing the semantics of the text description;

[0165] Step 32. Calculate the distance between the image tokens and the text tokens:

[0166] Evaluate their matching degree by calculating the semantic distance between the image tokens and the text tokens. For this purpose, the CLIP model can be used to map the image tokens and text tokens into a shared embedding space, thereby calculating the similarity between them;

[0167] The CLIP model: The CLIP model processes images and texts simultaneously, maps images and texts into a common semantic space, and given an image token and a text token the CLIP model calculates the similarity between them:

[0168]

[0169] where, and are the embedding vectors of the image token and the text token respectively, · represents the dot product of vectors, and ‖·‖ is the norm of the vector;

[0170] Step 33, Partial optimal transport calculation:

[0171] Since the number of image tokens and text tokens is usually not equal, partial optimal transport (Optimal Transport, OT) is needed to find the optimal alignment. The goal of partial optimal transport is to match image tokens with text tokens by minimizing the transport cost;

[0172] Cost matrix construction: Calculate a cost matrix C according to the similarity between image tokens and text tokens, where: each element C ij represents the transport cost between the image token and the text token and is usually constructed using the similarity between the image and the text as follows:

[0173]

[0174] Here, the cost matrix reflects the matching degree between image tokens and text tokens;

[0175] Partial optimal transport algorithm: Use the partial optimal transport algorithm to optimize the matching between image tokens and text tokens; this algorithm adjusts the pairing relationship between image tokens and text tokens through iterative optimization to minimize the overall transport cost of the matching; the goal of the optimization process is to solve an alignment matrix P, where: P ij represents the matching weight between the image token and the text token ;

[0176]

[0177] where, P ij represents the image token and the text token The matching probability, and satisfy the normalization constraint;

[0178]

[0179] By optimizing this objective, the best matching relationship between image tokens and text tokens is obtained;

[0180] Step 34. Semantic alignment:

[0181] Through the alignment matrix P optimized by partial optimal transport, the image tokens and text tokens are aligned according to this matrix, so as to realize their semantic matching. Specifically, the matching degree between image tokens and text tokens is determined by the weights in the matrix P;

[0182] For each pair of matching image tokens and text tokens, the model can further analyze their semantic relationships, such as associating image features and corresponding pathological description features;

[0183] Step 35. Output the aligned image tokens and text tokens;

[0184] Finally, the partial optimal transport module outputs the image tokens and text tokens after semantic alignment:

[0185]

[0186] Among them, T image_aligned and T text_aligned respectively represent the aligned visual prototype space and semantic prototype space; and respectively represent the aligned image tokens and text tokens;

[0187] Step 4: Feature transfer and application

[0188] The aligned feature distribution output by the semantic alignment module 150 is input into the cross-modal transfer application module 160 for the feature transfer task: for the input pathological image, its pathological type and pathological features can be obtained.

[0189] In the invention, the visual prototype module extracts local features of an image through a convolutional neural network. Through a quantization operation, the image features are discretized into an image token. The semantic prototype module obtains a text description of pathological features through GPT using a prompt statement. A text token is obtained using a BERT model. The partial optimal transport module 140 calculates a transport scheme between the image and text prototypes based on a cost matrix. The image and text are respectively mapped to a common semantic space using CLIP. The cross-modal transfer application module 160 uses the aligned features for cross-modal generation and classification tasks.

[0190] Through the above steps and modules, the image and text data can be efficiently aligned, supporting the alignment and transfer of multi-modal data in a shared semantic space.

[0191] The advantages of the present invention include:

[0192] 1. Efficient alignment: The matching accuracy of different modal data is significantly improved through the partial optimal transport algorithm.

[0193] 2. Flexible transfer: Supports the mutual transfer of visual and semantic features in multi-modal tasks.

[0194] 3. Wide application: Applicable to various scenarios such as multi-modal retrieval, generation, and classification.

[0195] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A text-guided ultrasound image prototype learning system based on partial optimal transmission, characterized in that: include: An image input module (100), used for receiving and preprocessing image data; A text generation module (110), used for generating text data; A visual prototype module (120), used for extracting features from image data and constructing an image prototype space; A semantic prototype module (130), used for extracting features from text data and constructing a text prototype space; A partial optimal transmission module (140), used for calculating a partial optimal transmission scheme between the image prototype space and the text prototype space, and realizing the alignment of the probability distributions of the two; A semantic alignment module (150), used for storing and outputting the aligned feature distribution; The cross-modal migration application module (160) is used to perform feature migration and application in multi-modal tasks.

2. The text-guided ultrasound image prototype learning system based on partial optimal transmission according to claim 1, characterized in that: The image input module (100) is used to receive pathological image input and transmit it to the visual prototype module for processing; The text generation module (110) is used to generate a text description according to the input prompt sentence and pass it to the semantic prototype module.

3. The text-guided ultrasound image prototype learning system based on partial optimal transmission according to claim 1, characterized in that: The visual prototype module (120) converts the input pathological image into an image token through VQGAN; The semantic prototype module (130) converts the generated text description into text tokens using BERT.

4. The text-guided ultrasound image prototype learning system based on partial optimal transmission according to claim 3, characterized in that: The semantic prototype module (130) is used to guide GPT to generate feature descriptions related to the pathology through prompt sentences, and use the BERT model to segment and embed the generated text to generate corresponding text tokens, wherein the text tokens represent semantic information in the image description; The GPT generates a pathology-related feature description including different manifestations of benign and malignant pathology ultrasound images, and the text description generated by the GPT includes a detailed description of the pathology feature.

5. The text-guided ultrasound image prototype learning system based on partial optimal transmission according to claim 1, characterized in that: The partial optimal transmission module (140) calculates the optimal transmission relationship between the image token and the text token and uses CLIP to perform semantic alignment on the image token and the text token; The partial optimal transmission module includes: An optimization algorithm for calculating the partial optimal transfer relationship between image tokens and text tokens; A computational module for executing the CLIP model and outputting the similarity between image tokens and text tokens.

6. A learning method for a text-guided ultrasound image prototype learning system based on partial optimal transmission as claimed in any one of claims 1 to 5, characterized in that: The alignment and migration of multimodal data is completed through the following steps: Step 1. Input data: Receive image data and text data as input: preprocess the input image, usually including denoising, normalization, and cropping steps; text data is generated by GPT through prompts and then passed to the text prototype learning module for processing; Step 2: Image feature extraction and prototype construction and text feature extraction and prototype construction: Among them: Image feature extraction and prototype construction include the following steps: Step 211: Input pathological ultrasound image: The input pathological ultrasound image is denoted as I, with a size of H×W×C. Where: H is the image height; W is the image width; C is the number of channels of the image; Step 212, VQGAN encoder: The input pathological ultrasound image is processed by the VQGAN encoder and converted into discrete image tokens. The VQGAN encoder maps the image to a high-dimensional feature space and uses a quantized method to represent the features as discrete vectors, each of which represents a local feature of the image. Step 213, encoder operation: The input image I is subjected to several layers of convolution, nonlinear activation and pooling operations to obtain a feature representation F enc ; F enc =CNN(I) Among them, CNN represents the convolution operation, F enc It is the feature representation of the image; The feature F enc Mapped to a discrete codeword space, the image features are discretized into an image token z through quantization operation i : z i =Quantize(F enc ) Among them, z i It is a quantized image token, belonging to a vector in a predefined discrete dictionary; Step 214: Generation of image token: After being processed by the VQGAN encoder, the input image I is mapped into multiple image tokens, denoted as: z1,z2,...,z n , each image token represents a local feature of the image; the image is decomposed into discrete prototype spaces, and the image token captures the global and local information of the image; Step 215: Generate image prototype space: The quantized image tokens constitute the image prototype space, and each image token is a discrete vector z i : Z={z1,z2,...,z n } Among them, Z is a set of image tokens, representing the discrete representation of the input pathological image I in the prototype space; Among them: text feature extraction and prototype construction include the following steps: Step 221, input the pathology description prompt sentence: Input a set of predefined pathology description prompt sentences, aiming to guide the generation of pathology feature descriptions; Step 222, text generation: Generate pathological description through GPT: The pathology description prompt sentence is input into the GPT model, and GPT generates a detailed description of the pathology features based on the input prompt; Step 223, convert text into token: The generated text description is encoded by the BERT model and then input into the BERT model. The BERT model will segment the text and convert each segment into a text token. The token will be a low-dimensional representation of the pathological feature description for subsequent modules to process. The process of the BERT model in step 223 is as follows: Accept the input text, segment the text, and then convert the segmented text into the corresponding token vector representation. Each word or subword will be mapped to a vector T of fixed dimension. i : T i =BERTEmbedding(T i ) Where: T i is the i-th token, T i It is the word embedding representation of token; Finally, the text prototype learning module outputs a set of text tokens T, which are the encoding results of BERT for the description of pathological features and will be used for semantic alignment with image tokens. Step 3: Partial optimal transmission, including the following steps: Step 31: Input image token and text token Image token: from the visual prototype module, representing a set of discrete feature vectors of pathological ultrasound images Where N is the number of image tokens, is the first image token, is the second image token, is the i-th image token, is the Nth image token, Represents the local features of the image; Text token: from the text prototype learning module, representing a set of discrete feature vectors describing the generated pathological features Where M is the number of text tokens, The first text token, is the second text token, is the jth text token, The Mth text token, Represents the semantics of the text description; Step 32: Calculate the distance between the image token and the text token: The matching degree is evaluated by calculating the semantic distance between the image token and the text token. To this end, the CLIP model is used to map the image token and the text token into a shared embedding space to calculate the similarity. The CLIP model: The CLIP model processes images and texts simultaneously, mapping images and texts to a common semantic space. and text token CLIP model calculates similarity: in, and are the embedding vectors of image token and text token respectively, · represents the dot product of the vector, ‖·‖ is the norm of the vector; Step 33, partial optimal transmission calculation: Since the number of image tokens and text tokens is usually not equal, partial optimal transfer is needed to find the optimal alignment. The goal of partial optimal transfer is to match image tokens and text tokens by minimizing the transfer cost. Cost matrix construction: Calculate a cost matrix C based on the similarity between image tokens and text tokens, where: each element C ij Represents an image token With text token The transmission cost between images and text is usually calculated using the similarity between images and text. To build: The cost matrix reflects the matching degree between image tokens and text tokens; Partially optimal transfer algorithm: Use the partially optimal transfer algorithm to optimize the matching between image tokens and text tokens; the goal of the optimization process is to solve an alignment matrix P, where: P ij Represents an image token and text token The matching weight of Among them, P ij Represents an image token and text token The matching probability of and satisfies the normalization constraint; By optimizing the target, the best matching relationship between image token and text token is obtained; Step 34: Semantic alignment: The alignment matrix P obtained by partial optimal transfer optimization is used to align image tokens and text tokens to achieve semantic matching. That is, the matching degree between image tokens and text tokens is determined by the weights in the matrix P. For each pair of matching image tokens and text tokens, the model further analyzes the semantic relationship; Step 35. Output the aligned image token and text token; Finally, some of the best transfer modules will output semantically aligned image tokens and text tokens: Among them, T image_aligned and T text_aligned Respectively represent the aligned visual prototype space and semantic prototype space; and Represent the aligned image token and text token respectively; Step 4: Feature Migration and Application The aligned feature distribution output by the semantic alignment module (150) is input to the cross-modal migration application module (160) for the feature migration task: the pathological type and pathological features can be obtained for the input pathological image.