Multi-modal medical image cross retrieval method and system based on comparative learning
Through the multimodal medical image cross-retrieval method based on contrast learning, medical images and text information of different modalities are converted into a single feature vector, solving the efficiency and accuracy of multimodal medical image retrieval, and achieving more accurate diagnosis and reducing the number of examinations.
Patent Information
- Application Number
- CN202510319358.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-22
AI Technical Summary
The existing multimodal medical image retrieval technology is difficult to efficiently and accurately match and retrieve medical images of different modalities, resulting in insufficient information during the clinical diagnosis process, increasing the number of patient examinations and medical costs.
Using a multimodal medical image cross-retrieval method based on contrast learning, the medical image and text information of different modalities are converted into a single feature vector through medical image encoder, image sequence feature aggregator, text encoder and multimodal contrast learning module, and aligned in the feature space to realize multimodal cross-retrieval from text to image and image to image.
It realizes efficient alignment of medical images and text information in different modalities, reduces the number of patient examinations, reduces medical costs, and provides more accurate diagnostic references.
Smart Images

Figure CN120353958A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly to a cross-modal medical image retrieval method and system. Background Art
[0002] In the current clinical diagnosis process, obtaining and analyzing medical images of different modalities (such as computed tomography (CT), magnetic resonance imaging (MRI), positron emission tomography (PET), whole slide images (WSI), etc.) has become an important technical means. Images acquired by different imaging techniques have their own characteristics and advantages. Medical images of different modalities can provide diagnostic information in different aspects. For example, MRI can provide information such as tissue structure, lesion location and size, CT can provide information such as blood vessels, bones and soft tissues, PET can provide information such as metabolic activity, and pathological images can directly reflect changes at the tissue and cell levels. Doctors can obtain more comprehensive and accurate diagnostic information by comprehensively analyzing the images acquired by different imaging techniques, and make more accurate diagnostic and treatment decisions.
[0003] In this process, the cross-modal retrieval technology of multi-modal medical images helps doctors find historical patients who are similar to the current patient in terms of image manifestations during clinical diagnosis, and include them in the diagnostic reference for the current patient. Especially when the examination items completed by the current patient are incomplete, the more complete examination items and the final diagnosis results completed by the historical patients can be included in the reference, so as to provide more comprehensive and accurate case data for doctors. This can not only help doctors comprehensively understand the disease situation from different perspectives and angles, assist doctors in making a diagnosis and formulating a treatment plan more quickly and accurately, but also help reduce the number of examinations received by patients and lower medical costs.
[0004] Medical image retrieval technology originated in the 1980s. The initial methods mainly retrieved based on text information and keywords. With the continuous development of computer technology and image processing algorithms, medical image retrieval technology has gradually developed from single text-based retrieval to content-based image retrieval. Traditional image retrieval methods usually rely on feature extraction and matching. However, due to the complexity, diversity and difference of medical images of different modalities themselves, as well as the dimension and scale of data, how to efficiently and accurately retrieve and match these images remains a challenging problem. Therefore, image retrieval methods based on multi-modal features have been proposed and received extensive attention. Among them, contrastive learning is an effective multi-modal feature learning method, which can map the image features of multiple modalities into a common feature space and retain the semantic information of different modality images.
[0005] In response to the need for analyzing multiple different modalities of medical images in the current clinical diagnosis process, this patent proposes a multi-modal medical image cross-retrieval method and system based on contrastive learning. The method can extract features from medical images of different modalities, convert different modalities of medical images such as single-scan images like DR, sequence images like CT and MRI, and whole large-sized pathological images into corresponding single feature vectors, and achieve alignment between features of different modalities of medical images and between medical images and natural language texts. On this basis, the image extraction method can achieve multi-modal cross-retrieval between text and image and between image and image, that is, it can retrieve the most similar medical image according to the input text information, and can also retrieve the medical image with the highest similarity, more similar content, and greater reference value within the same image modality and between different image modalities for a given medical image of a certain modality. The system can help doctors make more accurate diagnosis and treatment decisions, and contribute to reducing the number of examinations received by patients and lowering medical costs. Summary of the Invention
[0006] To this end, the present invention first proposes a multi-modal medical image cross-retrieval method based on contrastive learning. The composition of this method includes: a medical image encoder, an image sequence feature aggregator, a text encoder, a multi-modal contrastive learning module, and a multi-modal cross-retrieval database. Among them, the medical image encoder is used to extract features from single medical images of different modalities such as DR examinations, single CT images, MRI images, PET images, pathological image slices, etc., to obtain corresponding high-dimensional feature vectors; the image sequence feature aggregator is used to integrate the features of all single image features in the entire image sequence such as CT and MRI to obtain the entire image sequence features, or to integrate the features of all small image blocks in the whole pathological image to obtain the whole pathological image features; the text encoder is used to encode the description text information, diagnostic text information, or text information related to retrieval corresponding to the medical image to obtain corresponding text feature vectors; the multi-modal contrastive learning module performs contrastive learning among the multi-modal feature vectors output by the image encoder, the image sequence feature aggregator, and the text encoder, so as to realize the training of each module. Finally, through the optimized each module, medical image features and text features that are cross-modally aligned in the feature space can be obtained; the multi-modal cross-retrieval database is constructed based on the single image features, sequence image features, pathological image features, and text features, and is used to store the feature vectors of all medical images and text information for subsequent image and text retrieval operations.
[0007] The specific implementation steps of the multi-modal medical image cross-retrieval method based on contrastive learning are as follows:
[0008] 1) For medical images of different modalities, use the medical image encoder to extract features respectively to obtain corresponding high-dimensional feature vectors; for a medical image sequence containing multiple images, use the image sequence feature aggregator to aggregate the feature vectors of each image in the sequence obtained by the medical image encoder to obtain a feature vector representing the entire image sequence:
[0009] For a single scanned image such as a DR examination, use the medical image encoder to convert it into a high-dimensional feature vector representing the image. Taking single medical images of different sizes and resolutions as inputs, output feature vectors with the same dimension after encoding;
[0010] For a medical image sequence composed of multiple images and multiple sequences such as CT, MRI, and PET-CT, first use the medical image encoder to extract features from each image, encode the entire image sequence into a series of feature vectors with equal lengths, and then these feature vectors are encoded into an encoded vector representing the entire medical image sequence through the image sequence feature aggregator;
[0011] For a pathological image of huge size, first cut the pathological image into a large number of small image patches of the same size, use the medical image encoder to encode these image patches into feature vectors with equal lengths respectively, and finally use the image sequence feature aggregator to encode the feature vectors of all small image patches into a feature vector representing the entire pathological image;
[0012] 2) For image description text, diagnostic text, or text information related to retrieval, use the text encoder to encode it and convert it into a corresponding text feature vector;
[0013] 3) For any input medical image or text information, according to the methods in steps 1-4 above, use the medical image encoder, image sequence feature aggregator, and text encoder to process the input data to obtain corresponding feature vectors, and store them in the multi-modal cross-retrieval database;
[0014] 4) During the training process, based on the multi-modal contrast learning module, by aligning the medical image feature vectors and text feature vectors of different modalities in the multi-modal cross-retrieval database, train the image encoder, the text encoder, and the multi-modal contrast learning module respectively to obtain optimized versions of each module;
[0015] 5) During the retrieval process, the method can use any input medical image or text information as a query condition, extract features from the input medical image or text information through the medical image encoder, image sequence feature aggregator, and text encoder, and then, through the multi-modal contrast learning module, retrieve the most similar medical imaging data or text information from the multi-modal cross-retrieval database in the ways specified by the user, such as text-to-image, image-to-text, and one-modal image-to-the-same-or-another-modal image.
[0016] Based on a pre-trained cross-modal model and according to application requirements, the medical image encoder and text encoder can, for a single image, be based on a single cross-modal pre-trained model or multiple different cross-modal pre-trained models. The medical image encoder and text encoder can respectively include a single or multiple different cross-modal pre-trained models. The image encoder and text encoder in the selected cross-modal pre-trained model are respectively used to initialize the medical image encoder and text encoder to achieve feature extraction and encoding of different modalities and different types of medical images.
[0017] According to the types of medical images and application scenarios to be covered, the configuration, initialization method, and data encoding process of the medical image encoder and text encoder include:
[0018] 1) In the case of using only one cross-modal pre-trained model, the medical image encoder and text encoder are respectively directly configured and initialized using the image and text encoders of the pre-trained model. The image and text inputs will be processed using the medical image encoder and text encoder respectively, and the resulting feature vectors will be output.
[0019] 2) In the case of simultaneously selecting multiple cross-modal pre-trained models, all the selected pre-trained models will be used for the configuration and initialization of the medical image encoder and text encoder. The medical image encoder and text encoder will simultaneously include multiple different single encoders. In this case, it is necessary to specify the types of medical images applicable to each pre-trained model, and process the corresponding input data according to the specified applicable medical image types. For medical images and their corresponding text inputs, first, according to their types, select the pre-trained models specified by the user, and use the selected multiple pre-trained models to encode the input data respectively. Then, average the image features and text features output by all models to output the averaged image features and text feature vectors.
[0020] The image sequence feature aggregator converts a medical image sequence composed of multiple images into a corresponding feature vector based on a multi-instance learning framework and an attention mechanism. The image sequence feature aggregator takes the single-image feature vectors output by the medical image encoder as input, applies the self-attention mechanism to weight-aggregate all the input image feature vectors, and obtains the feature vector of the entire image sequence.
[0021] The multi-modal contrastive learning module includes two training processes: image contrastive learning between medical images and image-text contrastive learning between medical images and texts. Among them, the image contrastive learning includes three specific processes: intra-modal medical image contrastive learning, sequential image contrastive learning, and cross-modal medical image contrastive learning.
[0022] In the intra-modal medical image contrastive learning, the positive sample pairs are defined as multiple data samples obtained by subjecting the same picture to different feature transformations, and the negative sample pairs are defined as data samples from different pictures. The main training process of the contrastive learning is to reduce the distance between the positive sample pairs in the feature space and increase the distance between the negative sample pairs in the feature space.
[0023] For single-scan images such as DR examinations, the input of the intra-modal medical image contrastive learning is the image feature vector obtained through the medical image encoder; for image sequences or image matrices such as CT, MRI, and pathological images, the input of the intra-modal medical image contrastive learning is the image feature vector obtained by encoding each image through the medical image encoder and aggregating it through the image sequence feature aggregator.
[0024] In the training process of the intra-modal medical image contrastive learning, training will be carried out independently on each modality of medical images in sequence. For a selected modality of medical images, all the images of this modality are represented as {I1, I2, I3, …, I N}, where N represents the total number of images in this modality. For any one image I i After two random image transformations Two different transformed images are obtained and After that, through the image encoding process f, two image feature vectors are obtained and Randomly select one of them as the aligned target object q, then the other feature vector becomes the positive sample k of q + , and all the encoded vectors from other images become the negative sample k of q - , then for each selected alignment target q, the loss function is:
[0025]
[0026] Among them, k i represents all the encoded feature vectors, τ is a hyperparameter representing the temperature coefficient. For each image in {I1, I2, I3, …, I N}, two encoded vectors will be obtained through the above-mentioned random image transformation and feature encoding process. Therefore, the total number of encoded vectors k i is 2N. For the entire cross-modal medical image contrast learning process, its loss function is the sum of L q corresponding to all possible choices of q:
[0027] The operation object of the sequential image contrast learning in the image contrast learning is a medical image sequence or image matrix composed of multiple images, such as CT, MRI, pathological images, etc. In the sequential image contrast learning, for a given image sequence, its positive sample is defined as any single image in the image sequence, and its negative sample is defined as the images in other image sequences.
[0028] In the process of the sequential image contrast learning, each time an image sequence is selected as the target, all single images in the image sequence are feature-encoded through the medical image encoder to obtain an image feature sequence {f1, f2, …, f N}, and these single-image features are aggregated through the image sequence feature aggregator to obtain a single feature vector q representing the entire image sequence. For the target feature vector q, all the image feature vectors f i ∈{f1, f2, …, f N} are marked as positive samples k + , and the single image feature vectors from all other image sequences are marked as negative samples k - . The training process and loss function calculation method of the sequential image contrast learning are the same as those of the cross-modal medical image contrast learning described in claim 7, and its loss function is expressed as
[0029] In the cross-modal medical image contrast learning in the image contrast learning, alignment operations for contrast learning are performed on a patient examination basis. In the cross-modal medical image contrast learning, a positive sample pair is defined as a data sample of medical images of different modalities from the same patient, and a negative sample pair is defined as a data sample from different patients. During the process of the cross-modal medical image contrast learning, each time all the examination results of different modalities performed on the same patient during the same examination are selected, such as MRI, CT, pathological images, etc. The cross-modal medical image contrast learning process first completes the feature extraction of the target patient's examination images according to the image feature extraction steps described in claim 4. After that, all the image examination results of different modalities in this examination are mutually marked as positive sample pairs, and the images in other examinations from other patients are marked as negative sample pairs. Training is carried out according to the training process and loss function calculation method described in claim 7, and its loss function is expressed as
[0030] The operation object of the image-text contrast learning is medical images and corresponding text information. The image-text contrast learning takes the image and text features obtained by the medical image and text encoding steps as inputs, and through a cross-modal contrast learning training method including but not limited to CLIP, aligns the input medical image features with their corresponding text features, and expresses its loss function as L t2i 。
[0031] The total loss of the multi-modal contrast learning process is defined as L total =L img +L t2i ,where L img is the loss function of the same-modal medical image contrast learning, and L t2i is the loss function of the image-text contrast learning. During the training process, by minimizing L total , the training of the multi-modal contrast learning module is realized, so as to obtain medical image features and text features that retain key diagnostic information and are cross-modally aligned in the feature space.
[0032] The multi-modal medical image cross-retrieval method based on contrast learning supports the following different retrieval modes:
[0033] 1) Text-to-Image: Input text information as a query condition to retrieve the medical image in the multi-modal cross-retrieval database that is most similar to the text information. The retrieval process includes: First, use the text encoder to encode the input text information to obtain the corresponding text feature vector; then, through the multi-modal contrast learning module, retrieve the medical image feature vector in the multi-modal cross-retrieval database that is most similar to the text feature vector, and finally return the medical image corresponding to the medical image feature vector.
[0034] 2) Image of a Certain Type to Image of the Same Type: Input a medical image of a certain type as a query condition to retrieve other medical images of the same type in the multi-modal cross-retrieval database that are most similar to the query image;
[0035] 3) Image of a Certain Type to Image of a Different Type: Input a medical image of a certain type as a query condition to retrieve other medical images of different types in the multi-modal cross-retrieval database that are most similar to the query image.
[0036] All of the above retrieval processes use the above-mentioned medical image encoder, image sequence feature aggregator, and text encoder and their processing flows to extract and encode the features of the input query conditions, and compare them with the specified data in the multi-modal cross-retrieval database according to the cosine similarity of the feature vectors, and return one or more comparison results with the highest similarity, so as to realize the retrieval of medical images and text information.
[0037] Applying the multi-modal medical image cross-retrieval method based on contrast learning, it is characterized by including the following modules:
[0038] 1) Patient Data Management Module: This module acquires and manages the medical imaging data and related clinical information of patients, docks with the hospital's information system, obtains the basic information of patients, medical imaging examination records, and diagnostic result information from it, and saves them in the database of the system. The different-modal medical images stored in this module will be feature-extracted through the input encoding module described later, and the extracted encoded feature vectors will also be stored in the database for the image retrieval operation of Module 3 described later;
[0039] 2) Input Encoding Module: This module is based on the medical image encoder, image sequence feature aggregator, and text encoder, and is used to extract feature vectors according to different input data types;
[0040] 3) Image Retrieval Module: This module is used to realize cross-modal medical image cross-retrieval. This module uses the query condition feature vector encoded by the input encoding module, and according to the requirements, retrieves the target medical imaging data from the database according to the above retrieval mode and retrieval process;
[0041] 4) Retrieval operation and display module: This module consists of two parts: a retrieval operation module and a result display module. The retrieval operation module receives instructions from the user of the image retrieval system and retrieves the most similar historical images of the same or different modalities in the multi-modal cross-retrieval database according to the target image or target patient selected by the user, or retrieves the historical patients most similar to the clinical image manifestations of the target patient. The result display module realizes the display function of the retrieval results and provides image browsing and analysis tools, as well as a retrieval result clustering analysis tool to help doctors carry out diagnosis based on multi-modal medical images.
[0042] The technical effects to be achieved by the present invention are as follows: 1. The multi-modal medical image feature alignment method based on contrast learning proposed by the present invention can achieve
[0043] feature extraction for three different types of medical images, namely single images such as DR, sequence images such as CT and MRI, and pathological images, covering
[0044] multiple different modalities of medical images such as DR, CT, MRI, PET, and pathological images, and making similar images within the same modality, as well as images belonging to the same patient or having similar described organs or lesion sites between different modalities, have more similar feature representations, and making the medical images and their corresponding text information have more similar feature representations, so as to achieve feature alignment coding between different medical image modalities such as DR, CT, and MRI, and between images and text descriptions, providing a basis for subsequent
[0045] medical image cross-retrieval;
[0046] 2. The multi-modal medical image cross-retrieval system proposed by the present invention can retrieve relevant medical images within the same modality and different modalities through a certain modality of medical
[0047] image, and retrieve relevant medical images through text description information, so as to achieve missing image modality completion and reference patient retrieval operations. The retrieved medical images and the diagnostic results of historical patients can provide diagnostic references for doctors, thereby reducing the workload of doctors and helping to reduce the number of examinations received by patients. Brief Description of the Drawings
[0048] Figure 1 Multi-modal medical image cross-retrieval method architecture based on contrast learning; Detailed Embodiments
[0049] The following are the preferred embodiments of the present invention in combination with the accompanying drawings to further describe the technical solutions of the present invention, but the present invention is not limited to this embodiment.
[0050] The present invention proposes a multi-modal medical image feature alignment and retrieval method based on contrastive learning. Image features are extracted from medical images of different modalities including CT, MRI, PET, pathological images, etc., and text features are extracted from the descriptive texts corresponding to the medical images, and feature alignment is achieved between medical images of different modalities and between the images and their corresponding descriptive texts, so that the image examination results and text descriptions of patients with similar images and similar clinical diagnoses have closer distances in the feature space, thereby enabling cross-modal medical image cross-retrieval through similarity measurement between features.
[0051] The multi-modal medical image cross-retrieval method based on contrastive learning supports the following different retrieval modes:
[0052] 1) Text-to-image: Input text information as a query condition to retrieve the medical image in the multi-modal cross-retrieval database that is most similar to the text information. The retrieval process includes: First, use the text encoder to encode the input text information to obtain the corresponding text feature vector; then, through the multi-modal contrastive learning module, retrieve the medical image feature vector in the multi-modal cross-retrieval database that is most similar to the text feature vector, and finally return the medical image corresponding to the medical image feature vector.
[0053] 2) One type of image to the same type of image: Input a medical image of a certain type as a query condition to retrieve other medical images of the same type in the multi-modal cross-retrieval database that are most similar to the query image;
[0054] 3) One type of image to different types of images: Input a medical image of a certain type as a query condition to retrieve other medical images of different types in the multi-modal cross-retrieval database that are most similar to the query image.
[0055] All the above retrieval processes use the above-mentioned medical image encoder, image sequence feature aggregator, and text encoder and their processing flows to extract and encode the features of the input query conditions, and compare them with the specified data in the multi-modal cross-retrieval database according to the cosine similarity of the feature vectors, and return one or more comparison results with the highest similarity, thereby realizing the retrieval of medical images and text information.
[0056] 1. Method composition and implementation steps
[0057] The specific composition of the multi-modal medical image cross-retrieval method based on contrastive learning proposed by the present invention is as Figure 1As shown in the figure, it includes: a medical image encoder, an image sequence feature aggregator, a text encoder, a multi-modal contrast learning module, and a multi-modal cross-retrieval database. Among them, the medical image encoder is used to extract features from single medical images of different modalities such as DR examinations, single CT images, MRI images, PET images, and pathological image slices to obtain corresponding high-dimensional feature vectors; the image sequence feature aggregator is used to integrate the features of all single image features in the entire image sequence such as CT and MRI to obtain the entire image sequence feature, or to integrate the features of all small image patches in the entire pathological image to obtain the entire pathological image feature; the text encoder is used to encode the description text information, diagnosis text information, or text information related to retrieval corresponding to the medical image to obtain a corresponding text feature vector; the multi-modal contrast learning module performs contrast learning among the multi-modal feature vectors output by the image encoder, the image sequence feature aggregator, and the text encoder, so as to realize the training of each module. Finally, through each optimized module, medical image features and text features that are cross-modally aligned in the feature space can be obtained; the multi-modal cross-retrieval database is constructed based on the single image features, sequence image features, pathological image features, and text features, and is used to store the feature vectors of all medical images and text information for subsequent image and text retrieval operations.
[0058] The specific implementation steps of the multi-modal medical image cross-retrieval method are as follows:
[0059] 1) For medical images of different modalities, use the medical image encoder to extract features respectively to obtain corresponding high-dimensional feature vectors; for a medical image sequence containing multiple images, use the image sequence feature aggregator to aggregate the feature vectors of each image in the sequence obtained by the medical image encoder to obtain a feature vector representing the entire image sequence:
[0060] For single scanned images such as DR examinations, use the medical image encoder to convert them into high-dimensional feature vectors representing the images. Taking single medical images of different sizes and resolutions as inputs, output feature vectors with the same dimension after encoding;
[0061] For medical image sequences composed of multiple images and multiple sequences such as CT, MRI, and PET-CT, first use the medical image encoder to extract features from each image, encode the entire image sequence into a series of feature vectors with equal length, and then these feature vectors are encoded into an encoded vector representing the entire medical image sequence through the image sequence feature aggregator;
[0062] For extremely large pathological images, first cut the pathological image into a large number of small image patches of the same size, use the medical image encoder to encode these image patches into feature vectors of equal length respectively, and finally use the image sequence feature aggregator to encode the feature vectors of all small image patches into a feature vector representing the entire pathological image;
[0063] 2) For image description text, diagnostic text, or text information related to retrieval, use the text encoder for encoding to convert it into a corresponding text feature vector;
[0064] 3) For any input medical image or text information, according to the methods in steps 1-4 above, use the medical image encoder, image sequence feature aggregator, and text encoder to process the input data to obtain the corresponding feature vector, and store it in the multimodal cross-retrieval database;
[0065] 4) During the training process, based on the multimodal contrast learning module, by aligning the medical image feature vectors and text feature vectors of different modalities in the multimodal cross-retrieval database, train the image encoder, the text encoder, and the multimodal contrast learning module to obtain optimized versions of each module respectively;
[0066] 5) During the retrieval process, the method can use any input medical image or text information as a query condition, extract features from the input medical image or text information through the medical image encoder, image sequence feature aggregator, and text encoder, and then through the multimodal contrast learning module, retrieve the medical imaging data or text information most similar to the query image or text information from the multimodal cross-retrieval database in the ways specified by the user, such as from text to image, from image to text, from an image of one modality to an image of the same or another modality, etc.
[0067] 2. Medical Image Encoder and Text Encoder
[0068] The medical image encoder and text encoder are based on a pre-trained cross-modal model, and according to application requirements, can implement feature extraction and encoding of a single image based on a single cross-modal pre-trained model or multiple different cross-modal pre-trained models. The medical image encoder and text encoder can respectively include a single or multiple different cross-modal pre-trained models. The image encoder and text encoder in the selected cross-modal pre-trained model are respectively used to initialize the medical image encoder and text encoder to achieve feature extraction and encoding of different modalities and different types of medical images.
[0069] According to the types of medical images and application scenarios to be covered, the configuration, initialization method, and data encoding process of the medical image encoder and text encoder include:
[0070] 1) In the case of using only one cross-modal pre-training model, the medical image encoder and the text encoder are respectively configured and initialized directly using the image and text encoders of the pre-training model. The image and text inputs will be processed using the medical image encoder and the text encoder respectively, and the obtained feature vectors will be output.
[0071] 2) In the case of simultaneously selecting multiple cross-modal pre-training models, all the selected pre-training models will be used for the configuration and initialization of the medical image encoder and the text encoder. The medical image encoder and the text encoder will simultaneously contain multiple different single encoders. In this case, it is necessary to specify the type of medical image applicable to each pre-training model, and process the corresponding input data according to the specified applicable medical image type. For the medical image and its corresponding text input, first select the pre-training model specified by the user according to its type, and use the selected multiple pre-training models to encode the input data respectively. Then, average the image features and text features output by all models to output the averaged image feature and text feature vectors.
[0072] 3. Image sequence feature aggregator
[0073] The image sequence feature aggregator converts a medical image sequence composed of multiple images into a corresponding feature vector based on the multi-instance learning framework and the attention mechanism. The image sequence feature aggregator takes the single-image feature vectors output by the medical image encoder as input, and applies the self-attention mechanism to weight-aggregate all the input image feature vectors to obtain the feature vector of the entire image sequence.
[0074] 4. Multi-modal contrastive learning module and contrastive learning training process
[0075] The multi-modal contrastive learning module includes two training processes: image contrastive learning performed between medical images, and image-text contrastive learning performed between medical images and texts. Among them, the image contrastive learning includes three specific processes: intra-modal medical image contrastive learning, sequential image contrastive learning, and cross-modal medical image contrastive learning.
[0076] In the intra-modal medical image contrastive learning, the positive sample pairs are defined as multiple data samples obtained by subjecting the same picture to different feature transformations, and the negative sample pairs are defined as data samples from different pictures. The main training process of the contrastive learning is to reduce the distance between the positive sample pairs in the feature space and increase the distance between the negative sample pairs in the feature space.
[0077] For single-scan images such as DR examinations, the input for the same-modal medical image contrast learning is the image feature vector obtained through the medical image encoder; for image sequences or image matrices such as CT, MRI, and pathological images, the input for the same-modal medical image contrast learning is the image feature vector obtained by encoding each image therein through the medical image encoder and aggregating through the image sequence feature aggregator.
[0078] During the training process of the same-modal medical image contrast learning, training will be performed independently on each modality of medical images in sequence. For a selected modality of medical images, all images of this modality are represented as {I1, I2, I3, …, I N}, where N represents the total number of images in this modality. For any one image I i in it, after two random image transformations respectively, two different transformed images and are obtained, and then through the image encoding process f, two image feature vectors and are obtained. Randomly select one of them as the aligned target object q, then the other feature vector becomes the positive sample k of q + , and all the encoding vectors from other images become the negative sample k of q - . Then, for each selected alignment target q, the loss function is:
[0079]
[0080] where, k i represents all the encoding feature vectors, τ is a hyperparameter representing the temperature coefficient. For each image in {I1, I2, I3, …, I N}, two encoding vectors will be obtained through the above random image transformation and feature encoding process. Therefore, the total number of encoding vectors k i is 2N. For the entire same-modal medical image contrast learning process, its loss function is the sum of L q corresponding to all possible selections of q:
[0081] The operation object of the sequence image contrast learning in the image contrast learning is a medical image sequence or image matrix composed of multiple images, such as CT, MRI, pathological images, etc. In the sequence image contrast learning, for a given image sequence, its positive sample is defined as any single image in this image sequence, and its negative sample is defined as the images in other image sequences.
[0082] In the process of the sequence image contrast learning, each time an image sequence is selected as the target, all single images in the image sequence are feature-encoded by the medical image encoder to obtain an image feature sequence {f1, f2, …, f N}, and these single-image features are aggregated by the image sequence feature aggregator to obtain a single feature vector q representing the entire image sequence. For the target feature vector q, all image feature vectors f i ∈ {f1, f2, …, f N} are marked as positive samples k + , and the single image feature vectors from all other image sequences are marked as negative samples k - . The training process and loss function calculation method of the sequence image contrast learning are the same as those of the homogeneous-modal medical image contrast learning described in claim 7, and its loss function is expressed as
[0083] In the cross-modal medical image contrast learning in the image contrast learning, the alignment operation of the contrast learning is carried out in units of patient examinations. In the cross-modal medical image contrast learning, the positive sample pairs are defined as data samples of medical images of different modalities from the same patient, and the negative sample pairs are defined as data samples from different patients. In the process of the cross-modal medical image contrast learning, each time all different modality examination results of the same patient in the same examination are selected, such as MRI, CT, pathological images, etc. The cross-modal medical image contrast learning process first completes the feature extraction of the target patient's examination images according to the image feature extraction steps described in claim 4. After that, all different modality image examination results in this examination are mutually marked as positive sample pairs, and the images in other examinations of other patients are marked as negative sample pairs, and are trained according to the training process and loss function calculation method described in claim 7, and its loss function is expressed as
[0084] The operation object of the image-text contrast learning is medical images and corresponding text information. The image-text contrast learning takes the image and text features obtained by the medical image and text encoding steps as the input, and through a cross-modal contrast learning training method including but not limited to CLIP, aligns the input medical image features with their corresponding text features, and its loss function is expressed as L t2i .
[0085] The total loss of the multi-modal contrast learning process is defined as L total = L img + L t2i , where L img is the loss function of the homogeneous-modal medical image contrast learning, and L t2i is the loss function for image-text contrastive learning. During training, by minimizing L total , the training of the multi-modal contrastive learning module is achieved, thereby obtaining medical image features and text features that retain key diagnostic information and are cross-modally aligned in the feature space.
[0086] Based on the above multi-modal medical image cross-retrieval method, the present invention proposes a multi-modal medical image cross-retrieval system, which consists of:
[0087] 1) Patient data management module: This module acquires and manages the medical imaging data and relevant clinical information of patients, interfaces with the hospital's information system, obtains the basic information of patients, medical imaging examination records, and diagnostic result information therefrom, and stores them in the database of the system. The medical images of different modalities stored in this module will be subjected to feature extraction through the input encoding module described later, and the extracted encoded feature vectors will also be stored in the database for the image retrieval operation of module three described later;
[0088] 2) Input encoding module: This module is based on the medical image encoder, image sequence feature aggregator, and text encoder, and is used to extract feature vectors according to different input data types;
[0089] 3) Image retrieval module: This module is used to implement cross-modal medical image cross-retrieval. This module uses the query condition feature vector encoded by the input encoding module, and according to the requirements, retrieves the target medical imaging data from the database according to the retrieval mode and retrieval process described in claim 12;
[0090] 4) Retrieval operation and display module: This module includes two parts: a retrieval operation module and a result display module. The retrieval operation module accepts the instructions of the user of the image retrieval system, and according to the target image or target patient selected by the user, retrieves the same or different modality historical images that are most similar to the target image in the multi-modal cross-retrieval database, or retrieves the historical patients that are most similar to the clinical image manifestations of the target patient; the result display module realizes the display function of the retrieval results, and provides image browsing and analysis tools, as well as retrieval result clustering analysis tools to help doctors perform diagnosis based on multi-modal medical images.
[0091] 1. Patient data management module
[0092] The patient data management module includes a system data interface and a patient database. This module imports data such as patient information, diagnosis results, examination results, and multi-modal medical image examination results from the hospital information system through the data interface, stores them in the patient database, and provides the required data to subsequent system modules. When this module receives newly stored patient medical images, it directly transmits the images to the image feature encoding module and stores the encoded image feature vectors in the patient database for subsequent image retrieval operations.
[0093] 2. Input Encoding Module
[0094] The input module includes the multi-modal contrast learning module and the trained medical image encoder, image sequence feature aggregator, and text encoder obtained from the contrast learning training process, and is used to extract and encode the features of the input medical images and text information. When this module receives medical images of different modalities that need to be encoded, it will use the corresponding encoding module to encode the image features according to the category of the image (single image, sequence image, or pathological image) described, and return the encoded feature vectors.
[0095] 3. Image Retrieval Module
[0096] When the image retrieval module receives the input text or image as the retrieval condition, it first encodes its features through the input encoding module, and according to the retrieval operation and the needs of the user in the retrieval operation and display module, compares the encoded target feature vector with the specified modal image modal feature vectors stored in the patient database of the patient data management module. The features of the modality to be retrieved will be sorted according to their cosine similarity with the retrieval target image features, and these features and their corresponding patient information will be returned to the retrieval operation and display module. The above text-to-image retrieval process, a certain type of image-to-the-same type of image retrieval process, and a certain type of image-to-different types of image retrieval process correspond to the three alignment processes in the alignment method
[0097] 4. Retrieval Operation and Display Module
[0098] The retrieval operation and display module accepts the patient information and medical images provided by the user as the retrieval target, and provides the input medical images to the image retrieval module to obtain the retrieval results. This module includes two components: a retrieval operation module and a result display module.
[0099] The functions and execution processes provided by the retrieval operation module are as follows:
[0100] Text description - Image retrieval: The user inputs a text description information, and this module converts it into a feature vector through the text encoder, and uses the image retrieval module to retrieve the medical image most similar to the input text information in the multi-modal cross-retrieval database;
[0101] Same-modal or cross-modal image retrieval: This function receives an input image, converts it into a feature vector using the image feature encoding module, and uses the image retrieval module to retrieve several images with the highest similarity to the input image within the same or different image modalities as the input image through feature similarity;
[0102] Missing image modality completion: The user provides a patient with a missing medical image modality. This module retrieves the missing modality image with the highest similarity to the existing images through the above cross-modal image retrieval function based on the image modalities the patient has, and provides the retrieved image and its corresponding patient information to the user for reference;
[0103] Reference patient retrieval: The user selects a patient. This module successively performs same-modal image retrieval on each image modality the patient has, and finds the reference image and patient record with the highest similarity to the current patient and the current modality image. Finally, this module integrates the retrieval results of each modality to find several historical patients with the highest comprehensive similarity and the most reference value to the current patient, and provides them to the user for diagnostic reference.
[0104] The result display module receives the images and patient information returned by the retrieval operation module, and displays the target image of the query, the retrieved result images, as well as the patient information, all examination results, and diagnostic results corresponding to the target image and the result images to the user. The user can compare and analyze medical images of different modalities through the provided image browsing and analysis tools, including zooming in and out, window width and window level adjustment, marking and annotation, measurement, etc., and can save the analysis results to local or cloud for subsequent viewing and comparison. In addition, the result display module can perform clustering analysis on the medical image data obtained from cross-retrieval, compare the classification obtained according to the similarity of image features with the classification obtained according to the patient's diagnostic results, and obtain more comprehensive and accurate diagnostic information.
Claims
1. A multi-modal medical image cross-retrieval method based on contrastive learning, characterized in that, Using any input medical image or text information as a query condition, feature extraction is performed on the input medical image or text information through a medical image encoder, an image sequence feature aggregator, and a text encoder. Then, through a multi-modal contrast learning module, in the ways specified by the user, such as text-to-image, image-to-text, and one-modal image-to-the-same-or-another-modal image, the medical imaging data or text information most similar to the query image or text information is retrieved from the multi-modal cross-retrieval database; for any input medical image or text information, the input data is processed through the described medical image encoder, image sequence feature aggregator, and text encoder to obtain corresponding feature vectors, which are stored in the multi-modal cross-retrieval database; The composition of the method includes: 1) Medical image encoder: Feature extraction is performed on single medical images of different modalities including DR examinations, single CT images, MRI images, PET images, and pathological image slices to obtain corresponding high-dimensional feature vectors; 2) Image sequence feature aggregator: Feature integration is performed on all single-image features in the image sequence of DR examinations, single CT images, MRI images, PET images, and pathological image slices to obtain the entire image sequence feature, or feature integration is performed on all small image patch features in the entire pathological image to obtain the entire pathological image feature; for a medical image sequence composed of multiple images and multiple sequences, first, the medical image encoder is used to perform feature extraction on each image, encoding the entire image sequence into a series of feature vectors of equal length, and then these feature vectors are encoded into an encoding vector representing the entire medical image sequence through the image sequence feature aggregator; for a pathological image of huge size, first, the pathological image is cut into a large number of small image patches of the same size, the medical image encoder is used to encode these image patches into feature vectors of equal length respectively, and finally, the image sequence feature aggregator is used to encode the feature vectors of all small image patches into a feature vector representing the entire pathological image; 3) Text encoder: Encoding is performed on the description text information, diagnostic text information, or text information related to retrieval corresponding to the medical image to obtain corresponding text feature vectors; 4) Multi-modal contrast learning module: Contrast learning is performed among the multi-modal feature vectors output by the image encoder, image sequence feature aggregator, and text encoder, so as to realize the training of each module. Eventually, through each optimized module, medical image features and text features that are cross-modally aligned in the feature space can be obtained; 5) Multi-modal cross-retrieval database: A multi-modal cross-retrieval database constructed based on the single-image features, sequence image features, pathological image features, and text features, which is used to store the feature vectors of all medical images and text information for subsequent image and text retrieval operations.
2. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 1, wherein During the training process, based on the multi-modal contrastive learning module, by aligning the medical image feature vectors and text feature vectors of different modalities in the multi-modal cross-retrieval database, the training of the image encoder, the text encoder, and the multi-modal contrastive learning module is realized, and each of the optimized modules is obtained respectively.
3. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 2, wherein: The medical image encoder and the text encoder are based on a pre-trained cross-modal model, and there are two encoder setting schemes: using only a single cross-modal pre-trained model and simultaneously selecting multiple cross-modal pre-trained models: In the case of using only one cross-modal pre-trained model, the medical image encoder and the text encoder are respectively configured and initialized directly using the image and text encoders of the pre-trained model. The image and text inputs will be processed using the medical image encoder and the text encoder respectively, and the obtained feature vectors will be output. In the case of simultaneously selecting multiple cross-modal pre-trained models, all the selected pre-trained models will be used for the configuration and initialization of the medical image encoder and the text encoder. The medical image encoder and the text encoder will simultaneously include multiple different single encoders. At this time, the type of medical image applicable to each pre-trained model is specified, and the corresponding input data is processed respectively according to the specified applicable medical image type. For a medical image and its corresponding text input, first, according to its type, the pre-trained model specified by the user is selected, and the selected multiple pre-trained models are used to encode the input data respectively. Then, the image features and text features output by all the models are averaged respectively, and the averaged image features and text feature vectors are output.
4. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 1, wherein: The image sequence feature aggregator is realized based on the multi-instance learning framework and the attention mechanism. Taking the single-image feature vector output by the medical image encoder as the input, the self-attention mechanism is applied to weight-aggregate all the input image feature vectors to obtain the feature vector of the entire image sequence.
5. A multimodal medical image cross-retrieval method based on contrastive learning according to claim 1, characterized in that, The multi-modal contrastive learning module includes two training processes: image contrastive learning performed between medical images, and image-text contrastive learning performed between medical images and text; among them, the image contrastive learning includes three specific processes: same-modal medical image contrastive learning, sequence image contrastive learning, and cross-modal medical image contrastive learning.
6. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 5, wherein In the same-modal medical image contrastive learning, the positive sample pairs are defined as multiple data samples obtained by performing different feature transformations on the same picture, and the negative sample pairs are defined as data samples from different pictures; the training method of the contrastive learning is to reduce the distance between the positive sample pairs in the feature space and increase the distance between the negative sample pairs in the feature space. For single-shot scanned images such as DR examinations, the input for the same-modal medical image contrast learning is the image feature vector obtained through the medical image encoder; for image sequences or image matrices such as CT, MRI, and pathological images, the input for the same-modal medical image contrast learning is the image feature vector obtained by encoding each image therein through the medical image encoder and aggregating it through the image sequence feature aggregator; During the training process of the same-modal medical image contrast learning, training will be performed independently on each modality of medical images in sequence; for a selected modality of medical images, all images of this modality are represented as {I1, I2, I3, …, I N}, where N represents the total number of images in this modality. For any one image I i after two random image transformations two different transformed images are obtained and and then, through the image encoding process f, two image feature vectors are obtained and Randomly select one of them as the aligned target object q, then the other feature vector becomes the positive sample k of q + , and all encoding vectors from other images become the negative sample k of q - , then for each selected alignment target q, the loss function is: Among them, k i represents all the encoded feature vectors, τ is a hyperparameter representing the temperature coefficient. For each image in {I1, I2, I3, …, I N}, two encoded vectors will be obtained through the above random image transformation and feature encoding process. Therefore, the total number of encoded vectors k i is 2N; for the entire cross-modal medical image contrast learning process, its loss function is the sum of L q corresponding to all possible choices of q:
7. A multimodal medical image cross-retrieval method based on contrastive learning according to claim 5, characterized in that, The operation object of the sequence image contrast learning is a medical image sequence or image matrix composed of multiple images; In the sequence image contrast learning, for a given image sequence, its positive sample is defined as any single image in the image sequence, and its negative sample is defined as the images in other image sequences; During the process of the contrastive learning of the sequence images, each time an image sequence is selected as the target, all single images in the image sequence are feature-encoded through the medical image encoder to obtain an image feature sequence {f1, f2, …, f N}, and these single-image features are aggregated through the image sequence feature aggregator to obtain a single feature vector q representing the entire image sequence; for the target feature vector q, all the image feature vectors f i ∈ {f1, f2, …, f N} are marked as positive samples k + , and the single image feature vectors from all other image sequences are marked as negative samples k - ; the training process and loss function calculation method of the contrastive learning of the sequence images are the same as those of the same-modal medical image contrastive learning described in claim 7, and its loss function is expressed as 8. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 5, characterized in that The cross-modal medical images perform alignment operations for contrast learning on a patient examination basis; in the cross-modal medical image contrast learning, the positive sample pairs are defined as the data samples of medical images of different modalities from the same patient, and the negative sample pairs are defined as the data samples from different patients; During the process of cross-modal medical image contrastive learning, all different modality examination results of the same patient in the same examination are selected each time. The cross-modal medical image contrastive learning process first completes the feature extraction of the examination images of the target patient through the image feature extraction step. Then, all different modality image examination results in this examination are mutually marked as positive sample pairs, and the images in other examinations from other patients are marked as negative sample pairs. Training is carried out according to the training process and loss function calculation method, and its loss function is expressed as 9. The multimodal medical image cross-retrieval method based on contrastive learning according to claim 5, wherein The operation object of the image-text contrast learning is medical images and corresponding text information; the image-text contrast learning takes the image and text features obtained by the medical image and text encoding steps as inputs, and through a cross-modal contrast learning training method including but not limited to CLIP, aligns the input medical image features with their corresponding text features, and represents its loss function as L t2i .
10. The multimodal medical image cross-retrieval method based on contrastive learning according to claims 5-9, characterized in that, The total loss is defined as L total = L img + L t2i , where L img is the loss function for contrastive learning of homogeneous-modal medical images, and L t2i is the loss function for image-text contrastive learning; during the training process, by minimizing L total , the training of the multi-modal contrastive learning module is realized, so as to obtain medical image features and text features that retain key diagnostic information and are cross-modally aligned in the feature space.