Medical multi-modal information processing method and system
Multimodal feature fusion is performed by combining image encoder and text encoder with Transformer network, and using the comparative learning network to optimize feature matching, the problem of insufficient integration of multimodal data in the prior art is solved, and AI diagnosis with higher accuracy is achieved.
Patent Information
- Application Number
- CN202510504607.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-12
AI Technical Summary
Existing medical data processing methods mainly rely on single modal data, making it difficult to effectively integrate multi-dimensional medical information, resulting in incomplete or inaccurate AI diagnosis and lack of cross-modal information interaction and matching mechanisms.
Multimodal data features are extracted through image encoder and text encoder, and feature fusion is used using Transformer network, and multimodal feature matching is optimized with the comparative learning network to generate multimodal joint features to improve diagnostic accuracy.
It significantly improves the accuracy and clinical applicability of AI diagnosis, reduces the cost of data labeling, enhances the model's independent learning ability to relevance multi-source data, and reduces the rate of missed diagnosis.
Smart Images

Figure CN120473173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical information processing, and in particular to a method, device, and computer program product for processing medical multimodal information. Background Art
[0002] The rapid development of medical informatization has generated massive amounts of data in the medical field, including imaging data, text data, and laboratory test results. This data is of great significance for medical research and clinical decision-making. However, the processing and analysis of medical data still face many challenges. Traditional data processing methods often rely on a single type of data and fail to fully utilize multi-dimensional medical information. For example, in the diagnosis of pneumonia, doctors need to comprehensively consider a variety of information, including CT image descriptions provided by radiologists, laboratory reports, medical history, physical examination reports, and cardiopulmonary auscultation. The integration and analysis of this information is not only time-consuming but also difficult for interns to handle.
[0003] In the medical field, applications of multimodal data include medical image analysis and natural language processing. However, in practical applications, the inter-modal correlation of multimodal data is easily lost, which affects the accuracy of disease classification. In addition, existing multimodal fusion diagnosis solutions often separate image and text features and lack effective cross-modal information interaction and matching mechanisms. Therefore, how to effectively integrate and analyze multi-dimensional medical data, especially how to organically integrate imaging data with non-imaging data, has become an urgent problem to be solved. This is not only related to improving the accuracy and efficiency of disease diagnosis, but also has important implications for medical research and clinical decision-making.
[0004] Developing a processing method that can comprehensively consider the characteristics of multimodal data and achieve effective fusion of cross-modal information is of great value in promoting the development and application of medical artificial intelligence. Summary of the Invention
[0005] In order to solve the technical problem that current mainstream AI diagnostic methods mainly rely on single-modal data input and ignore the fact that doctors combine multiple information sources for diagnosis in actual clinical work, resulting in incomplete or inaccurate AI diagnosis, the present invention provides a multimodal medical information processing method that can integrate multimodal data to improve the accuracy and clinical practicality of AI diagnosis.
[0006] The present invention provides a multimodal medical information processing method, comprising the following steps: acquiring medical image data and non-image medical text data of a patient, wherein the medical image data includes a variable number of medical images of different modalities, and the non-image medical text data includes a variable number of medical texts of different categories; extracting features of each medical image data through an image encoder, and fusing the variable number of image features of different modalities into an image fusion feature; extracting features of each non-image medical text data through a text encoder, and fusing the variable number of text features of different categories into a text fusion feature; fusing the image fusion feature and the text fusion feature to generate a multimodal joint feature; inputting the multimodal joint feature and a given plurality of medical report features into a comparative learning network for comparative learning, respectively calculating the similarity between the multimodal joint feature and each medical report feature, and outputting the medical information corresponding to the medical report feature with the highest similarity as a multimodal medical information processing result.
[0007] Furthermore, the training process of the contrastive learning network includes: obtaining a batch of training medical image data and training non-image medical text data of patients, wherein the training medical image data of each patient includes a variable number of medical images of multiple different modalities, and the training non-image medical text data includes a variable number of medical texts of multiple different categories; for each patient's training medical image data, extracting image features of each modality through an image encoder, and fusing the variable number of image features of multiple different modalities into a fixed-length training image fusion feature; for each patient's training non-image medical text data, extracting text features of each category through a text encoder, The indefinite number of text features of different categories are fused into a fixed-length training text fusion feature; the training image fusion feature and the training text fusion feature of each patient are fused to generate the corresponding training multimodal joint feature; the training multimodal joint features of all patients in the same batch are cross-compared with the report features extracted from the medical reports corresponding to each patient, and the similarity between the training multimodal joint features of the same patient and its corresponding medical report features is made higher than the similarity between the training multimodal joint features of the patient and the medical report features of other patients through optimization; the above process is repeated for multiple batches of training until the comparative learning network converges.
[0008] Preferably, the fusing of an indefinite number of image features of multiple different modalities into a fixed-length training image fusion feature includes: determining the maximum number of modalities of the training medical image data of all patients in the current batch; setting the maximum length of the training image fusion feature based on the maximum number of modalities; and for the training medical image data of patients whose number of modalities is less than the maximum number of modalities, padding redundant feature vectors to the maximum length so that the length of the training image fusion feature of each patient is consistent.
[0009] Furthermore, the image feature fusion method also includes: compressing multiple training medical image data of different modalities of the same patient into feature maps of the same size through an image encoder; adding elements of the same spatial position in each feature map to generate a fixed-size fusion feature map as the fixed-length training image fusion feature, wherein the length of the fixed-length training image fusion feature is consistent with the size of the fusion feature map.
[0010] Preferably, the image feature fusion method also includes: determining the target number of modalities and setting a unified training image fusion feature length based on the target number of modalities; for training medical image data of patients whose number of modalities is less than the target number of modalities, filling redundant feature vectors to the unified training image fusion feature length so that the length of the training image fusion feature of all patients remains consistent with the target number of modalities.
[0011] Furthermore, the non-image medical text data includes at least two different types of text information, and the different types are selected from: patient symptom descriptions, medical history records, laboratory test indicators, vital sign parameters, functional test results, and social behavior and lifestyle records.
[0012] Preferably, the image fusion module and / or text fusion module realizes feature fusion through the Transformer network, including: constructing image features of different modalities or text features of different categories in the current batch into feature sequences respectively; filling redundant feature vectors into feature sequences of insufficient length according to the maximum length in the feature sequence to make the length of all feature sequences consistent; inputting feature sequences of consistent length into the Transformer network for fusion processing to generate fixed-length training image fusion features or fixed-length training text fusion features.
[0013] Furthermore, the cross-contrast learning is optimized in the following manner: based on the matching relationship between the trained multimodal joint features and the medical report features of different patients in the current batch, a contrastive learning loss is constructed to synchronously adjust the parameters of the image fusion module and the text fusion module so that the similarity between the trained multimodal joint features of the same patient and their medical report features is higher than the similarity with the medical report features of other patients.
[0014] Preferably, the objective function of the contrastive learning loss is configured to: maximize the normalized similarity between the training multimodal joint features and the corresponding medical report features of the same patient; and minimize the normalized similarity between the training multimodal joint features and the medical report features of different patients.
[0015] The beneficial effects of the present invention are: the present invention significantly improves the accuracy and clinical applicability of AI diagnosis through the dynamic fusion and comparative learning mechanism of multimodal medical data. Specifically: 1) supports an indefinite number of imaging modalities and non-imaging text inputs, and achieves feature length unification through padding, compression addition or Transformer network, which solves the diversity and inconsistency problems of clinical data and enables the model to flexibly adapt to patients with different examination conditions; 2) uses medical reports as supervisory signals for comparative learning, without the need for manual labeling of disease labels, greatly reducing data labeling costs, and enhancing the model's autonomous learning ability for the correlation of multi-source data; 3) integrates multimodal imaging features with multi-dimensional text information such as symptoms, medical history, and laboratory indicators, so that the diagnostic logic is closer to the doctor's comprehensive analysis process, improves the accuracy of the results, and effectively reduces the missed diagnosis rate; 4) optimizes the comparative loss function through normalized cosine similarity, enhances the matching accuracy of the patient's multimodal features and the corresponding medical report, and ensures that the diagnostic recommendations output by the model have higher clinical credibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The exemplary embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0017] Figure 1 This is a flowchart of an embodiment of a multimodal medical information processing method in this specification;
[0018] Figure 2 This is a schematic block diagram of an embodiment of a training method in a multimodal medical information processing method in this specification.
[0019] Figure 3 1 is a simplified block diagram illustrating an example apparatus that may be configured to perform one or more tasks described in embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0021] like Figure 1,2, a specific implementation of the multimodal medical information processing method 100 of the present invention includes the following steps: 102 obtaining the patient's medical image data and non-image medical text data, wherein the medical image data includes a variable number of medical images of different modalities (such as CT, MRI, X-ray, etc.), and the non-image medical text data includes a variable number of medical texts of different categories (such as patient symptom descriptions, medical history records, laboratory test indicators, etc.); 104 extracting the features of each medical image data through the image encoder 201, and fusing the variable number of image features of different modalities into image fusion features; 106 extracting the features of each non-image medical text data through the text encoder 202, and fusing the variable number of text features of different categories into text fusion features; 108 further fusing the image fusion features and the text fusion features to generate multimodal joint features; 110 inputting the multimodal joint features and the pre-stored multiple medical report features into a comparative learning network for similarity calculation, and finally outputting the medical information corresponding to the medical report feature with the highest similarity as the diagnosis result.
[0022] Specifically, the multimodal medical information processing method of this embodiment first obtains the patient's medical image data and non-image medical text data. Medical image data may include but is not limited to medical images of different modalities such as CT scan images, MRI images, X-ray images, ultrasound images, endoscopic images, and pathological slice images. These medical images may come from the same examination time point or different examination time points, and the number and types of medical image modalities of each patient may be different. For example, some patients may have images of three modalities: CT, MRI, and X-ray, while other patients may only have images of PET and pathological slices, or images of two modalities: MRI and X-ray.
[0023] Non-image medical text data includes, but is not limited to, patient symptom descriptions, medical histories, laboratory test results, vital sign parameters, functional test results, and records of social behavior and lifestyle. This text data may come from different medical record systems, and the number of text categories may vary for each patient. For example, some patients may have complete symptom descriptions, medical histories, and laboratory test results, while others may only have symptom descriptions and laboratory test results.
[0024] Symptom descriptions can be derived from the patient's chief complaint. For example, a patient may describe the location, nature (e.g., stabbing, dull), duration, and frequency of pain. These complaints can directly provide clues to potential medical conditions. A medical history should include past medical history, family medical history, and medication history. A history of chronic conditions such as diabetes and hypertension can indicate the risk of complications. A family history of premature heart disease may indicate a genetic risk for cardiovascular disease. A medical history is crucial for identifying a patient's high-risk factors. For example, patients with a history of stroke are at increased risk for recurrent stroke. Even when imaging does not reveal a clear lesion, doctors will consider the patient's medical history in screening. Laboratory tests include blood counts, liver function tests, and kidney function tests. For example, an elevated white blood cell count on a blood count may indicate infection, while elevated transaminases on a liver function test may indicate liver damage. Laboratory tests can provide biochemical evidence of disease. For example, when an infectious disease is suspected, imaging may only reveal areas of inflammation, while a blood count can further determine the likelihood and severity of infection. Vital signs can be derived from a physical examination, including temperature, blood pressure, and pulse. For example, a patient's blood pressure can influence the risk assessment for cardiovascular and cerebrovascular disease; an elevated body temperature often indicates infection or inflammation. Vital signs are crucial indicators for doctors to assess the severity of a patient's condition. For example, in an emergency setting, hypertension accompanied by a severe headache may indicate a cerebrovascular emergency, necessitating immediate action. Functional testing results can be derived from functional assessments, such as electrocardiograms (ECGs) and pulmonary function tests. For example, an ECG can be used to detect arrhythmias, while a pulmonary function test can assess the severity of lung diseases such as asthma or COPD. Functional assessments can reveal the functional status of organs that are not easily detected on imaging. For example, some cardiac abnormalities are not obvious on CT or MRI but can be better captured on an ECG. Social behavior and lifestyle records, such as smoking history, alcohol consumption history, and occupational exposures (such as long-term exposure to chemicals), can also be used. Lifestyle and environmental factors can significantly influence diagnosis. For example, patients with a long-term smoking history have a higher risk of lung cancer, and occupational exposures may be associated with specific types of lung or skin diseases.
[0025] Non-image medical text data contains at least two different types of text information, selected from the following categories: patient symptom descriptions (such as the nature, location, and duration of pain), medical history records (including past medical history, family genetic history, and medication use), laboratory test indicators (such as blood routine, liver function data), vital sign parameters (such as blood pressure, body temperature, pulse), functional test results (such as electrocardiogram abnormalities, lung function indicators), and social behavior and lifestyle records (such as smoking history, occupational exposure history). For example, a patient's text data may contain two combinations of information at the same time - such as the symptom description of "paroxysmal upper abdominal pain with acid reflux" and the laboratory test result of "Helicobacter pylori positive"; or it may contain three combinations of information - such as the chief complaint of "chest pain radiating to the left arm", the family history of "father's 55-year-old myocardial infarction", and the biochemical indicator of "low-density lipoprotein cholesterol (LDL-C) 5.2mmol / L". These combined information strengthens the diagnostic basis through multi-dimensional association. For example, the occupational exposure record of "long-term exposure to asbestos" combined with the symptoms of "chest pain and difficulty breathing" can indicate the possibility of mesothelioma, while the laboratory data of "elevated CRP level" superimposed with the vital sign of "body temperature 39°C" can assist in judging the severity of the infectious disease.
[0026] The training process of the contrastive learning network includes: obtaining a batch of training medical image data and training non-image medical text data of patients, wherein the training medical image data of each patient includes a variable number of multiple medical images of different modalities, and the training non-image medical text data includes a variable number of multiple medical texts of different categories; for each patient's training medical image data, extracting the image features of each modality through the image encoder 201, and fusing the image features of the variable number of multiple different modalities into a fixed-length training image fusion feature; for each patient's training non-image medical text data, extracting the text features of each category through the text encoder 202, and fusing the number of Multiple text features of varying categories are fused into fixed-length training text fusion features; the training image fusion features and training text fusion features of each patient are fused to generate corresponding training multimodal joint features 206; the training multimodal joint features of all patients in the same batch are cross-compared with the report features 207 extracted from the medical reports 209 corresponding to each patient, and optimization is performed to ensure that the similarity between the training multimodal joint features of the same patient and their corresponding medical report features is higher than the similarity between the training multimodal joint features of the same patient and the medical report features of other patients; the above process is repeated for multiple batches of training until the contrastive learning network converges. The fixed length can be different between different batches.
[0027] Training a contrastive learning network requires acquiring a large amount of patient training data, including both medical image data and non-image medical text data. This training data should cover a wide range of common diseases and health conditions to ensure the model can learn characteristic representations of different diseases. Each patient's training data includes a variable number of medical images of multiple modalities and a variable number of medical texts from multiple categories, consistent with real-world clinical scenarios, as different patients may undergo different types of examinations.
[0028] like Figure 2 As shown, various types of image data are used Figure 1 The “Image encoder” (image encoder 201) shown performs feature extraction. Figure 2 The process of inputting various image types into the image encoder is indicated by dashed lines. In practice, the number and type of images vary. During training, the network is trained using multimodal images of various types. During testing, any number and type of images can be used. For example, a plain scan CT image can be used to generate a single encoder output feature, or both CT and PET images can be used to generate two image encoding features. The image encoder can be a variety of neural network structures, such as ViT, CNN, and Mamba.
[0029] Because different patients have different medical image modalities, the feature vector dimensions of a single image after passing through the image encoder are consistent. However, due to the inconsistent number of input images, the total number of image features is also inconsistent. Therefore, it is necessary to fuse the image features of multiple different modalities into a fixed-length training image fusion feature.
[0030] The fusion method can be to determine the maximum number of modalities for training medical image data for all patients in the current batch, set a maximum length for the training image fusion features based on the maximum number of modalities, and pad redundant feature vectors for training medical image data of patients with fewer modalities than the maximum number of modalities to the maximum length, so that the length of the training image fusion features for each patient is consistent. The redundant feature vectors can be zero vectors, average vectors, or other padded vectors. The feature sequences of consistent length are input into a Transformer network for fusion processing to generate fixed-length training image fusion features.
[0031] For example, if the largest number of patients in the current batch has medical images of four modalities, and the feature vector dimension of each modality is 2048, the maximum length of the training image fusion feature is 8192 (4 × 2048). For patients with only two modalities, two 2048-dimensional zero vectors need to be padded to make the training image fusion feature length also reach 8192.
[0032] In another embodiment, the fusion method is to compress the training medical image data of multiple different modalities of the same patient into feature maps of the same size through an image encoder, add the elements of the same spatial position in each feature map, and generate a fixed-size fusion feature map as a fixed-length training image fusion feature, wherein the length of the fixed-length training image fusion feature is consistent with the size of the fusion feature map.
[0033] For example, if the feature map size of each modality medical image after passing through the image encoder is 7×7×2048, the feature maps of different modalities can be added or averaged in the channel dimension to obtain a 7×7×2048 fused feature map, which can then be flattened into a training image fused feature of length 100352 (7×7×2048). The feature sequence of the same length is input into the Transformer network for fusion processing to generate a fixed-length training image fused feature.
[0034] In another embodiment, the fusion method determines a target number of modalities and sets a unified training image fusion feature length based on the target number of modalities. For training medical image data of patients whose number of modalities is less than the target number of modalities, redundant feature vectors are padded to the unified training image fusion feature length so that the length of the training image fusion feature of all patients is consistent with the target number of modalities. The redundant feature vector can be a zero vector, an average vector, or other padded vector.
[0035] For example, the target number of modalities can be set to 3, and the feature vector dimension of each modality can be set to 2048. The unified training image fusion feature length is 6144 (3 × 2048). For patients with only two modalities, a 2048-dimensional zero vector is required to pad the training image fusion feature length to 6144. For patients with four modalities, the three most important modalities can be selected, or the features of all four modalities can be reduced in dimension to achieve a training image fusion feature length of 6144.
[0036] These padding operations or fusion operations use specifically defined blank features to ensure that the image feature sequences in each batch have the same length. Image feature sequences of the same length are input into the transformer network to extract feature information. The result of this is that no matter how the number of input images changes, the output feature length after processing by the transformer network will remain consistent. Through the processing process of the above-mentioned image fusion module, the fusion of different types of image input information can be effectively achieved, and a comprehensive feature expression can be generated, thereby solving the problem of inconsistent feature lengths caused by the uncertain number of input images. This method not only improves the flexibility of model processing, but also enhances the model's adaptability and parsing capabilities for multi-source image data. After passing through the image fusion module, a feature expression that fused various types of image input information is obtained.
[0037] In another embodiment, fusing the indefinite number of image features of multiple different modalities into fixed-length training image fusion features includes: constructing the image features of different modalities in the current batch into feature sequences respectively; filling redundant feature vectors into feature sequences of insufficient length according to the maximum length of the feature sequences to make the lengths of all feature sequences consistent; and inputting feature sequences of consistent length into the Transformer network for fusion processing to generate fixed-length training image fusion features.
[0038] Similarly, Figure 1 The process of inputting various types of non-imaging data into the image encoder is represented by dotted lines. In practice, the types and quantities of non-imaging data vary, depending on the type and number of examinations a patient undergoes. For each patient's training non-image medical text data, a text encoder extracts text features of each category and fuses an indefinite number of text features from multiple categories into a fixed-length training text fusion feature. The text encoder can be a pretrained natural language processing model, such as BERT, GPT, Llama, Qwen, and other text models. These models have been pretrained on large-scale text datasets and can effectively extract semantic features from medical text. For each category of medical text, the text encoder outputs a feature vector of fixed dimensions, such as 768. During training, the network collects multiple categories of text input for training, and during testing, text of any category and quantity can be used for testing.
[0039] The method of fusing text features of different categories is similar to fusing image features of different modalities. Techniques such as padding, weighted averaging, or attention mechanism can be used to ensure that the length of the training text fusion features of all patients is consistent.
[0040] Because different patients have different medical text categories, the feature vector dimensions of a single medical text are consistent after passing through the text encoder 202. However, due to the inconsistent number of input medical texts, the total number of medical text features is also inconsistent. Therefore, it is necessary to fuse multiple medical text features of different categories into a fixed-length training text fusion feature.
[0041] The fusion method can be to determine the maximum number of categories for the training medical text data of all patients in the current batch, set the maximum length of the training text fusion features based on the maximum number of categories, and pad the redundant feature vectors of the training medical text data of patients with fewer categories than the maximum number of categories to the maximum length so that the length of the training text fusion features of each patient is consistent. The redundant feature vector can be a zero vector, an average vector, or other padded vector. The feature sequence of the same length is input into the Transformer network for fusion processing to generate a fixed-length training image fusion feature.
[0042] In another embodiment, the fusion method is to compress multiple different categories of training medical text data from the same patient into feature maps of the same size using a text encoder, add the elements at the same spatial position in each feature map, and generate a fixed-size fused feature map as a fixed-length training text fusion feature, where the length of the fixed-length training text fusion feature is consistent with the size of the fused feature map. The feature sequence of the same length is input into a Transformer network for fusion processing to generate a fixed-length training text fusion feature.
[0043] In another embodiment, the fusion method is to determine the number of target categories and set a unified training image text feature length based on the target number of categories. For training medical text data of patients whose number of categories is less than the target number of categories, redundant feature vectors are padded to the unified training text fusion feature length, so that the length of the training text fusion feature of all patients is consistent with the target number of categories. The redundant feature vector can be a zero vector, an average vector, or other padded vector.
[0044] These padding operations or fusion operations use specifically defined blank features to ensure that the text feature sequences in each batch have the same length. Text feature sequences of the same length are input into the transformer network to extract feature information. The result of this is that no matter how the amount of input text changes, the output feature length after processing by the transformer network will remain consistent. Through the processing process of the above-mentioned text fusion module, the fusion of text input information of different categories can be effectively achieved, and a comprehensive feature expression can be generated, thereby solving the problem of inconsistent feature lengths caused by the uncertain number of input medical text categories. This method not only improves the flexibility of model processing, but also enhances the model's adaptability and parsing capabilities for multi-source text data. After passing through the text fusion module, a feature expression that fused the text input information of each category is obtained.
[0045] In another embodiment, the indefinite number of multiple different categories of text features are fused into a fixed-length training text fusion feature, including: constructing the text features of different categories in the current batch into feature sequences respectively; filling the feature sequences of insufficient length with redundant feature vectors according to the maximum length of the feature sequences to make the lengths of all feature sequences consistent; and inputting the feature sequences of consistent length into the Transformer network for fusion processing to generate the fixed-length training text fusion feature.
[0046] Specifically, the image features of different modalities or text features of different categories in the current batch are first constructed into feature sequences. For image features, the feature vector of each modality can be regarded as an element in the sequence; for text features, the feature vector of each type of text can be regarded as an element in the sequence.
[0047] Since different patients have different numbers of modalities or categories, the lengths of feature sequences are also different. In order for the Transformer network to process these sequences, it is necessary to fill the feature sequences of insufficient length with redundant feature vectors based on the maximum length in the feature sequence so that all feature sequences have the same length. Redundant feature vectors can be zero vectors or special padding vectors. Feature sequences of consistent length are input into the Transformer network for fusion processing. The Transformer network contains multiple self-attention layers and feedforward neural network layers, which can capture the relationship between different elements in the sequence and generate context-aware feature representations. The output of the Transformer network is a sequence of the same length as the input sequence. This sequence can be converted into a fixed-dimensional vector through average pooling, maximum pooling, or the use of special tags, and used as a fixed-length training image fusion feature or a fixed-length training text fusion feature.
[0048] Next, the training image fusion features and training text fusion features of each patient are fused to generate the corresponding training multimodal joint features. The fusion method can be simple feature concatenation or more complex multimodal fusion techniques such as cross-attention mechanism or multimodal Transformer.
[0049] The trained multimodal joint features for all patients in the same batch are then cross-referenced with the report features extracted from each patient's corresponding medical report. Medical report 209 can be a document such as a doctor's diagnosis report, treatment plan, or prognosis assessment. Its features are extracted through text encoder 208 to obtain report features. The text encoder used to extract features from medical report 209 can be a network such as BERT or a large model QWEN.
[0050] The goal of cross-contrast learning is to make the similarity between the trained multimodal joint features of the same patient and their corresponding medical report features higher than the similarity between the trained multimodal joint features of the same patient and the medical report features of other patients, thereby achieving effective cross-modal feature learning. Contrast learning is a self-supervised learning method whose core idea is to make the feature representations of similar samples similar and the feature representations of dissimilar samples dissimilar. In this method, contrastive learning is used to optimize the multimodal medical information processing model, making the multimodal joint features of the same patient similar to their corresponding medical report features and dissimilar to the medical report features of other patients.
[0051] For example, for a batch of N patients, an N×N similarity matrix can be constructed, where the element in the i-th row and j-th column represents the similarity between the training multimodal joint features of the i-th patient and the medical report features of the j-th patient. Ideally, the diagonal elements of the similarity matrix (i.e., the similarity between the training multimodal joint features of the same patient and their corresponding medical report features) should be much larger than the off-diagonal elements (i.e., the similarity between different patients).
[0052] By optimizing the contrastive learning loss function, the model learns to match a patient's multimodal medical data with their corresponding medical reports. This allows the model to find the most suitable medical report or report features for a new patient during the inference phase. During the inference phase, a medical report can include sentences describing the treatment results, such as "tumor present" and "tumor not found." The report features are then compared with the multimodal joint features. The one with the highest similarity indicates the specific treatment result, including the classification outcome.
[0053] Repeat the above training process for multiple batches until the contrastive learning network converges. Convergence can be indicated by the contrastive learning loss on the validation set no longer decreasing, or the matching accuracy on the validation set reaching a preset threshold.
[0054] In one embodiment, the specific goals of comparative learning are as follows: the normalized cosine distance between the training multimodal joint features and the reported features extracted from medical reports for the same patient is minimized, while the normalized cosine distance between the training multimodal joint features and the reported features extracted from medical reports for different patients is maximized. By learning from big data, the network can automatically distinguish between different types of patients, thereby gaining diagnostic capabilities for different disease types.
[0055] To ensure that image features and text features are compared at the same scale, the extracted features (training multimodal joint features and report features extracted from medical reports) are usually L2 normalized before calculating the above loss. Assuming x is an unnormalized feature vector, its normalized feature vector x can be obtained by the following formula:
[0056] in,
[0057] After normalizing both the training multimodal joint features and the report features extracted from the medical report, the cosine distance between the two features is calculated as follows: where x and y represent the training multimodal joint features and report features extracted from medical reports, respectively.
[0058] In one embodiment, the objective function of the contrastive learning loss is configured to maximize the normalized similarity between the trained multimodal joint features and the corresponding medical report features for the same patient, and minimize the normalized similarity between the trained multimodal joint features and the medical report features for different patients. The normalized similarity can be cosine similarity, dot product similarity, or other similarity metrics. By optimizing this objective function, the model can learn to match a patient's multimodal medical data with their corresponding medical reports.
[0059] For example, we can use the InfoNCE loss function as a contrastive learning loss:
[0060] L=-log[exp(sim(f_i,r_i) / τ) / Σ_j exp(sim(f_i,r_j) / τ)]
[0061] where fi is the trained multimodal joint feature of the i-th patient, ri is the medical report feature of the i-th patient, sim(·,·) is the similarity function (e.g., cosine similarity), and τ is the temperature parameter that controls the smoothness of the distribution.
[0062] By minimizing this loss function, the model can make the similarity between the training multimodal joint features of the same patient and their medical report features (numerator) as large as possible relative to the sum of the similarities with the medical report features of all patients (denominator), thereby achieving the goal of contrastive learning.
[0063] During the training process, contrastive learning loss will synchronously adjust the parameters of the image fusion module and the text fusion module, so that the entire model can effectively fuse medical information of different modalities and match it with medical reports.
[0064] In another embodiment, Figure 1 , 2, a specific embodiment of the multimodal medical information processing method of the present invention includes the following steps: 102 obtaining the patient's medical image data and non-image medical text data, wherein the medical image data includes an indefinite number of multiple medical images of different modalities (such as CT, MRI, X-ray, PET, etc.), and the non-image medical text data includes an indefinite number of multiple medical texts of different categories (such as patient symptom descriptions, medical history records, laboratory test indicators such as blood routine, etc.); 104 extracting the features of each medical image data through the image encoder 201, and combining the indefinite number of medical image data through the image fusion module 203 The image features of multiple different modalities are fused into image fusion features; 106 the features of each non-image medical text data are extracted through the text encoder 202, and the text features of the indefinite number of different categories are fused into text fusion features through the text fusion module 204; 108 the image fusion features and the text fusion features are further fused through the image-text fusion module 205 to generate multimodal joint features; 110 the multimodal joint features and the pre-stored multiple medical report features are input into the comparative learning network for similarity calculation, and finally the medical information corresponding to the medical report feature with the highest similarity is output as the diagnosis result.
[0065] To simplify the explanation, the training operations are depicted and described herein in a particular order. However, it should be understood that the training operations can occur in a different order, simultaneously, and / or with other operations not presented or described herein. Furthermore, it should be noted that not all operations depicted and described herein may be included in the training process, and not all illustrated operations need to be performed.
[0066] The systems, methods, and / or instrumentalities described herein may be implemented using one or more processors, one or more storage devices, and / or other suitable accessory devices (eg, display devices, communication devices, input / output devices, etc.). Figure 3300 is a block diagram illustrating an exemplary apparatus 300 that can be configured to perform the tasks described herein. As shown, the apparatus 300 may include a processor (e.g., one or more processors) 302, which may be a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or any other circuit or processor capable of performing the functions described herein. The apparatus 300 may further include communication circuitry 304, a memory 306, a mass storage device 308, an input device 310, and / or a communication link 312 (e.g., a communication bus) through which one or more of the components shown in the figure can exchange information.
[0067] Communication circuitry 304 may be configured to transmit and receive information using one or more communication protocols (e.g., TCP / IP) and one or more communication networks, including a local area network (LAN), a wide area network (WAN), the Internet, or wireless data networks (e.g., Wi-Fi, 3G, 4G / LTE, or 5G networks). Memory 306 may include a storage medium (e.g., a non-transitory storage medium) configured to store machine-readable instructions that, when executed, cause processor 302 to perform one or more of the functions described herein. Examples of machine-readable media may include volatile or non-volatile memory, including, but not limited to, semiconductor memory (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), flash memory, and the like. Mass storage device 308 may include one or more disks, such as one or more internal hard disks, one or more removable disks, one or more magneto-optical disks, one or more CD-ROM or DVD-ROM disks, and the like, on which instructions and / or data may be stored to facilitate the operation of processor 302. The input device 310 may include a keyboard, a mouse, a voice-controlled input device, a touch-sensitive input device (eg, a touch screen), etc. for receiving user input to the apparatus 300 .
[0068] It should be noted that the apparatus 300 can operate as a standalone device or can be connected (e.g., networked or clustered) with other computing devices to perform the functions described herein. Figure 3 Only one instance of each component is shown in the figure, but those skilled in the art will understand that the device 300 may include multiple instances of one or more components shown in the figure.
[0069] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0070] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A multimodal medical information processing method, characterized in that: The method comprises: Acquiring medical image data and non-image medical text data of a patient, wherein the medical image data includes a variable number of medical images of different modalities, and the non-image medical text data includes a variable number of medical texts of different categories; Extracting features of each medical image data through an image encoder, and fusing the image features of the indefinite number of different modalities into an image fusion feature; Extracting features of each non-image medical text data through a text encoder, and fusing the indefinite number of text features of different categories into a text fusion feature; fusing the image fusion feature and the text fusion feature to generate a multimodal joint feature; The multimodal joint feature and a given plurality of medical report features are input into a comparative learning network for comparative learning, and the similarity between the multimodal joint feature and each medical report feature is calculated respectively, and the medical information corresponding to the medical report feature with the highest similarity is output as the multimodal medical information processing result.
2. The method according to claim 1, wherein in, The training process of the contrastive learning network includes: Obtaining a batch of training medical image data and training non-image medical text data of patients, wherein the training medical image data of each patient includes a variable number of medical images of multiple different modalities, and the training non-image medical text data includes a variable number of medical texts of multiple different categories; For each patient's training medical image data, extracting image features of each modality through an image encoder, and fusing the indefinite number of image features of multiple different modalities into a fixed-length training image fusion feature; For each patient's training non-image medical text data, extract text features of each category through a text encoder, and fuse the indefinite number of text features of multiple different categories into a fixed-length training text fusion feature; Fuse the training image fusion features and training text fusion features of each patient to generate the corresponding training multimodal joint features; Cross-comparison learning is performed on the training multimodal joint features of all patients in the same batch and the report features extracted from the medical reports corresponding to each patient. Through optimization, the similarity between the training multimodal joint features of the same patient and its corresponding medical report features is higher than the similarity between the training multimodal joint features of the patient and the medical report features of other patients; The above process is repeated for multiple batches of training until the contrastive learning network converges.
3. The method according to claim 2, wherein The step of fusing the indefinite number of image features of different modalities into a fixed-length training image fusion feature comprises: Determine the maximum number of modalities for training medical image data for all patients in the current batch; Setting a maximum length of the training image fusion feature based on the maximum number of modalities; For training medical image data of patients whose number of modalities is less than the maximum number of modalities, redundant feature vectors are padded to the maximum length so that the length of the training image fusion features of each patient is consistent.
4. The method according to claim 2, wherein The step of fusing the indefinite number of image features of different modalities into a fixed-length training image fusion feature comprises: The training medical image data of multiple different modalities of the same patient are compressed into feature maps of the same size through an image encoder; The elements at the same spatial position in each feature map are added together to generate a fixed-size fusion feature map as the fixed-length training image fusion feature, wherein the length of the fixed-length training image fusion feature is consistent with the size of the fusion feature map.
5. The method according to claim 2, wherein The step of fusing the indefinite number of image features of different modalities into a fixed-length training image fusion feature comprises: Determining the target number of modalities and setting a unified training image fusion feature length based on the target number of modalities; For training medical image data of patients whose number of modalities is less than the target number of modalities, padding redundant feature vectors to the unified training image fusion feature length, The length of the fusion features of the training images of all patients is made consistent with the number of target modalities.
6. The method according to claim 1, wherein The non-image medical text data includes at least two different types of text information, and the different types are selected from: patient symptom descriptions, medical history records, laboratory test indicators, vital sign parameters, functional test results, and social behavior and lifestyle records.
7. The method according to claim 2, wherein The method comprises fusing the indefinite number of image features of different modalities into a fixed-length training image fusion feature or fusing the indefinite number of text features of different categories into a fixed-length training text fusion feature, including: Construct the image features of different modalities or text features of different categories in the current batch into feature sequences respectively; According to the maximum length of the feature sequences, redundant feature vectors are filled into the feature sequences with insufficient length to make the lengths of all feature sequences consistent; The feature sequences with the same length are input into the Transformer network for fusion processing to generate fixed-length training image fusion features or fixed-length training text fusion features.
8. The method according to claim 2, wherein The cross-contrast learning is optimized in the following way: Based on the matching relationship between the trained multimodal joint features and the medical report features of different patients in the current batch, a contrastive learning loss is constructed to synchronously adjust the parameters of the image fusion module and the text fusion module so that the similarity between the trained multimodal joint features of the same patient and their medical report features is higher than the similarity with the medical report features of other patients.
9. The method according to claim 8, wherein The objective function of the contrastive learning loss is configured as: Maximizing the normalized similarity between the training multimodal joint features and the corresponding medical report features of the same patient; And minimize the normalized similarity between the trained multimodal joint features and medical report features of different patients.
10. A computer program product comprising instructions, which, when executed on a computer, causes the computer to perform the method according to any one of claims 1 to 9.