Multi-semantic-granularity cross-modal pre-training method
By performing self-supervised pre-training in the matching of medical imaging and diagnostic reports, and using technologies such as triple extraction and multi-task comparison learning, the performance degradation and gradient problems of existing medical imaging cross-modal pre-trained models are solved, achieving higher analysis accuracy and resource savings.
Patent Information
- Application Number
- CN202510238639.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The existing medical imaging cross-modal pre-trained models have problems such as performance decline, increased training data demand, and disappearance or explosion of gradients in the application of the medical field.
A multi-semantic particle-size cross-modal pre-training method is proposed, and self-supervised pre-training is performed through the matching of medical imaging and diagnostic reports. The key information correlation between medical imaging slices and corresponding diagnostic texts is constructed using triple extraction, multi-task comparison learning and cross-modal attention mechanisms.
It improves the accuracy and interpretability of medical imaging analysis, enhances the accuracy and accuracy of subsequent model training, saves resources, and avoids the problems of gradient disappearance and explosion.
Smart Images

Figure CN120182779A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross-modal pre-training, and in particular to a cross-modal pre-training method with multiple semantic granularities. Background Art
[0002] Pre-training is an important technical strategy in deep learning, aiming to pre-train a model on large-scale data so that it can learn general feature representations, thereby providing good initial parameters for subsequent downstream tasks. This process can significantly reduce the consumption of training time and computing resources for downstream tasks and, to a certain extent, alleviate problems such as gradient vanishing or explosion.
[0003] Pre-trained models can be mainly divided into three categories: visual models, natural language processing models, and multi-modal models. Among them, typical representatives of multi-modal models include CLIP and DALL·E, etc. However, these models are not trained based on medical image-text pair data. If their parameters are directly transferred to deep learning tasks in the medical field, the following problems may occur: model performance degradation, increased training data requirements, gradient vanishing or gradient explosion, etc.
[0004] Hospitals store a large amount of paired medical image and text data. These text data mainly contain the diagnosis and description of medical images, which play a crucial guiding role in interpreting the pathological features presented by the image data. These data provide a solid practical basis for cross-media pre-training. Research results in the field of natural image processing have confirmed that the matching of text and image is an efficient cross-modal pre-training method. Based on the existing research results, the present invention proposes a cross-modal pre-training method with multiple semantic granularities, a self-supervised pre-training method based on the matching of medical images and diagnostic reports, and constructs the relevance of key information between medical image slices and corresponding diagnostic texts through contrastive learning technology. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: to make up for the deficiencies in the current field of cross-modal pre-training of medical images, propose a cross-modal pre-training method with multiple semantic granularities, achieve enhanced training accuracy and accuracy of subsequent models, save resources, and avoid problems of gradient vanishing and explosion. The aim is to improve the accuracy and interpretability of medical image analysis by integrating medical knowledge bases, structured triple extraction, multi-task contrastive learning, and cross-modal attention mechanisms. This method is applicable to lesion detection, disease typing and grading diagnosis, and is compatible with various medical image modalities such as CT and MRI.
[0006] The technical solution of the present invention is as follows: A cross-modal pre-training method with multiple semantic granularities, the steps are as follows:
[0007] Step (1), extract triples for text preprocessing, and construct a 3D visual encoder for medical image processing;
[0008] Extract {entity, position, exist} triples from medical diagnosis reports, where entity represents the disease type, position represents the disease location, and exist represents the presence or absence of the disease; among them, entity and exist are combined as the text category marker [CLS] input, all triples are used as text input, and the text encoder extracts the text feature T and the text category marker feature t cls ;
[0009] After cascading the medical image with the corresponding visual category marker [CLS], it is sent into the 3D visual encoder to obtain the corresponding visual feature V and visual category marker feature v cls ;
[0010] Step (2) Construct a visual-text marker contrast loss module;
[0011] Introduce a visual-text contrast learning loss function to constrain the visual category marker feature and the text category marker feature for preliminary alignment, and the visual-text matching module and the masked language modeling module perform matching based on this preliminary alignment result;
[0012] Step (3) Construct a visual-text matching module;
[0013] Visual-text matching is a binary classification task, supervised by cross-entropy loss for training, and the parameters of the 3D visual encoder and the text encoder are updated backward. The visual feature and the text feature are input, and the two perform contrast learning at a coarse semantic granularity in the visual-text matching module; the positive sample is the correctly matched image-text pair in the batch, and the negative sample is the cross-mismatched image-text pair in the batch;
[0014] Step (4) Construct a masked language modeling module;
[0015] The masked language modeling module predicts the masked text information through the visual feature and the masked text feature, and updates the 3D visual encoder and the text encoder backward.
[0016] The visual-text marker contrast loss module uses a linear projection layer g * (·) to transform the visual category feature v cls and the text category feature t cls to the same dimension, and calculates the cosine similarity s(v, t) between the visual category feature v cls and the text category feature t cls :
[0017] s(v, t) = g v (v cls ) · g t (tcls )
[0018] During the contrastive learning process, the correctly matched image-text pairs in the same batch are used as positive samples, and the cross-mismatched image-text pairs are used as negative samples; the visual-text contrastive learning loss function includes the contrastive loss from vision to text and the contrastive loss from text to vision:
[0019]
[0020] where N is the number of positive samples in a batch, i and j are used to index the medical images and medical diagnosis reports in the batch, τ is a hyperparameter, L v2t represents the contrastive loss function from image to text, L t2v represents the contrastive loss function from text to image, and s represents the cosine similarity.
[0021] In the masked language modeling module, the text features are randomly masked with a probability not higher than 45% to obtain masked text features, which are used as the Query input of the cross-attention model. The Key and Value come from the visual features; the masked text features and the unmasked visual features are input here for fine semantic granularity alignment; the output of the cross-attention model is used to calculate the loss with the original text features to supervise the learning of the 3D visual encoder and the text encoder to update the parameters.
[0022] Advantages of the present invention: Applying triple extraction can effectively obtain the high-level semantics of medical reports. The graphic-text contrast module enables vision and text to be aligned at a low semantic granularity, and the masked language modeling module enables vision and text to be matched at a higher semantic granularity, thus completing cross-modal contrast at multiple semantic granularities. When the visual encoder and the text encoder obtained by applying the present invention are used for downstream task training, the learning efficiency can be effectively improved, and the problems of gradient disappearance or gradient explosion can be avoided. Brief Description of the Drawings
[0023] Figure 1 is the overall framework diagram of the cross-modal pre-training method for multiple semantic granularities;
[0024] Figure 2 is the schematic diagram of triple extraction. (a) is the report of the left circumflex and left main trunk walls, and (b) is the report of the left anterior descending branch and left circumflex;
[0025] Figure 3 is the schematic diagram of the masked language modeling module;
[0026] Figure 4 is the schematic diagram of the visual-text matching module. Detailed Embodiments
[0027] The following further illustrates the detailed embodiments of the present invention in conjunction with the drawings and technical solutions.
[0028] A cross-modal pre-training method with multiple semantic granularities is as follows:
[0029] Step (1): Extract triples for text preprocessing and construct a 3D visual encoder for medical image processing;
[0030] Use the RadGraph tool to extract {entity, position, exist} from medical diagnostic reports, representing disease type, location, and presence respectively, where entity and exist are combined as the text category label [CLS] input, and all triples are used as text inputs. Use ClinicalBERT to extract text features T and text category label features t cls .
[0031] Since medical images contain the three-dimensional structural information of human organs, the present invention uses a 3D visual encoder to capture three-dimensional spatial cues and retain the continuity of anatomical structures. For medical images, after concatenating them with the corresponding [CLS] category labels, they are fed into the 3D visual encoder to obtain the corresponding visual features V and visual category label features v cls .
[0032] Step (2): Construct a visual-text label contrast loss module;
[0033] To promote the representational unity of images and texts in the common embedding space, introduce a vision-text contrastive learning loss function (VTC) to constrain the visual category label features and text category label features for preliminary alignment. Specifically, use a linear projection layer g * (·) to transform the visual category label features v cls and the text category label features t cls to the same dimension, and calculate the cosine similarity s(v, t) between the two features:
[0034] s(v, t) = g v (v cls ) · g t (t cls )
[0035] During the contrastive learning process, use the correctly matched image-text pairs in the same batch as positive samples and the cross-mismatched image-text pairs as negative samples. The contrast loss function here includes visual-to-text contrast loss and text-to-visual contrast loss:
[0036]
[0037] Where N is the number of positive samples in a batch, i and j are used to index the medical images and medical diagnostic reports in the batch, τ is a hyperparameter, and L v2t represents the image-to-text contrast loss function, and L t2v represents the text-to-image contrast loss function, and s represents the cosine similarity.
[0038] The obtained vision-image pairs play a guiding role in subsequent alignment, and subsequent comparison work is carried out based on them.
[0039] Step (3) Construct a vision-text matching module;
[0040] Vision-text matching is a binary classification task, supervised by cross-entropy loss for training, and the parameters of the 3D vision encoder and text encoder are updated in reverse; input visual features and text features, and contrastive learning with coarse semantic granularity is performed here. Positive samples are correctly matched image-text pairs in the batch, and negative samples are cross-mismatched image-text pairs in the batch.
[0041] Step (4) Construct a masked language modeling module;
[0042] The masked language modeling module aims to use visual features and masked text features to predict the masked text information, and update the 3D vision encoder and text encoder in reverse. Specifically, with a probability not higher than 45%, such as 15%, the text features are randomly masked as the input of the cross-attention model Query, and Key and Value come from visual features. Input masked text features and unmasked visual features, and fine-grained semantic alignment is performed here. Calculate the loss between the output of the cross-attention model and the original text features, and supervise the learning to update the parameters of the 3D vision encoder and text encoder, which can use visual information to complete the masked text.
[0043] Through the above steps, cross-modal pre-training applied to medical images can be achieved.
[0044] As Figure 1 shown, the cross-modal pre-training method framework with multi-semantic granularity is generally composed of a text encoder, a vision encoder, masked language modeling, and vision-text matching. Since there are many redundant words in the imaging diagnosis report that are not required in the deep learning process, it is first input into the text processing module for text cleaning. Among them, the disease type and whether the disease type exists are input as classification markers, all triples are used as text inputs, and the ClinicalBERT encoder is used to extract the text feature T and the text category marker feature t cls . Similarly, tokenize the medical images, and after concatenating the visual category marker [CLS], send them into the 3D vision encoder to obtain the visual feature V and the visual category marker feature v clsThe visual features and text category features are initially compared, and then fine-grained comparison is carried out through the masked language modeling module and the vision-text matching module respectively.
[0045] As Figure 2 shown, since there is a lot of text in the imaging report that is not required in the deep learning process, it is first input into the text processing module for text cleaning. The RadGraph tool is used to extract {entity, position, exist} from the medical diagnosis report, and entity and exist among them are used as classification markers for input, and all triples are used as text input.
[0046] As Figure 3 shown, the text features are randomly masked as the Query input of the cross-attention model with a probability of 15%, and the Key and Value come from the visual features. The loss is calculated between the output of the cross-attention model and the original text features for supervision.
[0047] As Figure 4 shown, vision-text matching is a binary classification task, and the model is trained using cross-entropy loss. The positive samples are the correctly matched image-text pairs in the batch, and the negative samples are the cross-mismatched image-text pairs in the batch.
[0048] The pre-trained model of the present invention is applied to the medical image segmentation task for a comparative experiment. The ImageCAS dataset contains 1000 three-dimensional coronary angiography (CTA) images with a size range of 512×512×(206 - 275). In the experiment, the method of the present invention, Clip, and ImageNet were used to compare the performance of various segmentation methods. As shown in Table 1, the experimental results show that applying the visual encoder and text encoder parameters obtained by this method to the segmentation task has obvious advantages in indicators such as DSC, NSD, and ASD compared with other pre-trained models.
[0049] Table 1 Data comparison of different pre-trained models applied to the segmentation task
[0050]
Claims
1. A multi-semantic granularity cross-modal pre-training method, characterized in that: Here are the steps: Step (1), extracting triples for text preprocessing and constructing a 3D visual encoder for medical image processing; Extract {entity, position, exist} triples from medical diagnosis reports, where entity represents the disease type, position represents the disease location, and exist represents whether the disease exists; entity and exist are combined as text category tag [CLS] input, all triples are used as text input, and the text encoder extracts text features T and text category tag features t cls ; After the medical image is concatenated with the corresponding visual category label [CLS], it is sent to the 3D visual encoder to obtain the corresponding visual feature V and visual category label feature v cls ; Step (2) constructing a visual-textual labeling contrast loss module; The visual-text contrast learning loss function is introduced to constrain the visual category label features and the text category label features to be preliminarily aligned. The visual-text matching module and the masked language modeling module perform matching based on this preliminary alignment result. Step (3) constructing a visual-text matching module; Visual-text matching is a binary classification task. It uses cross-entropy loss to supervise training and reversely update the parameters of the 3D visual encoder and text encoder. Visual features and text features are input, and the two are compared and learned at a coarse semantic granularity in the visual-text matching module. Positive samples are correctly matched image-text pairs in the batch, and negative samples are cross-mismatched image-text pairs in the batch. Step (4) constructing a masked language modeling module; The masked language modeling module predicts the masked text information through visual features and masked text features, and updates the 3D visual encoder and text encoder in reverse.
2. The multi-semantic granularity cross-modal pre-training method according to claim 1, characterized in that: The visual-textual labeling contrast loss module uses a linear projection layer g * (·) The visual category feature v cls and text category feature t cls Convert to the same dimension and calculate the visual category feature v cls and text category feature t cls The cosine similarity s(v,t) between them is: s(v,t)=g v (v cls )·g t (t cls ) In the contrastive learning process, the correctly matched image-text pairs in the same batch are used as positive samples, and the cross-mismatched image-text pairs are used as negative samples; the visual-text contrastive learning loss function includes the contrastive loss from vision to text and the contrastive loss from text to vision: Where N is the number of positive samples in a batch, i and j are used to index the medical images and medical diagnosis reports in the batch, τ is a hyperparameter, and L v2t represents the contrast loss function from image to text, L t2v represents the contrast loss function from text to image, and s represents the cosine similarity.
3. The multi-semantic granularity cross-modal pre-training method according to claim 1, characterized in that: In the masked language modeling module, text features are randomly masked with a probability not higher than 45% to obtain masked text features as the query input of the cross-attention model, and the Key and Value come from the visual features; the masked text features and the unmasked visual features are input, and the two are aligned at a fine semantic granularity; the output of the cross-attention model is compared with the original text features to calculate the loss, and the parameters of the 3D visual encoder and the text encoder are updated through supervised learning.
Citation Information
Cited By
Training method and device of open vocabulary detection model and electronic equipment
CN120726615A