Medical vision-language pre-training method based on multi-view and text combination

By combining the positive lateral view feature integration and alignment module, the problem of existing methods ignoring lateral view is solved, and more comprehensive image and report feature extraction is achieved, improving the accuracy and generalization ability of the model in downstream tasks.

CN120339754APending Publication Date: 2025-07-18CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510483798.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing medical vision-language pre-training methods mainly rely on the upright view of chest X-rays, ignoring the importance of lateral view, resulting in insufficient understanding of certain pathological features by the model in practical applications, especially those diseases that are not easy to observe in the upright view but are more obvious in the lateral view.

Method used

The medical vision-language pre-training method based on the combination of multi-view and text is adopted to build a general medical vision-language model, combine the positive lateral view feature integration module and the positive lateral feature alignment module, and extract key pathological patches using the positive lateral view feature integration module, and restore the loss function through cross-modal alignment loss function and mask report to ensure the information consistency and integrity of different views and text reports.

Benefits of technology

It improves the accuracy of downstream tasks, can extract images and reports more fully, enhances the model's understanding of different views, improves the accuracy and generalization capabilities of downstream tasks, and especially shows good performance under low data volume conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339754A_ABST
    Figure CN120339754A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of medical vision-language pre-training, and relates to a medical vision-language pre-training method based on multi-view and text combination, which comprises the following steps of: preprocessing chest X-ray image-report data; inputting the preprocessed image into a view encoder to obtain local features VF and VL and global representations gF and gL of a positive view and a side view, wherein the local features VF and VL of the positive view and the local features VL of the side view and the global representations gF and gL of the positive view and the side view are compared with the report sequence XT of the mask and the report sequence # imgabs0 # of the mask; respectively inputting the sequence XT and # imgabs1 # into a report encoder to obtain a local feature T, a global representation gT and a mask report representation # imgabs2 #, and inputting VF, VL, T, gF, gL, gT and # imgabs3 # into a positive side view feature integration module to obtain a mask report generation text T'and a prediction probability P thereof; inputting the VF, the VL and the T into a positive and lateral feature alignment module to obtain fine-grained representations F and L of positive and lateral views; calculating a loss function value according to P, F, L, gF, gL and gT, and updating model parameters according to the loss function value until a pre-trained medical vision-language general model is obtained; according to the method, the lateral view is introduced into pre-training, so that the diagnosis accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical vision-language pre-training in artificial intelligence, and particularly relates to a medical vision-language pre-training method based on the combination of multi-view and text. Background Art

[0002] In recent decades, deep learning technology has significantly promoted the understanding of medical images. However, there has always been a problem of data shortage in the field of medical imaging, mainly due to the following reasons: First, annotating medical images requires more professional practitioners compared to natural images, which is both expensive and time-consuming; Second, it is difficult to scale up data for some rare cases; Third, factors such as ethical privacy make it impossible to aggregate and publicly disclose data. When the number of samples in a medical dataset for a specific field is small, it cannot provide enough information for the deep learning model to learn medical image representations. Therefore, people hope to learn more general feature representations on relevant datasets through pre-training tasks and transfer them to specific fields to perform better in specific tasks.

[0003] In the medical field, there are often paired diagnostic reports for image pictures. Researchers have noticed that text reports written by experienced doctors usually contain rich domain knowledge and are relatively easy to obtain. Therefore, using paired images and reports for vision-language pre-training has become a promising solution and has attracted extensive attention in the medical field. This paradigm requires no additional manual annotation and uses words and sentences in free text reports as supervision to integrate domain knowledge and guide the neural network to learn general medical image representations. The trained pre-training model can be used as a feature extractor to directly extract feature embeddings or be fine-tuned to downstream tasks to improve the performance of downstream tasks such as classification, segmentation, retrieval, and report generation, and has now become an important method for medical image analysis.

[0004] Existing medical VLP methods mainly use two paradigms, contrastive learning and masked reconstruction, to learn image and text representations. Contrastive learning enhances the model's ability to recognize high-level features of diseases by comparing the similarities and differences between different samples. Masked reconstruction enhances the model's ability to understand details by predicting the occluded content.

[0005] The lateral view can provide additional information about the lesion location and anatomical structure, which is crucial for a comprehensive understanding and accurate diagnosis of diseases. It is worth noting that in clinical practice, doctors often rely on the combined information of the frontal and lateral views of radiological scans for accurate assessment.

[0006] However, existing methods mainly rely on the frontal view of chest X-rays during the pre-training process, often neglecting the importance of the lateral view. This limitation of discarding the lateral view may lead to insufficient understanding of certain pathological features by the model in practical applications, especially those diseases that are not easily observable in the frontal view but are more obvious in the lateral view. Summary of the Invention

[0007] To solve the above problems of the existing technology, the present invention adopts a medical vision-language pre-training method based on the combination of multi-views and text, which is characterized by including:

[0008] S1. Construct a medical vision-language general model, which includes: a view encoder, a report encoder, a frontal-lateral view feature integration module, and a frontal-lateral feature alignment module;

[0009] S2. Obtain chest X-ray image-report data, preprocess the chest X-ray image-report data to obtain a preprocessed chest X-ray image, a report sequence X T and a masked report sequence The preprocessed chest X-ray image includes a frontal image X F and a lateral image X L ;

[0010] S3. Input the preprocessed chest X-ray image into the view encoder to obtain local features V F , V L and global representations g F , g L ; of the frontal view and the lateral view; Input the report sequence X T and the masked report sequence into the report encoder respectively to obtain local features T of the report sequence and global representation g T as well as a masked report representation

[0011] S4. Input the local features V F , V L , T of the frontal view, the lateral view, and the report sequence, and the global representations g F , g L , g T as well as the masked report representation into the frontal-lateral view feature integration module to obtain a masked report generation text T' and its prediction probability P;

[0012] S5. Input the local features V F , V L , T of the frontal view, the lateral view, and the report sequence into the frontal-lateral feature alignment module to obtain fine-grained representations F, L of the frontal view and the lateral view;

[0013] S6. According to the predicted probability P, the fine-grained representations F and L of the frontal view and the lateral view, and the global representations g of the frontal view, the lateral view, and the report sequence F , g L , g T Calculate the overall loss function value, update the model parameters according to the overall loss function value, and when the overall loss function value is minimized, obtain the pre-trained medical vision-language general model.

[0014] Beneficial effects:

[0015] 1. Aiming at the problem that the existing medical vision-language pre-training method ignores the lateral view, resulting in the model failing to fully learn the information contained in the image and the report, the present invention introduces the lateral view into the pre-training, which can more fully extract the features of the image and the report, thereby improving the accuracy of downstream tasks; 2. The present invention uses the front and lateral view feature integration module to select the key pathological patches in the chest X-ray images of different views for restoring the masked report, enabling the model to have a comprehensive understanding of different views, while strengthening the joint understanding of the model for the image and the report, so as to more fully extract the features of the image and the report and improve the accuracy of downstream tasks; 3. The front and lateral feature alignment module used in the present invention ensures the synergistic effect of different views and the integrity of the information related to the report, so as to more fully extract the image features and improve the accuracy of downstream tasks; 4. The cross-modal alignment loss function of the present invention enables the frontal and lateral views and the text report to maintain overall consistency at the high-order semantic level and can capture meaningful disease-level features; 5. The present invention combines the masked report restoration loss function, the front and lateral view fine-grained representation consistency loss function, and the cross-modal alignment loss function, enabling the model to not only learn the fine-grained features of different-angle views and reports, but also maintain consistency at the high-order semantics, and showing good generalization ability in a variety of downstream tasks. Brief description of the drawings

[0016] Figure 1 It is a flowchart of a medical vision-language pre-training method based on the joint of multi-views and text provided by an embodiment of the present invention;

[0017] Figure 2 It is a structural diagram of a medical vision-language pre-training method based on the joint of multi-views and text provided by an embodiment of the present invention;

[0018] Figure 3 It is a result diagram on the downstream linear probing classification task provided by an embodiment of the present invention;

[0019] Figure 4 It is a result diagram on the downstream segmentation task provided by an embodiment of the present invention. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] As Figure 1 、 Figure 2 shown, the present invention adopts a medical vision-language pre-training method based on the combination of multi-views and text, including:

[0022] S1. Construct a medical vision-language general model, which includes: a view encoder, a report encoder, an anteroposterior view feature integration module, and an anteroposterior feature alignment module;

[0023] S2. Obtain chest X-ray image-report data, preprocess the chest X-ray image-report data to obtain the preprocessed chest X-ray image, report sequence X T and the masked report sequence The preprocessed chest X-ray image includes an anteroposterior image X F and a lateral image X L ;

[0024] Use the chest X-ray image-report dataset MIMIC for pre-training. The MIMIC dataset includes multiple chest X-ray image-report data, and each chest X-ray image-report data includes: anteroposterior and lateral chest X-ray images and a text report.

[0025] Specifically, the preprocessing of the chest X-ray image-report data includes:

[0026] Randomly crop the anteroposterior and lateral chest X-ray images, and unify the image size to 224×224 to obtain the preprocessed anteroposterior image X F and lateral image X L ; Use regular matching to remove special symbols in the text report, and use the tokenizer of BioClinicalBERT to tokenize the text report after removing special symbols and convert it into a sequence to obtain the report sequence X T .

[0027] According to the importance of medical entity nouns and their frequencies of appearance in radiology reports, 44 diseases and symptoms related to chest X-ray images, such as atelectasis, cardiomegaly, etc., were selected from Medical Subject Headings (MESH) to obtain a medical entity set. Medical entity words in the text report after removing special symbols were masked with a probability of 75%, and the remaining words were masked with a probability of 65%. The masked text report was tokenized to obtain a masked report sequence

[0028] S3. Input the preprocessed chest X-ray image into the view encoder to obtain local features V of the frontal view and the lateral view F 、V L and the global representation g F 、g L ; Input the report sequence X T and the masked report sequence into the report encoder respectively to obtain the local feature T of the report sequence and the global representation g T and the masked report representation

[0029] The view encoder includes: a frontal view encoder and a lateral view encoder The view encoder processes the preprocessed chest X-ray image including:

[0030] Input the frontal image X F and the lateral image X L into the corresponding frontal view encoder and the lateral view encoder to obtain the local feature of the frontal view global representation and the local feature of the lateral view lateral view global representation where N is the number of patches of the local feature of the frontal view, and d is the feature dimension

[0031] The frontal view encoder and the lateral view encoder are both ViT-B / 16; specifically, ViT-B / 16 processes an image by: dividing the image into 16×16 patches, and linearly mapping each patch into a feature vector; adding a special [CLS] token at the beginning of the feature vector of each patch for aggregating global information; after the feature vectors of the patches with the [CLS] token added pass through multiple layers of Transformer in ViT-B / 16, the output features are obtained; the features corresponding to the feature vectors of the patches in the output features are local features, and the feature corresponding to the [CLS] token is the global representation; where the [CLS] token is a classification token, which is a special token used in the BERT model.

[0032] The report encoder processes the report sequence X T and the masked report sequence as follows:

[0033] Using BioClinicalBERT as the report encoder E txt , the report sequence X T is input into the report encoder E txt for encoding to obtain local features T = {t1, t2,..., t M} ∈ R M×d and the global representation g T ; the masked report sequence is input into the report encoder E txt to obtain the masked report representation where M is the number of tokens (words) in the report sequence and d is the feature dimension.

[0034] BioClinicalBERT as the report encoder E txt processes the report sequence X T as follows: The report sequence X T is tokenized and converted into a token sequence, and a special [CLS] token is added at the beginning of the sequence; after the token sequence with the special [CLS] token added passes through multiple layers of Transformer in BioClinicalBERT, the token features output by the token sequence in the last layer of Transformer are used as local features, and the output feature of the [CLS] token is the global representation.

[0035] The masked report sequence is processed in the same way.

[0036] S4. Input the local features V of the frontal view, lateral view, and report sequence F , V L , T, and the global representation g F , g L , g T and the masked report representation into the frontal-lateral view feature integration module to obtain the masked report generation text T' and its prediction probability P;

[0037] The frontal-lateral view feature integration module includes: an attention layer, an aggregation layer, a report decoder, and a classifier; specifically, the frontal-lateral view feature integration module processes the local features V F , V L , T, and the global representation g F , g L , g T and the masked report representation as follows:

[0038] S41. Concatenate the local features V of the frontal view and lateral view F , V L to obtain the concatenated feature Concatenate the global representations g of the frontal view and lateral view F , g L to obtain the concatenated feature g of the global representation V ; where v i is the i-th patch of the concatenated feature;

[0039] S42. Input the concatenated feature V, the concatenated feature g V and the global representation g of the report sequence T into the attention layer to obtain the image-report cross-attention score i of each patch v of the concatenated feature V and the self-attention score of the image

[0040] Specifically, for each patch, first use the cross-attention mechanism to calculate the image-report cross-attention score for the frontal and lateral view features and report features Then perform self-attention calculation within the view features to obtain the self-attention score of the image The calculation is as follows:

[0041]

[0042] where Norm represents normalizing the score to the range of 0-1, t g , v gRepresents the overall information representation of the report and the image, respectively, for the global representation g of the report sequence T and the concatenated feature g of the global representation V obtained by average pooling.

[0043] S43. According to the cross-attention score of the image-report and the self-attention score of the image Calculate the final score of each patch v of the concatenated feature i and select the top N patches ranked by score k ;

[0044] The final score of each patch where the final score s of each patch i is calculated as follows:

[0045]

[0046] Select the top N patches ranked by score k and form a refined patch sequence from the selected patches representing the relatively important regions in the anterior-posterior and lateral views; where N K =α*2N is the number of important patches selected and N K <2N, α∈(0,1) is the proportion of the selected patches in the total number of patches.

[0047] S44. Input the selected patches into the aggregation layer to aggregate N K important patches to obtain an aggregated visual patch sequence where is the j-th patch, N R is the number of patches in the aggregated visual patch sequence, and j represents the index of the patch in the aggregated visual patch sequence;

[0048] The patch aggregation layer is used to adaptively aggregate patches with similar semantics to generate a more compact and semantically complete visual feature representation, thereby retaining the complementary information of the anterior-posterior and lateral views and strengthening their global consistency. Specifically, the patch aggregation layer processes the selected patches including:

[0049]

[0050] where W ij is the weight matrix of the patch aggregation layer and N R is the number of patches after aggregation (NR <N K )。

[0051] S45. Input the aggregated visual patch sequence and the mask report representation into the report decoder to obtain the masked report generated text T';

[0052] After obtaining the aggregated visual patch sequence containing the key information in the PA and LAT views , it is used to perform the masked report reconstruction task. In this process, the visual representation (i.e., the aggregated visual patch sequence) guides the generation of the masked text by providing rich context clues, helping the model generate a more accurate and context-related medical report. This not only helps capture the important details in the image but also strengthens the interaction between the visual and text modalities. Inputting the mask report representation and the aggregated visual patch sequence into the report decoder together can more effectively fuse the multi-view representation and the mask report representation.

[0053] Specifically, the report decoder consists of one layer of multi-head cross-attention layer and 4 layers of Transformer blocks;

[0054] The processing of the aggregated visual patch sequence and the mask report representation by the report decoder includes:

[0055] Taking the mask report representation as the query Q, the aggregated visual patch sequence as the key (K) and value (V), after being calculated by the multi-head cross-attention layer, it is then sent into the Transformer block, and the interaction process is as follows:

[0056]

[0057]

[0058] where T' is the masked report generated text, is the i-th token of the masked report generated text, H is the token length of the masked report generated text, LN(·) represents layer normalization, and CrossAtt(·) is the multi-head cross-attention layer.

[0059] S46. Input the masked report generated text T' into the classifier to obtain the prediction probability P of each token in the masked report generated text T'.

[0060] For each Use an MLP classifier to predict its original tokens. The process of mask report recovery can be expressed as follows:

[0061]

[0062] where |V| is the number of words in the BioClinicalBERT vocabulary, is the i-th token in the text generated for the mask report is the probability of the j-th word in the BioClinicalBERT vocabulary. The BioClinicalBERT vocabulary can be obtained by loading the tokenizer of the corresponding BioClinicalBERT model.

[0063] S5. Input the local features V F 、V L 、T of the frontal view, lateral view, and report sequence into the frontal-lateral feature alignment module to obtain the fine-grained representations F and L of the frontal view and lateral view;

[0064] The frontal view and lateral view provide observations of the same anatomical structure from different angles, which makes them both complementary and somewhat consistent in information expression. To strengthen the information consistency between different views and enhance the mutual cooperation between the frontal and lateral views, while mining the unique features of each view, the frontal-lateral feature alignment module ensures that the representations of the same instance in different views are as consistent as possible.

[0065] The frontal-lateral feature alignment module includes: a frontal feature alignment module and a lateral feature alignment module; the processing of the local features of the frontal view and the local features of the lateral view by the frontal-lateral feature alignment module includes:

[0066] S51. Use the local feature V F of the frontal view as the value vector and key vector, and use the local feature T of the report sequence as the query vector; input the query vector, key vector, and value vector into the frontal feature alignment module to obtain the fine-grained representation F of the frontal view containing pathological information;

[0067] Using the local representation of the report as the query vector Q, relevant representations can be extracted from the local representations of the frontal and lateral views respectively, which can capture the pathological fine-grained features at different angles in the image, and the pathological information described by these features is consistent.

[0068] The frontal feature alignment module consists of one cross-attention layer and 2 transformer blocks. The extraction process of the fine-grained frontal view representation is as follows:

[0069]

[0070]

[0071] where M is the token length of the local feature T of the report sequence, LN(·) represents layer normalization, CrossAtt(·) is the multi-head cross-attention layer, and f i,j is the element at the i-th row and j-th column of the fine-grained frontal view representation F.

[0072] S52. Use the local feature V of the lateral view L as the value vector and the key vector, and use the local feature T of the text sequence as the query vector; input the query vector, the key vector, and the value vector into the lateral feature alignment module to obtain the fine-grained representation of the lateral view containing pathological information where l i,j is the element at the i-th row and j-th column of the fine-grained lateral view representation L.

[0073] S6. According to the prediction probability P, the fine-grained representations F and L of the frontal view and the lateral view, and the global representations g F 、g L 、g T of the frontal view, the lateral view, and the report sequence, calculate the value of the overall loss function, update the model parameters according to the value of the overall loss function, and when the value of the overall loss function is minimized, obtain the pre-trained medical vision-language general model.

[0074] Calculating the value of the overall loss function includes:

[0075] S61. Calculate the mask report recovery loss function value L according to the prediction probability P of the text T′ generated from the masked report MLM ;

[0076]

[0077] where is the one-hot vocabulary distribution, that is, the true label of the i-th token being the j-th vocabulary in the BioClinicalBERT vocabulary, the label of the ground-truth token is 1, and k is the index of the vocabulary in the BioClinicalBERT vocabulary.

[0078] S62. Calculate the fine-grained representation consistency loss function value L of the frontal and lateral views according to the fine-grained representations F and L of the frontal and lateral views FL ;

[0079] Since the fine-grained representations of the extracted frontal and lateral views share similar semantic content, model the similarity between the fine-grained features of the extracted frontal and lateral views to maintain the consistency between the fine-grained characterizations of the frontal and lateral views, and the loss function value LFL It is defined as follows:

[0080]

[0081] Among them, f i,j is the element at the i-th row and j-th column of the fine-grained representation F of the frontal view, l i,j is the element at the i-th row and j-th column of the fine-grained representation L of the lateral view, τ4 is the temperature parameter for maintaining consistency between the fine-grained representations of the frontal and lateral views, sim(·,·) is the similarity function, N and M are the number of rows and columns of the fine-grained representations of the frontal and lateral views respectively, and k is the index of the number of columns.

[0082] S63. Calculate the cross-modal alignment loss function value according to the global representations g F 、g L 、g T ;

[0083] The present invention hopes that the model can not only learn the fine-grained information of the pathological region, but also ensure the consistency between the features of the frontal and lateral views and the text features at the high-level semantic level, while retaining the unique information of each view. Therefore, cross-modal alignment is performed from the perspectives of global features and prototype features. The specific steps include:

[0084] S631. Calculate the global alignment loss function value L F 、g L 、g T according to the global representations g GLOBAL of the frontal view, lateral view, and report sequence;

[0085] Calculating the global alignment loss function value L GLOBAL includes:

[0086] Calculate the cosine similarity L F 、g T of the frontal view-report and the cosine similarity L f2t of the report-frontal view according to the global representations g t2f :

[0087]

[0088] Among them, is the contrastive loss from image to text, is the contrastive loss from text to image, sim(·,·) is the similarity function, which is calculated by dot product, B is the number of chest X-ray image-report data, τ1 is the instance-level temperature hyperparameter, and the loss function value of the frontal view-report alignment is calculated as follows:

[0089]

[0090] Similarly, based on the global representations g of the frontal view and the report sequence L 、g T calculate the cosine similarity L between the frontal view - report l2t and the cosine similarity L between the report - frontal view t2l ; based on the cosine similarities L l2t and L t2l calculate the loss function value L of the frontal view - report alignment LT ;

[0091] Concatenate the global representations g of the frontal view and the lateral view F 、g L and project the concatenated features to obtain the feature g (F,L) ; based on the feature g (F,L) and the global representation g T calculate the loss function value L of the alignment between the combined global representation of the frontal and lateral views and the report (F,L)T ;

[0092] Specifically, to utilize the complementarity between the frontal and lateral views, a method for aligning joint representations is introduced; concatenate the global representations g of the frontal view and the lateral view F 、g L to obtain the joint representation (i.e., ), and then project the joint representation to an appropriate dimension using the linear layer f(F,L) Finally, obtain the loss function value L of the alignment between the combined global representation of the frontal and lateral views and the report (F,L)T which is:

[0093]

[0094] where τ1 is the temperature parameter.

[0095] Combining the above losses, the global alignment loss function value L GLOBAL is:

[0096]

[0097] S632. Calculate the prototype alignment loss function value L based on the global representations g of the frontal view, lateral view, and report sequence F 、g L 、g T ; PROTO ;

[0098] Calculating the prototype alignment loss function value includes:

[0099] Construct a trainable prototype set O = {o1,...,o J}; where o j∈R d is a trainable cross-modal prototype, and J is the number of prototypes; this set of prototypes is obtained by a linearly layer with random initialization, and the linearly layer can be regarded as a matrix, and each row of the matrix can be regarded as a prototype vector.

[0100] Adopt the iterative Sinkhorn-Knopp clustering algorithm to assign each feature representation T in g to the clusters corresponding to J cross-modal prototypes, and obtain each feature representation and the soft clustering assignment code of each cross-modal prototype o j

[0101] Specifically, first input each feature representation into the linearly layer composed of the prototype set O, and the linearly layer maps the report representation to the prototype space to obtain each feature representation and the initial similarity between each prototype o j ; adopt the iterative Sinkhorn-Knopp algorithm to convert the initial similarity between each feature representation and each prototype o j into a doubly stochastic matrix, and each element in this matrix is the soft clustering assignment code of the feature representation and the cross-modal prototype o j indicating the soft assignment probability between each feature representation and each prototype; according to the soft clustering assignment code assign each feature representation to the corresponding cluster.

[0102] Calculate the cosine similarity between each feature representation F in the global representation g L of the front view and the side view respectively and each cross-modal prototype o , and calculate the probability j between each feature representation and each cross-modal prototype o j according to the cosine similarity The calculation is as follows:

[0103]

[0104] where is the probability of the i-th global representation of the front view and the j-th prototype in the prototype set, ​​The probability of the i-th global representation in the lateral view and the j-th prototype in the prototype set, where sim(·,·) represents calculating the cosine similarity between the two, τ3 is the temperature parameter for prototype alignment, and j′ is the index of the prototype.

[0105] To guide the learning of frontal and lateral view representations and provide more accurate semantic labels for images, the reported assignment code c is adopted i as the pseudo-label for training the frontal view representation and the lateral view representation This enables the model to learn the high-level semantic consistency between images and text, thereby improving the accuracy of cross-modal alignment and enhancing the ability to capture meaningful disease features. The prototype alignment loss is defined as follows:

[0106]

[0107] S633. Combining the global alignment loss function L GLOBAL and the prototype alignment loss function L PROTO , the cross-modal alignment loss function L ALIGN is obtained as follows:

[0108] L ALIGN = λ1L GLOBAL + λ2L PROTO

[0109] S64. Combining the masked report recovery loss function L MLM , the consistency loss function L FL and the cross-modal alignment loss function, the overall loss function L is obtained as follows:

[0110] L = λ3L MLM + λ4L ALIGN + λ5L FL

[0111] where λ1, λ2, λ3, λ4, λ5 are weight hyperparameters used to balance the weights associated with each loss.

[0112] After the pre-training is completed, the weights of the pre-trained medical vision-language general model are used for downstream tasks such as classification, segmentation, and object detection.

[0113] Taking the classification task as an example, a medical dataset is obtained, and a classification model is trained using the medical dataset and the pre-trained medical vision-language general model to obtain a trained classification model; classification is performed according to the trained classification model.

[0114] In one embodiment, the dataset for the downstream task is a chest X-ray image dataset. The frontal view encoder of the pre-trained medical vision-language general model is loaded, and the parameters of its backbone network are frozen to ensure that these parameters will not be updated during training. The last fully connected layer of the frontal view encoder is replaced with a linear classification layer, and the output dimension of this linear classification layer matches the number of classes of the downstream task to obtain a classification model. The classification model is trained according to the chest X-ray image dataset. During the training process, only the parameters of the newly added linear classification layer are updated, while the parameters of the backbone network are kept frozen to obtain a trained classification model. The trained classification model can be used to classify the chest X-ray images to be tested to obtain classification results. The chest X-ray image dataset is the CheXpert dataset, which contains approximately 224,000 chest X-ray films, and each image may be labeled with 5 types of lesions, namely atelectasis, cardiomegaly, consolidation, edema, and pleural effusion.

[0115] In the above manner, the model can utilize the general features extracted by the pre-trained medical vision-language general model to quickly adapt to the downstream task.

[0116] The results of the present invention are compared with those of previous pre-training methods (Randominit, ImageNet init, ConVIRT, CLoRIA, MedKLIP, KAD, MLIP, MGCA, MRM) on downstream datasets (CXR14, CheXpert, RSNA, COVIDx). When the model performs downstream tasks, the backbone remains frozen, and the results are as Figure 2 、 Figure 3 shown; among them, Randominit is random initialization, ImageNet init means using the model weights pre-trained on the large-scale image dataset ImageNet as the initial weights, ConVIRT is a contrastive vision-text representation learning method, CLoRIA is a contrastive language-image relationship alignment method, MedKLIP is a medical knowledge-enhanced language-image pre-training method, KAD is a knowledge-enhanced alignment method, MLIP is a multi-modal language-image pre-training method, MGCA is multi-granularity contrastive alignment, and MRM is a masked region modeling method; Figure 3 The results on four dataset tasks are reported, which are basically better than the existing methods, demonstrating the effectiveness of the framework of the present invention. When fine-tuning with a 1% data ratio, the accuracy ACC on the dataset COVIDx is 3.2% better than that of the second-best MRM, indicating that the model of the present invention still has excellent discrimination ability in the scenario of extremely low labeled data volume. Figure 4The performance of the model was evaluated on the segmentation and detection tasks of three datasets (SIIM, RSNA, RSNA). Notably, under the condition of a low annotation data volume of 1%, the model of the present invention was significantly superior to the existing methods in both the segmentation and object detection tasks, which verified the data efficiency of the model of the present invention when transferred to related tasks that require the extraction of fine-grained visual features. In particular, when using 100% of the data, the segmentation performance of the model of the present invention on SIIM was 6.3% higher than that of MRM, which indicated that the model of the present invention learned high-quality low-level visual features (such as shape, contour, and texture) during pre-training.

[0117] In the above embodiments, the objectives, technical solutions, and advantages of the present invention have been further described in detail. It should be understood that the above embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A medical vision-language pre-training method based on the combination of multi-view and text, characterized in that Including: S1. Construct a medical vision-language general model, which includes a view encoder, a report encoder, an anteroposterior view feature integration module, and an anteroposterior feature alignment module; S2. Obtain chest X-ray image-report data, preprocess the chest X-ray image-report data to obtain the preprocessed chest X-ray image and the report sequence X T and the report sequence of the mask The preprocessed chest X-ray image includes the frontal image X F and the lateral image X L ; S3. Input the preprocessed chest X-ray image into the view encoder to obtain the local features V of the frontal view and the lateral view F , V L and the global representation g F , g L ; Input the report sequence X T and the masked report sequence into the report encoder respectively to obtain the local feature T of the report sequence and the global representation g T and the masked report representation S4. Input the local features V of the frontal view, lateral view, and report sequence F , V L , T and the global representation g F , g L , g T and the masked report representation into the frontal and lateral view feature integration module to obtain the masked report generation text T' and its predicted probability P; S5. Input the local features V F and V L of the frontal view, lateral view, and report sequence into the frontal-lateral feature alignment module to obtain the fine-grained representations F and L of the frontal view and lateral view; S6. According to the predicted probability P, the fine-grained representations F and L of the frontal view and the lateral view, and the global representations g of the frontal view, the lateral view, and the report sequence F , g L , g T Calculate the overall loss function value, update the model parameters according to the overall loss function value, and when the overall loss function value is minimized, obtain the pre-trained medical vision-language general model.

2. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 1, characterized in that The view encoder includes an anterior view encoder and a lateral view encoder; The view encoder processes the preprocessed chest X-ray images, including: inputting the frontal image X F and the lateral image X L into the corresponding frontal view encoder and lateral view encoder respectively, to obtain the local feature V F and the global representation g F of the frontal view, as well as the local feature V L and the global representation g L of the lateral view; The anterior view encoder processes the anterior image as follows: segment the anterior image into multiple patches, linearly map each patch into a feature vector; add a special [CLS] token at the beginning of the feature vector of each patch; use a multi-layer Transformer to process the feature vectors of the patches with the [CLS] token added to obtain the output features; among the output features, the feature corresponding to the feature vector of the patch is the local feature, and the feature corresponding to the [CLS] token is the global representation; where the [CLS] token is a classification marker and the patch is a patch.

3. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 1, characterized in that The front and side view feature integration module includes: an attention layer, an aggregation layer, a report decoder, and a classifier; the front and side view feature integration module processes the local features V F 、V L 、T, and the global representations g F 、g L 、g T and the masked report representation and the processing includes: S41. Concatenate the local features V F and V L of the front view and the side view to obtain the concatenated feature Concatenate the global representations g F and g L of the front view and the side view to obtain the concatenated feature g V ; where v i is the i-th patch of the concatenated feature; where patch is a patch S42. Input the splicing feature V, the splicing feature g V and the global representation g of the report sequence T into the attention layer to obtain the image-report cross-attention scores for each patch v of the splicing feature V i and the self-attention scores of the image ​ S43. Calculate the final score of each patch v of the concatenated feature according to the cross-attention score between the image and the report and the self-attention score of the image to select the top N patches with the highest scores; where N i is the number of selected patches; k k ​​ S44. Input the selected patches into the aggregation layer for aggregation to obtain an aggregated visual patch sequence S45. Aggregate the visual patch sequence and the mask report representation as the input to the report decoder to obtain the masked report generation text T′; S46. Input the masked report generation text T′ into the classifier to obtain the prediction probability of each token in the masked report generation text T′; where the token is a word.

4. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 1, characterized in that, The anteroposterior and lateral feature alignment module includes: an anteroposterior feature alignment module and a lateral feature alignment module; the anteroposterior and lateral feature alignment module processes the local features V F , V L , T and includes: S51. Take the local feature V of the frontal view F as the value vector and the key vector, and take the local feature T of the report sequence as the query vector; input the query vector, the key vector, and the value vector into the frontal feature alignment module to obtain the frontal view fine-grained representation F; S52. Take the local feature V of the side view L as the value vector and the key vector, and take the local feature T of the text sequence as the query vector; input the query vector, the key vector, and the value vector into the side view feature alignment module to obtain the side view fine-grained representation L.

5. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 1, characterized in that, Calculating the overall loss function value includes: S61. Calculate the mask report recovery loss function value L according to the predicted probability P of generating the text T′ from the mask report MLM ; S62. Calculate the consistency loss function value L of the fine-grained representations of the front and side views based on the fine-grained representations F and L of the front and side views FL ; S63. Calculate the cross-modal alignment loss function value according to the global representation g F and g L and g T ​ S64. Combine the mask report to recover the loss function value L MLM , the consistency loss function value L FL and the cross-modal alignment loss function value to obtain the overall loss function value.

6. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 5, characterized in that, Calculate the loss function value L of the mask report recovery MLM including: where |V| is the number of words in the BioClinicalBERT vocabulary, is the true label that the i-th token of the report text T is the j-th word in the BioClinicalBERT vocabulary, is the probability that the i-th token of the masked report generated text T′ is the j-th word in the BioClinicalBERT vocabulary, H is the number of tokens in the report text T and the masked report generated text T′, and k is the index of the word in the BioClinicalBERT vocabulary.

7. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 5, characterized in that Calculate the consistency loss function value \(L\) of the fine-grained representations of the front and side views FL including: where f i,j is the element in the \(i\)-th row and \(j\)-th column of the frontal view fine-grained representation \(F\), \(l i,j is the element in the \(i\)-th row and \(j\)-th column of the lateral view fine-grained representation \(L\), \(\tau_4\) is the temperature parameter for maintaining consistency between the frontal and lateral view fine-grained representations, \(sim(\cdot,\cdot)\) is the similarity function, \(N\) and \(M\) are the number of rows and columns of the frontal and lateral view fine-grained representations respectively, and \(k\) is the index of the number of columns.

8. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 5, characterized in that, Calculating the cross-modal alignment loss function value includes: S631. Calculate the global alignment loss function value L based on the global representations g of the frontal view, lateral view, and report sequence F g L g T GLOBAL ;​ S632. Calculate the prototype alignment loss function value L based on the global representations g of the anterior-posterior view, lateral view, and report sequence F 、g L 、g T ; PROTO ; S633. Combine the global alignment loss function value L GLOBAL and the prototype alignment loss function value L PROTO to obtain the cross-modal alignment loss function value.

9. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 8, characterized in that Calculate the value L of the global alignment loss function GLOBAL including: Based on the global representation g of the frontal view and the report sequence F 、g T Calculate the cosine similarity L between the frontal view and the report f2t and the cosine similarity L between the report and the frontal view t2f , and based on the cosine similarities L f2t and L t2f Calculate the loss function value L for the alignment between the frontal view and the report FT ; According to the global representation g of the lateral view and the report sequence L 、g T Calculate the cosine similarity L between the lateral view and the report l2t and the cosine similarity L between the report and the lateral view t2l , according to the cosine similarities L l2t and L t2l Calculate the loss function value L of the lateral view-report alignment LT ; Concatenate the global representations g of the anterior-posterior view and the lateral view F , g L , and project the concatenated features to obtain the feature g (F ,L) ; According to the feature g (F,L) and the global representation g T , calculate the alignment loss function value L of the joint global representation of the anterior-posterior view and the lateral view and the report (F,L)T ; Combined with the loss function value L FT , the loss function value L LT , and the alignment loss function value L (F,L)T , the global alignment loss function value L GLOBAL is obtained.

10. A medical vision-language pre-training method based on the combination of multi-view and text according to claim 8, characterized in that, Calculate the value of the prototype alignment loss function L PROTO including: Construct a set O of trainable cross-modal prototypes O = {o1,..., o J}; where o j is the j-th trainable cross-modal prototype, and J is the number of prototypes; Use the iterative Sinkhorn-Knopp clustering algorithm to assign each feature representation T in the global representation g to the clusters corresponding to J cross-modal prototypes, obtaining the soft clustering assignment codes for each feature representation and each cross-modal prototype o j ​ Calculate the global representations g F and g L for each feature representation between each cross-modal prototype o j and calculate the cosine similarity between each feature representation and each cross-modal prototype o j to obtain the probability According to probability and the soft clustering assignment code Calculate the prototype alignment loss function value.

Citation Information

Cited By

  • Image privacy risk grading method and device based on cross-layer attention aggregation and ordinal constraint

    CN121121251A

  • Bone tumor fine-grained classification model training and classification method and device

    CN121708408A

  • A bone tumor fine-grained classification model training, classification method and device

    CN121708408B