Artificial Intelligence Model Training Method for Chest Imaging, Electronic Device, and Storage Medium
By using the mean teacher model for multi-view distillation and contrast learning in the chest imaging artificial intelligence model, the limitations of existing models in utilizing multi-view information and processing diversified report descriptions are solved, and higher visual and language comprehension and zero-sample learning abilities are achieved.
Patent Information
- Application Number
- CN202510467393.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing basic chest imaging models have limitations when utilizing multi-perspective information and processing diverse medical report descriptions, and it is difficult to fully explore the complex correlation and differential information between different perspectives, resulting in limitations in diagnostic accuracy and comprehensive understanding of the lesions.
By designing a chest image artificial intelligence model training method, multi-view distillation is performed using the mean teacher model, combining contrast learning and cross-modal contrast loss, it effectively utilizes the multi-view information of chest image and the diversified description of medical reports to extract and integrate features under different perspectives and text descriptions.
The visual and language comprehension ability of the model is improved, the image and text comprehension ability of the same disease is enhanced, and excellent zero-sample learning ability in downstream tasks is achieved.
Smart Images

Figure CN120014386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and specifically relates to a method for training an artificial intelligence model for chest imaging, an electronic device, and a storage medium. Background Art
[0002] In recent years, the basic model has emerged in the field of artificial intelligence. Its aim is to build a pre-trained model with wide applicability and strong representation ability. By training on a large amount of data, it learns general knowledge and feature representations, and then can be fine-tuned on various downstream tasks to meet specific needs. The research and development of the basic model for chest imaging is in line with this trend. It is committed to using a large amount of chest imaging data and related clinical information for pre-training, so as to learn the deep semantic features of chest imaging, the commonalities and differences of different diseases in images, etc. Once such a basic model is established, it can be quickly adapted and optimized for various specific application scenarios such as chest disease diagnosis, disease severity assessment, and disease progression prediction. It is expected to significantly improve the accuracy, efficiency, and automation of chest imaging analysis, provide strong technical support for the precise medical treatment of chest diseases, and greatly promote the intelligent development process in the field of medical imaging.
[0003] Although certain progress has been made in the research on chest imaging basic models, there are still many defects and deficiencies. Regarding the utilization of multi-view information in chest X-rays, traditional research often focuses only on a single view or simply splices and fuses multi-view information; this approach fails to fully explore the complex internal relationships and differential information between different views, ignoring the uniqueness and complementarity of each view's images in disease manifestation. For example, the posteroanterior chest X-ray and the lateral chest X-ray have different emphases in showing certain lung lesions. Existing methods are difficult to accurately extract and effectively integrate the subtle differences in lesion characteristics from these different views, resulting in limitations in diagnostic accuracy and comprehensive understanding of the lesions. Moreover, during the model training process, there is a lack of an efficient and targeted knowledge transfer mechanism for multi-view data, which cannot well guide the model to learn discriminative and representative features from different-view data, thereby affecting the model's ability to handle complex situations of chest diseases. In addition, in the current field of medical imaging basic model research, there are obvious deficiencies in dealing with the situation where there are different description reports for the same disease in different image-text pairs. When different medical staff write disease reports, due to differences in personal habits, professional backgrounds, and clinical experiences, the descriptions of the same disease will vary significantly in terms of vocabulary usage, sentence structure, and level of detail. For example, for pulmonary inflammation, some reports may focus on describing the size and location of the inflammatory area, while others will mention more about the imaging features of the inflammation or the relevant symptoms of the patient. Traditional models are difficult to perform unified semantic understanding and structured processing on these diverse and scattered description information, unable to extract the common core disease information, and thus, when using image-report pairs for model training, they are easily interfered by description differences, resulting in inaccurate and incomplete disease manifestations learned by the model. Summary of the Invention
[0004] Based on this, the present invention provides a method for training an artificial intelligence model for chest imaging, an electronic device, and a storage medium, which at least solve one problem in the prior art.
[0005] In a first aspect, the present invention provides a method for training an artificial intelligence model for chest imaging, which includes the following steps:
[0006] Obtain a chest imaging data set and its corresponding report text data set, where the chest imaging data set includes an image pair composed of a first-view chest X-ray and a second-view chest X-ray, and the report text data set includes a text pair composed of an original report text and a polished text derived from the original report text;
[0007] Train a mean teacher model using the chest imaging dataset and its corresponding report text dataset. The mean teacher model includes a student model and a teacher model. The student model includes a student image encoder and a student text encoder, and the teacher model includes a teacher image encoder and a teacher text encoder.
[0008] Perform contrastive learning by comparing the first-view chest X-ray features extracted by the student image encoder based on the image pair with the second-view chest X-ray features extracted by the teacher image encoder based on the image pair, and also compare the second-view chest X-ray features extracted by the student image encoder based on the image pair with the first-view chest X-ray features extracted by the teacher image encoder based on the image pair.
[0009] Perform contrastive learning by comparing the original report text features extracted by the student text encoder based on the text pair with the polished text features extracted by the teacher text encoder based on the text pair, and also compare the polished text features extracted by the student text encoder based on the text pair with the original report text features extracted by the teacher text encoder based on the text pair.
[0010] Perform contrastive learning by comparing the first-view chest X-ray features and / or second-view chest X-ray features extracted by the student image encoder based on the image pair with the original report text features and / or polished text features extracted by the teacher text encoder based on the text pair, and also compare the original report text features and / or polished text features extracted by the student text encoder based on the text pair with the first-view chest X-ray features and / or second-view chest X-ray features extracted by the teacher image encoder based on the image pair.
[0011] Update the parameters of the mean teacher model according to the total loss function.
[0012] It should be noted that the image pair composed of the first-view chest X-ray and the second-view chest X-ray can be an image pair composed of a posteroanterior chest X-ray and its corresponding lateral chest X-ray, or an image pair composed of a posteroanterior chest X-ray or a lateral chest X-ray and its enhanced image obtained by image enhancement. There is a corresponding relationship between the original report text and the posteroanterior chest X-ray or the lateral chest X-ray, generally including a literal description of the interpretation of the posteroanterior chest X-ray or the lateral chest X-ray. The polished text can be obtained by optimizing the original report text in the form of a question-and-answer using a large language model.
[0013] In some optional embodiments, the first-view chest X-ray and the second-view chest X-ray have been standardized. The standardization process may include image cropping and normalization operations to ensure that chest X-rays from different sources and shooting conditions have a unified size and gray level range for subsequent model processing.
[0014] In some alternative embodiments, the total loss function is a weighted function of multi-view image contrast loss, multi-view text contrast loss, cross-modal image-text contrast loss, and cross-modal instance similarity consistency loss, as shown in the following formula:
[0015]
[0016] Wherein, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal image-text contrast loss, represents the cross-modal instance similarity consistency loss; , , , respectively represent weight coefficients.
[0017] In some alternative embodiments, the contrast loss in the multi-view image contrast loss, multi-view text contrast loss, and cross-modal image-text contrast loss is calculated by the following formula:
[0018]
[0019] Wherein, , and respectively represent features of different modalities (image modality or text modality), represents the temperature hyperparameter, and respectively represent features of different modalities in a batch, n represents the number of samples in a batch, T represents the matrix transpose operation.
[0020] In some alternative embodiments, the multi-view image contrast loss is calculated by the following formula:
[0021]
[0022] Wherein, represents the multi-view image contrast loss, ([[]] , ) and ( , ) respectively represent the features of the student model and the teacher model under the first view and the second view.
[0023] In some alternative embodiments, the multi-view text contrast loss is calculated by the following formula:
[0024]
[0025] Among them, represents the multi-view text comparison loss, ( , ) and ( , ) respectively represent the features of the student model and the teacher model under the original report text and the polished text.
[0026] In some alternative embodiments, the cross-modal image-text comparison loss is calculated by the following formula:
[0027]
[0028] Among them,
[0029]
[0030] Among them, represents the cross-modal image-text comparison loss.
[0031] In some alternative embodiments, the cross-modal instance similarity consistency loss is calculated by the following formula:
[0032]
[0033] Among them,
[0034]
[0035]
[0036] Among them, represents the cross-modal instance similarity consistency loss; represents the similarity matrix of the first-view chest X-ray features, represents the similarity matrix of the original report text features; represents the features of the specified modality, represents the temperature hyperparameter; represents the similarity of the feature similarity matrix after removing the diagonal.
[0037] In a second aspect, the present invention provides an electronic device, which includes:
[0038] At least one processor;
[0039] And a memory communicatively connected to the at least one processor;
[0040] Wherein, the memory stores instructions that, when executed by the at least one processor, implement the chest imaging artificial intelligence model training method as described above.
[0041] In a third aspect, the present invention provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the chest imaging artificial intelligence model training method described above.
[0042] Due to the above technical solutions, the embodiments of the present invention have at least the following beneficial effects:
[0043] Effectively utilize multi-view information of chest images for pre-training, improving the model's visual and language understanding capabilities;
[0044] Correlation learning can effectively enhance the model's ability to understand images and texts of the same disease;
[0045] Effectively train the model and can be transferred to downstream tasks, demonstrating excellent zero-shot learning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic flowchart of the chest imaging artificial intelligence model training method in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the concept of the present invention and the technical effects produced, so as to fully elaborate the purpose, solution and effects of the present invention.
[0048] The embodiments of the present invention provide a chest imaging artificial intelligence model training method, which mainly includes the steps of data preprocessing, feature extraction, multi-view distillation, and model parameter optimization.
[0049] In the data preprocessing step, a chest image dataset containing posteroanterior chest radiographs and lateral chest radiographs is collected, and each chest radiograph is standardized, including image cropping and normalization operations, to ensure that chest radiographs under different sources and shooting conditions have a unified size and gray level range for subsequent model processing. If there is only a posteroanterior chest radiograph, a certain image enhancement is performed on the posteroanterior chest radiograph, and it is used as an image pair with the original image. For the original report text corresponding to the chest radiograph, a large language model is used to optimize it in the form of questions and answers to obtain a polished version of the text (polished text) and form a text pair with the original report text.
[0050] In the feature extraction step, an image encoder is designed to extract the features of posteroanterior chest radiographs and lateral chest radiographs. The posteroanterior chest radiograph and the lateral chest radiograph are input into the image encoder as an image pair to obtain the features of the posteroanterior chest radiograph and the lateral chest radiograph. For the medical text report corresponding to the image, the Tokenizer in the BERT model is used to convert the words in the text into vector representations, and then the BERT model processes the text vector sequence to extract the semantic features of the text.
[0051] In the multi-view distillation step, a multi-view distillation strategy of the Mean Teacher (MT) model is adopted, that is, both the image encoder and the text encoder have a teacher model with momentum update. The architectures of the teacher model and the student model in the Mean Teacher model are the same to achieve self-distillation. Specifically, image pairs are simultaneously fed into the student image encoder and the teacher image encoder. The frontal image features obtained by the student encoder are used for contrastive learning with the lateral image features obtained by the teacher model, and at the same time, the lateral image features obtained by the student encoder are used for contrastive learning with the frontal image features obtained by the teacher model. This can effectively utilize the teacher model to guide the student model to better learn multi-view features. Text pairs are simultaneously fed into the student text encoder and the teacher text encoder. The original text features obtained by the student encoder are used for contrastive learning with the polished text features obtained by the teacher model, and at the same time, the polished text features obtained by the student encoder are used for contrastive learning with the original text features obtained by the teacher model. In addition, the same is true for cross-modal learning. The frontal image features and lateral image features obtained by the student image encoder are respectively used for contrastive learning with the original text features and polished text features obtained by the teacher text encoder, and the polished text features and original text features obtained by the student text encoder are respectively used for contrastive learning with the frontal image features and lateral image features obtained by the teacher image encoder.
[0052] In the model parameter optimization step, an image correlation matrix is constructed by calculating the cosine similarity between different image feature vectors. This matrix reflects the similarity degree of visual features between different chest images. For example, images with similar lesion types or imaging patterns have higher similarity values in the matrix. Similarly, a correlation matrix of text features is calculated. A text correlation matrix is constructed using the similarity metric between text semantic feature vectors. This matrix reflects the semantic content association between different medical text reports. For example, different description texts of the same disease will have a certain correlation in the matrix. A modality association graph is constructed based on the correlation matrices of images and texts. By calculating the correlation consistency loss between the two modalities, the model's ability to understand images and texts of the same disease can be effectively enhanced. Through the backpropagation algorithm, based on the designed loss functions, including multi-view image contrast loss, multi-view text contrast loss, cross-modal image-text contrast loss, and cross-modal instance similarity consistency loss, the parameters of the entire model are then optimized and trained. During the training process, the model parameters are continuously adjusted to minimize the loss function value, enabling the student model to effectively acquire cross-modal knowledge from the teacher model and improve the model's performance in various downstream tasks.
[0053] Figure 1 The flowchart of the method for training an artificial intelligence model for chest images according to an embodiment of the present invention is shown. Feeding the image-text data into the Mean Teacher model for training includes the following steps:
[0054] Step S1: Feed the image pair and the text pair into the image encoder and the text encoder respectively to obtain corresponding features;
[0055] Step S2: Align the frontal image features and the lateral image features, and at the same time align the original report text and the polished text features; through the way of contrastive learning, use the contrastive loss to complete the feature alignment process; finally, align the image features and the text features pairwise, and this process is also completed through the contrastive loss; among them, the calculation process of the contrastive loss is as follows:
[0056]
[0057] Among them, 、 and respectively represent features of different modalities (image modality or text modality), represents the temperature hyperparameter, and respectively represent features of different modalities in a batch, n represents the number of samples in a batch, T represents the matrix transpose operation;
[0058] Step S3: Construct a correlation graph between the image features and the text features, and calculate the consistency loss.
[0059] Specifically, step S1 may include step S11 and step S12.
[0060] Step S11: Feed the frontal image and the lateral image as paired data into the student image encoder and the teacher image encoder in the mean teacher model to obtain four features:
[0061]
[0062] Among them, and respectively represent the frontal image and the lateral image; if there is only the frontal image, then represents the after rotation, flipping and other enhancements. and respectively represent the frontal image features and the lateral image features obtained by the student image encoder. and respectively represent the frontal image features and the lateral image features obtained by the teacher image encoder.
[0063] Step S12: Feed the original report text and the polished text as paired data into the student text encoder of the mean teacher model and the teacher text encoder Four features are obtained as follows:
[0064]
[0065] Among them, and represent the original report text and the polished text respectively; and represent the features of the original report text and the polished text obtained by the student text encoder respectively, and represent the features of the original text and the polished text obtained by the teacher text encoder respectively.
[0066] Step S2 may include steps S21 to S23.
[0067] Step S21: Align the image features obtained in step S11 using contrastive learning, where the positive sample is the current image pair and the negative sample is other image features. The specific process is as follows:
[0068]
[0069] Among them, represents the multi-view image contrast loss, ([[]] , ) and ([[]] , ) represent the features of the student model and the teacher model under different views (the first view and the second view) respectively.
[0070] Step S22: Align the text features obtained in step S12 using contrastive learning, where the positive sample is the current text pair and the negative sample is other image features. The specific process is as follows:
[0071]
[0072] Among them, represents the multi-view text contrast loss, ([[]] , ) and ([[]] , ) represent the features of the student model and the teacher model under the original report text and the polished text respectively.
[0073] Step S23: Align the image features obtained in step S11 and the text features obtained in step S12 using contrastive learning, where the positive sample is the current image-text pair feature, the negative sample of the image feature is the remaining text features, and the negative sample of the text feature is other image features. The specific process is as follows:
[0074]
[0075] Among them, represents the cross-modal contrast loss between image and text.
[0076] Step S3 may include steps S31 to S33.
[0077] Step S31: Process the image features obtained in step S11, and calculate the correlation matrix of the image features under the same batch. The specific process is as follows:
[0078]
[0079] Among them, represents the features of the specified modality, represents the temperature hyperparameter. Since the diagonal is the dot product of the feature and itself, after eliminating the diagonal, the final correlation matrix can be obtained as shown in the following formula:
[0080]
[0081] Among them, represents the similarity of the feature similarity matrix after removing the diagonal.
[0082] Step S32: Process the text features obtained in step S12, and calculate the correlation matrix of the image features under the same batch. The calculation method is the same as that of the image.
[0083] Step S33: Use KL-Divergence to perform consistency learning on the image feature correlation matrix and the text feature correlation matrix. The specific process is as follows:
[0084]
[0085] Among them, represents the cross-modal instance similarity consistency loss, represents the similarity matrix of the first-view chest X-ray features, represents the similarity matrix of the original report text features.
[0086] Finally, the total loss function of the model can be expressed by the following formula:
[0087]
[0088] Among them, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal contrast loss between image and text, Denotes the cross-modal instance similarity consistency loss; 、 、 、 Denote the weight coefficients respectively.
[0089] Example 1
[0090] According to the method provided by the present invention, pre-training was carried out on the MIMIC-CXR dataset, and performance tests of zero-shot classification were carried out on the Vindr-CXR dataset and the Chexpert5×200 dataset. The Vindr-CXR dataset contains 18,000 chest images, and its label distribution is 22 local labels and 6 global labels, which are marked by experienced radiologists. In this example, the image size was adjusted to 248×248 and randomly cropped to 224×224. The CheXpert5×200 dataset is a subset of the CheXpert dataset and is used for the classification of 5 diseases, with 200 images in each category. In this example, the same processing method was used for its preprocessing.
[0091] At the same time, the method of this example was compared with the state-of-the-art method CXR-CLIP. From the objective evaluation index, a comparison experiment was carried out using the AUC evaluation index. The experimental results are shown in Table 1.
[0092] Table 1 AUC results of zero-shot classification on the Vindr-CXR dataset and Chexpert5×200
[0093]
[0094] It can be seen that on the Vindr-CXR dataset, the AUC of the method in this example increased by 6.3% compared with the CXR-CLIP of the ResNet-50 version and by 6.7% compared with the CXR-CLIP of the Swin-Transformer version. On the Chexpert5×200 dataset, the AUC of the method in this example increased by 3.0% compared with the CXR-CLIP of the ResNet-50 version and by 2.5% compared with the CXR-CLIP of the Swin-Transformer version.
[0095] Example 2
[0096] According to the method provided by the present invention, pre-training was carried out on the MIMIC-CXR dataset, and performance tests of cross-modal image-text retrieval were carried out on the MIMIC-CXR dataset.
[0097] Meanwhile, the method of this embodiment was compared with the state-of-the-art method CXR-CLIP. From the objective evaluation indicators, the recall rates R@1, R@5, and R@10 were used as evaluation indicators for the comparative experiment. The experimental results are shown in Table 2.
[0098] Table 2 Cross-modal retrieval results of text and image in MIMIC-CXR dataset
[0099]
[0100] It can be seen that the method of this embodiment is the highest in all three indicators of R@1, R@5, and R@10, which are 2.6%, 6.1%, and 6.6% higher than those of the ResNet-50 version of CXR-CLIP, and 2.0%, 5.7%, and 5.3% higher than those of the Swin-Transformer version of CXR-CLIP, respectively.
[0101] As mentioned above, it is only a preferred embodiment of the present invention. The present invention is not limited to the above-mentioned implementation manners. As long as the same or equivalent means are used to achieve the technical effects of the present invention, they should fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and changes can be made to its technical solutions and / or implementation manners.
Claims
1. A chest image artificial intelligence model training method, characterized in that: The following steps are involved: Acquire a chest image dataset and a corresponding report text dataset, wherein the chest image dataset includes an image pair consisting of a first-view chest X-ray and a second-view chest X-ray, and the report text dataset includes a text pair consisting of an original report text and a polished text derived from the original report text; Using the chest image dataset and its corresponding report text dataset to train a mean teacher model, the mean teacher model includes a student model and a teacher model, wherein the student model includes a student image encoder and a student text encoder, and the teacher model includes a teacher image encoder and a teacher text encoder; Compare and learn the chest X-ray features of the first view extracted by the student image encoder based on the image pair with the chest X-ray features of the second view extracted by the teacher image encoder based on the image pair, and compare and learn the chest X-ray features of the second view extracted by the student image encoder based on the image pair with the chest X-ray features of the first view extracted by the teacher image encoder based on the image pair; Comparatively learning the original report text features extracted by the student text encoder based on the text pair and the polished text features extracted by the teacher text encoder based on the text pair, and comparatively learning the polished text features extracted by the student text encoder based on the text pair and the original report text features extracted by the teacher text encoder based on the text pair; Compare and learn the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the student image encoder based on the image pair with the original report text features and / or the polished text features extracted by the teacher text encoder based on the text pair, and compare and learn the original report text features and / or the polished text features extracted by the student text encoder based on the text pair with the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the teacher image encoder based on the image pair; The parameters of the mean teacher model are updated according to the total loss function.
2. The method according to claim 1, characterized in that The first-view chest X-ray and the second-view chest X-ray have been standardized.
3. The method according to claim 1, characterized in that The total loss function is a weighted function of multi-view image contrast loss, multi-view text contrast loss, image-text cross-modal contrast loss, and cross-modal instance similarity consistency loss, as shown in the following formula: in, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal contrast loss between images and texts, represents the consistent loss of cross-modal instance similarity; , , , They represent weight coefficients respectively.
4. The method according to claim 3, characterized in that The contrast loss in the multi-view image contrast loss, the multi-view text contrast loss, and the image-text cross-modal contrast loss is calculated by the following formula: in, , and Represent the characteristics of different modes, represents the temperature hyperparameter, and Respectively represent the characteristics of different modes in a batch, n represents the number of samples in a batch, T Represents a matrix transpose operation.
5. The method according to claim 4, characterized in that The multi-view image contrast loss is calculated by the following formula: in, represents the multi-view image contrast loss, ( , )and( , ) represent the characteristics of the student model and the teacher model in the first perspective and the second perspective respectively.
6. The method according to claim 4, characterized in that The multi-view text contrast loss is calculated by the following formula: in, represents the multi-view text contrast loss, ( , )and( , ) represent the features of the student model and the teacher model under the original report text and the polished text respectively.
7. The method according to claim 4, characterized in that The image-text cross-modal contrast loss is calculated by the following formula: in, in, Represents the cross-modal contrast loss between images and text.
8. The method according to claim 4, characterized in that The cross-modal instance similarity consistency loss is calculated as follows: in, in, represents the consistent loss of cross-modal instance similarity; Represents the similarity matrix of the first-person chest radiograph features, A similarity matrix representing the text features of the original report; Represents the characteristics of the specified mode, represents the temperature hyperparameter; Represents the similarity of the feature similarity matrix after removing the diagonal.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions, which, when executed by at least one processor, implement the chest imaging artificial intelligence model training method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed by the processor, the chest imaging artificial intelligence model training method as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Text processing method and device, electronic equipment and storage medium
CN117952071A
Semi-supervised nasopharyngeal carcinoma segmentation method based on image text contrast learning
CN118297960A