Chest image artificial intelligence model training method, electronic equipment and storage medium

By using the mean teacher model for multi-view distillation and comparison learning in the chest imaging artificial intelligence model, the problem of insufficient utilization of multi-view information and diversified report descriptions in the prior art is solved, and a more accurate and comprehensive chest imaging diagnosis and lesion understanding is achieved.

CN120014386AActive Publication Date: 2025-05-16NANCHANG UNIV

Patent Information

Application Number
CN202510467393.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing basic chest imaging models have shortcomings in using multi-perspective information and processing diversified medical report descriptions. It is difficult to fully explore the complex intrinsic correlations and different information between different perspectives, and it is difficult to unify the description information of different medical personnel, resulting in limitations in diagnostic accuracy and comprehensive understanding of the lesions.

Method used

By designing a chest image artificial intelligence model training method, multi-view distillation is performed using the mean teacher model, combining contrast learning and cross-modal learning, the multi-view information of chest image and medical report text information are effectively used to extract and integrate feature information under different perspectives and descriptions.

Benefits of technology

It improves the visual and language understanding ability of the model, enhances the image and text understanding ability of the same disease, realizes effective feature extraction and integration, and improves the accuracy and comprehensiveness of the model in diagnosis and lesion understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014386A_ABST
    Figure CN120014386A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, in particular to a chest image artificial intelligence model training method, electronic equipment and a storage medium. According to the invention, a student image encoder and a teacher image encoder are used to carry out comparative learning on extracted first visual angle chest radiograph features and / or second visual angle chest radiograph features based on an image pair; respectively carrying out comparative learning on the extracted original report text features and / or retouching text features based on texts by a student text encoder and a teacher text encoder, and respectively carrying out comparative learning on the original report text features and / or retouching text features and the first visual angle chest radiograph features and / or the second visual angle chest radiograph features; and the consistency between the prediction results of the student model and the teacher model is measured through the correlation matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a chest imaging artificial intelligence model training method, electronic equipment and storage medium. Background Art

[0002] In recent years, basic models have emerged in the field of artificial intelligence. They aim to build pre-trained models with broad applicability and strong representation capabilities. By training on large-scale data, general knowledge and feature representations can be learned, and then fine-tuned on various downstream tasks to meet specific needs. The development of the basic model for chest imaging is in line with this trend. It is committed to using massive chest imaging data and related clinical information for pre-training, so that it can learn the deep semantic features of chest imaging, the commonalities and differences of different diseases in imaging, etc. Once such a basic model is established, it can be quickly adapted and optimized for a variety of specific application scenarios such as chest disease diagnosis, disease severity assessment, and disease progression prediction. It is expected to significantly improve the accuracy, efficiency, and automation of chest image analysis, provide strong technical support for precision medicine for chest diseases, and greatly promote the development of intelligentization in the field of medical imaging.

[0003] Although the research on the basic model of chest imaging has made some progress, there are still many defects and shortcomings. For the use of multi-view information of chest X-rays, traditional research often focuses on a single view or simply splices and fuses multi-view information; this method fails to fully explore the complex internal correlation and difference information between different viewpoints, and ignores the uniqueness and complementarity of each viewpoint image in disease representation. For example, the frontal chest X-ray and lateral chest X-ray have different focuses on the display of certain lung lesions. Existing methods are difficult to accurately extract and effectively integrate the subtle differences in the characteristics of lesions under these different viewpoints, resulting in limitations in diagnostic accuracy and comprehensive understanding of lesions. Moreover, in the process of model training, there is a lack of efficient and targeted knowledge transfer mechanism for multi-view data, which cannot well guide the model to learn discriminative and representative features from different viewpoint data, thereby affecting the model's ability to cope with complex chest diseases. In addition, in the current research field of basic models of medical imaging, there are obvious deficiencies in handling the situation where different descriptions and reports of the same disease exist in different image-text pairs. When different medical personnel write disease reports, due to differences in personal habits, professional backgrounds, and clinical experience, the descriptions of the same disease may vary significantly in terms of vocabulary usage, sentence structure, and level of detail. For example, for lung inflammation, some reports may focus on describing the size and location of the inflammation area, while others may mention more imaging features of inflammation or related symptoms of the patient. Traditional models find it difficult to perform unified semantic understanding and structural processing on these diverse and scattered descriptive information, and are unable to extract the common core disease information. As a result, when using image-report pairs for model training, they are easily disturbed by differences in descriptions, resulting in the disease representation learned by the model being inaccurate and incomplete. Summary of the invention

[0004] Based on this, the present invention provides a chest image artificial intelligence model training method, electronic device and storage medium, which at least solve one problem in the prior art.

[0005] In a first aspect, the present invention provides a chest image artificial intelligence model training method, which comprises the following steps: Acquire a chest image dataset and a corresponding report text dataset, wherein the chest image dataset includes an image pair consisting of a first-view chest X-ray and a second-view chest X-ray, and the report text dataset includes a text pair consisting of an original report text and a polished text derived from the original report text; Using the chest image dataset and its corresponding report text dataset to train a mean teacher model, the mean teacher model includes a student model and a teacher model, wherein the student model includes a student image encoder and a student text encoder, and the teacher model includes a teacher image encoder and a teacher text encoder; Compare and learn the chest X-ray features of the first view extracted by the student image encoder based on the image pair with the chest X-ray features of the second view extracted by the teacher image encoder based on the image pair, and compare and learn the chest X-ray features of the second view extracted by the student image encoder based on the image pair with the chest X-ray features of the first view extracted by the teacher image encoder based on the image pair; Comparatively learning the original report text features extracted by the student text encoder based on the text pair and the polished text features extracted by the teacher text encoder based on the text pair, and comparatively learning the polished text features extracted by the student text encoder based on the text pair and the original report text features extracted by the teacher text encoder based on the text pair; Compare and learn the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the student image encoder based on the image pair with the original report text features and / or the polished text features extracted by the teacher text encoder based on the text pair, and compare and learn the original report text features and / or the polished text features extracted by the student text encoder based on the text pair with the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the teacher image encoder based on the image pair; The parameters of the mean teacher model are updated according to the total loss function.

[0006] It should be noted that the image pair formed by the first-view chest radiograph and the second-view chest radiograph can be an image pair formed by an anteroposterior chest radiograph and its corresponding lateral chest radiograph, or an image pair formed by an anteroposterior chest radiograph or a lateral chest radiograph and an enhanced image obtained by image enhancement. There is a corresponding relationship between the original report text and the anteroposterior chest radiograph or the lateral chest radiograph, which generally includes a textual description of the interpretation of the anteroposterior chest radiograph or the lateral chest radiograph. The polished text can be obtained by optimizing the original report text in the form of question and answer using a large language model.

[0007] In some optional embodiments, the first-view chest radiograph and the second-view chest radiograph are subjected to standardization processing, which may include image cropping and normalization operations to ensure that chest radiographs from different sources and shooting conditions have a uniform size and grayscale range for subsequent model processing.

[0008] In some optional embodiments, the total loss function is a weighted function of multi-view image contrast loss, multi-view text contrast loss, image-text cross-modal contrast loss, and cross-modal instance similarity consistency loss, as shown in the following formula: in, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal contrast loss between images and texts, represents the consistent loss of cross-modal instance similarity; , , , They represent weight coefficients respectively.

[0009] In some optional embodiments, the contrast loss in the multi-view image contrast loss, the multi-view text contrast loss, and the image-text cross-modal contrast loss is calculated by the following formula: in, , and Respectively represent the features of different modalities (image modality or text modality), represents the temperature hyperparameter, and Respectively represent the characteristics of different modes in a batch, n represents the number of samples in a batch, T Represents a matrix transpose operation.

[0010] In some optional embodiments, the multi-view image contrast loss is calculated by the following formula: in, represents the multi-view image contrast loss, ( , )and( , ) represent the characteristics of the student model and the teacher model in the first perspective and the second perspective respectively.

[0011] In some optional embodiments, the multi-view text contrast loss is calculated by the following formula: in, represents the multi-view text contrast loss, ( , )and( , ) represent the features of the student model and the teacher model under the original report text and the polished text respectively.

[0012] In some optional embodiments, the image-text cross-modal contrast loss is calculated by the following formula: in, in, Represents the cross-modal contrast loss between images and text.

[0013] In some optional embodiments, the cross-modal instance similarity consistency loss is calculated by the following formula: in, in, represents the consistent loss of cross-modal instance similarity; Represents the similarity matrix of the first-person chest radiograph features, A similarity matrix representing the text features of the original report; represents the characteristics of the specified mode, represents the temperature hyperparameter; Represents the similarity of the feature similarity matrix after removing the diagonal.

[0014] In a second aspect, the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; Among them, the memory stores instructions, and when the instructions are executed by at least one processor, they implement the chest imaging artificial intelligence model training method as described above.

[0015] In a third aspect, the present invention provides a computer-readable storage medium storing instructions, which, when executed by a processor, implement the chest imaging artificial intelligence model training method as described above.

[0016] Due to the adoption of the above technical solution, the embodiments of the present invention have at least the following beneficial effects: Effectively utilize multi-view information of chest images for pre-training, improving the model's visual and language understanding capabilities; Correlation learning can effectively enhance the model's ability to understand images and texts of the same disease; The model is effectively trained and can be transferred to downstream tasks, demonstrating excellent zero-sample learning capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the process of chest image artificial intelligence model training method in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following is a clear and complete description of the concept of the present invention and the technical effects produced, so as to fully explain the purpose, scheme and effect of the present invention.

[0019] An embodiment of the present invention provides a chest imaging artificial intelligence model training method, which mainly includes the steps of data preprocessing, feature extraction, multi-view distillation, and model parameter optimization.

[0020] In the data preprocessing step, a chest imaging dataset including AP and lateral chest X-rays is collected, and each chest X-ray is standardized, including image cropping and normalization operations, to ensure that chest X-rays from different sources and shooting conditions have a uniform size and grayscale range for subsequent model processing. If there is only an AP chest X-ray, the AP chest X-ray is enhanced to a certain extent and used as an image pair with the original image. For the original report text corresponding to the chest X-ray, a large language model is used to optimize it in the form of question and answer, and a polished version of the text (polished text) is obtained and made into a text pair with the original report text.

[0021] In the feature extraction step, an image encoder is designed to extract the features of the anteroposterior and lateral chest radiographs. The anteroposterior and lateral chest radiographs are input as image pairs into the image encoder to obtain the features of the anteroposterior and lateral chest radiographs. For the medical text report corresponding to the image, the Tokenizer in the BERT model is used to convert the words in the text into vector representations, and then the text vector sequence is processed by the BERT model to extract the semantic features of the text.

[0022] In the multi-view distillation step, the multi-view distillation strategy of the Mean Teacher (MT) model is adopted, that is, both the image encoder and the text encoder have a momentum-updated teacher model. The architecture of the teacher model and the student model in the Mean Teacher model is the same to achieve self-distillation. Specifically, the image pair is sent to the student image encoder and the teacher image encoder at the same time, and the front image features obtained by the student encoder are compared with the side image features obtained by the teacher model. At the same time, the side image features obtained by the student encoder are compared with the front image features obtained by the teacher model. In this way, the teacher model can be effectively used to guide the student model to better learn multi-view features. The text pair is sent to the student text encoder and the teacher text encoder at the same time, and the original text features obtained by the student encoder are compared with the polished text features obtained by the teacher model. At the same time, the polished text features obtained by the student encoder are compared with the original text features obtained by the teacher model. In addition, the same is true for cross-modal learning. The frontal image features and side image features obtained by the student image encoder are compared with the original text features and polished text features obtained by the teacher text encoder. The polished text features and original text features obtained by the student text encoder are also compared with the frontal image features and side image features obtained by the teacher image encoder.

[0023] In the model parameter optimization step, the image correlation matrix is ​​constructed by calculating the cosine similarity between different image feature vectors. The matrix reflects the similarity between different chest images in visual features. For example, images with similar lesion types or image patterns have higher similarity values ​​in the matrix. The correlation matrix between text features is also calculated. The text correlation matrix is ​​constructed using the similarity measure between text semantic feature vectors. The matrix reflects the association between different medical text reports in semantic content. For example, different description texts of the same disease will have a certain correlation in the matrix. The modal association graph is constructed based on the image and text correlation matrix. By calculating the correlation consistency loss between the two modalities, the model's ability to understand images and texts of the same disease can be effectively enhanced. Through the back propagation algorithm, based on the designed loss functions, including multi-view image contrast loss, multi-view text contrast loss, image-text cross-modal contrast loss, and cross-modal instance similarity consistency loss, the parameters of the entire model are optimized and trained. During the training process, the model parameters are continuously adjusted to minimize the loss function value, so that the student model can effectively acquire cross-modal knowledge from the teacher model and improve the performance of the model in various downstream tasks.

[0024] Figure 1 The process of the chest image artificial intelligence model training method in an embodiment of the present invention is shown, and inputting graphic data into the mean teacher model for training includes the following steps: Step S1, sending the image pair and the text pair to the image encoder and the text encoder respectively to obtain corresponding features; Step S2: align the features of the frontal image and the lateral image, and align the features of the original report text and the polished text; complete the feature alignment process by contrastive learning and using contrastive loss; finally, align the image features and text features in pairs, also by using contrastive loss; wherein, the calculation process of contrastive loss is as follows: in, , and Respectively represent the features of different modalities (image modality or text modality), represents the temperature hyperparameter, and Respectively represent the characteristics of different modes in a batch, n represents the number of samples in a batch, T Represents a matrix transpose operation; Step S3: construct a correlation graph between image features and text features, and calculate consistency loss.

[0025] Specifically, step S1 may include step S11 and step S12.

[0026] Step S11: Send the frontal image and the lateral image as paired data to the student image encoder of the mean teacher model at the same time and the teacher image encoder In this paper, four features are obtained: in, and Represent the frontal image and lateral image respectively; if there is only the frontal image, then Indicates that after rotation, flipping and other enhancements . and They respectively represent the frontal image features and lateral image features obtained by the student image encoder. and They respectively represent the frontal image features and side image features obtained by the teacher image encoder.

[0027] Step S12: The original report text and the polished text are fed into the student text encoder of the mean teacher model as paired data at the same time. and teacher text encoder In this paper, four features are obtained: in, , Represents the original report text and the polished text respectively; and They represent the original report text features and polished text features obtained by the student text encoder, and They represent the original text features and polished text features obtained by the teacher text encoder respectively.

[0028] Step S2 may include steps S21 to S23.

[0029] Step S21: Align the image features obtained in step S11 using contrastive learning, where the positive sample is the current image pair and the negative sample is other image features. The specific process is as follows: in, represents the multi-view image contrast loss, ( , )and( , ) represent the characteristics of the student model and the teacher model under different perspectives (first perspective and second perspective), respectively.

[0030] Step S22: align the text features obtained in step S12 using contrastive learning, where the positive sample is the current text pair and the negative sample is other image features. The specific process is as follows: in, represents the multi-view text contrast loss, ( , )and( , ) represent the features of the student model and the teacher model under the original report text and the polished text respectively.

[0031] Step S23: align the image features obtained in step S11 and the text features obtained in step S12 by contrastive learning, wherein the positive sample is the current image-text pair feature, the negative sample of the image feature is the remaining text features, and the negative sample of the text feature is the other image features. The specific process is as follows: in, Represents the cross-modal contrast loss between images and text.

[0032] Step S3 may include steps S31 to S33.

[0033] Step S31: Process the image features obtained in step S11 and calculate the correlation matrix of the image features in the same batch. The specific process is as follows: in, represents the characteristics of the specified mode, Represents the temperature hyperparameter. Since the diagonal is the dot product of the feature and itself, after eliminating the diagonal, the final correlation matrix can be obtained, as shown in the following formula: in, Represents the similarity of the feature similarity matrix after removing the diagonal.

[0034] Step S32: Process the text features obtained in step S12, and calculate the correlation matrix of the image features in the same batch, using the same calculation method as that of the images.

[0035] Step S33: Use KL-Divergence to perform consistency learning on the image feature correlation matrix and the text feature correlation matrix. The specific process is as follows: in, represents the cross-modal instance similarity consistency loss, Represents the similarity matrix of the first-person chest radiograph features, A similarity matrix representing the features of the original report text.

[0036] Finally, the total loss function of the model can be expressed as follows: in, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal contrast loss between images and texts, represents the consistent loss of cross-modal instance similarity; , , , They represent weight coefficients respectively.

[0037] Example 1 According to the method provided by the present invention, pre-training was performed on the MIMIC-CXR dataset, and performance testing of zero-sample classification was performed on the Vindr-CXR dataset and the CheXpert5×200 dataset. The Vindr-CXR dataset contains 18,000 chest images, whose labels are distributed as 22 local labels and 6 global labels, and are labeled by experienced radiologists. In this embodiment, the image size is adjusted to 248×248 and randomly cropped to 224×224. The CheXpert5×200 dataset is a subset of the CheXpert dataset, which is used for 5 disease classifications, each category has 200 images, and the same processing method is used to pre-process it in this embodiment.

[0038] At the same time, the method of this embodiment is compared with the most advanced method CXR-CLIP, and the comparison is made from the perspective of objective evaluation indicators. A comparative experiment is performed using the AUC evaluation indicator. The experimental results are shown in Table 1.

[0039] Table 1 AUC results of Vindr-CXR dataset and CheXpert5×200 zero-shot classification It can be seen that on the Vindr-CXR dataset, the AUC of the method of this embodiment is 6.3% higher than that of the ResNet-50 version of CXR-CLIP and 6.7% higher than that of the Swin-Transformer version of CXR-CLIP. On the CheXpert5×200 dataset, the AUC of the method of this embodiment is 3.0% higher than that of the ResNet-50 version of CXR-CLIP and 2.5% higher than that of the Swin-Transformer version of CXR-CLIP.

[0040] Example 2 According to the method provided by the present invention, pre-training was performed on the MIMIC-CXR dataset, and the performance test of cross-modal image-text retrieval was performed on the MIMIC-CXR dataset.

[0041] At the same time, the method of this embodiment is compared with the most advanced method CXR-CLIP in terms of objective evaluation indicators. The recall rates R@1, R@5 and R@10 are used as evaluation indicators for comparative experiments. The experimental results are shown in Table 2.

[0042] Table 2 Cross-modal image and text retrieval results of the MIMIC-CXR dataset It can be seen that the method of this embodiment is at the highest level in the three indicators of R@1, R@5, and R@10, which are 2.6%, 6.1%, and 6.6% higher than the ResNet-50 version of CXR-CLIP, and 2.0%, 5.7%, and 5.3% higher than the Swin-Transformer version of CXR-CLIP.

[0043] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same or equivalent means, it should belong to the protection scope of the present invention. Within the protection scope of the present invention, its technical scheme and / or implementation method can have various modifications and changes.

Claims

1. A chest image artificial intelligence model training method, characterized in that: The following steps are involved: Acquire a chest image dataset and a corresponding report text dataset, wherein the chest image dataset includes an image pair consisting of a first-view chest X-ray and a second-view chest X-ray, and the report text dataset includes a text pair consisting of an original report text and a polished text derived from the original report text; Using the chest image dataset and its corresponding report text dataset to train a mean teacher model, the mean teacher model includes a student model and a teacher model, wherein the student model includes a student image encoder and a student text encoder, and the teacher model includes a teacher image encoder and a teacher text encoder; Compare and learn the chest X-ray features of the first view extracted by the student image encoder based on the image pair with the chest X-ray features of the second view extracted by the teacher image encoder based on the image pair, and compare and learn the chest X-ray features of the second view extracted by the student image encoder based on the image pair with the chest X-ray features of the first view extracted by the teacher image encoder based on the image pair; Comparatively learning the original report text features extracted by the student text encoder based on the text pair and the polished text features extracted by the teacher text encoder based on the text pair, and comparatively learning the polished text features extracted by the student text encoder based on the text pair and the original report text features extracted by the teacher text encoder based on the text pair; Compare and learn the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the student image encoder based on the image pair with the original report text features and / or the polished text features extracted by the teacher text encoder based on the text pair, and compare and learn the original report text features and / or the polished text features extracted by the student text encoder based on the text pair with the first-view chest X-ray features and / or the second-view chest X-ray features extracted by the teacher image encoder based on the image pair; The parameters of the mean teacher model are updated according to the total loss function.

2. The method according to claim 1, characterized in that The first-view chest X-ray and the second-view chest X-ray have been standardized.

3. The method according to claim 1, characterized in that The total loss function is a weighted function of multi-view image contrast loss, multi-view text contrast loss, image-text cross-modal contrast loss, and cross-modal instance similarity consistency loss, as shown in the following formula: in, represents the total loss, represents the multi-view image contrast loss, represents the multi-view text contrast loss, represents the cross-modal contrast loss between images and texts, represents the consistent loss of cross-modal instance similarity; , , , They represent weight coefficients respectively.

4. The method according to claim 3, characterized in that The contrast loss in the multi-view image contrast loss, the multi-view text contrast loss, and the image-text cross-modal contrast loss is calculated by the following formula: in, , and Represent the characteristics of different modes, represents the temperature hyperparameter, and Respectively represent the characteristics of different modes in a batch, n represents the number of samples in a batch, T Represents a matrix transpose operation.

5. The method according to claim 4, characterized in that The multi-view image contrast loss is calculated by the following formula: in, represents the multi-view image contrast loss, ( , )and( , ) represent the characteristics of the student model and the teacher model in the first perspective and the second perspective respectively.

6. The method according to claim 4, characterized in that The multi-view text contrast loss is calculated by the following formula: in, represents the multi-view text contrast loss, ( , )and( , ) represent the features of the student model and the teacher model under the original report text and the polished text respectively.

7. The method according to claim 4, characterized in that The image-text cross-modal contrast loss is calculated by the following formula: in, in, Represents the cross-modal contrast loss between images and text.

8. The method according to claim 4, characterized in that The cross-modal instance similarity consistency loss is calculated as follows: in, in, represents the consistent loss of cross-modal instance similarity; Represents the similarity matrix of the first-person chest radiograph features, A similarity matrix representing the text features of the original report; represents the characteristics of the specified mode, represents the temperature hyperparameter; Represents the similarity of the feature similarity matrix after removing the diagonal.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions, which, when executed by at least one processor, implement the chest imaging artificial intelligence model training method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed by the processor, the chest imaging artificial intelligence model training method as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Text processing method and device, electronic equipment and storage medium

    CN117952071A

  • Semi-supervised nasopharyngeal carcinoma segmentation method based on image text contrast learning

    CN118297960A

  • Text data processing method and device and text data detection method and device

    CN119621984A

  • Method for pre-training vision-language transformer, and artificial intelligence system comprising vision-language transformer pre-trained through same method

    WO2024186177A1

Cited By

  • Crop all-weather monitoring and disaster assessment method based on cross-modal self-distillation

    CN122090284A