Method, apparatus and electronic device for determining coronary artery CT calcium score

Through the multi-class segmentation model and self-attention mechanism training coronary CT calcification integral method, the problems of high calculation costs and insufficient accuracy in the existing technology are solved, and efficient and accurate calculation of coronary CT calcification integral is achieved.

CN119559158BActive Publication Date: 2025-07-11FUWAI HOSPITAL CHINESE ACAD OF MEDICAL SCI & PEKING UNION MEDICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411827482.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-07-11
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

The existing calculation method for coronary CT calcification integral has the problem of high optimization cost, and the false positive rate and robustness of the model are insufficient, resulting in insufficient accuracy.

Method used

The coronary CT image is processed by using a multi-class segmentation model, and the target mask is obtained through characterization learning encoder and multi-layer perceptron module training. Combined with self-attention mechanism and contrast learning technology, the dependence on the annotated data is reduced and the robustness and accuracy of the model is improved.

Benefits of technology

It is achieved to improve the accuracy and consistency of coronary CT calcification integral calculation under low labeled data volume, reduce training costs, and improve the flexibility and stability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559158B_ABST
    Figure CN119559158B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and electronic device for determining coronary CT calcification score. Among them, the method includes: obtaining a coronary CT image to be processed; processing the coronary CT image with a multi-class segmentation model to obtain a target mask, where the multi-class segmentation model is trained based on a coronary CT feature map for a calcification binary segmentation task and a calcification multi-class segmentation task, and the framework of the multi-class segmentation model includes a feature learning encoder, a binary classification task branch and a multi-class classification task branch; determining the coronary CT calcification score corresponding to the coronary CT image according to the target mask. The present application solves the technical problem of the high optimization cost of the calculation method of coronary CT calcification score in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the medical field, and more particularly, to a method, apparatus, and electronic device for determining coronary CT calcium scores. Background Technique

[0002] Cardiovascular diseases have become one of the major health challenges worldwide. One of its main forms is coronary artery disease, which is specifically manifested as atherosclerotic plaques inside the coronary arteries causing blood flow obstruction, thereby triggering serious complications such as myocardial ischemia, angina pectoris, and myocardial infarction. Therefore, early detection and assessment of the degree of atherosclerosis are crucial for preventing the development of cardiovascular diseases. In clinical practice, coronary CT (computed tomography) calcium score has become an important non-invasive diagnostic tool, providing an objective method for assessing the risk of coronary artery diseases. In recent years, due to the development and implementation of deep learning, current automatic calculation schemes for coronary CT calcium scores have been continuously proposed and applied. However, the accuracy of these technologies still has certain problems, and the optimization cost is relatively high.

[0003] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, and electronic device for determining coronary CT calcium scores to at least solve the technical problem of relatively high optimization cost in the calculation method of coronary CT calcium scores in related technologies.

[0005] According to one aspect of the embodiments of this application, a method for determining coronary CT calcium scores is provided, including: obtaining a coronary CT image to be processed; processing the coronary CT image using a multi-class segmentation model to obtain a target mask, where the multi-class segmentation model is trained based on coronary CT feature maps for a binary calcium segmentation task and a multi-class calcium segmentation task, and the framework of the multi-class segmentation model includes a representation learning encoder, a binary classification task branch, and a multi-class classification task branch; determining the coronary CT calcium score corresponding to the coronary CT image based on the target mask.

[0006] Optionally, the multi-class segmentation model is trained in the following manner: obtaining historical coronary CT images, where the historical coronary CT images include calcium category annotation information, and the calcium category annotation information is used to indicate whether there is calcium in the historical coronary CT images; determining the corresponding historical matrix data of the historical coronary CT images; randomly dividing the historical matrix data to obtain a first data set, and determining the data in the first data set that contains calcium category annotation information and the annotation information indicates the presence of calcium as a second data set; training the representation learning encoder based on the first data set and the second data set; constructing a multi-class segmentation model based on the representation learning encoder.

[0007] Optionally, training the representation learning encoder based on the first data set and the second data set includes: inputting the first data set into the framework model where the representation learning encoder to be trained is located to obtain a first feature representation and a second feature representation; determining a first loss value based on the first feature representation and the second feature representation; when the change rate of the first loss value is less than a first threshold, inputting the second data set into the framework model to obtain a third feature representation and a fourth feature representation; determining a second loss value based on the third feature representation and the fourth feature representation; when the change rate of the second loss value is less than a second threshold, obtaining the trained representation learning encoder from the framework model, where the first threshold is greater than the second threshold.

[0008] Optionally, the framework model includes a first encoder, a second encoder, a first multi-layer perceptron module, and a second multi-layer perceptron module. Inputting the first data set into the framework model where the representation learning encoder to be trained is located to obtain a first feature representation and a second feature representation includes: performing two data augmentation processes on the first data set to obtain a first image pair, where when the first image pair is from the same image of different data augmentation processes, the first image pair is a positive sample, and when the first image pair is from different images, the first image pair is a negative sample; inputting the first image pair into the first encoder and the second encoder respectively to obtain a first image feature representation and a second image feature representation; inputting the first image feature representation and the second image feature representation into the first multi-layer perceptron module and the second multi-layer perceptron module respectively to obtain a first feature representation and a second feature representation.

[0009] Optionally, inputting the second data set into the framework model to obtain a third feature representation and a fourth feature representation includes: performing two data augmentation processes on the second data set to obtain a second image pair, where when the second image pair is from the same image of different data augmentation processes, the second image pair is a positive sample, and when the second image pair is from different images, the second image pair is a negative sample; inputting the second image pair into the first encoder and the second encoder respectively to obtain a third image feature representation and a fourth image feature representation; inputting the third image feature representation and the fourth image feature representation into the first multi-layer perceptron module and the second multi-layer perceptron module respectively to obtain a third feature representation and a fourth feature representation.

[0010] Optionally, a multi-class segmentation model is constructed based on a representation learning encoder, including: obtaining a third dataset composed of image patches corresponding to the second dataset and calcification segmentation mask pairs corresponding to the second dataset, where the calcification segmentation mask pairs include multi-class segmentation masks and binary segmentation masks; extracting features from the third dataset through the representation learning encoder to obtain coronary CT feature maps; determining a first output based on the coronary CT feature maps and the calcification binary segmentation task in the multi-class segmentation model; determining a second output based on the coronary CT feature maps and the calcification multi-class segmentation task in the multi-class segmentation model; determining a first loss function based on the first output and the binary segmentation masks in the second dataset, and determining a second loss function based on the second output and the multi-class segmentation masks in the second dataset; determining a total loss function based on the first loss function and the second loss function, adjusting the parameters of the total loss function, and stopping training until the total loss function meets a preset condition to obtain a trained multi-class segmentation model.

[0011] Optionally, the calcification binary segmentation task includes a first adapter, a first decoder, and a calcification binary segmentation head, and the calcification multi-class segmentation task includes a second adapter, a second decoder, and a calcification multi-class segmentation head. Determining a first output based on the coronary CT feature maps and the calcification binary segmentation task in the multi-class segmentation model, and determining a second output based on the coronary CT feature maps and the calcification multi-class segmentation task in the multi-class segmentation model includes: inputting the coronary CT feature maps into the first adapter and the second adapter respectively for feature extraction to obtain a first coronary CT feature map and a second coronary CT feature map; inputting the first coronary CT feature map into the second adapter for feature extraction, and fusing the feature extraction result with the second coronary CT feature map to obtain a third coronary CT feature map; inputting the first coronary CT feature map into the first decoder to obtain a fourth coronary CT feature map; inputting the fourth coronary CT feature map into the calcification binary segmentation head to obtain a first output; inputting the third coronary CT feature map into the second decoder to obtain a fifth coronary CT feature map; inputting the fifth coronary CT feature map into the calcification multi-class segmentation head to obtain a second output.

[0012] Optionally, when the representation learning encoder is Resnet, the method further includes: dividing the coronary CT feature maps into multiple windows, and determining the correlation between the features of each window through a self-attention mechanism, and determining the correlation between the features of different windows.

[0013] Optionally, determining the coronary CT calcium score corresponding to the coronary CT image according to the target mask includes: performing pixel processing on the target mask to obtain target pixels; determining the density score corresponding to each pixel in the target pixels, and determining the maximum density score of each layer of the coronary CT image; determining the calcium score of each layer of the coronary CT image according to the maximum density score of each layer and the calcium area of each layer, where the calcium area is determined according to the number of pixels in the calcium region of each layer and the pixel pitch of the coronary CT image; determining the coronary CT calcium score according to the calcium score of each layer and the slice thickness of the coronary CT image.

[0014] According to another aspect of the embodiments of the present application, there is also provided a device for determining a coronary CT calcium score, including: an acquisition module, configured to acquire a coronary CT image to be processed; a classification module, configured to process the coronary CT image by using a multi-class segmentation model to obtain a target mask, where the multi-class segmentation model is trained based on a coronary CT feature map for a calcification binary segmentation task and a calcification multi-class segmentation task, and the framework of the multi-class segmentation model includes a feature learning encoder, a binary classification task branch, and a multi-class classification task branch; a determination module, configured to determine the coronary CT calcium score corresponding to the coronary CT image according to the target mask.

[0015] According to still another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the memory is configured to store program instructions; the processor is connected to the memory and is configured to execute to implement the above method for determining a coronary CT calcium score.

[0016] According to yet another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program, where the device where the non-volatile storage medium is located executes the above method for determining a coronary CT calcium score by running the computer program.

[0017] According to yet another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, where the computer instructions implement the above method for determining a coronary CT calcium score when executed by a processor.

[0018] In the embodiments of the present application, by acquiring a coronary CT image to be processed; processing the coronary CT image by using a multi-class segmentation model to obtain a target mask, where the multi-class segmentation model is trained based on a coronary CT feature map for a calcification binary segmentation task and a calcification multi-class segmentation task, and the framework of the multi-class segmentation model includes a feature learning encoder, a binary classification task branch, and a multi-class classification task branch; and determining the coronary CT calcium score corresponding to the coronary CT image according to the target mask, the technical problem of high optimization cost in the calculation method of the coronary CT calcium score in the related art can be solved. Description of the Drawings

[0019] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a flowchart of a method for determining the coronary CT calcium score according to an embodiment of the present application;

[0021] Figure 2 is a schematic flowchart of a process for training a representation learning encoder according to an embodiment of the present application;

[0022] Figure 3 is a schematic flowchart of a process for training a multi-class segmentation model according to an embodiment of the present application;

[0023] Figure 4 is a structural diagram of a device for determining the coronary CT calcium score according to an embodiment of the present application;

[0024] Figure 5 is a hardware structure block diagram of a computer terminal for implementing the method for determining the coronary CT calcium score according to an embodiment of the present application. Detailed implementation manners

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] The information collected in the embodiments of this application is information and data that have been authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards of the relevant regions, necessary confidentiality measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for users to choose to authorize or reject the automated decision-making results. If the user chooses to reject, the expert decision-making process will be entered.

[0028] With the rise of deep learning, computer vision, and artificial intelligence technologies, researchers have begun to explore the application of these advanced technologies to coronary CT calcium scoring. By training a neural network, automated detection, segmentation, and classification of calcified regions can be achieved, thereby reducing manual intervention and improving the consistency and accuracy of analysis. For example, in the field of medical imaging, the U-Net (U-Net) based on CNN (Convolutional Neural Network) is widely used for the segmentation of lesions or organs, and the model structure is simple and easy to train. To solve specific problems, variants based on the U-Net model have also been widely applied, such as U-Net++, UNet3+, etc. In addition, the Transformer architecture originating from the field of natural language processing has gradually become popular. Compared with convolutional neural networks, the self-attention mechanism based on transformers can capture the long-range dependencies of images and obtain more accurate results. However, due to the extremely large number of pixels in images, directly learning self-attention weights in units of image pixels will result in a very high time complexity for the algorithm. Therefore, Transformer structures combined with the characteristics of images, such as Swin Transformer, and TransUnet that combines CNN and Transformer have also been developed.

[0029] In addition, due to the high cost of obtaining labeled data, especially in the field of medical imaging, unsupervised training methods based on contrastive learning have also developed rapidly and become popular. Relevant frameworks include MOCO, SimCLR, SimSiam, BYOL, SwAV, etc. Contrastive learning does not require labeled data but mines the internal representations based on the similarity of images. A large amount of unlabeled data can be used for pre-training, and then fine-tuned on a small amount of labeled data to achieve even better results than supervised models.

[0030] There are the following two mainstream algorithm frameworks in the related art, and the specific disadvantages are as follows: (1) Learning a binary segmentation model or a multi-class segmentation model based on a convolutional neural network. This method lacks long-distance dependencies, that is, the existing convolutional neural network cannot capture the correlation of pixels with longer distances, and can only capture local image information. The ability to capture global image information is limited, resulting in many false positives on the aorta; (2) Using a rough segmentation model and a false positive removal model, the output of the first model is the input of the second model. This method makes the model training not flexible enough, and often requires joint optimization of the two models. In addition, due to different imaging devices and parameters, the trained model often does not perform robustly in actual applications, that is, the model generalization ability is insufficient. In the case of many image interference factors, the model stability is insufficient. To solve this problem, a large amount of high-quality data needs to be collected and labeled, which will result in too high time and labor costs.

[0031] The automated calculation technology of coronary artery calcium score based on coronary CT can greatly improve the efficiency of radiologists manually annotating calcified lesions and improve the consistency of results. However, the existing coronary artery calcium score technology has problems such as false positives and insufficient robustness, which will lead to insufficient accuracy; and the existing model training is not flexible enough, which will lead to higher optimization costs. To solve the above problems, the embodiments of the present application provide a method for determining the coronary CT calcium score.

[0032] Figure 1 It is a flowchart of a method for determining the coronary CT calcium score according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0033] Step S202, obtaining a coronary CT image to be processed.

[0034] Step S204, processing the coronary CT image with a multi-class segmentation model to obtain a target mask, where the multi-class segmentation model is trained based on the coronary CT feature map for the binary calcification segmentation task and the multi-class calcification segmentation task. The framework of the multi-class segmentation model includes a representation learning encoder, a binary classification task branch, and a multi-classification task branch.

[0035] In the above step S204, applying the trained multi-class segmentation model to a new coronary CT image (i.e., the above-mentioned coronary CT image to be processed), the model will automatically identify and segment different category calcified regions in the image. Through the segmentation result output by the model, the target mask (mask) of each category calcified region can be obtained.

[0036] Step S206, determining the coronary CT calcium score corresponding to the coronary CT image according to the target mask.

[0037] In the above step S206, the coronary artery CT calcium score is an index used to quantitatively evaluate the overall calcification of the coronary arteries. It is scored based on the area and density of the calcified regions detected in the CT image, reflecting the severity of coronary artery disease and the level of the patient's cardiovascular risk. The target mask is obtained from the coronary artery CT image through a multi-class segmentation model, which identifies the calcified regions of different classes in the image. Using the target mask, the area and density information of each calcified region can be accurately extracted. Substituting this information into the calculation formula of the calcium score, the calcium score of the coronary artery CT image can be obtained.

[0038] In step S202 of the above method for determining the coronary artery CT calcium score, after obtaining the coronary artery CT image to be processed, the method further includes: determining a preset CT value range for the coronary artery CT image, where the preset CT value range includes a preset maximum value and a preset minimum value; replacing the CT values in the coronary artery CT image that are less than or equal to the preset minimum value with the preset minimum value, and replacing the CT values in the coronary artery CT image that are greater than or equal to the preset maximum value with the preset maximum value to obtain a processed coronary artery CT image; resampling the processed coronary artery CT image to obtain the matrix data of the coronary artery CT image.

[0039] In the embodiments of the present application, the CT value (Hounsfield Units, HU) is a unit for describing the degree of X-ray absorption of different tissues in the CT image. In the coronary artery CT image, different CT value ranges usually correspond to different tissue types. In order to better highlight the calcified regions or other regions of interest and reduce the influence of irrelevant tissues, a preset CT value range, that is, a preset maximum value and a preset minimum value, needs to be set. In an optional embodiment, the preset CT value range can be [-800, 1200], where -800 is the above-mentioned preset minimum value and 1200 is the above-mentioned preset maximum value. The CT values in the coronary artery CT image that are less than or equal to -800 are unified to -800, and those greater than or equal to 1200 are unified to 1200. This can improve the contrast of the CT image and reduce the image of non-target regions. Resampling the preprocessed coronary artery CT image as described above, resampling is the process of changing the pixel spacing (or called resolution) of the image to a new value. After resampling, the coronary artery CT image will be converted into matrix data with a specified pixel spacing. For example, bilinear interpolation is used to resample the coronary artery CT image to set the axial spacing of the CT image to 1.5 mm to obtain the matrix data of the coronary artery CT image.

[0040] In the above method for determining the coronary artery CT calcium score, multiple segmentation models are trained as follows: Obtain historical coronary artery CT images, where the historical coronary artery CT images include calcium category annotation information, and the calcium category annotation information is used to indicate whether there is calcium in the historical coronary artery CT image and the location where the calcium is located; Determine the historical matrix data corresponding to the historical coronary artery CT image; Randomly divide the historical matrix data to obtain a first data set, and determine the data in the first data set that contains calcium category annotation information and the annotation information indicates the presence of calcium as the second data set; Train a representation learning encoder based on the first data set and the second data set; Construct a multiple segmentation model based on the representation learning encoder.

[0041] In the embodiments of the present application, it is necessary to first collect a large number of historical coronary artery CT images, which contain key information such as calcium regions. By annotating these historical coronary artery CT images, calcium category annotation information is obtained. For example, a radiologist annotates the calcium category on the aorta in the historical coronary artery CT image. There are a total of 7 vascular position categories (left main, anterior descending branch, circumflex branch, right coronary artery, posterior descending branch, ascending aorta, descending aorta). When performing model training later, an auxiliary position information can be given to the model to guide the model to distinguish calcium in different positions. Among them, the annotation ratio is 10-30%. After annotation, the calcium region and the background region can be distinguished, so as to obtain the segmentation mask of the calcium region corresponding to the historical coronary artery CT image. The segmentation mask can be a grayscale image. It should be noted that when performing annotation, binary segmentation annotation is performed simultaneously during multi-class segmentation mask annotation, that is, the branch artery and the aorta.

[0042] After annotating historical coronary artery images and then performing preprocessing, for example, restricting the CT values in the historical coronary artery CT images to [-800, 1200], and resampling the preprocessed historical coronary artery CT images to obtain the historical matrix data corresponding to the historical coronary artery CT images. Randomly divide the historical matrix data into image blocks with a size of [48, 128, 128] to obtain the first dataset, which can be represented by dataset A. Among them, the historical matrix data includes labeled data and unlabeled data, and the random division means that the number and position of the image blocks are both random. It should be noted that the size of the above image blocks is only an example and does not represent a limitation. The specific size can be adjusted according to the actual situation. In the first dataset, the image blocks with labeled information and the labeled information being calcified are used as the second dataset, which can be represented by dataset B. Train the representation learning encoder through the obtained first dataset and second dataset. For example, unsupervised training can be carried out through contrastive learning to obtain the representation learning encoder. Train the multi-head supervised multi-task model according to the trained representation learning encoder, so as to construct a multi-class segmentation model. Using contrastive learning for pre-training reduces the need for labeled data and can obtain better results with less fine labeling. In addition, a large number of data augmentation techniques are used in contrastive learning, making the model more robust. The following is an explanation.

[0043] In the above steps, training the representation learning encoder according to the first dataset and the second dataset includes: inputting the first dataset into the framework model where the representation learning encoder to be trained is located to obtain the first feature representation and the second feature representation; determining the first loss value based on the first feature representation and the second feature representation; when the change rate of the first loss value is less than the first threshold, inputting the second dataset into the framework model to obtain the third feature representation and the fourth feature representation; determining the second loss value based on the third feature representation and the fourth feature representation; when the change rate of the second loss value is less than the second threshold, obtaining the trained representation learning encoder from the framework model, where the first threshold is greater than the second threshold.

[0044] In the above steps, the framework model includes a first encoder, a second encoder, a first multi-layer perceptron module, and a second multi-layer perceptron module. The first dataset is input into the framework model where the representation learning encoder to be trained is located, and a first feature representation and a second feature representation are obtained, including: the first dataset is subjected to two data augmentation processes to obtain a first pair of images. Among them, when the first pair of images is from the same image patch of different data augmentation processes, the first pair of images is a positive sample; when the first pair of images is from different image patches, the first pair of images is a negative sample; the first pair of images is respectively input into the first encoder and the second encoder to obtain a first image feature representation and a second image feature representation; the first image feature representation and the second image feature representation are respectively input into the first multi-layer perceptron module and the second multi-layer perceptron module to obtain a first feature representation and a second feature representation.

[0045] In the above steps, the second dataset is input into the framework model to obtain a third feature representation and a fourth feature representation, including: the second dataset is subjected to two data augmentation processes to obtain a second pair of images. Among them, when the second pair of images is from the same image patch of different data augmentation processes, the second pair of images is a positive sample; when the second pair of images is from different image patches, the second pair of images is a negative sample; the second pair of images is respectively input into the first encoder and the second encoder to obtain a third image feature representation and a fourth image feature representation; the third image feature representation and the fourth image feature representation are respectively input into the first multi-layer perceptron module and the second multi-layer perceptron module to obtain a third feature representation and a fourth feature representation.

[0046] In the embodiment of the present application, Figure 2 is a schematic flowchart of a process for training a representation learning encoder in an embodiment of the present application. The following combines Figure 2 to illustrate the above process of training the representation learning encoder. The model structure where the representation learning encoder is located adopts the MoCo framework, and the encoder is not restricted (Resnet or Transformer). The MoCo framework includes a first encoder, a second encoder, and two multi-layer perceptron (MLP) modules. In Figure 2 the first multi-layer perceptron module is connected to the first encoder, and the second multi-layer perceptron module is connected to the second encoder. The input is historical coronary CT images, and then a pair of images is obtained through two data augmentations A and B. If this pair of images is from the same image with different data augmentation schemes, then this pair of images is a positive sample; if a pair of images is from different images, then this pair of images is a negative sample. Among them, the data augmentation schemes include translation, rotation, brightness shift, and scaling. This pair of images respectively pass through the first encoder and the second encoder to obtain a first image feature representation and a second image feature representation, and then respectively pass through the first and second multi-layer perceptron modules to obtain a first feature representation and a second feature representation.

[0047] Specifically, in the training stage, first use dataset A (i.e., the first dataset) as the input of the MoCo framework model. After two data augmentations (the above-mentioned data augmentation A and data augmentation B), a first image pair is obtained. The first image pair is respectively input into the first encoder and the second encoder to obtain a first image feature representation and a second image feature representation, and then they are respectively input into the first multi-layer perceptron module and the second multi-layer perceptron module to obtain a first feature representation and a second feature representation. Contrastive loss is used for training, and a first loss value is determined based on the first feature representation and the second feature representation. In the above process, the formula of the contrastive loss function is as follows:

[0048]

[0049] where q represents the first feature representation, k + represents the second feature representation when the input first image pair is a positive sample, and k _ represents the second feature representation when the input first image pair is a negative sample, is a temperature hyperparameter, Lloss represents the contrastive loss function. The numerator part of this loss function is the similarity of the positive sample pairs, and the denominator is the similarity of all sample pairs, and it is normalized using softmax, enabling the model to learn the feature representation of the image to improve the similarity of the positive samples.

[0050] When the first loss value gradually converges such that the change rate in 5 consecutive epochs (one generation of training) is less than 0.1% (i.e., the above-mentioned first threshold), then use dataset B (i.e., the second dataset) as the input of the MoCo framework model to make the model pay more attention to the features of calcification. Specifically, after two data augmentations of dataset B, a second image pair is obtained. The second image pair is respectively input into the first encoder and the second encoder to obtain a third image feature representation and a fourth image feature representation. The third image feature representation and the fourth image feature representation are respectively input into the first multi-layer perceptron module and the second multi-layer perceptron module to obtain a third feature representation and a fourth feature representation. Similarly, contrastive loss is used for training, and a second loss value is determined based on the third feature representation and the fourth feature representation. When the second loss gradually converges such that the change rate in 5 consecutive epochs is less than 0.001% (i.e., the above-mentioned second threshold), the pre-training ends, and a trained representation learning encoder is obtained from the MoCo framework model. This representation learning encoder can be the trained first encoder.

[0051] In the above steps, a multi-class segmentation model is constructed based on the representation learning encoder, including: obtaining a third dataset composed of image patches corresponding to the second dataset and pairs of calcification segmentation masks corresponding to the second dataset, where the pair of calcification segmentation masks includes a multi-class segmentation mask and a binary segmentation mask; extracting features from the third dataset through the representation learning encoder to obtain coronary CT feature maps; determining a first output based on the coronary CT feature maps and the calcification binary segmentation task in the multi-class segmentation model; determining a second output based on the coronary CT feature maps and the calcification multi-class segmentation task in the multi-class segmentation model; determining a first loss function based on the first output and the binary segmentation mask in the second dataset, and determining a second loss function based on the second output and the multi-class segmentation mask in the second dataset; determining a total loss function based on the first loss function and the second loss function, adjusting the parameters of the total loss function until the total loss function meets the preset conditions, and then stopping the training to obtain a trained multi-class segmentation model.

[0052] In the above steps, the calcification binary segmentation task includes a first adapter, a first decoder, and a calcification binary segmentation head, and the calcification multi-class segmentation task includes a second adapter, a second decoder, and a calcification multi-class segmentation head. Determining a first output based on the coronary CT feature maps and the calcification binary segmentation task in the multi-class segmentation model, and determining a second output based on the coronary CT feature maps and the calcification multi-class segmentation task in the multi-class segmentation model includes: respectively inputting the coronary CT feature maps into the first adapter and the second adapter for feature extraction to obtain a first coronary CT feature map and a second coronary CT feature map; inputting the first coronary CT feature map into the second adapter for feature extraction, and fusing the feature extraction result with the second coronary CT feature map to obtain a third coronary CT feature map; inputting the first coronary CT feature map into the first decoder to obtain a fourth coronary CT feature map; inputting the fourth coronary CT feature map into the calcification binary segmentation head to obtain a first output; inputting the third coronary CT feature map into the second decoder to obtain a fifth coronary CT feature map; inputting the fifth coronary CT feature map into the calcification multi-class segmentation head to obtain a second output.

[0053] In the embodiment of the present application, the data blocks corresponding to the second dataset and the pairs of calcification segmentation masks are combined to form a third dataset, which can be represented by dataset C. The pair of calcification segmentation masks includes a multi-class segmentation mask and a binary segmentation mask. In the embodiment of the present application, multi-class segmentation means dividing into 7 vascular categories (namely left main coronary artery, left anterior descending artery, left circumflex artery, right coronary artery, posterior descending artery, ascending aorta, descending aorta); binary segmentation is to divide the above 7 vascular categories into 2 categories, including branch arteries and aorta. Taking dataset C as the input of the multi-class segmentation model to be trained, the following data augmentations are added during training: random cropping, flipping, rotation, cutmix, mixup, random erasing, so that the model can learn more robust information with a small amount of labeled data.

[0054] To train a multi-class segmentation model, a supervised model framework is constructed. When the representation learning encoder is a Transformer, as Figure 3 shown, Figure 3 is a schematic flowchart of a method for training a multi-class segmentation model according to an embodiment of the present application. The following combines Figure 3 to describe the above process of training the multi-class segmentation model. The trained representation learning encoder is used as the backbone network of a supervised model (i.e., the multi-class segmentation model). According to two tasks (calcification binary classification and calcification multi-class classification), it is split into two branches of networks, each of which consists of an adapter, a decoder, and a segmentation head. Specifically, the first network where the calcification binary classification segmentation task is located includes a first adapter, a first decoder, and a calcification binary class segmentation head. The second network where the calcification multi-class segmentation task is located includes a second adapter, a second decoder, and a calcification multi-class segmentation head. During training, the parameters of the representation learning encoder are frozen, and only the adapter, decoder, and segmentation head are learned to adapt to downstream tasks, while preventing catastrophic forgetting of the pre-trained model, reducing the training parameters, and improving the training efficiency.

[0055] In Figure 3 , the representation learning encoder extracts features from the coronary CT data in the third dataset to obtain a coronary CT feature map. The coronary CT feature map output by the representation learning encoder is input into the adapter. The adapter architecture includes: a normalization layer (Layer Norm), a first fully connected layer (Linear1), an activation layer (ReLU layer), and a second fully connected layer (Linear2). After the first adapter processes the coronary CT feature map (including feature extraction and conversion into the same format to facilitate feature fusion), it obtains a first coronary CT feature map and transmits it to the first decoder. After the second adapter processes the coronary CT feature map, it obtains a second coronary CT feature map and transmits it to the second decoder. At the same time, the first adapter transmits the first coronary CT feature map to the second adapter. The second adapter extracts features from the first coronary CT feature map and adds and fuses the feature extraction results with the second coronary CT feature map to obtain a third coronary CT feature map.

[0056] Among them, the decoder can be a Transformer or CNN architecture. The decoder receives the coronary artery feature maps transmitted by the adapter through the upsampling layer (i.e., the first decoder receives the first coronary artery CT feature map, and the second decoder receives the third coronary artery CT feature map), gradually restores the coronary artery feature maps to the original image size, and there is a residual connection between the decoders of each level and the corresponding encoders to fuse the detailed information of the low level and the semantic information of the high level, so as to obtain a finely learned segmentation mask. Specifically, the first coronary artery CT feature map is input into the first decoder to obtain the fourth coronary artery CT feature map, and the third coronary artery CT feature map is input into the second decoder to obtain the fifth coronary artery CT feature map. The fourth coronary artery CT feature map is input into the calcification binary segmentation head to obtain the first output; the fifth coronary artery CT feature map is input into the calcification multi-class segmentation head to obtain the second output, where the output1 (i.e., the first output above) output by the binary segmentation head is a binary segmentation mask with a size of [1, 48, 128, 128], and the output of the multi-class segmentation head is a multi-class segmentation mask output2 (i.e., the second output above) with a size of [8, 48, 128, 128]. It should be noted that the size of the mask can also be adjusted according to actual needs.

[0057] The output1 and the manually annotated binary segmentation mask are input into the binary cross-entropy loss function L1 (i.e., the first loss function above), and the output2 and the manually annotated multi-class segmentation mask are input into the multi-class loss function L2 (i.e., the second loss function above). Since the multi-class head is an 8-classification (7 vascular classes plus background) task, and the actual occurrence probabilities of the 7 vascular classes are different, there will be a serious problem of class imbalance. To solve this problem, a multi-class loss function for this segmentation head is designed, and a penalty factor is introduced based on prior knowledge on the basis of the weighted cross-entropy loss. According to the vascular distribution of the coronary artery, the confusion difficulty between different blood vessels can be scored (scores from extremely easy to confuse to not easy to confuse increase from low to high), so as to obtain a KxK relationship matrix (K is the number of classes), and it is normalized. When calculating the WCE (Weighted Cross Entropy) loss, the annotated mask is processed into a one-hot encoding. Different from the original multi-class WCE that only calculates the loss of the pixels in the channel where the Label class is located, the designed loss function will calculate the loss of K channels, so it has more guiding properties for the training direction of the model. The formulas of L1 and L2 are as follows:

[0058]

[0059] Among them, N is the number of pixels in the segmentation mask, K is the number of classes, w is the weight of K classes (a 1*k vector), y iDenotes the manually annotated calcification segmentation mask, p i Denotes the output of the segmentation head. The total loss L is the weighted average of the losses (L1, L2) of the binary segmentation head and the multi-class segmentation head. In the early stage of training, the weight ɑ of L1 is 1 and the weight of L2 is 0, enabling the model to first focus on the calcification area. When the threshold of the binary classification head reaches a relatively high value (dice is 0.8), it assists the multi-task network in learning. And at this time, the weight of L1 is set to 0.1 and the weight of L2 is set to 0.9, focusing on training the multi-class segmentation head, and finally obtaining a trained multi-class segmentation model.

[0060] It should be noted that the adopted representation learning encoder is the backbone network of a supervised model. By setting 2 task branches, each branch includes an adapter, a decoder, and a segmentation head. During training, the representation learning encoder is frozen, and the adapter, decoder, and segmentation head are learned, reducing the training parameters. While improving the training efficiency, it prevents the catastrophic forgetting of the representation learning encoder. The multi-class segmentation model has two segmentation heads, one is a binary segmentation head and the other is a multi-class segmentation head. By establishing the connection between joint training and calcification classification, the training difficulty of the multi-class segmentation model can be reduced, the convergence of the model can be accelerated, and the classification error rate can be reduced. The above steps enable the feature information obtained from the binary segmentation task to be fused into the multi-classification task, enabling the multi-classification task to focus on the features of the binary segmentation task, making the model more robust.

[0061] In the above steps, when the representation learning encoder is Resnet, the method further includes: dividing the coronary CT feature map into multiple windows, and determining the correlation between the features of each window through the self-attention mechanism, as well as determining the correlation between the features of different windows.

[0062] In the embodiment of the present application, the difference between the multi-class segmentation model when the representation learning encoder is Resnet and the multi-class segmentation model when the representation learning encoder is Transformer lies in: between the representation learning encoder and the first adapter and the second adapter, there is a swin transformer block module. After the swin transformer block module processes the coronary CT feature map, the processed coronary CT feature map is input into the first adapter (Adapter1) and the second adapter (Adapter1). The self-attention mechanism of the swin transformer block can capture long-range dependencies, enabling the model to pay attention to more overall position information and improving the accuracy of calcification classification.

[0063] In step S206 of the above method for determining the coronary artery CT calcification score, determining the coronary artery CT calcification score corresponding to the coronary artery CT image based on the target mask includes: performing pixel processing on the target mask to obtain target pixels; determining the density score corresponding to each pixel in the target pixels, and determining the maximum density score of each layer of the coronary artery CT image; determining the calcification score of each layer of the coronary artery CT image based on the maximum density score of each layer and the calcification area of each layer, where the calcification area is determined based on the number of calcified region pixels of each layer and the pixel spacing of the coronary artery CT image; and determining the coronary artery CT calcification score based on the calcification score of each layer and the slice spacing of the coronary artery CT image.

[0064] In the above step, performing pixel processing on the target mask to obtain target pixels includes: assigning the pixels on the ascending aorta and descending aorta in the target mask to 0; removing the pixels with HU less than the first value on the coronary artery CT image in the target mask; dividing the target mask into lesions according to connected components, and removing the lesions with a volume less than the second value or greater than the third value, where the second value is less than the third value.

[0065] In the embodiment of the present application, classifying the coronary artery CT image by using a trained multi-class segmentation model, using the output2 corresponding to the coronary artery CT image as the above target mask, and performing subsequent processing, including:

[0066] (1) Removing the calcifications on the ascending aorta and descending aorta in the target mask, that is, assigning the pixels with the class numbers of the ascending aorta and descending aorta in the target mask to 0;

[0067] (2) Removing (i.e., assigning to 0) the pixels with HU values less than the first value on the corresponding original coronary artery CT image in the target mask. For example, the first value can be 130, but this value is only for illustration and does not represent a limitation, and can be adjusted according to the actual situation;

[0068] (3) Dividing the target mask into lesions according to connected components, and removing the lesions with a volume less than the second value or greater than the third value. The second value can be 5, and the third value can be 1,000,000 for example. It should be noted that the above second value and third value are only for illustration and do not represent a limitation, and can be adjusted according to the actual situation.

[0069] After performing the above processing on the target mask, target pixels will be obtained, and the density score corresponding to each pixel in the target pixels will be determined. Specifically, pixels are assigned density scores according to their HU values. Among them, HU values in the range of 130-199 are 1 point, 200-299 are 2 points, 300-399 are 3 points, and above 400 are 4 points. Multiply the maximum density score of each layer of the coronary CT image by the calcification area of each layer to determine the calcification score of each layer of the coronary CT image. Among them, the calcification area is equal to the product of the number of pixels in the calcification region of each layer and the pixel spacing of the coronary CT image. Finally, add up the calcification score values of all layers and multiply by the slice spacing (spacing) of the coronary CT image to obtain the overall coronary CT calcification score.

[0070] The method for determining the coronary CT calcification score provided by the embodiments of the present application can be applied to the early screening system for cardiovascular diseases in hospitals and physical examination institutions. Patients only need to undergo a non-invasive chest or coronary CT plain scan, and then a disease grade classification can be quickly obtained according to this technology, and early intervention and treatment can be carried out according to this type. Specifically, in the CT image analysis system, when the CT image of a patient is uploaded to the system, the segmentation of the calcification score and the calculation of the calcification score value can be automatically performed, and the calcification score on each coronary artery can be further subdivided. Further, the patient type can be classified according to the calcification score and recorded in the report.

[0071] Figure 4 It is a structural diagram of a device for determining the coronary CT calcification score according to an embodiment of the present application, as Figure 4 shown. The device includes:

[0072] An acquisition module 40, configured to acquire a coronary CT image to be processed;

[0073] A classification module 42, configured to process the coronary CT image by using a multi-class segmentation model to obtain a target mask. Among them, the multi-class segmentation model is trained based on the coronary CT feature map for the calcification binary segmentation task and the calcification multi-class segmentation task. The framework of the multi-class segmentation model includes a feature learning encoder, a binary classification task branch, and a multi-class classification task branch;

[0074] A determination module 44, configured to determine the coronary CT calcification score corresponding to the coronary CT image according to the target mask.

[0075] Through the acquisition module 40, the classification module 42, and the determination module 44 in the above device for determining the coronary CT calcification score, the purpose of accurately evaluating the coronary disease risk is achieved, thereby achieving the technical effect of improving the accuracy of the calcification score calculation, and further solving the technical problem of the high optimization cost of the calculation method of the coronary CT calcification score in the related technology.

[0076] In the above-mentioned coronary CT calcification score determination device, a training module 46 is further included, which is used to train a multi-class segmentation model. Specifically, the training module is used to obtain historical coronary CT images, where the historical coronary CT images include calcification category annotation information, and the calcification category annotation information is used to indicate whether there is calcification in the historical coronary CT images; determine the historical matrix data corresponding to the historical coronary CT images; randomly divide the historical matrix data to obtain a first data set, and determine the data in the first data set that includes the calcification category annotation information and the annotation information indicates the presence of calcification as the second data set; train a representation learning encoder based on the first data set and the second data set; and construct a multi-class segmentation model based on the representation learning encoder.

[0077] In the training module of the above-mentioned coronary CT calcification score determination device, the training module is further used to input the first data set into the framework model where the representation learning encoder to be trained is located, and obtain a first feature representation and a second feature representation; determine a first loss value based on the first feature representation and the second feature representation; when the change rate of the first loss value is less than a first threshold, input the second data set into the framework model, and obtain a third feature representation and a fourth feature representation; determine a second loss value based on the third feature representation and the fourth feature representation; and when the change rate of the second loss value is less than a second threshold, obtain the trained representation learning encoder from the framework model, where the first threshold is greater than the second threshold.

[0078] In the training module of the above-mentioned coronary CT calcification score determination device, the framework model includes a first encoder, a second encoder, a first multi-layer perceptron module, and a second multi-layer perceptron module. The training module is further used to perform two data augmentation processes on the first data set to obtain a first image pair. Among them, when the first image pair comes from the same image of different data augmentation processes, the first image pair is a positive sample, and when the first image pair comes from different images, the first image pair is a negative sample; input the first image pair into the first encoder and the second encoder respectively, and obtain a first image feature representation and a second image feature representation; and input the first image feature representation and the second image feature representation into the first multi-layer perceptron module and the second multi-layer perceptron module respectively, and obtain a first feature representation and a second feature representation.

[0079] In the training module of the above-mentioned coronary CT calcification score determination device, the training module is further configured to perform two types of data augmentation processing on the second data set to obtain a second image pair. Among them, when the second image pair is derived from the same image with different data augmentation processing, the second image pair is a positive sample; when the second image pair is derived from different images, the second image pair is a negative sample. The second image pair is respectively input into the first encoder and the second encoder to obtain a third image feature representation and a fourth image feature representation. The third image feature representation and the fourth image feature representation are respectively input into the first multi-layer perceptron module and the second multi-layer perceptron module to obtain a third feature representation and a fourth feature representation.

[0080] In the training module of the above-mentioned coronary CT calcification score determination device, the training module is further configured to obtain a third data set composed of an image patch corresponding to the second data set and a calcification segmentation mask pair corresponding to the second data set. The calcification segmentation mask pair includes a multi-class segmentation mask and a binary segmentation mask. The third data set is subjected to feature extraction through a representation learning encoder to obtain a coronary CT feature map. According to the coronary CT feature map and the calcification binary segmentation task in the multi-class segmentation model, a first output is determined. According to the coronary CT feature map and the calcification multi-class segmentation task in the multi-class segmentation model, a second output is determined. According to the first output and the binary segmentation mask in the second data set, a first loss function is determined, and according to the second output and the multi-class segmentation mask in the second data set, a second loss function is determined. According to the first loss function and the second loss function, a total loss function is determined, and the parameters of the total loss function are adjusted until the total loss function meets the preset conditions, and then the training is stopped to obtain a trained multi-class segmentation model.

[0081] In the training module of the above-mentioned coronary CT calcification score determination device, the calcification binary segmentation task includes a first adapter, a first decoder, and a calcification binary segmentation head. The calcification multi-class segmentation task includes a second adapter, a second decoder, and a calcification multi-class segmentation head. The training module is further configured to respectively input the coronary CT feature map into the first adapter and the second adapter for feature extraction to obtain a first coronary CT feature map and a second coronary CT feature map. The first coronary CT feature map is input into the second adapter for feature extraction, and the result of the feature extraction is fused with the second coronary CT feature map to obtain a third coronary CT feature map. The first coronary CT feature map is input into the first decoder to obtain a fourth coronary CT feature map. The fourth coronary CT feature map is input into the calcification binary segmentation head to obtain a first output. The third coronary CT feature map is input into the second decoder to obtain a fifth coronary CT feature map. The fifth coronary CT feature map is input into the calcification multi-class segmentation head to obtain a second output.

[0082] In the training module of the above-mentioned coronary CT calcium score determination device, when the representation learning encoder is Resnet, the training module is further configured to divide the coronary CT feature map into multiple windows, determine the correlation between the features of each window through the self-attention mechanism, and determine the correlation between the features of different windows.

[0083] In the determination module of the above-mentioned coronary CT calcium score determination device, the determination module is further configured to perform pixel processing on the target mask to obtain target pixels; determine the density score corresponding to each pixel in the target pixels, and determine the maximum density score of each layer of the coronary CT image; determine the calcium score of each layer of the coronary CT image according to the maximum density score of each layer and the calcium area of each layer, where the calcium area is determined according to the number of calcium region pixels in each layer and the pixel pitch of the coronary CT image; determine the coronary CT calcium score according to the calcium score of each layer and the layer spacing of the coronary CT image.

[0084] It should be noted that Figure 4 The shown coronary CT calcium score determination device is used to execute Figure 1 The shown coronary CT calcium score determination method, so the relevant explanations in the above-mentioned coronary CT calcium score determination method also apply to this coronary CT calcium score determination device, and will not be repeated here.

[0085] The method embodiment for determining the coronary CT calcium score provided by the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Figure 5 The hardware structure block diagram of a computer terminal for implementing the method for determining the coronary CT calcium score is shown. As Figure 5 shown, the computer terminal 10 may include one or more processors (processors may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA, shown as 102a, 102b,..., 102n in the figure), a memory 104 for storing data, and a transmission module 106 for communication functions connected by wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 5 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 5 shown, or have a different configuration from Figure 5 shown.

[0086] It should be noted that one or more of the above-mentioned processors and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated, in whole or in part, into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0087] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the method for determining the coronary CT calcium score in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned method for determining the coronary CT calcium score. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0088] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0089] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.

[0090] It should be noted here that in some alternative embodiments, the above-mentioned Figure 5 shown computer terminal can include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 5 is only an example of a specific specific instance and is intended to show the types of components that can exist in the above-mentioned computer terminal.

[0091] Under the above operating environment, an embodiment of the present application provides a method for determining the coronary CT calcium score. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0092] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. Among them, the memory is used to store program instructions; the processor is connected to the memory and is used to execute the method for determining the coronary CT calcium score described above.

[0093] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program. Among them, the device where the non-volatile storage medium is located executes the method for determining the coronary CT calcium score described above by running the computer program.

[0094] An embodiment of the present application also provides a computer program product, including computer instructions, which implement the steps of the method for determining the coronary CT calcium score in each embodiment of the present application when executed by a processor.

[0095] An embodiment of the present application also provides a computer program, which implements the steps of the method for determining the coronary CT calcium score in each embodiment of the present application when executed by a processor.

[0096] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0097] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0098] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0099] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0100] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical disks and other various media that can store program codes.

[0102] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for determining the coronary artery CT calcium score, characterized in that Including: Obtain the coronary CT image to be processed; Process the coronary CT image using a multi-class segmentation model to obtain a target mask. Among them, the multi-class segmentation model is trained based on the coronary CT feature map for the calcification binary segmentation task and the calcification multi-class segmentation task. The framework of the multi-class segmentation model includes a feature learning encoder, a binary classification task branch, and a multi-class classification task branch; Determine the coronary CT calcification score corresponding to the coronary CT image according to the target mask; The multi-class segmentation model is trained in the following way: Obtain historical coronary CT images, where the historical coronary CT images include calcification category annotation information, and the calcification category annotation information is used to indicate whether there is calcification in the historical coronary CT images; Determine the corresponding historical matrix data of the historical coronary CT images; Randomly divide the historical matrix data to obtain a first data set, and determine the data in the first data set that contains the calcification category annotation information and the annotation information is "with calcification" as the second data set; Train the feature learning encoder according to the first data set and the second data set; Construct the multi-class segmentation model according to the feature learning encoder; Construct the multi-class segmentation model according to the feature learning encoder, including: Obtain a third data set composed of the image blocks corresponding to the second data set and the calcification segmentation mask pairs corresponding to the second data set, where the calcification segmentation mask pairs include a multi-class segmentation mask and a binary segmentation mask; Extract features from the third data set through the feature learning encoder to obtain a coronary CT feature map; Determine a first output according to the coronary CT feature map and the calcification binary segmentation task in the multi-class segmentation model; Determine a second output according to the coronary CT feature map and the calcification multi-class segmentation task in the multi-class segmentation model; Determine a first loss function according to the first output and the binary segmentation mask in the second data set, and determine a second loss function according to the second output and the multi-class segmentation mask in the second data set; Determine the total loss function according to the first loss function and the second loss function, adjust the parameters of the total loss function, and stop training until the total loss function meets the preset conditions to obtain the trained multi-class segmentation model.

2. The method according to claim 1, wherein Training the feature learning encoder according to the first data set and the second data set includes: Input the first data set into the framework model where the feature learning encoder to be trained is located to obtain a first feature representation and a second feature representation; Determine a first loss value according to the first feature representation and the second feature representation; When the change rate of the first loss value is less than a first threshold, input the second data set into the framework model to obtain a third feature representation and a fourth feature representation; Determine a second loss value according to the third feature representation and the fourth feature representation; When the change rate of the second loss value is less than a second threshold, obtain the trained feature learning encoder from the framework model, where the first threshold is greater than the second threshold.

3. The method according to claim 2, characterized in that, The framework model includes a first encoder, a second encoder, a first multi-layer perceptron module, and a second multi-layer perceptron module. Inputting the first data set into the framework model where the representation learning encoder to be trained is located, to obtain a first feature representation and a second feature representation, including: Performing two data augmentation processes on the first data set to obtain a first image pair. Among them, when the first image pair is derived from the same image under different data augmentation processes, the first image pair is a positive sample; when the first image pair is derived from different images, the first image pair is a negative sample; Inputting the first image pair into the first encoder and the second encoder respectively to obtain a first image feature representation and a second image feature representation; Inputting the first image feature representation and the second image feature representation into the first multi-layer perceptron module and the second multi-layer perceptron module respectively to obtain the first feature representation and the second feature representation.

4. The method according to claim 3, wherein Inputting the second data set into the framework model to obtain a third feature representation and a fourth feature representation, including: Performing two data augmentation processes on the second data set to obtain a second image pair. Among them, when the second image pair is derived from the same image under different data augmentation processes, the second image pair is a positive sample; when the second image pair is derived from different images, the second image pair is a negative sample; Inputting the second image pair into the first encoder and the second encoder respectively to obtain a third image feature representation and a fourth image feature representation; Inputting the third image feature representation and the fourth image feature representation into the first multi-layer perceptron module and the second multi-layer perceptron module respectively to obtain the third feature representation and the fourth feature representation.

5. The method according to claim 1, wherein The calcification binary segmentation task includes a first adapter, a first decoder, and a calcification binary segmentation head. The calcification multi-class segmentation task includes a second adapter, a second decoder, and a calcification multi-class segmentation head. Determining a first output according to the coronary CT feature map and the calcification binary segmentation task in the multi-class segmentation model, and determining a second output according to the coronary CT feature map and the calcification multi-class segmentation task in the multi-class segmentation model, including: Inputting the coronary CT feature map into the first adapter and the second adapter respectively for feature extraction to obtain a first coronary CT feature map and a second coronary CT feature map; Inputting the first coronary CT feature map into the second adapter for feature extraction, and fusing the feature extraction result with the second coronary CT feature map to obtain a third coronary CT feature map; Inputting the first coronary CT feature map into the first decoder to obtain a fourth coronary CT feature map; Inputting the fourth coronary CT feature map into the calcification binary segmentation head to obtain the first output; Inputting the third coronary CT feature map into the second decoder to obtain a fifth coronary CT feature map; inputting the fifth coronary CT feature map into the calcification multi-class segmentation head to obtain the second output.

6. The method according to claim 5, wherein In the case where the representation learning encoder is Resnet, the method further includes: Divide the coronary CT feature map into multiple windows, and determine the correlation between the features of each window and the correlation between the features of different windows through the self-attention mechanism.

7. The method according to claim 1, characterized in that Determine the coronary CT calcium score corresponding to the coronary CT image according to the target mask, including: Perform pixel processing on the target mask to obtain target pixels; Determine the density score corresponding to each pixel in the target pixels, and determine the maximum density score of each layer of the coronary CT image; Determine the calcium score of each layer of the coronary CT image according to the maximum density score of each layer and the calcium area of each layer, wherein the calcium area is determined according to the number of pixels in the calcium region of each layer and the pixel spacing of the coronary CT image; Determine the coronary CT calcium score according to the calcium score of each layer and the slice spacing of the coronary CT image.

8. An apparatus for determining the coronary CT calcium score, characterized in that, Including: An acquisition module for acquiring a coronary CT image to be processed; A classification module for processing the coronary CT image by using a multi-class segmentation model to obtain a target mask, wherein the multi-class segmentation model is trained based on a coronary CT feature map for a two-class calcium segmentation task and a multi-class calcium segmentation task, and the framework of the multi-class segmentation model includes a representation learning encoder, a binary classification task branch, and a multi-class classification task branch; A determination module for determining the coronary CT calcium score corresponding to the coronary CT image according to the target mask; The multi-class segmentation model is trained in the following manner: acquire historical coronary CT images, wherein the historical coronary CT images include calcium category annotation information for indicating whether there is calcium in the historical coronary CT images; determine the corresponding historical matrix data of the historical coronary CT images; randomly divide the historical matrix data to obtain a first data set, and determine the data in the first data set that contains the calcium category annotation information and the annotation information indicates the presence of calcium as a second data set; train the representation learning encoder according to the first data set and the second data set; construct the multi-class segmentation model according to the representation learning encoder; Constructing the multi-class segmentation model based on the representation learning encoder includes: obtaining a third data set composed of image patches corresponding to the second data set and calcification segmentation mask pairs corresponding to the second data set, where the calcification segmentation mask pairs include multi-class segmentation masks and binary segmentation masks; extracting features from the third data set through the representation learning encoder to obtain coronary CT feature maps; determining a first output based on the coronary CT feature maps and the calcification binary segmentation task in the multi-class segmentation model; determining a second output based on the coronary CT feature maps and the calcification multi-class segmentation task in the multi-class segmentation model; determining a first loss function based on the first output and the binary segmentation masks in the second data set, and determining a second loss function based on the second output and the multi-class segmentation masks in the second data set; determining a total loss function based on the first loss function and the second loss function, adjusting the parameters of the total loss function, and stopping training until the total loss function meets a preset condition to obtain a trained multi-class segmentation model.

9. An electronic device, characterized in that, Including: A memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is used to execute the method for determining the coronary CT calcification score according to any one of claims 1 to 7.

10. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, where the device where the non-volatile storage medium is located executes the method for determining the coronary CT calcification score according to any one of claims 1 to 7 by running the computer program.

11. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method for determining the coronary CT calcification score according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Antibiotic resistance gene attribute labeling framework based on causal inference theory

    CN118380050A

  • Method and System for Assessing Vessel Obstruction Based on Machine Learning

    US20190318476A1