A method and device for training a coronary artery vulnerable plaque segmentation model based on multi-modal semi-supervised learning
Patent Information
- Application Number
- CN202510985918.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-07-17
AI Technical Summary
[0056]通过使用本申请的基于多模态半监督学习的冠状动脉易损斑块分割模型训练方法所训练的基于多模态半监督学习的冠状动脉易损斑块分割模型可克服传统依赖于手动设计损失函数的配准方法的缺点,且手动设计的损失函数往往无法满足复杂的多模态配准场景需求,相比于基于监督学习的传统深度学习配准方法,强化学习模型能够通过与环境交互自我调整与优化,算法可以根据不同图像集的特点动态调整其参数和策略,从而提高配准精度,通过迭代式的策略更新,可以在每一步不断调整和优化配准结果,提升整体精度。
Smart Images

Figure CN120876551B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a training method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, a training device for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, and a recognition method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning. Background Technology
[0002] Existing technologies employ feature extraction-based radiomics methods and multimodal image diagnosis to identify vulnerable plaques. They calculate key morphological parameters of vulnerable plaques, such as fibrous cap thickness and lipid core angle, based on image data, and assess plaque status by referencing relevant mechanical parameters. Chen Duanduan invented a multimodal image-based method for assessing vulnerable vascular plaques. First, it automatically segments the vascular lumen and outer wall from IV-OCT images and automatically extracts 2D morphological and mechanical parameters. Then, it reconstructs DSA images obtained from C-arm images containing multiple angles using computer vision methods through backprojection, obtaining an original 3D model and extracting the vascular centerline. Based on the obtained 3D vascular skeleton line and the vascular lumen obtained from IV-OCT segmentation, it constructs the spatial relationships between them for image fusion and registration. Based on the combined patient-specific 3D vascular model, it performs biomechanical calculations. Finally, based on all morphological and key mechanical parameters, it constructs a vulnerable plaque risk assessment model.
[0003] Deep learning-based object detection and segmentation techniques are also widely used in vulnerable patch detection. Compared to radiomics methods, deep learning methods can uncover deeper features of images. However, traditional deep learning models require a large amount of finely labeled data, which is often difficult to collect in real-world medical scenarios, thus limiting their widespread application.
[0004] By combining an attention model-based approach with a multi-task neural network, integrating top-down attention model noise removal with multi-task neural network classification and segmentation, this method addresses the low recall, accuracy, and overlap rates in vulnerable patch detection in OCT images, achieving efficient and accurate vulnerable patch detection. However, most current deep learning-based vulnerable patch detection methods rely on single-modality images. Single-modality images only provide specific types of information and may not be sufficient to comprehensively assess the composition and stability of patches. CT modal imaging techniques may produce artifacts, further obscuring image interpretation and affecting the accurate detection of vulnerable patches. Therefore, combining multimodal imaging can more accurately assess the state and risk of patches. A major challenge in multimodal imaging is the registration problem between different modalities, especially the registration of 2D and 3D images. Traditional registration methods often struggle to handle registration in complex scenarios involving different dimensions and modalities.
[0005] In clinical applications, invasive intravascular imaging data acquisition carries high risks and is expensive, thus its application scope and data volume are far less than that of CTA images. To address the challenge of limited, precisely labeled clinical data, we aim to utilize a semi-supervised training framework to solve the problem of segmenting vulnerable plaques in the coronary arteries.
[0006] Therefore, it is desirable to have a technical solution to overcome or at least mitigate one of the aforementioned defects of the prior art.
[0007] Application content
[0008] The purpose of this application is to provide a training method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning to overcome or at least mitigate one of the above-mentioned defects in the prior art.
[0009] To achieve the above objectives, this application provides a method for training a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning. The method includes:
[0010] Acquire a CTA coronary artery image set and a vascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the vascular lumen image set includes multiple vascular lumen images. In this case, a CTA coronary artery image and a vascular lumen image belong to the same patient and form a pair of images.
[0011] Each pair of paired images is preprocessed to obtain preprocessed paired images;
[0012] Construct a registration model and a coronary artery vulnerable plaque segmentation model based on a semi-supervised learning framework;
[0013] The registration model is trained by using preprocessed paired images from each group to obtain a trained registration model.
[0014] Each group of preprocessed paired images is input into a trained registration model to obtain a multimodal image registered by the registration model.
[0015] The coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework is trained using multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework.
[0016] Optionally, the step of preprocessing each pair of paired images to obtain preprocessed paired images includes:
[0017] The CTA coronary artery images in each paired image group were processed as follows:
[0018] The CT values in the CTA coronary artery images are compressed to 0-255 HU, and then the compressed CTA coronary artery images are normalized to between 0 and 1.
[0019] The normalized CTA coronary artery images were processed using a deep learning-based segmentation network to extract the left and right coronary artery branches, thus obtaining an optimized segmentation mask.
[0020] The key point information of the vessel centerline of the CTA coronary vessel image and the vessel lumen image after each segmentation mask is obtained respectively.
[0021] Optionally, the compression of CT values in CTA coronary artery images is performed using the following formula:
[0022]
[0023] Where ww represents the window width, wl represents the window level, x is the CT value before transformation, and y is the CT value after transformation.
[0024] Optionally, training the registration model using preprocessed paired images to obtain the trained registration model includes:
[0025] A rigid registration algorithm is used to register the preprocessed CTA coronary artery image and the intravascular lumen image, so that the two modalities can be aligned. During the registration process, the CTA image is used as the reference image and the intravascular lumen image is used as the floating image, and two registration matrices are obtained respectively. Then, the registration matrices are applied to obtain the transformed intravascular lumen image.
[0026] Fine registration is performed on the transformed intravascular images using reinforcement learning-based flexible registration.
[0027] Optionally, the registration model includes a planner, an executor, and an evaluator;
[0028] The fine registration of the transformed intravascular image using reinforcement learning-based flexible registration includes:
[0029] Based on the transformed intravascular image and the preprocessed CTA coronary artery image, obtain the preprocessed CTA coronary artery image with three cluster labels and the intravascular image with three cluster labels.
[0030] The fixed and moving images with three cluster labels are iteratively registered until the number of iterations is reached or convergence is achieved. Each iteration includes the following steps: registration step, reward value calculation step, evaluator evaluation and planner update step, and registration loss function update step.
[0031] Optionally, in the reward value calculation step, the reward value is calculated using the Dice index of its segmentation result, and the calculation method is as follows:
[0032]
[0033] I F For fixed images; I M For moving images; U F Preprocessed CTA coronary artery images with three cluster labels; U M An image of the vascular lumen with three cluster labels; Let be the cumulative deformation field at time t; Let be the cumulative deformation field at time t-1.
[0034] Optionally, the value function formula used in the evaluator evaluation and planner update steps is as follows:
[0035]
[0036] Among them, s t p represents the state at a given time t. t The plan for a given time t; k ψ For the planner's model parameters; Q θ (s t ,p t ) is the Q function of the soft Bellman residual fitting commentator;
[0037] The loss function of the planner in the evaluator evaluation and planner update steps is as follows:
[0038]
[0039] The loss function used for updating the actuator parameters is as follows:
[0040]
[0041] Optionally, the training process of the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework includes three stages. In the first stage, supervised learning is performed on a small number of paired CTA and intravascular lumen images to obtain a teacher model. The teacher model contains three branch networks: two feature extraction networks for extracting features from different modalities of images, and a multimodal feature extraction network for fusing multimodal information. It is also equipped with three different segmentation heads to output vulnerable plaque segmentation results. The three segmentation heads do not use shared network parameters, and each calculates the loss and updates the parameters using the gold standard. In order to achieve alignment of multimodal information in the training network, consistency regularization constraints are applied to the features output by the three heads.
[0042] In the second stage, the CTA branch network in the teacher model trained in the first stage is used to infer pseudo-labels from the unlabeled CTA data and initialize the pseudo-label memory. In this step, the pseudo-labels are generated only once. In the third stage of collaborative training between the teacher and student models, the pseudo-labels will be gradually corrected as the training process progresses. Therefore, the accuracy of unlabeled data labeling can be improved in each iteration.
[0043] The third stage combines labeled and unlabeled CTA data with the teacher model to train the student model. First, unlabeled and labeled data are combined as training data to train the student model. When the model is input with unlabeled data, it first performs forward inference through the teacher model to obtain the latest predicted output. The current output is then combined with the stored historical labels to perform label fusion, resulting in an updated pseudo-label, which is then updated. The consistency regularization loss is calculated between the output of the student model and the output of the teacher model. When the model is input with labeled data, the gold standard is only used for calculating the loss function of the student model.
[0044] This application also provides a training device for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, the training device comprising:
[0045] The image set acquisition module is used to acquire a CTA coronary artery image set and a vascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the vascular lumen image set includes multiple vascular lumen images. In this case, a CTA coronary artery image and a vascular lumen image belong to the same patient and form a pair of images.
[0046] The preprocessing module is used to preprocess each pair of paired images to obtain preprocessed paired images.
[0047] The model building module is used to build a registration model and a coronary artery vulnerable plaque segmentation model based on a semi-supervised learning framework.
[0048] A registration model training module is used to train the registration model using preprocessed paired images to obtain a trained registration model.
[0049] A multimodal image acquisition module is used to input each group of preprocessed paired images into a trained registration model, thereby acquiring a multimodal image registered by the registration model.
[0050] The segmentation model training module is used to train the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework using multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework.
[0051] This application also provides a method for identifying vulnerable plaques in coronary arteries based on multimodal semi-supervised learning, the method comprising:
[0052] The registration model trained by the multimodal semi-supervised learning coronary artery vulnerable plaque segmentation model training method described above, and the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework are obtained. The trained registration model and the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework constitute the segmentation model.
[0053] Acquire the image to be recognized;
[0054] The image to be identified is input into the segmentation model to obtain the recognition result.
[0055] This application has the following advantages:
[0056] The coronary artery vulnerable plaque segmentation model trained using the multimodal semi-supervised learning training method of this application overcomes the shortcomings of traditional registration methods that rely on manually designed loss functions. Moreover, manually designed loss functions often cannot meet the needs of complex multimodal registration scenarios. Compared with traditional deep learning registration methods based on supervised learning, reinforcement learning models can self-adjust and optimize through interaction with the environment. The algorithm can dynamically adjust its parameters and strategies according to the characteristics of different image sets, thereby improving registration accuracy. Through iterative strategy updates, the registration results can be continuously adjusted and optimized at each step, improving the overall accuracy.
[0057] Using a small amount of finely labeled CTA and intravascular lumen paired data (with vulnerable plaque region of interest boxes and their classification labels for vascular segments) and a large amount of unlabeled paired data, a segmentation model (Teacher Model) with weak generalization ability was trained on the small amount of paired labeled data. The Teacher Model was then used to label the unlabeled paired data with pseudo-labels. The unlabeled data (with pseudo-labels) and the labeled data were combined into a larger training set to train the Student Model. During the training process, the pseudo-labels generated by the Teacher Model were gradually optimized to supervise the training of the Student Model on the unlabeled data. KL divergence was used to constrain the distribution consistency of the output features to improve the model's generalization ability. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating a training method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, according to an embodiment of this application.
[0059] Figure 2 This is a schematic diagram of the reinforcement learning-based flexible registration proposed in this application.
[0060] Figure 3 This is a schematic diagram of the registration environment for this application.
[0061] Figure 4 This is the semi-supervised training flowchart of this application.
[0062] Figure 5 This is a flowchart of the collaborative training process for teacher and student models in this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0064] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this application.
[0065] like Figure 1 The training method for the coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, as shown, includes:
[0066] Step 1: Acquire CTA coronary artery image set and intravascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the intravascular lumen image set includes multiple intravascular lumen images. In this case, one CTA coronary artery image and one of the intravascular lumen images belong to the same patient, forming a pair of images.
[0067] Step 2: Preprocess each pair of images to obtain preprocessed paired images;
[0068] Step 3: Construct the registration model and the coronary artery vulnerable plaque segmentation model based on a semi-supervised learning framework;
[0069] Step 4: Train the registration model using the preprocessed paired images of each group to obtain the trained registration model;
[0070] Step 5: Input the preprocessed paired images of each group into the trained registration model to obtain the multimodal images registered by the registration model;
[0071] Step 6: Train the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework using the multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework.
[0072] The CTA image coronary artery segmentation results are extracted based on deep learning methods. The vessel centerline and vessel bifurcation points are obtained based on the segmentation results as rigid registration key points. A flexible registration method based on reinforcement learning is used to further align the multi-frame sequence of CTA vessel tomography and intravascular multimodal coronary artery images.
[0073] Using a small amount of finely labeled CTA and intravascular lumen paired data (with vulnerable plaque region of interest boxes and their classification labels for vascular segments) and a large amount of unlabeled paired data, a segmentation model (Teacher Model) with weak generalization ability was trained on the small amount of paired labeled data. The Teacher Model was then used to label the unlabeled paired data with pseudo-labels. The unlabeled data (with pseudo-labels) and the labeled data were combined into a larger training set to train the Student Model. During the training process, the pseudo-labels generated by the Teacher Model were gradually optimized to supervise the training of the Student Model on the unlabeled data. KL divergence was used to constrain the distribution consistency of the output features to improve the model's generalization ability.
[0074] In this embodiment, the step of preprocessing each pair of paired images to obtain preprocessed paired images includes:
[0075] The CTA coronary artery images in each paired image group were processed as follows:
[0076] The CT values in CTA coronary artery images are compressed to 0–255 HU, and then normalized to between 0 and 1. Specifically, CTA images of the heart region include many parts, such as the atria, ventricles, pulmonary arteries, bones, and coronary arteries. These areas have different densities, resulting in different HU values on the CTA image. However, the coronary arteries and surrounding tissues show little difference on CTA. To increase contrast, the window level and width are adjusted. The CT values of coronary arteries in CTA are generally between 0 and 600 HU. A window level of 300 HU and a window width of 600 HU are used for data preprocessing, calculated as follows:
[0077] The compression of CT values in CTA coronary angiography images is achieved using the following formula:
[0078]
[0079] Where ww represents the window width, wl represents the window level, x is the CT value before transformation, and y is the CT value after transformation.
[0080] The normalized CTA coronary artery images were processed using a deep learning-based segmentation network to extract the left and right coronary artery branches, resulting in an optimized segmentation mask. Specifically, deep learning was used to extract the left and right coronary arteries and their branches. The oversegmentation results were optimized using a Hessian matrix. First, the multi-scale Hessian matrix of the initial segmentation mask was calculated. Then, the vessel similarity response (e.g., the Frangi response) was calculated based on the eigenvalues to enhance the coronary structure. The initial segmentation mask was multiplied by the vessel similarity response to suppress low-response regions. Finally, low-response regions below a set threshold were removed to obtain the optimized segmentation mask.
[0081] Key point information of the vessel centerline is obtained from the CTA coronary artery image and the intravascular lumen image after each segmentation mask. Specifically, the CTA vessel segmentation result is a three-dimensional binary mask, where the vessel region is marked as the foreground with a value of 1, and the background is marked as 0. A skeletonization algorithm is used to reduce the segmented vessel region to a single-pixel-width "skeleton." A distance-transform-based method can be used, calculating the distance from each pixel to the nearest background pixel and retaining the pixel with the maximum local distance as the skeleton. Vessel skeletonization may produce some pseudo-branches or spikes, which are post-processed using a pseudo-branch removal algorithm. The skeletonization result can be represented as a graph structure, where nodes are key points on the skeleton (such as branch points or endpoints), and edges are skeleton ends connecting these nodes. The geometric features (length, diameter) of each vessel are extracted, and shorter branch structures are removed using a thresholding method.
[0082] In this embodiment, training the registration model using preprocessed paired images to obtain the trained registration model includes:
[0083] A rigid registration algorithm is used to register the preprocessed CTA coronary artery image and the intravascular lumen image, so that the two modalities can be aligned. During the registration process, the CTA image is used as the reference image and the intravascular lumen image is used as the floating image, resulting in two registration matrices. The registration matrices are then applied to obtain the transformed intravascular lumen image. In this embodiment, the geometric transformation registration matrix includes rotation and translation operations in three-dimensional space, with a total of 6 degrees of freedom, and the transformation formula is constructed as shown in the following equation.
[0084]
[0085] Where α, β, and γ represent rotations around x, y, and z, respectively, and a, b, and c represent translations along x, y, and z, respectively. By selecting a suitable similarity calculation function, the optimal matching quantity is solved using optimization algorithms such as gradient descent, Newton's method, random search, genetic algorithm, and simulated annealing, thus obtaining the optimal solution α for the parameters. * β * γ* x * y * z * .
[0086] Fine registration is performed on the transformed intravascular images using reinforcement learning-based flexible registration.
[0087] In this embodiment, the registration model includes a planner, an actor, and an evaluator; specifically, the planner, actor, and evaluator have scientific learning parameters ψ, φ, and φ, respectively. The planner's role is to generate a high-level plan in the low-dimensional latent space to guide the actor in generating the high-dimensional deformation field.
[0088] See Figure 2 The fine registration of the transformed intravascular image using reinforcement learning-based flexible registration includes:
[0089] Based on the transformed intravascular image and the preprocessed CTA coronary artery image, obtain the preprocessed CTA coronary artery image with three cluster labels and the intravascular image with three cluster labels.
[0090] The fixed and moving images with three cluster labels are iteratively registered until the number of iterations is reached or convergence is achieved. Each iteration includes the following steps: registration step, reward value calculation step, evaluator evaluation and planner update step, and registration loss function update step.
[0091] In this embodiment, leveraging the sequential nature of reinforcement learning, we decompose large deformation registration into T steps instead of predicting the deformation field of the final transformation all at once.
[0092] At time step t, action a t It is the planner k ψ and actuator π φ Based on fixed image I F and moving images in the middle The deformation field at time t during generation.
[0093] It is a t The deformation field at the previous time step t-1 The generated data can be represented as:
[0094]
[0095] At each time step, the cumulative deformation field is used. Successively moving image I M Deformation Then deform the image Used as the moving image in the next time step t+1.
[0096]
[0097] Where G represents the similarity measure between the fixed image and the deformed image (including but not limited to the sum of squared differences, normalized mutual information, and negative normalized cross-correlation). R represents the smoothing constraint term of the deformed field to prevent large abrupt changes in the deformed field.
[0098] In this embodiment, Figure 3 This section presents an overview of the stepwise registration environment, which contains only a pair of fixed and moving images. Three cluster labels are generated using the K-means unsupervised segmentation method, resulting in segmented images. While the generated segmentation images cannot accurately segment every anatomical structure, each voxel can be assigned to a virtual anatomical structure for reward calculation. The reward is calculated using the Dice metric of the segmentation results, as follows:
[0099]
[0100] I F For fixed images; I M For moving images; U F Preprocessed CTA coronary artery images with three cluster labels; U M An image of the vascular lumen with three cluster labels; Let be the cumulative deformation field at time t; Let be the cumulative deformation field at time t-1.
[0101] In this embodiment, the present application adopts the following meta policy:
[0102] Given the state s at time t t A stochastic scheme can be modeled as a subspace of a deformable field, within which a low-dimensional vector p is obtained. t The actions produced by the actuator are actually planned by p. t The determined deformation field is used to guide the transformation of the floating image. The model parameters for the planner and executor are defined as k, respectively. ψ and π φ randomized plan p t The sampling can be represented as: p t ~k ψ (p t |s t Then, through the actuator, p t Decoded as a deformable field in a high-dimensional space: a t =πφ (a t |p t The overall optimization objective can be expressed as:
[0103]
[0104] Where α is the temperature parameter, used to control the entropy H (the policy in state s). t Entropy (a measure of the randomness of action distribution) and reward r t The balance; ρ (k,π) It is k ψ (p t |s t ) and π φ (a t |p t The trajectory distribution of ).
[0105] To effectively improve learning efficiency and reduce learning difficulty, the commentator Q... θ The evaluation is of plan p t Instead of action a t Low-dimensional plan p t The downsampled vector is concatenated to the commentator and outputs the soft-Q function Q. θ (s t ,p t The current state plan value is estimated. In this scheme, a commentator is used to evaluate the planner, where the reward and soft Q-value, calculated based on the Dice difference between the current step t and the previous step t-1, are used to iteratively guide the improvement of the stochastic policy. The classic Soft Actor-Critic (SAC) is an off-policy algorithm based on a maximum entropy framework, combining the Actor-Critic architecture and entropy regularization techniques. Unlike traditional policy gradient methods, SAC not only attempts to maximize the expected reward but also the policy entropy, leading to a more uniform action distribution under the same policy reward, thereby enhancing exploration ability and algorithm robustness, and avoiding getting trapped in local optima. In this application, the Stochastic Planner-Actor-Critic (SPAC) algorithm is adopted, where the planner learns the policy to generate the plan, and the commentator's Q-function Q is fitted by minimizing the soft Bellman residual. θ (s t ,p t ), which uses converted data sampled from playback buffer pool D.
[0106] In this embodiment, the value function formula used in the evaluator evaluation and planner update steps is as follows:
[0107]
[0108] Among them, s t p represents the state at a given time t. t The plan for a given time t; k ψ For the planner's model parameters; Q θ (s t ,p t ) is the Q function of the soft Bellman residual fitting commentator;
[0109] Use target network To stabilize training, the parameters It is obtained through the exponential moving average of the critic network parameters. The update method is J can be optimized using the stochastic gradient descent algorithm. Q (θ).
[0110] Since the evaluator is based on the planner's output, the optimization process will also influence the planner's decisions. Therefore, we can minimize the KL divergence between the policy and the Boltzmann distribution guided by the Q function. The loss function of the planner in the evaluator evaluation and planner update steps is as follows:
[0111]
[0112] In this embodiment, the planner and executor learn based on unsupervised registration.
[0113] The similarity between fixed images (CTA images) and deformed images (vascular lumen images) is learned using locally normalized cross-correlation:
[0114]
[0115] A higher value indicates better registration and alignment. Additionally, to produce a smooth deformation field, a total variational regularizer is used to smooth the spatial gradient of the deformation field.
[0116]
[0117] The loss function used for updating the actuator parameters is as follows:
[0118]
[0119] In this embodiment, the semi-supervised learning framework is divided into three stages. The first stage first performs supervised learning on a small amount of labeled data with paired CTA and intravascular lumen images to train a Teacher model, which includes three network modules: the Backbone network is a feature extraction module, which can use CNN feature extraction networks such as ResNet and DenseNet; the FPN network is a feature pyramid, which upsamples the small-sized feature map and adds it to the large-sized feature map to obtain a feature map that integrates multi-scale information, and then inputs it into the RPN network (Region Generation Network) to extract candidate boxes. The output candidate boxes and the gold standard box are used to calculate the box loss, and the output class cross-entropy loss is also calculated.
[0120] See Figure 4 Phase 1: Supervised learning is performed on a small number of paired CTA and intravascular images to obtain the Teacher Model. The Teacher Model contains three branch networks: two feature extraction networks for extracting features from different modalities of the images, and a multimodal feature extraction network for fusing multimodal information. It also includes three different segmentation heads for outputting vulnerable plaque segmentation results. The three segmentation heads do not share network parameters; each calculates its loss using the gold standard and updates its parameters. Furthermore, to achieve alignment of multimodal information during network training, consistency regularization constraints are applied to the features output by the three heads.
[0121] Phase 2: Using the CTA branch network in the teacher model trained in Phase 1, inference is performed on the unlabeled CTA data to obtain pseudo-labels and the pseudo-label memory is initialized. In this step, the pseudo-labels are generated only once. In the third phase of the collaborative training of the teacher and student models, the pseudo-labels will be gradually corrected as the training process progresses. Therefore, the accuracy of unlabeled data labeling can be improved in each iteration.
[0122] Phase 3: Training the student model by combining labeled and unlabeled CTA data with the teacher model, such as... Figure 5As shown, the student model is first trained by combining unlabeled and labeled data. When the model receives unlabeled data as input, it first performs forward inference through the teacher model to obtain the latest predicted output. This current output is then fused with historical labels stored in the Pseudo Labels Memory to obtain updated pseudo-labels, which are then updated in the Pseudo Labels Memory. A consistency regularization loss is calculated between the student model's output and the teacher model's output. This aims to guide the student model's parameter updates using the teacher model's parameters and output distribution. When the model receives labeled data as input, the gold standard is only used to calculate the student model's loss function, ensuring that the student model also possesses strong supervision capabilities.
[0123] The consistency regularization loss used in this scheme is used to constrain the consistency of the output feature distribution of different models. It can be calculated using methods such as L2 distance loss function, KL divergence or cosine similarity.
[0124] Loss function based on L2 distance regularization:
[0125]
[0126] Loss function based on KL divergence regularization:
[0127]
[0128] Cosine similarity-based regularized loss function:
[0129]
[0130] During the training of the Student Model, data augmentation can be performed on the labeled data, such as translation, rotation, inversion, and Gaussian blur, to increase the robustness of the model.
[0131] This application also provides a training device for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning. The training device includes an image set acquisition module, a preprocessing module, a model construction module, a registration model training module, a multimodal image acquisition module, and a segmentation model training module.
[0132] The image set acquisition module is used to acquire CTA coronary artery image set and vascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the vascular lumen image set includes multiple vascular lumen images. In this case, a CTA coronary artery image and a vascular lumen image belong to the same patient and form a pair of images.
[0133] The preprocessing module is used to preprocess each pair of images separately to obtain preprocessed paired images;
[0134] The model building module is used to build registration models and coronary artery vulnerable plaque segmentation models based on a semi-supervised learning framework;
[0135] The registration model training module is used to train the registration model using each group of preprocessed paired images, thereby obtaining the trained registration model;
[0136] The multimodal image acquisition module is used to input each group of preprocessed paired images into a trained registration model, thereby obtaining multimodal images registered by the registration model.
[0137] The segmentation model training module is used to train the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework using multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A training method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, characterized in that, The training method for the coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning includes: Acquire a CTA coronary artery image set and a vascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the vascular lumen image set includes multiple vascular lumen images. In this case, a CTA coronary artery image and a vascular lumen image belong to the same patient and form a pair of images. Each pair of paired images is preprocessed to obtain preprocessed paired images; Construct a registration model and a coronary artery vulnerable plaque segmentation model based on a semi-supervised learning framework; The registration model is trained by using each group of preprocessed paired images to obtain the trained registration model. Each group of preprocessed paired images is input into a trained registration model to obtain a multimodal image registered by the registration model. The coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework is trained using multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework. The training process of the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework includes three stages. In the first stage, a small number of CTA and intravascular lumen images are paired for supervised learning to obtain a teacher model. The teacher model contains three branch networks: two feature extraction networks for extracting features from different modalities of images, and a multimodal feature extraction network for fusing multimodal information. It is also equipped with three different segmentation heads to output vulnerable plaque segmentation results. The three segmentation heads do not use shared network parameters, and each calculates the loss and updates the parameters using the gold standard. In order to achieve alignment of multimodal information in the training network, consistency regularization constraints are applied to the features output by the three heads. In the second stage, the CTA branch network in the teacher model trained in the first stage is used to infer pseudo-labels from the unlabeled CTA data and initialize the pseudo-label memory. In this step, the pseudo-labels are generated only once. In the third stage of collaborative training between the teacher and student models, the pseudo-labels will be gradually corrected as the training process progresses. Therefore, the accuracy of unlabeled data labeling can be improved in each iteration. The third stage combines labeled and unlabeled CTA data with the teacher model to train the student model. First, unlabeled and labeled data are combined as training data to train the student model. When the model is input with unlabeled data, it first performs forward inference through the teacher model to obtain the latest predicted output. The current output is then combined with the stored historical labels to perform label fusion, resulting in an updated pseudo-label, which is then updated. The consistency regularization loss is calculated between the output of the student model and the output of the teacher model. When the model is input with labeled data, the gold standard is only used for calculating the loss function of the student model.
2. The method for training a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning as described in claim 1, characterized in that, The step of preprocessing each pair of images to obtain preprocessed paired images includes: The CTA coronary artery images in each paired image group were processed as follows: The CT values in the CTA coronary artery images are compressed to 0~255HU, and then the compressed CTA coronary artery images are normalized to between 0 and 1. The left and right coronary artery branches were extracted from the normalized CTA coronary artery images using a deep learning-based segmentation network to obtain an optimized segmentation mask. The key point information of the vessel centerline of the CTA coronary vessel image and the vessel lumen image after each segmentation mask is obtained respectively.
3. The method for training a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning as described in claim 2, characterized in that, The compression of CT values in CTA coronary angiography images is achieved using the following formula: ; in, Indicates window width. The window level is represented by x, where x is the CT value before transformation and y is the CT value after transformation.
4. The method for training a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning as described in claim 3, characterized in that, The step of training the registration model using preprocessed paired images to obtain the trained registration model includes: A rigid registration algorithm is used to register the preprocessed CTA coronary artery image and the intravascular lumen image, so that the two modalities can be aligned. During the registration process, the CTA image is used as the reference image and the intravascular lumen image is used as the floating image, and two registration matrices are obtained respectively. Then, the registration matrices are applied to obtain the transformed intravascular lumen image. Fine registration is performed on the transformed intravascular images using reinforcement learning-based flexible registration.
5. The method for training a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning as described in claim 4, characterized in that, The registration model includes a planner, an executor, and an evaluator; The fine registration of the transformed intravascular image using reinforcement learning-based flexible registration includes: Based on the transformed intravascular image and the preprocessed CTA coronary artery image, obtain the preprocessed CTA coronary artery image with three cluster labels and the intravascular image with three cluster labels. The fixed and moving images with three cluster labels are iteratively registered until the number of iterations is reached or convergence is achieved. Each iteration includes the following steps: registration step, reward value calculation step, evaluator evaluation and planner update step, and registration loss function update step.
6. A training device for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning, used in the training method for a coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning as described in any one of claims 1 to 5, characterized in that, The training device for the coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning includes: The image set acquisition module is used to acquire a CTA coronary artery image set and a vascular lumen image set. The CTA coronary artery image set includes multiple CTA coronary artery images, and the vascular lumen image set includes multiple vascular lumen images. In this case, a CTA coronary artery image and a vascular lumen image belong to the same patient and form a pair of images. The preprocessing module is used to preprocess each pair of paired images to obtain preprocessed paired images. The model building module is used to build a registration model and a coronary artery vulnerable plaque segmentation model based on a semi-supervised learning framework. A registration model training module is used to train the registration model using preprocessed paired images to obtain a trained registration model. A multimodal image acquisition module is used to input each group of preprocessed paired images into a trained registration model, thereby acquiring a multimodal image registered by the registration model. The segmentation model training module is used to train the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework using multimodal images registered by each group of registration models, thereby obtaining the trained coronary artery vulnerable plaque segmentation model based on the semi-supervised learning framework.
7. A recognition method for a segmentation model of vulnerable plaques in coronary arteries based on multimodal semi-supervised learning, characterized in that, The identification method used in the coronary artery vulnerable plaque segmentation model based on multimodal semi-supervised learning includes: The registration model and the coronary artery vulnerable plaque segmentation model based on the semi-supervised learning training method described in any one of claims 1 to 5 are obtained, and the training model is composed of the registration model and the semi-supervised learning framework. Acquire the image to be recognized; The image to be identified is input into the segmentation model to obtain the recognition result.
Citation Information
Patent Citations
Coronary plaque type determination method and device, electronic equipment and storage medium
CN116664938A