A blood vessel OCT image training method and device based on a DCN cross-modal fusion network and structure-blood flow coupling modeling, and a blood vessel OCT image generation method and device
By using DCN cross-modal fusion network and structure-blood flow coupling modeling, combined with CTA and OCT images and hemodynamic parameters, high-resolution OCT images are generated, which solves the problems of high invasiveness of OCT imaging and insufficient resolution of CTA, and realizes non-invasive and efficient plaque feature display and diagnostic assistance.
Patent Information
- Application Number
- CN202510922700.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-07-04
AI Technical Summary
In existing technologies, OCT imaging is highly invasive, expensive, and has insufficient spatial resolution in the diagnosis of coronary artery stenosis and vulnerable plaques, while CTA imaging cannot clearly display key plaque features.
A method based on DCN cross-modal fusion network and structure-flow coupling modeling is adopted. By acquiring and preprocessing coronary CTA and OCT images and combining hemodynamic parameters, a cross-modal fusion network and an OCT generation network are trained to generate high-resolution OCT images.
It enables the non-invasive generation of high-resolution OCT images, which can clearly display key plaque features such as fibrous cap thickness and microcalcifications, replacing some invasive OCT examinations and assisting doctors in assessing plaque vulnerability and guiding interventional treatment.
Smart Images

Figure CN120807442B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method for training vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling, a method for generating vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling, and a device for training vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling. Background Technology
[0002] In the diagnosis and analysis of coronary artery stenosis and vulnerable plaques, OCT imaging is the "gold standard" for assessing the microstructure of coronary and cerebral blood vessels, but it has significant limitations. OCT imaging requires invasive catheter manipulation, posing risks of vascular injury, thrombosis, and infection, and is poorly tolerated by patients, especially the elderly or those with comorbidities. Furthermore, a single OCT test is expensive and requires a specialized interventional team. Computed tomography angiography (CTA) is a recognized non-invasive imaging method for diagnosing coronary arteries and is currently an important clinical method for early screening and definitive diagnosis of coronary artery disease. It can be used to identify coronary plaques and assess the degree of vascular stenosis. Although CTA is non-invasive, its spatial resolution is insufficient, and it cannot clearly display key plaque features such as fibrous cap thickness and microcalcifications.
[0003] Therefore, it is desirable to have a technical solution to overcome or at least mitigate one of the aforementioned defects of the prior art.
[0004] Application content
[0005] The purpose of this application is to provide a vascular OCT image training method based on DCN cross-modal fusion network and structure-blood flow coupling modeling to overcome or at least mitigate one of the above-mentioned defects of the prior art.
[0006] To achieve the above objectives, this application provides a method for training vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling. The method includes:
[0007] Acquire coronary CTA and OCT images of multiple groups of patients, with the spatial location of the OCT images initially aligned with that of the coronary CTA images;
[0008] Each of the CTA images is preprocessed to obtain image data marked with blood vessel segmentation results and plaque locations;
[0009] Each of the OCT images is preprocessed to obtain preprocessed OCT image data;
[0010] Image data belonging to the same patient, labeled with vascular segmentation results and plaque locations, and preprocessed OCT image data are registered to obtain each registered image data.
[0011] Obtain hemodynamic parameters of the same patient; obtain cross-modal fusion network and OCT generation network based on deformable convolution;
[0012] The deformable convolution-based cross-modal fusion network and the OCT generation network are trained using the registered image data and hemodynamic parameters, respectively, to obtain the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
[0013] Optionally, acquiring coronary CTA and OCT images of multiple patients, wherein the spatial location of the OCT images is initially aligned with that of the coronary CTA images, includes:
[0014] Collect N sets of paired coronary CTA and OCT images. The CTA data are high-resolution CT angiography covering the target vessel and include enhanced information on the vessel lumen, plaque, and vessel wall. The OCT data should be intravascular OCT images of the same patient as the CTA to ensure initial spatial alignment with the CTA.
[0015] Optionally, the preprocessing of each of the CTA images to obtain image data labeled with vessel segmentation results and plaque locations includes:
[0016] Obtain a pre-trained segmentation model;
[0017] Each of the CTA images is input into a pre-trained segmentation model to obtain mask image data after segmenting blood vessels and plaques;
[0018] Noise is removed from the mask image data by combining morphological operations, while preserving the topology of branch vessels;
[0019] Using the blood vessel centerline as a reference, a unified coordinate system origin and direction matrix are established for the mask image data, thereby obtaining image data marked with blood vessel segmentation results and plaque locations.
[0020] Optionally, the registration of image data with labeled blood vessel segmentation results and plaque locations belonging to the same patient and preprocessed OCT image data to obtain registered image data includes:
[0021] The OCT image is converted from polar coordinates to Cartesian coordinates based on the duct marker points. Then, motion artifact compensation and coordinate system transformation are used to correct the image, thereby obtaining the corrected OCT image.
[0022] Optionally, the process of registering the image data marked with blood vessel segmentation results and plaque locations with the preprocessed OCT image data to obtain registered image data includes:
[0023] The corrected OCT images were matched with the corresponding CTA locations by identifying vascular branch points and calcified plaque positions, and the registration was optimized using a mutual information maximization algorithm.
[0024] This application also provides a method for generating vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling. The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling includes:
[0025] Obtain the deformable convolution-based cross-modal fusion network and the trained OCT generation network after training using the vascular OCT image training method based on DCN cross-modal fusion network and structure-blood flow coupling modeling as described above;
[0026] Acquire coronary CTA images and hemodynamic parameters of the patient to be generated;
[0027] The coronary CTA images of the patient to be generated and the hemodynamic parameters are input into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network.
[0028] The CTA features of the cross-modal fusion network are input into the trained OCT generation network to obtain a cross-sectional OCT image containing information about the blood vessel wall and plaque.
[0029] Optionally, the step of inputting the coronary CTA image of the patient to be generated and the hemodynamic parameters into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network includes:
[0030] The coronary CTA images of the patient to be generated are encoded into CTA feature maps using a 3D deformable convolutional network;
[0031] The coronary CTA images of the patient to be generated are preprocessed to obtain image data marked with vessel segmentation results and plaque locations;
[0032] Image data labeled with blood vessel segmentation results and plaque locations are encoded into a heatmap, which is then multiplied channel-by-channel with the CTA feature map to obtain the first fusion feature.
[0033] Extracting dynamic features based on hemodynamic parameters;
[0034] The CTA features of the cross-modal fusion network are obtained by fusing dynamic features with the first fusion feature using a multi-head attention mechanism.
[0035] Optionally, the trained OCT generation network includes a structural similarity loss function, which is as follows:
[0036] L smi =α·‖y gen -y real ||1+(1-α)·(1-SSIM(y)|| gen ,y real ))
[0037] Where α is the weighting coefficient, SSIM is the structural similarity index, and L smi Let y be the structural similarity loss function. gen It is the generated OCT image, y real These are real OCT images;
[0038] The trained OCT generation network further includes a structure-blood flow coupling loss, which is as follows:
[0039]
[0040] Where, θ lipid M represents the lipid core curvature in the OC image. loss-wss For low WSS mask, L align It is the structure-blood flow loss term.
[0041] Optionally, the final loss function of the trained OCT generation network is as follows:
[0042] L all =L smi +γL align
[0043] Among them, L all For the final loss function, L smi Let L be the structural similarity loss function; γ = 0.5; align This is due to structural blood flow coupling loss.
[0044] This application also provides a vascular OCT image training device based on DCN cross-modal fusion network and structure-blood flow coupling modeling, wherein the vascular OCT image training device based on DCN cross-modal fusion network and structure-blood flow coupling modeling includes:
[0045] The image acquisition module is used to acquire multiple sets of coronary CTA images and OCT images based on the same patient, wherein the spatial position of the OCT images is initially aligned with the coronary CTA images;
[0046] The CTA image preprocessing module is used to preprocess each of the CTA images to obtain image data marked with blood vessel segmentation results and plaque locations.
[0047] The OCT image preprocessing module is used to preprocess each of the OCT images to obtain preprocessed OCT image data.
[0048] The registration module is used to register image data of the same patient with marked blood vessel segmentation results and plaque locations, as well as preprocessed OCT image data, to obtain registered image data.
[0049] A hemodynamic parameter acquisition module, which is used to acquire hemodynamic parameters of the same patient;
[0050] The model acquisition module is used to acquire a cross-modal fusion network based on deformable convolution and an OCT generation network;
[0051] The training module is used to train the deformable convolution-based cross-modal fusion network and the OCT generation network using the registered image data and hemodynamic parameters respectively, so as to obtain the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
[0052] This application generates high-resolution OCT images by fusing anatomical features from CTA imaging with CFD (computed tomography) hemodynamic parameters of blood vessels, and assesses plaque vulnerability. This method can replace some invasive OCT examinations, assisting physicians in assessing plaque vulnerability and guiding interventional treatment. Attached Figure Description
[0053] Figure 1 This is a flowchart illustrating a method for training vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling according to an embodiment of this application.
[0054] Figure 2 This is a schematic diagram of deformable convolution in this application.
[0055] Figure 3 This is a schematic diagram of the DCN-based cross-modal fusion network of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are some, but not all, embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0057] In the description of this application, it should be understood that the terms "center", "longitudinal", "lateral", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the scope of protection of this application.
[0058] like Figure 1 The vascular OCT image training method based on DCN cross-modal fusion network and structure-flow coupling modeling shown includes:
[0059] Step 1: Acquire coronary CTA and OCT images of multiple patients, with the spatial location of the OCT images initially aligned with that of the coronary CTA images;
[0060] Step 2: Preprocess each of the CTA images to obtain image data marked with blood vessel segmentation results and plaque locations;
[0061] Step 3: Preprocess each of the OCT images to obtain preprocessed OCT image data;
[0062] Step 4: Register the image data of the same patient, which are marked with the results of vessel segmentation and the location of plaques, with the preprocessed OCT image data to obtain the registered image data;
[0063] Step 5: Obtain hemodynamic parameters for the same patient;
[0064] Step 6: Obtain the cross-modal fusion network based on deformable convolution and the OCT generation network;
[0065] Step 7: Train the deformable convolution-based cross-modal fusion network and the OCT generation network using the registered image data and hemodynamic parameters respectively, thereby obtaining the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
[0066] In this embodiment, acquiring multiple sets of coronary CTA and OCT images based on the same patient, wherein the spatial location of the OCT images is initially aligned with that of the coronary CTA images, includes:
[0067] Collect N sets of paired coronary CTA and OCT images. The CTA data are high-resolution CT angiography covering the target vessel and include enhanced information on the vessel lumen, plaque, and vessel wall. The OCT data should be intravascular OCT images of the same patient as the CTA to ensure initial spatial alignment with the CTA.
[0068] For example, there could be data from 100 patient groups, each group including at least one coronary CTA image (when there are more than two, the two coronary CTA images have different time sequences) and at least one OCT image (when there are more than two, the two OCT images have different time sequences). The spatial location of the OCT images for each patient group is initially aligned with the coronary CTA images. It can be understood that if there are more than two images in time sequence, they can also be aligned in time sequence.
[0069] In this embodiment, acquiring coronary CTA and OCT images of multiple patients, and initially aligning the spatial location of the OCT images with the coronary CTA images, includes:
[0070] Obtain a pre-trained segmentation model;
[0071] Each of the CTA images is input into a pre-trained segmentation model to obtain mask image data after segmenting blood vessels and plaques;
[0072] Noise is removed from the mask image data by combining morphological operations, while preserving the topology of branch vessels;
[0073] Using the blood vessel centerline as a reference, a unified coordinate system origin and direction matrix are established for the mask image data, thereby obtaining image data marked with blood vessel segmentation results and plaque locations.
[0074] In this embodiment, the preprocessing of each of the OCT images to obtain the preprocessed OCT image data includes:
[0075] The OCT image is converted from polar coordinates to Cartesian coordinates based on catheter markers (such as radio frequency signals from the OCT imaging catheter). Then, motion artifact compensation and coordinate system transformation are used to correct the image, thereby obtaining the corrected OCT image.
[0076] In this embodiment, the registration of image data belonging to the same patient, marked with vessel segmentation results and plaque locations, and preprocessed OCT image data to obtain registered image data includes:
[0077] The corrected OCT images were matched with the corresponding CTA locations using vascular branching points and calcified plaque locations, and the registration was optimized using the mutual information maximization algorithm.
[0078] In this embodiment, CTA images are segmented into vessels and plaques using models including but not limited to 3D U-Net++, and the centerline and lumen-wall boundary are extracted. Then, morphological operations (such as erosion and expansion) are used to remove noise while preserving the topology of branch vessels. Finally, the origin and direction matrix of the coordinate system are unified using the vessel centerline as a reference. OCT images are transformed from polar coordinates to Cartesian coordinates based on catheter markers (such as the radiofrequency signal of the OCT imaging catheter), and motion artifact compensation and coordinate system transformation are used to correct the image. Finally, the corresponding positions of the OCT cross-section and CTA are matched using vessel branching points and calcified plaque locations, and the registration is optimized using a mutual information maximization algorithm.
[0079] This application also provides a method for generating vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling. The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-blood flow coupling modeling includes:
[0080] Obtain the deformable convolution-based cross-modal fusion network and the trained OCT generation network after training using the vascular OCT image training method based on DCN cross-modal fusion network and structure-blood flow coupling modeling as described above;
[0081] Acquire coronary CTA images and hemodynamic parameters of the patient to be generated;
[0082] The coronary CTA images of the patient to be generated and the hemodynamic parameters are input into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network.
[0083] The CTA features of the cross-modal fusion network are input into the trained OCT generation network to obtain a cross-sectional OCT image containing information about the blood vessel wall and plaque.
[0084] In this embodiment, the step of inputting the coronary CTA image of the patient to be generated and the hemodynamic parameters into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network includes:
[0085] The coronary CTA images of the patient to be generated are encoded into CTA feature maps using a 3D deformable convolutional network;
[0086] The coronary CTA images of the patient to be generated are preprocessed to obtain image data marked with vessel segmentation results and plaque locations;
[0087] Image data labeled with blood vessel segmentation results and plaque locations are encoded into a heatmap, which is then multiplied channel-by-channel with the CTA feature map to obtain the first fusion feature.
[0088] Extracting dynamic features based on hemodynamic parameters;
[0089] The CTA features of the cross-modal fusion network are obtained by fusing dynamic features with the first fusion feature using a multi-head attention mechanism.
[0090] Specifically, the inputs to the deformable convolution-based cross-modal fusion network of this application include: CTA images, vessel segmentation results, plaque locations, and hemodynamic distributions calculated by CFD.
[0091] To adapt to the diversity of blood vessel tortuosity and plaque morphology, and to improve the feature extraction capability of complex anatomical structures, a 3D deformable convolutional network is used to encode CTA image blocks (64×64×64) into CTA feature maps. Let the input CTA image be I. CTA ∈R H ×W×D After passing through 3D deformable convolutional layers:
[0092] F CTA =DCN(I CTA ;θ DCN )
[0093] Where, θ DCN It is the deformable convolution parameter, F CTA This is the output feature map. In the deformable convolution formula:
[0094]
[0095] Where p represents the coordinates of a certain position on the output feature map (e.g., (x, y, z) in a 3D image), corresponding to the center position of the region to be sampled on the input feature map, F out (p) represents the output value at that position. k Δp is the standard offset of the k-th sampling point within the convolution kernel. k It is a learnable offset used to dynamically adjust the position of each sampling point, w k These are weight parameters.
[0096] Δp k It is the core of deformable convolution, allowing the model to adaptively adjust the sampling position to adapt to complex geometries (such as tortuous blood vessels), generated from the input feature map and a single-layer convolution, such as... Figure 2 As shown.
[0097] The CTA segmentation results (vascular mask, plaque location) are encoded into a heatmap and multiplied channel-by-channel with the feature map to enhance the response of the vessel wall and plaque region and suppress background noise.
[0098] F CTA =F CTA ⊙(Sigmoid(M vessel +H plaque ))
[0099] Among them, M vessel It is a vascular mask, H plaque This represents the location of the center point of the plaque.
[0100] To incorporate hemodynamic parameters calculated by CFD, the wall shear force and blood flow velocity distribution simulated by CFD were first extracted using a 3D convolutional network (two layers of 3D convolution, channel count: 2→32→64, activation function: LeakyReLU). Since the CFD features and CTA images were already registered and aligned, they were fused using a multi-head attention mechanism to enhance the ability to capture different modal relationships, resulting in further enhanced CTA features. Finally, the enhanced CTA features were input into an OCT generation network to generate a reasonable OCT map.
[0101] Multi-head attention is used to enhance the model's ability to capture relationships between different modalities:
[0102] Attention(Q,K,V)=Concat(head1,…,head h W o
[0103] The calculation for each attention head is as follows:
[0104]
[0105] Where Q is a linear mapping of CTA features, and the number of channels in Q is defined as C. q K and V represent the CFD features. h is the number of attention heads; this invention sets 8 or 16 for optimal results. d k =C q / h represents the dimension of each head. W i Q , W i V W represents the learnable projection matrix. o This indicates the output projection matrix.
[0106] To enhance the response to local hemodynamic characteristics, a spatial modulation factor α∈R is introduced. B×1×H×W×D :
[0107] α=σ(Conv3D([Q;K]))
[0108] Where σ is the Sigmoid function, the final fused feature is:
[0109] F fusion =Q + α⊙Attention(Q,K,V)
[0110] In this embodiment, the trained OCT generation network is based on an OCT generation network including but not limited to a conditional diffusion model (or other diffusion models). The input is the CTA features of the cross-modal fusion network, and the final output is a cross-sectional OCT map containing information about the blood vessel wall and plaque.
[0111] To stabilize the OCT image generation process and improve generation quality, this invention employs two loss terms—image similarity loss and structure-blood flow coupling loss—for constraint. The image similarity loss consists of image structural similarity and pixel-level L1 norm. The L1 norm constrains the absolute difference in pixel values between the generated image and the real image, preventing excessive blurring or deviation from the basic grayscale distribution. The structural similarity loss constrains the continuity between the generated image and the real OCT image in key structures such as vessel wall layering and plaque edges.
[0112] The trained OCT generation network includes a structural similarity loss function, which is as follows:
[0113] L smi =α·‖y gen -y real ||1+(1-α)·(1-SSIM(y)|| gen ,y real ))
[0114] Where α is the weighting coefficient and SSIM is the structural similarity index, multiple experiments show that a smaller α value (less than 0.3) results in better continuity and image quality in the OCT image, but also introduces problems such as missing positions and content distortion. A larger α value (greater than 0.5) results in more complete content, but poorer continuity; this invention sets α to 0.35 as the final parameter; y gen It is the generated OCT image, y real It is a real OCT image;
[0115] To enhance the pathological guidance of hemodynamic parameters in OCT image generation, a structure-blood flow coupling loss was designed to define the association rules between hemodynamic parameters (such as low WSS regions) and plaque vulnerability. The loss function constraint was used to better generate the lipid core curvature and fibrous cap thickness of OCT.
[0116] Generate a low WSS mask:
[0117] M loss-WSS =Π(WSS<1.5Pa)∈{0,1}
[0118] In this embodiment, if the CFD simulation shows a localized low WSS, then the lipid core curvature (θ) in the corresponding generated OCT image is... lipid (≥180°), otherwise apply L2 norm penalty.
[0119]
[0120] Where, θ lipid M represents the lipid core curvature in the OCT image. loss-wss For low WSS mask, L align It is the structure-blood flow loss term;
[0121] The final loss term for the network is as follows, γ = 0.5:
[0122] L all =L smi +γL align
[0123] In this embodiment, training employs a two-stage approach. The first stage uses one-third of the data to train a cross-modal fusion network based on deformable convolutions and an OCT generation network, optimizing only pixel-level L1 and SSIM losses. In the second stage (fine-tuning), a structural blood flow coupling loss is introduced, and all data is used for training. A gradient pruning strategy is employed during training to prevent generation mode collapse.
[0124] Two post-processing operations are performed on the generated OCT images: first, based on the confidence map of the generated images, nonlocal mean denoising is used to eliminate generated noise; second, morphological refinement is performed on the vascular intima boundary to improve the visualization accuracy of the fibrous cap thickness.
[0125] This invention proposes a vascular OCT image generation technology based on Deformable Convolution Network (DCN) cross-modal fusion network and structure-flow coupling modeling. Through cross-modal feature fusion and physical constraint modeling, it realizes cross-scale and cross-functional high-fidelity image generation from CTA to OCT.
[0126] The overall process of this invention is as follows: Figure 1 As shown, paired CTA (3D images) and OCT cross-sectional images (intravascular OCT images from the same patient as the CTA) were collected. First, 3D U-Net++ was used to segment the CTA images into vessels and plaques, extracting the vessel centerline, lumen-wall structure, and plaque structure. Simultaneously, motion artifact compensation and coordinate system transformation were performed on the paired OCT images to correct the images. Then, the corresponding positions of the OCT cross-sections and CTA images were matched using vessel bifurcation points and calcified plaque locations, and the multimodal registration results were optimized using a mutual information maximization algorithm. The next step was to calculate the CFD hemodynamic parameters of the vessels, including wall shear force distribution and blood flow velocity distribution.
[0127] The input to the cross-modal fusion network based on deformable convolution includes: CTA image patches, vessel segmentation results, plaque locations, and CFD-calculated hemodynamic distributions. The specific process involves using a 3D deformable convolutional network to encode CTA image patches into CTA feature maps, then encoding the CTA segmentation results (vessel mask, plaque location) into heatmaps, and multiplying them channel-by-channel with the CTA feature maps. To incorporate CFD-calculated hemodynamic parameters, CFD-simulated wall shear force and blood flow velocity distributions are input as additional channels to the cross-modal fusion network. Finally, an OCT generation network generates OCT cross-sectional images, including the three layers of the vessel wall (intima, media, and adventitia) and plaque details. During training, two loss functions control the model: in addition to the image similarity loss between the real OCT image and the generated OCT image, a correlation rule is defined between hemodynamic parameters (such as low WSS regions) and plaque vulnerability. The structure-blood flow loss function constrains the lipid core curvature and fibrous cap thickness of the generated OCT.
[0128] The vascular OCT image generation method based on DCN cross-modal fusion network and structure-blood flow coupling modeling in this application includes DCN-based cross-modal fusion and blood flow-structure joint modeling. Deformable convolution is embedded in the fusion network to dynamically adjust the receptive field according to the curvature of CTA vessels, generating multi-scale features and solving the deformation problem of tortuous vessels. In addition, the low wall shear force region simulated by CFD is associated with the lipid core distribution of OCT, and the pathological rationality of the generated image is improved by using a physical constraint loss function.
[0129] This application also provides a vascular OCT image training device based on DCN cross-modal fusion network and structure-flow coupling modeling. The vascular OCT image training device based on DCN cross-modal fusion network and structure-flow coupling modeling includes an image acquisition module, a CTA image preprocessing module, an OCT image preprocessing module, a registration module, a dynamic parameter acquisition module, a model acquisition module, and a training module; wherein,
[0130] The image acquisition module is used to acquire coronary CTA images and OCT images of multiple patients, and the spatial position of the OCT image is initially aligned with the coronary CTA image;
[0131] The CTA image preprocessing module is used to preprocess each of the CTA images to obtain image data marked with blood vessel segmentation results and plaque locations;
[0132] The OCT image preprocessing module is used to preprocess each of the OCT images to obtain preprocessed OCT image data.
[0133] The registration module is used to register image data belonging to the same patient, which is marked with blood vessel segmentation results and plaque location, as well as preprocessed OCT image data, to obtain registered image data.
[0134] The dynamic parameter acquisition module is used to acquire hemodynamic parameters of the same patient.
[0135] The model acquisition module is used to acquire a cross-modal fusion network based on deformable convolution and an OCT generation network;
[0136] The training module is used to train the deformable convolution-based cross-modal fusion network and the OCT generation network using the registered image data and hemodynamic parameters, respectively, so as to obtain the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for training vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling, characterized in that, The vascular OCT image training method based on DCN cross-modal fusion network and structure-blood flow coupling modeling includes: Acquire coronary CTA and OCT images of multiple groups of patients, with the spatial location of the OCT images initially aligned with that of the coronary CTA images; Each of the CTA images is preprocessed to obtain image data marked with blood vessel segmentation results and plaque locations; Each of the OCT images is preprocessed to obtain preprocessed OCT image data; Image data belonging to the same patient, labeled with vascular segmentation results and plaque locations, and preprocessed OCT image data are registered to obtain each registered image data. Obtain hemodynamic parameters from the same patient; Obtain a cross-modal fusion network and an OCT generation network based on deformable convolution; The deformable convolution-based cross-modal fusion network and the OCT generation network are trained using the registered image data and hemodynamic parameters, respectively, to obtain the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
2. The vascular OCT image training method based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 1, characterized in that, The acquisition of coronary CTA and OCT images from multiple groups of patients, wherein the spatial location of the OCT images is initially aligned with that of the coronary CTA images, includes: Collect N sets of paired coronary CTA and OCT images. The CTA data are high-resolution CT angiography covering the target vessel and include enhanced information on the vessel lumen, plaque, and vessel wall. The OCT data should be intravascular OCT images of the same patient as the CTA to ensure initial spatial alignment with the CTA.
3. The vascular OCT image training method based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 2, characterized in that, The preprocessing of each of the CTA images to obtain image data marked with vessel segmentation results and plaque locations includes: Obtain a pre-trained segmentation model; Each of the CTA images is input into a pre-trained segmentation model to obtain mask image data after segmenting blood vessels and plaques; Noise is removed from the mask image data by combining morphological operations, while preserving the topology of branch vessels; Using the blood vessel centerline as a reference, a unified coordinate system origin and direction matrix are established for the mask image data, thereby obtaining image data marked with blood vessel segmentation results and plaque locations.
4. The vascular OCT image training method based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 3, characterized in that, The preprocessing of each of the OCT images to obtain the preprocessed OCT image data includes: The OCT image is converted from polar coordinates to Cartesian coordinates based on the duct marker points. Then, motion artifact compensation and coordinate system transformation are used to correct the image, thereby obtaining the corrected OCT image.
5. The vascular OCT image training method based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 4, characterized in that, The process of registering image data belonging to the same patient, marked with vascular segmentation results and plaque locations, with preprocessed OCT image data to obtain registered image data includes: The corrected OCT images were matched with the corresponding CTA locations by identifying vascular branch points and calcified plaque positions, and the registration was optimized using a mutual information maximization algorithm.
6. A method for generating vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling, characterized in that, The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling includes: Obtain the deformable convolution-based cross-modal fusion network and the trained OCT generation network after training by the vascular OCT image training method based on DCN cross-modal fusion network and structure-blood flow coupling modeling as described in any one of claims 1 to 5; Acquire coronary CTA images and hemodynamic parameters of the patient to be generated; The coronary CTA images of the patient to be generated and the hemodynamic parameters are input into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network. The CTA features of the cross-modal fusion network are input into the trained OCT generation network to obtain a cross-sectional OCT image containing information about the blood vessel wall and plaque.
7. The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 6, characterized in that, The step of inputting the coronary CTA image of the patient to be generated and the hemodynamic parameters into the trained cross-modal fusion network based on deformable convolution to obtain the CTA features of the cross-modal fusion network includes: The coronary CTA images of the patient to be generated are encoded into CTA feature maps using a 3D deformable convolutional network; The coronary CTA images of the patient to be generated are preprocessed to obtain image data marked with vessel segmentation results and plaque locations; Image data labeled with blood vessel segmentation results and plaque locations are encoded into a heatmap, which is then multiplied channel-by-channel with the CTA feature map to obtain the first fusion feature. Extracting dynamic features based on hemodynamic parameters; The CTA features of the cross-modal fusion network are obtained by fusing dynamic features with the first fusion feature using a multi-head attention mechanism.
8. The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 7, characterized in that, The trained OCT generation network includes a structural similarity loss function, which is as follows: L smi =α·‖y gen -y real ‖1+(1-α)·(1-SSIM(y gen ,y real )); among them, α is the weighting coefficient, SSIM is the structural similarity index, and L smi Here, y represents the structural similarity loss function. gen It is the generated OCT image, y real These are real OCT images; The trained OCT generation network further includes a structure-blood flow coupling loss, which is as follows: Where, θ lipid M represents the lipid core curvature in the OC image. loss-wss For low WSS mask, L align It is the structure-blood flow loss term.
9. The method for generating vascular OCT images based on DCN cross-modal fusion network and structure-flow coupling modeling as described in claim 8, characterized in that, The final loss function of the trained OCT generation network is as follows: L all =L smi +γL align Among them, L all For the final loss function, L smi Let L be the structural similarity loss function; γ = 0.5; align This is the structure-blood flow coupling loss.
10. A vascular OCT image training device based on DCN cross-modal fusion network and structure-blood flow coupling modeling, characterized in that, The vascular OCT image training device based on DCN cross-modal fusion network and structure-flow coupling modeling includes: The image acquisition module is used to acquire coronary CTA images and OCT images of multiple patients, wherein the spatial position of the OCT image is initially aligned with the coronary CTA image; The CTA image preprocessing module is used to preprocess each of the CTA images to obtain image data marked with blood vessel segmentation results and plaque locations. The OCT image preprocessing module is used to preprocess each of the OCT images to obtain preprocessed OCT image data. The registration module is used to register image data belonging to the same patient, which is marked with blood vessel segmentation results and plaque locations, and preprocessed OCT image data, thereby obtaining registered image data. A hemodynamic parameter acquisition module, which is used to acquire hemodynamic parameters of the same patient; The model acquisition module is used to acquire a cross-modal fusion network based on deformable convolution and an OCT generation network; The training module is used to train the deformable convolution-based cross-modal fusion network and the OCT generation network using the registered image data and hemodynamic parameters respectively, so as to obtain the trained deformable convolution-based cross-modal fusion network and the trained OCT generation network.
Citation Information
Patent Citations
Deep learning-based acute cerebral apoplexy early-stage intelligent screening and early-warning method and system
CN119889670A
Coronary angiography image blood vessel segmentation system based on multi-scale feature fusion
CN120219751A