Multimodal collaborative enhancement and dynamic alignment of vascular image segmentation method and device

Through multimodal synergistic enhancement and dynamic alignment of vascular image segmentation methods, multimodal medical image fusion and difference problems are solved, significantly improving the accuracy and robustness of vascular image segmentation, especially in complex and small blood vessel segmentation tasks.

CN119888241BActive Publication Date: 2025-05-23XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386253.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-05-23
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art faces the problems of effective fusion of multimodal data and intermodal differences when processing multimodal medical images, resulting in limited accuracy and generalization ability of vascular image segmentation models in complex vascular structure analysis.

Method used

The vascular image segmentation method with multimodal coordinated enhancement and dynamic alignment is adopted. By acquiring the 3D vascular original images of different modes, segmenting and feature fusion, the pre-trained diffusion model is used for denoising and enhancement processing, and a modal embedding map is generated. Then, a 3D UNet encoder containing parallel multi-scale branch structure is used to extract multi-scale low-level features, and global feature enhancement and dynamic alignment are performed through graph neural networks to generate alignment high-level features. Finally, the synergistic enhancement of multimodal features is achieved by fusion by calculating feature information entropy and adaptive generation of dynamic weights.

Benefits of technology

It significantly improves the accuracy and robustness of vascular image segmentation, especially in complex and tiny blood vessel segmentation tasks, and can capture the morphological characteristics of tiny blood vessels more efficiently and reduce the impact of noise interference and mode inconsistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888241B_ABST
    Figure CN119888241B_ABST
Patent Text Reader

Abstract

The present invention provides a vascular image segmentation method and device with multi-modal collaborative enhancement and dynamic alignment, and relates to the technical field of deep learning and image processing. The present invention obtains a set of 3D vascular original images of different modalities, and after segmentation and fusion of modal feature labels, uses a diffusion model to perform denoising and enhancement to generate a modal embedding graph; the modal embedding graph is input into an improved 3D UNet encoder to extract multi-scale low-level features; low-level features are globally enhanced and dynamically aligned based on a graph neural network to generate aligned high-level features; modal collaborative features are obtained by calculating the feature information entropy of aligned high-level features and performing dynamic weight fusion; based on contrast learning of positive and negative sample pairs to optimize the difference representation between modal collaborative features, fusion features are obtained; the fusion features are input into a decoder to generate vascular segmentation results. The present invention can efficiently capture the morphological features of small blood vessels of different modalities, and significantly improve the segmentation accuracy and robustness of multi-modal vascular images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and image processing technology, and in particular to a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method and device. Background Art

[0002] Deep neural networks have made breakthrough progress in 3D medical image vascular segmentation, especially the introduction of the UNet architecture, which achieves accurate pixel-level prediction of medical images through the encoder-decoder structure and skip connections.

[0003] However, existing methods face two core challenges when processing multimodal medical images: effective fusion of multimodal data and differences between modalities, which limit the accuracy and generalization ability of medical image segmentation models in complex vascular structure analysis. Single-modal data only utilizes part of the image information. For example, CTA may miss non-calcified plaques, and MRA may misjudge blood flow signals. Direct fusion will lead to loss of detail information. Moreover, the noise distribution of different modalities is different, which will affect the stability of feature extraction. Furthermore, the performance of the same vascular structure in different modalities is significantly different (such as high-density blood vessels in CTA and bright signals in MRA), which makes cross-modal feature alignment difficult.

[0004] In view of this, the applicant filed this application after studying the existing technology. Summary of the invention

[0005] The present invention aims to provide a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method and device to efficiently integrate multi-modal data and improve the accuracy and robustness of vascular image segmentation, especially in the segmentation tasks of complex and small blood vessels.

[0006] In order to solve the above technical problems, the present invention is implemented through the following technical solutions:

[0007] A multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method, comprising:

[0008] S1, obtain 3D vascular original image sets of different modalities;

[0009] S2, segmenting each image of the 3D vascular original image set, fusing the local patches obtained by segmentation with the small blocks marked with corresponding modal features, and then using the pre-trained diffusion model to denoise and enhance the marked fused images to generate a modal embedding map;

[0010] S3, inputting the modality embedding graph into a 3D UNet encoder having a parallel multi-scale branch structure to extract multi-scale low-level features;

[0011] S4, based on the graph neural network, performing global feature enhancement and dynamic alignment on the multi-scale low-level features to generate aligned high-level features;

[0012] S5, calculating the feature information entropy of the aligned high-level features and adaptively generating dynamic weights for fusion to achieve collaborative enhancement of multimodal features and obtain modal collaborative features;

[0013] S6, based on the comparative learning of positive samples and negative samples, optimizing the inter-modal difference representation of the modal collaborative feature, and fusing the optimized features to obtain a fused feature;

[0014] S7, inputting the fused features into a decoder to generate a final blood vessel segmentation result.

[0015] Preferably, each image of the 3D vascular original image set is segmented, and the local patches obtained by segmentation are fused with small blocks marked with corresponding modality features, specifically:

[0016] Each image of the 3D blood vessel original image set is divided into a plurality of local patches after linear layer mapping;

[0017] Each local patch corresponds to a small block of the same size marked with modal features, namely, a modal feature marked block; the feature labels of the modal feature marked block are obtained by continuous training and optimization of the pre-trained model to adapt to the features of a specific image modality and capture the specific information of the image modality;

[0018] Each local patch is fused with the corresponding modality feature labeled block through feature fusion, and the formula is:

[0019] ;

[0020] in, Indicates a local patch The fused patch; [ ] represents the feature concatenation operation; the Encoder function represents the conversion of the modal feature marker block s into the corresponding local patch Feature vectors of the same dimension;

[0021] The merged patch After stitching, a labeled fused image is obtained.

[0022] Preferably, the denoising and enhancement steps of the diffusion model are as follows:

[0023] The labeled fused image, i.e., the labeled fused image, is input into the diffusion model, and a noise image is generated by gradually adding Gaussian noise at each time step t, so that the data distribution of the image gradually tends to an isotropic Gaussian distribution, and the expression is:

[0024] ;

[0025] in, Mark the data of the fused image for the tth step; is the noise scaling factor; is random Gaussian noise; It means the mean is 0 and the variance is Multidimensional normal distribution of;

[0026] Using Neural Network Learn a denoising function to gradually remove the noise of the noisy image, restore the modality-specific vascular image, and obtain the modality embedding map , the expression is:

[0027] ;

[0028]

[0029] in, is the labeled fused image after being processed by the diffusion model; m represents the modal index of the image; θ is the model parameter of the neural network; Diffusion is the reconstruction function of the diffusion model.

[0030] Preferably, the 3D UNet encoder introduces parallel multi-scale branches, and extracts features from high-resolution, medium-resolution and low-resolution input data respectively through multiple parallel encoding branches, so as to enhance the modeling ability of the model for vascular structures of different scales;

[0031] The structure of each encoding branch includes a 3D convolution layer, a batch normalization layer, and an activation function. The local features are extracted through the 3D convolution layer of each encoding branch, and the nonlinear expression ability is enhanced by the ReLU activation function to obtain low-level features at different scales.

[0032] When the 3D UNet encoder is used for feature extraction, a feature pyramid mechanism is used to perform cross-scale fusion of low-level features of different scales to enhance the hierarchical representation capability of vascular structures in images; wherein the feature pyramid mechanism uses cross-scale residual connections to establish information transfer channels between features of different resolutions, so that high-resolution features can be optimized under the guidance of low-resolution features while maintaining the integrity of detail information.

[0033] Preferably, it also includes: when extracting features through the 3D UNet encoder, a variable receptive field mechanism is introduced, and the feature map is compressed through a maximum pooling layer to ensure that the model can simultaneously focus on encoding branches of different scales and improve the modeling capability of the vascular topology structure;

[0034] The low-level features extracted by the 3D UNet encoder are input into a position encoder to add position information to each low-level feature; wherein the position encoder is generated by sine and cosine functions to capture the structural information inside the data and help the model understand the relative position of each element in the sequence.

[0035] Preferably, S4 is specifically:

[0036] The multi-scale low-level features As graph nodes, the relationships between modalities are used as edges to construct the feature relationship graph G;

[0037] The features of the feature relationship graph G are propagated and updated through the message passing mechanism of the graph neural network. The formula is:

[0038] ;

[0039] in, For the The feature representation of the layer; A is the adjacency matrix of the feature relationship graph; is an aggregate function; is a trainable parameter; is the activation function; m represents the modal index of the image;

[0040] Based on the inconsistency of feature distribution of different modalities, by calculating the mean μ and standard deviation σ of each modal feature, modal normalization is used to adjust the feature distribution to dynamically align the features of different modalities and generate aligned high-level features. The formula is:

[0041] ;

[0042] in, To align high-level features; and is a learnable parameter; Propagate updated features to graph neural networks; Features The corresponding mean; Features The corresponding standard deviation.

[0043] Preferably, it also includes: using a feature alignment loss function to continuously optimize the difference between the aligned high-level features and the reference modality features to enhance the consistency of cross-modality features, the formula is:

[0044] ;

[0045] in, To align high-level features; is the reference modal feature; is the feature alignment loss function.

[0046] Preferably, S5 is specifically:

[0047] Compute aligned high-level features for each modality The information entropy of is used to measure the uncertainty of the modal characteristics. The formula is:

[0048] ;

[0049] in, is the information entropy of m-modal features; Representation alignment high-level features Probability distribution in feature space;

[0050] Based on the information entropy calculated , using the exponential normalization strategy to generate the dynamic weights of each modality through a single-layer multi-layer perceptron:

[0051] ;

[0052] in, is the normalized weight of the m mode; for Information entropy of modal features; m, Represents the modality index of the image;

[0053] Align high-level features of different modalities based on calculated dynamic weights Perform weighted fusion to generate the final modal collaborative features. The formula is:

[0054] ;

[0055] in, is the modal collaboration feature.

[0056] Preferably, the S6 is specifically:

[0057] According to the modality collaborative features, positive samples and negative samples are constructed through comparative learning; wherein the positive samples are vascular image features based on the same anatomical structure in different modalities; and the negative samples are vascular image features of different anatomical parts or different patients;

[0058] Adopting contrast loss function based on cosine similarity , measure the similarity between different sample pairs, and continuously compare the loss to optimize the difference representation to strengthen the feature aggregation ability of positive sample pairs and enhance the feature differentiation ability of negative sample pairs. The formula is:

[0059] ;

[0060] Where P represents the set of positive sample pairs; N represents the set of negative sample pairs; i, j, and k represent the index variables of the feature layer respectively; sim is the cosine similarity function, and τ is the temperature parameter used to control the degree of aggregation of feature distribution; is the modal synergy feature; is the modal collaborative feature of the i-th layer;

[0061] The features after loss optimization are fused and generated by a trainable mapping function to obtain the fused features , the formula is:

[0062] ( ) ;

[0063] in, is the mapping function, The present invention also provides a multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation device, which is characterized by comprising:

[0064] An acquisition unit, used for acquiring a set of 3D vascular original images of different modalities;

[0065] A modality embedding unit is used to segment each image of the 3D vascular original image set, fuse the local patches obtained by segmentation with the small blocks marked with corresponding modality features, and then use a pre-trained diffusion model to denoise and enhance the marked fused images to generate a modality embedding graph;

[0066] A multi-scale encoding unit, used for inputting the modal embedding graph into a 3DUNet encoder having a parallel multi-scale branch structure to extract multi-scale low-level features;

[0067] A GNN enhancement unit, used for performing global feature enhancement and dynamic alignment on the multi-scale low-level features based on a graph neural network to generate aligned high-level features;

[0068] A collaborative generation unit, configured to calculate the feature information entropy of the aligned high-level features and adaptively generate dynamic weights for fusion, so as to achieve collaborative enhancement of multimodal features and obtain modal collaborative features;

[0069] A contrastive learning unit, used for optimizing the inter-modal difference representation of the modal collaborative feature based on contrastive learning of positive samples and negative samples, and fusing the optimized features to obtain a fused feature;

[0070] The output unit is used to input the fusion feature into a decoder to generate a final blood vessel segmentation result.

[0071] The present invention also provides a vascular image segmentation device with multimodal collaborative enhancement and dynamic alignment, characterized in that it includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program can be executed by the processor to implement a vascular image segmentation method with multimodal collaborative enhancement and dynamic alignment as described above.

[0072] The present invention also provides a computer-readable storage medium, characterized in that computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method as described above is implemented.

[0073] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0074] The present invention uses a modality labeling method and a pre-trained diffusion model to perform modal feature labeling, denoising and enhancement on the input multimodal 3D vascular original image to generate a modal embedding graph, so as to effectively capture modality-specific information, make up for the differences between different imaging modalities, provide high-quality data input for subsequent vascular analysis tasks, and enhance the model's ability to process multimodal data.

[0075] The present invention adopts a 3D UNet encoder with a parallel multi-scale branch structure, which enables it to process input data of different resolutions simultaneously to enhance the expression ability of vascular structure. Through feature extraction at different scales, the present invention can ensure that the model can capture the fine structural information of local blood vessels while retaining the global features of the overall vascular topology.

[0076] The present invention adopts an adaptive dynamic weight coordination mechanism to dynamically adjust the contribution of modal features in the fusion process to ensure that the final fused features have stronger robustness and lower noise interference.

[0077] The present invention fully exploits the complementary information of multimodal data through multi-scale feature extraction, global feature modeling of graph neural networks, adaptive dynamic weight coordination, and differential optimization of contrastive learning, overcomes the limitations of traditional single-modal methods in small blood vessel segmentation, and significantly improves the segmentation accuracy and robustness of vascular images, especially in the segmentation tasks of complex and small blood vessels. Compared with the existing technology, the present invention can more efficiently capture the morphological characteristics of small blood vessels, reduce the influence of noise interference and inconsistency between modalities, and generate accurate and detailed segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0079] Figure 1 A schematic diagram of a blood vessel image segmentation method with multimodal collaborative enhancement and dynamic feature alignment provided in Example 1.

[0080] Figure 2 A schematic diagram of the structure of the blood vessel image segmentation method with multimodal collaborative enhancement and dynamic feature alignment provided in Example 1.

[0081] Figure 3 This is a schematic diagram of the structure of the modal embedding module provided in Example 1.

[0082] Figure 4 This is a visualization diagram of the segmentation result provided in Example 1.

[0083] Figure 5 A schematic diagram of a blood vessel image segmentation device with multimodal collaborative enhancement and dynamic feature alignment provided in Example 2.

[0084] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0085] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0086] Embodiment 1

[0087] Embodiment 1 of the present invention provides a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method, which can be implemented by a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation device (hereinafter referred to as the vascular image segmentation device), and in particular, is executed by one or more processors in the vascular image segmentation device.

[0088] In this embodiment, the vascular image segmentation device may be an electronic device equipped with a processor, the processor having a computer program of the multimodal collaborative enhancement and dynamic alignment vascular image segmentation method and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0089] like Figure 1-Figure 2 As shown, a multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method includes steps S1 to S7.

[0090] S1, obtain 3D vascular original image sets of different modalities.

[0091] Specifically, the 3D vascular original image sets of different modalities include but are not limited to: Computed Tomography Angiography (CTA) and Magnetic Resonance Angiography (MRA).

[0092] S2, segmenting each image of the 3D vascular original image set, fusing the local patches obtained by segmentation with the small blocks marked with corresponding modal features, and then using the pre-trained diffusion model to denoise and enhance the marked fused images to generate a modal embedding map.

[0093] like Figure 3 As shown in the figure, taking CTA images and MRA images (i.e., original images) as examples, the modality difference labeling method is used to map each image of the 3D vascular original image set through a linear layer and then segment it, and the segmented local patches are fused with small blocks labeled with corresponding modality features, and then the pre-trained diffusion model is used to denoise and enhance the fused and labeled images.

[0094] In this embodiment, the 3D blood vessel original image is expanded into a corresponding one-dimensional vector after being mapped through a linear layer, so as to facilitate image segmentation.

[0095] Each local patch contains local area information of the vascular image; each local patch corresponds to a small block of the same size marked with modal features, namely, a modal feature marked block. The feature labels of the modal feature marked blocks are obtained by continuous training and optimization of the pre-trained model to adapt to the features of a specific image modality and capture the specific information of the image modality.

[0096] Specifically, each original image of the 3D vascular original image set is divided into several local patches, each local patch introduces a feature mark related to the original image modality, and each local patch is fused with the corresponding modality feature mark block through feature fusion. The formula is:

[0097] ;

[0098] in, Indicates a local patch The fused patch; [ ] indicates the feature concatenation operation; Encoder(s) indicates the conversion of the modal feature marker block s into the corresponding local patch Feature vectors of the same dimension.

[0099] The merged patch After stitching, a labeled fused image is obtained.

[0100] Next, the diffusion model gradually adds noise to the labeled fusion image and reversely denoises it to generate a modality embedding graph.

[0101] In this embodiment, the diffusion model is a deep learning model based on probability generation. Its basic principle is to convert the original data into a noise distribution by gradually adding noise, and to restore it to a high-quality image through a reverse process of denoising. The specific process of this step is as follows:

[0102] Forward diffusion process: For the input labeled fusion image data, Gaussian noise is gradually added in each time step t, so that the data distribution gradually tends to an isotropic Gaussian distribution, which is specifically expressed as:

[0103] ;

[0104] in, is the labeled fused image data of the t-th step, t=1,2,...T, T is the total number of diffusion steps; is the noise scaling factor; is random Gaussian noise; It means the mean is 0 and the variance is The multidimensional normal distribution of represents independent and identically distributed standard Gaussian noise. As t increases, The noise ratio gradually decreases, and finally the image distribution tends to the isotropic standard Gaussian distribution.

[0105] Reverse denoising process: through neural network Learn a denoising function to gradually remove noise and restore modality-specific vascular images to obtain a modality embedding map , the expression is:

[0106] ;

[0107] The denoising network can adopt a U-Net structure and combine the input of modality-specific conditions (such as feature embeddings of CTA or MRA) to improve the quality and preserve the modality-specific information.

[0108] Generate modality embedding map: The labeled fused image data processed by the diffusion model contains modality-specific enhanced features, which are recorded as modality embedding map , where m represents the modality index. For example, m=1 represents CTA and m=2 represents MRA. It not only removes the noise in the original image, but also enhances the vascular structure, making subsequent tasks (such as segmentation, registration or cross-modal fusion) more stable and reliable.

[0109] This step uses the noise addition and reverse denoising mechanism of the diffusion model to make the generated modal embedding graph It not only retains the high-resolution vascular structure of the original image, but also eliminates interference factors between different modalities, providing high-quality data input for subsequent vascular analysis tasks.

[0110] S3, input the modality embedding graph into a 3D UNet encoder with a parallel multi-scale branch structure to extract multi-scale low-level features.

[0111] In this step, the modality is embedded into the graph The improved 3D UNet encoder is input to extract multi-scale low-level features of each image modality. The improved 3D UNet encoder includes multiple parallel encoding branches to process input data of different resolutions and generate a multi-scale low-level feature set.

[0112] In this embodiment, in order to meet the multi-scale feature expression requirements of medical images, the 3D U-Net encoder structure is improved as follows:

[0113] Introducing parallel multi-scale branches: Through multiple parallel encoding branches, feature extraction is performed on high-resolution, medium-resolution and low-resolution input data respectively to enhance the modeling ability of vascular structures at different scales. The high-resolution branch is used to capture subtle vascular details, the medium-resolution branch is used to process medium-scale vascular structures, and the low-resolution branch is used to capture the global vascular layout and larger-scale structures.

[0114] Each branch uses 3D convolutional layers, batch normalization (BN) layers, and activation functions (ReLU) for feature extraction and generates corresponding low-level features of different scales. Local features are extracted through the 3D convolutional layer of each encoding branch, and the nonlinear expression ability is enhanced through the ReLU activation function.

[0115] Multi-scale feature fusion: In the feature extraction process, the feature pyramid mechanism is used to cross-scale fuse features of different scales to enhance the hierarchical representation of vascular structures. Specifically, cross-scale residual connections are used to establish information transmission channels between features of different resolutions, so that high-resolution features can be optimized under the guidance of low-resolution features (that is, in the process of high-resolution feature optimization, the contextual information of low-resolution features is fused to improve the accuracy of vascular boundary positioning), while maintaining the integrity of detail information and avoiding the gradient vanishing problem.

[0116] Variable receptive field design: By introducing the adaptive receptive field mechanism, the model can focus on both large-scale trunk vessels and small-scale vascular branches at the same time, thereby improving the modeling ability of vascular topology.

[0117] In summary, the specific feature extraction process of this step is: embed the modality into the graph The improved 3DUNet encoder is input, and each multi-scale branch uses 3D convolution to extract local features, and performs preliminary processing on the modal embedding map to obtain basic feature representation; then, feature extraction is performed based on high, medium, and low resolutions to obtain low-level features of three scales, and cross-scale residual connections are used to cross-scale fuse features of different scales to form a multi-scale low-level feature set.

[0118] S4, based on the graph neural network, performs global feature enhancement and dynamic alignment on the multi-scale low-level features to generate aligned high-level features.

[0119] Furthermore, before each low-level feature is input into the graph neural network, position information can be added through a position encoder, and then the multi-scale low-level features with added position information are input into the graph neural network for global feature enhancement and dynamic alignment. The position encoder is generated by sine and cosine functions, aiming to add position information to each input feature, thereby helping the model understand the relative position of each element in the sequence. Through position encoding, the input feature not only contains the original image information, but also introduces the sequential relationship of the data, so that the model can capture the structural information inside the data.

[0120] Specifically, the multi-scale low-level features The global feature enhancement module based on graph neural network (GNN) is input, and the aligned high-level features of each modality are generated through the dynamic alignment layer.

[0121] In this embodiment, a feature relationship graph (Feature RelationshipGraph) is constructed based on the graph neural network GNN to capture the global dependency between modalities, and a feature propagation mechanism is used to extract vascular image features with global information. The dynamic alignment layer adjusts the consistency of modal features based on feature distribution to improve the alignment of cross-modal features and the expression of global information. The specific steps are as follows:

[0122] Global feature enhancement based on GNN: First, GNN is used to construct a feature relationship graph G=(V,E), where the node set V represents low-level features of different scales, and the edge set E is the cross-scale and cross-modal feature relationship established by calculating feature similarity.

[0123] After the feature relationship graph is constructed, the message passing mechanism is used to update the features. The update rules are as follows:

[0124] ;

[0125] Among them, among them, For the The feature representation of the layer; A is the adjacency matrix of the feature relationship graph; is an aggregate function; is a trainable parameter; is the activation function. Through multi-layer GNN operations, vascular image features with cross-modal global information are gradually extracted.

[0126] Dynamic alignment layer: In view of the inconsistency of feature distribution of different modalities, a dynamic alignment mechanism is used to normalize and optimize the features. First, the mean μ and standard deviation σ of each modal feature are calculated, and the feature distribution is adjusted using modal normalization, which is expressed as:

[0127] ;

[0128] in, To align high-level features; and is a learnable parameter; Propagate updated features to graph neural networks; Features The corresponding mean; Features The corresponding standard deviation.

[0129] In addition, a feature alignment loss function is introduced to continuously optimize the difference between the aligned high-level features and the reference modality features to enhance the consistency of cross-modal features. The formula is:

[0130] ;

[0131] in, To align high-level features; is the reference modal feature; is the feature alignment loss function. By optimizing this loss function, the consistency of cross-modal features can be further enhanced.

[0132] Output high-level features: After GNN feature enhancement, dynamic alignment layer and alignment loss optimization, the aligned high-level features are finally obtained .

[0133] High-level features after alignment It contains cross-modal enhancement information and optimizes the consistency between modalities, and can be used for subsequent medical image analysis tasks, such as vascular image segmentation, registration, and classification.

[0134] S5, by calculating the feature information entropy of the aligned high-level features and adaptively generating dynamic weights for fusion, the collaborative enhancement of multimodal features is achieved to obtain modal collaborative features.

[0135] In this embodiment, the alignment high-level features The modal collaborative features are generated by fusing them through an adaptive weight mechanism based on information entropy. The adaptive weight mechanism calculates the feature information entropy to measure the information richness and stability of each modal feature, and generates dynamic weights to achieve synergistic enhancement of different modal features. The specific steps are as follows:

[0136] Calculate information entropy: First, align high-level features of each modality The information entropy is calculated to measure the uncertainty of the modal feature. The calculation formula of information entropy is as follows:

[0137] ;

[0138] in, is the information entropy of m-modal features; Represents alignment of high-level features Probability distribution in feature space.

[0139] The calculation method of information entropy includes normalization processing to ensure the comparability of different modal features. In the calculation results of information entropy: lower information entropy means that the modal feature information is more stable and has a higher confidence level; higher information entropy means that the modal feature contains more uncertainty and may be affected by noise or modal differences.

[0140] Generate adaptive dynamic weights: based on calculated information entropy , using the exponential normalization strategy, a single-layer multi-layer perceptron (MLP) is used to generate the dynamic weights of each modality. The formula is as follows:

[0141] ;

[0142] in, is the normalized weight of the m mode, ensuring that the sum of the weights of all modal features is 1; for Information entropy of modal features.

[0143] This method can ensure that: the modality with lower information entropy is given a higher weight, so that the modality with stable information contributes more in the fusion; the modality with higher information entropy is given a lower weight to reduce the interference of noise information on the final modality collaborative features.

[0144] Compute modal collaborative features: Use the calculated dynamic weights to align high-level features of different modalities Perform weighted fusion to generate the final modality collaborative features , the formula is:

[0145] ;

[0146] in, The modal collaborative feature has the information advantages of each modal feature and realizes feature complementarity and collaborative enhancement between modalities through an adaptive weight mechanism controlled by information entropy.

[0147] S6, based on the comparative learning of positive samples and negative samples, the inter-modal difference representation of the modal collaborative feature is optimized, and the optimized features are fused to obtain a fused feature.

[0148] Specifically, based on contrastive learning, the modal collaborative features Optimized for differentially enhanced fusion features The contrastive learning constructs positive samples and negative samples to explore the deep differences between different modalities to improve segmentation accuracy.

[0149] This step is used to enhance the feature expression ability of cross-modal vascular images. By mining the common information and difference information between different modalities through comparative learning, on the basis of optimizing the modality collaborative features, the complementarity and distinguishing ability of cross-modal features are further strengthened, thereby improving the accuracy of vascular image segmentation tasks.

[0150] First, the construction of positive sample pairs is based on the image features of the same anatomical structure in different modalities. For example, the CTA image and the corresponding vascular area in the MRA image of the same patient constitute a positive sample pair. Since the positive samples come from the same anatomical part, they should have a high similarity in the feature space. Therefore, one of the optimization goals of contrastive learning is to increase the feature consistency between positive samples so that the features of the same anatomical structure remain highly comparable between different modalities.

[0151] In contrast, negative sample pairs are composed of vascular image features from different anatomical parts or different patients. For example, a patient's cerebrovascular CTA image and his lower limb vascular CTA image, or MRA images and CTA images from different patients. Since these images do not belong to the same anatomical region or different individuals, they should maintain a large degree of distinction in the feature space. Therefore, the second optimization goal of contrastive learning is to reduce the similarity between negative samples, so that the model can automatically learn how to maintain the independence of modal information and avoid feature confusion between different modalities or individuals.

[0152] In order to achieve the above optimization goals, an optimization strategy based on contrast loss is adopted to minimize the feature distance between positive samples and maximize the feature distance between negative samples. Specifically, a contrast loss function based on cosine similarity is adopted. , to measure the similarity between different sample pairs, and by adjusting the optimization direction of the loss function, guide the model to strengthen the feature aggregation ability of positive sample pairs during training, while enhancing the feature differentiation ability of negative sample pairs. It is expressed as:

[0153] ;

[0154] Among them, P represents the set of positive sample pairs, N represents the set of negative sample pairs, i, j, and k represent the index variables of the feature layer respectively; sim is the cosine similarity function; τ is the temperature parameter used to control the degree of aggregation of feature distribution.

[0155] By minimizing the contrastive loss function , so that the distance between positive samples in the feature space is reduced, while the distance between negative samples in the feature space is increased, thereby improving the distinguishing ability of the fused features. The features after loss optimization are fused through residual connections and a trainable mapping function Generate and obtain optimized fusion features :

[0156] ( ) ;

[0157] in, is the mapping function, are trainable parameters of the mapping function.

[0158] Fusion Features The output of this step can be used for subsequent medical image analysis tasks.

[0159] S7, inputting the fused features into a decoder to generate a final blood vessel segmentation result.

[0160] In this step, the fusion feature The decoder is input to generate the final blood vessel segmentation result S. The decoder efficiently outputs the pixel-level blood vessel image segmentation result through upsampling and convolution operations.

[0161] Specifically, the feature map resolution can be restored by trilinear interpolation upsampling, and then the number of channels can be adjusted using a 1x1x1 convolution kernel to output the segmentation probability map, and then the Softmax function can be applied to generate the final blood vessel segmentation result S.

[0162] In another preferred embodiment, the method further comprises: optimizing the accuracy of multimodal vascular image segmentation using a combined loss function L, which is mathematically expressed as:

[0163] ;

[0164] ;

[0165] in, is the weighted cross entropy loss; is the contrast loss function based on cosine similarity; is the true value; is the predicted value; γ is the balance coefficient; represents the category label, C is the total number of labeled categories; n N, is a variable for the number of pixels of the vascular image, and N is the total number of pixels of the vascular image.

[0166] like Figure 4The visualization of the vascular image segmentation results of the CT and MR modalities shown in the figure. Among them, GroundTruth refers to the real annotation or reference segmentation result, which is used as a benchmark for evaluating the segmentation performance; Ours is the segmentation result obtained using the model proposed in this study.

[0167] The following are several commonly used medical image segmentation models:

[0168] 3DUnet, which is a 3D image segmentation model based on convolutional neural network (CNN), is widely used in medical image segmentation tasks.

[0169] SegResNetVAE, a SegResNet model combined with a variational autoencoder (VAE), improves the robustness of segmentation by enhancing generation capabilities.

[0170] SegResNet, a segmentation network based on the ResNet architecture, uses deep residual learning to improve image segmentation performance.

[0171] Swin UNETR, which combines Swin Transformer and convolutional neural network (CNN), captures long-range dependencies through Swin Transformer and uses CNN for local feature extraction.

[0172] CSNet3D (vessel) is a 3D convolutional neural network (CNN) model designed for vascular structure analysis in medical images.

[0173] like Figure 4 As shown in the figure, the red frame area of ​​the Ground Truth shows the small and easily overlooked blood vessels between different blood vessels. In the CT and MR modalities, our proposed model successfully focused on these small blood vessels and performed correct segmentation. Other models either ignored these blood vessels or incorrectly segmented them into other types. This shows that our proposed model can pay more attention to the details in the image through multi-scale feature extraction, dynamic alignment and contrast learning, thereby achieving a more accurate segmentation effect.

[0174] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0175] The present invention uses a modality labeling method and a pre-trained diffusion model to perform modality feature labeling, denoising and enhancement on the input multimodal 3D vascular original image to generate a modality embedding graph, so as to effectively capture modality-specific information, compensate for the differences between different imaging modalities, and enhance the model's ability to process multimodal data.

[0176] The present invention introduces a fusion method based on multimodal collaborative enhancement and dynamic feature alignment. Through multi-scale feature extraction, global modeling of graph neural networks, dynamic weight collaboration, and differential optimization of contrastive learning, the complementary information of multimodal data is fully mined, and the limitations of traditional single-modal methods in small blood vessel segmentation are overcome. The segmentation accuracy and robustness are significantly improved, especially in the segmentation tasks of complex and small blood vessels. Compared with the existing technology, the present invention can more efficiently capture the morphological characteristics of small blood vessels, reduce the influence of noise interference and inconsistency between modalities, and generate accurate and detailed segmentation results.

[0177] Embodiment 2

[0178] like Figure 5 As shown, the second embodiment of the present invention also provides a multi-modal collaborative enhancement and dynamic alignment vascular image segmentation device, which is mainly composed of multiple functional units, and sequentially completes the modality embedding, feature extraction, modality alignment, feature fusion and decoding segmentation of vascular images to achieve efficient and accurate vascular image segmentation tasks. The device includes the following functional units:

[0179] An acquisition unit, used for acquiring a set of 3D vascular original images of different modalities;

[0180] A modality embedding unit is used to segment each image of the 3D vascular original image set, fuse the local patches obtained by segmentation with the small blocks marked with corresponding modality features, and then use a pre-trained diffusion model to denoise and enhance the marked fused images to generate a modality embedding graph;

[0181] A multi-scale encoding unit, used for inputting the modal embedding graph into a 3DUNet encoder having a parallel multi-scale branch structure to extract multi-scale low-level features;

[0182] A GNN enhancement unit, used for performing global feature enhancement and dynamic alignment on the multi-scale low-level features based on a graph neural network to generate aligned high-level features;

[0183] A collaborative generation unit, configured to calculate the feature information entropy of the aligned high-level features and adaptively generate dynamic weights for fusion, so as to achieve collaborative enhancement of multimodal features and obtain modal collaborative features;

[0184] A contrastive learning unit, used for optimizing the inter-modal difference representation of the modal collaborative feature based on contrastive learning of positive samples and negative samples, and fusing the optimized features to obtain a fused feature;

[0185] The output unit is used to input the fusion feature into a decoder to generate a final blood vessel segmentation result.

[0186] Specifically, the modality embedding unit removes noise from the image while retaining the modality-specific feature information through the process of modality labeling, gradually adding noise, and reverse denoising, thereby generating a modality embedding graph containing modality features. The three-dimensional vascular original image set includes but is not limited to computed tomography angiography (CTA) and magnetic resonance imaging angiography (MRA), ensuring that the segmentation method can be applied to a variety of medical imaging data.

[0187] The multi-scale encoding unit adopts a parallel multi-scale branch structure, which enables it to process input data of different resolutions simultaneously to enhance the expression of vascular structure and generate a multi-scale low-level feature set. By extracting features at different scales, the unit can ensure that the model can capture the fine structural information of local blood vessels while retaining the global characteristics of the overall vascular topology.

[0188] The GNN enhancement unit builds a feature relationship graph through a global feature enhancement module based on a graph neural network (GNN) to capture the global dependencies between modalities. In addition, the unit also contains a dynamic alignment layer to adjust the distribution of features of different modalities so that they are highly consistent before fusion. Through the processing of this unit, the features of images of different modalities can be more closely aligned, thereby improving the fusion stability of cross-modal features.

[0189] The collaborative generation unit first calculates the information entropy of each modal feature to measure the uncertainty of the feature, and dynamically adjusts the contribution of the modal feature in the fusion process accordingly. Modal features with higher information stability are given higher weights, while modal features with greater information uncertainty are weakened to ensure that the final fused features have stronger robustness and lower noise interference.

[0190] The contrastive learning unit learns the deep feature differences between modalities by constructing positive and negative sample pairs, ensuring that similar anatomical structures of different modalities are more consistently expressed in the feature space, while enhancing the discrimination between modalities. The positive sample pairs are composed of features of the same anatomical part in different modalities, while the negative sample pairs are composed of features of different individuals or different anatomical parts. Through contrastive loss optimization, the positive sample features are made closer, while the negative sample features maintain discrimination, thereby generating difference-enhanced fusion features to improve the accuracy and generalization ability of segmentation tasks.

[0191] The output unit inputs the optimized fusion features into the decoder to efficiently generate the final blood vessel segmentation result, and adopts a skip connection mechanism to ensure that the global information of high-level features is combined with the local detail information of low-level features, so that the final segmentation result has both the integrity of the global structure and the clarity of the blood vessel edges.

[0192] Embodiment 3

[0193] The third embodiment of the present invention also provides a vascular image segmentation device with multimodal collaborative enhancement and dynamic alignment, which includes a memory and a processor, wherein a computer program is stored in the memory, and the computer program can be executed by the processor to implement the vascular image segmentation method with multimodal collaborative enhancement and dynamic alignment as described above.

[0194] Embodiment 4

[0195] The fourth embodiment of the present invention further provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method as described above is implemented.

[0196] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0197] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0198] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0199] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0200] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0201] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0202] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0203] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method, characterized in that: include: S1, obtain 3D vascular original image sets of different modalities; S2, segmenting each image of the 3D vascular original image set, fusing the local patches obtained by segmentation with the small blocks marked with corresponding modal features, and then using the pre-trained diffusion model to denoise and enhance the marked fused images to generate a modal embedding map; S3, inputting the modality embedding graph into a 3D UNet encoder having a parallel multi-scale branch structure to extract multi-scale low-level features; S4, based on the graph neural network, the multi-scale low-level features are globally enhanced and dynamically aligned to generate aligned high-level features, specifically: The multi-scale low-level features As graph nodes, the relationships between modalities are used as edges to construct the feature relationship graph G; The features of the feature relationship graph G are propagated and updated through the message passing mechanism of the graph neural network. The formula is: ; in, For the The feature representation of the layer; A is the adjacency matrix of the feature relationship graph; is an aggregate function; is a trainable parameter; is the activation function; m represents the modal index of the image; Based on the inconsistency of feature distribution of different modalities, by calculating the mean μ and standard deviation σ of each modal feature, modal normalization is used to adjust the feature distribution to dynamically align the features of different modalities and generate aligned high-level features. The formula is: ; in, To align high-level features; and is a learnable parameter; Propagate updated features to graph neural networks; Features The corresponding mean; Features The corresponding standard deviation; S5, by calculating the feature information entropy of the aligned high-level features and adaptively generating dynamic weights for fusion, to achieve collaborative enhancement of multimodal features, modal collaborative features are obtained, specifically: Compute aligned high-level features for each modality The information entropy of is used to measure the uncertainty of the modal characteristics. The formula is: ; in, is the information entropy of m-modal features; Representation alignment high-level features Probability distribution in feature space; Based on the information entropy calculated , using the exponential normalization strategy to generate the dynamic weights of each modality through a single-layer multi-layer perceptron: ; in, is the normalized weight of the m mode; for Information entropy of modal features; m, Represents the modality index of the image; Align high-level features of different modalities based on calculated dynamic weights Perform weighted fusion to generate the final modal collaborative features. The formula is: ; in, is the modal synergy feature; S6, based on the comparative learning of positive samples and negative samples, optimizing the inter-modal difference representation of the modal collaborative feature, and fusing the optimized features to obtain a fused feature; S7, inputting the fused features into a decoder to generate a final blood vessel segmentation result.

2. The multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method according to claim 1, characterized in that: Each image of the 3D vascular original image set is segmented, and the local patches obtained by segmentation are fused with small blocks marked with corresponding modality features, specifically: Each image of the 3D blood vessel original image set is divided into a number of local patches after linear layer mapping; Each local patch corresponds to a small block of the same size marked with modal features, namely, a modal feature marked block; the feature labels of the modal feature marked block are obtained by continuous training and optimization of the pre-trained model to adapt to the features of a specific image modality and capture the specific information of the image modality; Each local patch is fused with the corresponding modality feature labeled block through feature fusion, and the formula is: ; in, Indicates a local patch The fused patch; [ ] represents the feature concatenation operation; the Encoder function represents the conversion of the modal feature marker block s into the corresponding local patch Feature vectors of the same dimension; The merged patch After stitching, a labeled fused image is obtained.

3. The multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method according to claim 2, characterized in that: The denoising and enhancement steps of the diffusion model are as follows: The labeled fused image, i.e., the labeled fused image, is input into the diffusion model, and a noise image is generated by gradually adding Gaussian noise at each time step t, so that the data distribution of the image gradually tends to an isotropic Gaussian distribution, and the expression is: ; in, Mark the data of the fused image for the tth step; is the noise scaling factor; is random Gaussian noise; It means the mean is 0 and the variance is Multidimensional normal distribution of; Using Neural Network Learn a denoising function to gradually remove the noise of the noisy image, restore the modality-specific vascular image, and obtain the modality embedding map , the expression is: ; in, is the labeled fused image after being processed by the diffusion model; m represents the modal index of the image; θ is the model parameter of the neural network; Diffusion is the reconstruction function of the diffusion model.

4. The multi-modal collaborative enhancement and dynamic alignment vascular image segmentation method according to claim 1, characterized in that: The 3D UNet encoder introduces parallel multi-scale branches, and extracts features from high-resolution, medium-resolution and low-resolution input data respectively through multiple parallel encoding branches, so as to enhance the model's modeling ability for vascular structures of different scales; The structure of each encoding branch includes a 3D convolution layer, a batch normalization layer, and an activation function. The local features are extracted through the 3D convolution layer of each encoding branch, and the nonlinear expression ability is enhanced by the ReLU activation function to obtain low-level features at different scales. When the 3D UNet encoder is used for feature extraction, a feature pyramid mechanism is used to perform cross-scale fusion of low-level features of different scales to enhance the hierarchical representation capability of vascular structures in images; wherein the feature pyramid mechanism uses cross-scale residual connections to establish information transfer channels between features of different resolutions, so that high-resolution features can be optimized under the guidance of low-resolution features while maintaining the integrity of detail information.

5. The multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method according to claim 4, characterized in that: Also includes: When performing feature extraction through the 3D UNet encoder, a variable receptive field mechanism is introduced, and the feature map is compressed through the maximum pooling layer to ensure that the model can simultaneously focus on encoding branches of different scales and improve the modeling ability of the vascular topology structure; The low-level features extracted by the 3D UNet encoder are input into a position encoder to add position information to each low-level feature; wherein the position encoder is generated by sine and cosine functions to capture the structural information inside the data and help the model understand the relative position of each element in the sequence.

6. The multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method according to claim 1, characterized in that: It also includes: using a feature alignment loss function to continuously optimize the difference between the aligned high-level features and the reference modality features to enhance the consistency of cross-modality features, the formula is: ; in, To align high-level features; is the reference modal feature; is the feature alignment loss function; m represents the modality index of the image.

7. The multi-modal collaborative enhancement and dynamic alignment blood vessel image segmentation method according to claim 1, characterized in that: The S6 is specifically: According to the modality collaborative features, positive samples and negative samples are constructed through comparative learning; wherein the positive samples are vascular image features based on the same anatomical structure in different modalities; and the negative samples are vascular image features of different anatomical parts or different patients; Adopting the contrast loss function based on cosine similarity , measure the similarity between different sample pairs, and continuously compare the loss to optimize the difference representation to strengthen the feature aggregation ability of positive sample pairs and enhance the feature differentiation ability of negative sample pairs. The formula is: ; Where P represents the set of positive sample pairs; N represents the set of negative sample pairs; i, j, and k represent the index variables of the feature layer respectively; sim is the cosine similarity function, and τ is the temperature parameter used to control the degree of aggregation of feature distribution; is the modal synergy feature; is the modal collaborative feature of the i-th layer; The features after loss optimization are fused and generated by a trainable mapping function to obtain the fused features , the formula is: ( ) ; in, is the mapping function, are trainable parameters of the mapping function.

8. A multi-modal collaborative enhancement and dynamic alignment vascular image segmentation device, characterized in that: include: An acquisition unit, used for acquiring a set of 3D vascular original images of different modalities; A modality embedding unit is used to segment each image of the 3D vascular original image set, fuse the local patches obtained by segmentation with the small blocks marked with corresponding modality features, and then use a pre-trained diffusion model to denoise and enhance the marked fused images to generate a modality embedding graph; A multi-scale encoding unit, used for inputting the modality embedding graph into a 3D UNet encoder having a parallel multi-scale branch structure to extract multi-scale low-level features; The GNN enhancement unit is used to perform global feature enhancement and dynamic alignment on the multi-scale low-level features based on the graph neural network to generate aligned high-level features, specifically: The multi-scale low-level features As graph nodes, the relationships between modalities are used as edges to construct the feature relationship graph G; The features of the feature relationship graph G are propagated and updated through the message passing mechanism of the graph neural network. The formula is: ; in, For the The feature representation of the layer; A is the adjacency matrix of the feature relationship graph; is an aggregate function; is a trainable parameter; is the activation function; m represents the modal index of the image; Based on the inconsistency of feature distribution of different modalities, by calculating the mean μ and standard deviation σ of each modal feature, modal normalization is used to adjust the feature distribution to dynamically align the features of different modalities and generate aligned high-level features. The formula is: ; in, To align high-level features; and is a learnable parameter; Propagate updated features to graph neural networks; Features The corresponding mean; Features The corresponding standard deviation; The collaborative generation unit is used to calculate the feature information entropy of the aligned high-level features and adaptively generate dynamic weights for fusion to achieve collaborative enhancement of multimodal features, and obtain the modal collaborative features as follows: Compute aligned high-level features for each modality The information entropy of is used to measure the uncertainty of the modal characteristics. The formula is: ; in, is the information entropy of m-modal features; Representation alignment high-level features Probability distribution in feature space; Based on the information entropy calculated , using the exponential normalization strategy to generate the dynamic weights of each modality through a single-layer multi-layer perceptron: ; in, is the normalized weight of the m mode; for Information entropy of modal features; m, Represents the modality index of the image; Align high-level features of different modalities based on calculated dynamic weights Perform weighted fusion to generate the final modal collaborative features. The formula is: ; in, is the modal synergy feature; A contrastive learning unit, used for optimizing the inter-modal difference representation of the modal collaborative feature based on contrastive learning of positive samples and negative samples, and fusing the optimized features to obtain a fused feature; The output unit is used to input the fusion feature into a decoder to generate a final blood vessel segmentation result.

Citation Information

Patent Citations

  • Color eye fundus image blood vessel segmentation method and device

    CN117611603A

  • MRI (Magnetic Resonance Imaging) brain tumor segmentation method based on multi-scale feature fusion of improved U-Net

    CN117876399A