Vascular Image Segmentation Method, Device and Equipment Based on Differential Learning Model

Through the vascular image segmentation method based on the differential learning model, the differential characteristics of multimodal images are integrated, and the problems of insufficient multimodal data fusion and differential processing in the prior art are solved, which improves the accuracy and robustness of vascular segmentation, especially in the segmentation task of tiny blood vessels.

CN119832252BActive Publication Date: 2025-05-27XIAMEN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309179.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-05-27
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

The existing vascular image segmentation methods have significant shortcomings in multimodal data fusion and differential processing, resulting in limited segmentation accuracy and robustness, especially in the segmentation of complex and small blood vessels.

Method used

The vascular image segmentation method based on the differential learning model is adopted, and the difference between modes is integrated, and the differential features of multimodal images are extracted and fused through the modal difference marking module, 3D UNet encoder, position encoder, differential learning module and Transformer encoder module.

Benefits of technology

Improve the accuracy and robustness of vascular segmentation, especially in complex and tiny blood vessel segmentation tasks, the morphological characteristics of tiny blood vessels can be captured more accurately and the effects of noise interference and mode differences can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832252B_ABST
    Figure CN119832252B_ABST
Patent Text Reader

Abstract

The vascular image segmentation method, device and equipment based on the differential learning model provided by the present invention relate to the fields of neural networks and image processing. The present invention obtains a 3D vascular original image set of different modalities and inputs it into a vascular image segmentation model. In the segmentation model, a modality difference marking method is used to obtain a reconstructed image embedded with modality features; the reconstructed image is input into a 3D UNet encoder to extract low-level features; according to the low-level features, high-level features are generated through a position encoder and a self-attention mechanism; the high-level features are unfolded into one-dimensional vectors through a linear layer and optimized by a Transformer encoder to generate optimized features; then, the features optimized by the Transformer are used to calculate the difference features between modalities by element-wise subtraction, so as to obtain difference fusion features; the difference fusion features are reshaped by the Transformer encoder to obtain the final vascular segmentation result. The present invention can accurately capture the morphological features of small blood vessels and effectively improve the accuracy of vascular image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network learning and image processing. Specifically, it relates to a method, device, and equipment for vascular image segmentation based on a differential learning model. Background Art

[0002] Deep neural networks have been widely applied in the field of medical image processing. Especially in the vascular segmentation of 3D medical images, such as CT and MR angiography image analysis, it helps doctors to more accurately make disease diagnoses and formulate treatment plans. In recent years, UNet has made significant breakthroughs in medical image segmentation. Through end-to-end training for pixel-by-pixel prediction, UNet introduced skip connections between the encoder and the decoder, integrating low-resolution features into high-resolution features, which greatly improved the segmentation ability.

[0003] However, existing vascular image segmentation methods have significant deficiencies in multi-modal data fusion and differential processing. Traditional methods mainly rely on a single imaging modality (such as CTA or MRA), failing to fully utilize the complementarity of multi-modal data, resulting in limited segmentation accuracy and robustness. Especially in the segmentation of complex and tiny blood vessels, single-modal cannot comprehensively capture vascular features, affecting the accuracy of the segmentation results. In addition, existing methods have not effectively processed the differences between different modalities (such as visual features, resolution, and noise differences, etc.), resulting in inconsistent segmentation results for the same vascular structure under different modalities, affecting the overall performance and application effect of the segmentation model.

[0004] In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention

[0005] The present invention aims to provide a method, device, and equipment for vascular image segmentation based on a differential learning model to effectively integrate the differences between modalities, improve the accuracy and robustness of vascular segmentation, especially in the segmentation tasks of complex and tiny blood vessels.

[0006] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0007] A method for vascular image segmentation based on a differential learning model, comprising:

[0008] S1. Obtain a set of 3D vascular original images of different modalities and input them into a vascular image segmentation model based on differential learning; wherein, the vascular image segmentation model includes a modality difference marking module, a 3D UNet encoder, a position encoder, a differential learning module, and a Transformer encoder module;

[0009] S2. In the modal difference marking module, using the modal difference marking method, fuse and reconstruct the local patches obtained by splitting each image with the small cut blocks marked with corresponding modal features to obtain a reconstructed image embedded with modal features;

[0010] S3. Input the reconstructed image into the 3D UNet encoder to extract the low-level features of each image modality;

[0011] S4. Add position information to the low-level features of each image modality through the position encoder, and capture the global information of the input features by combining the self-attention mechanism to generate the high-level features corresponding to the image modality;

[0012] S5. The high-level features are expanded into corresponding one-dimensional vectors through a linear layer, and the normalization layer, multi-head self-attention mechanism and multi-layer perceptron of the Transformer encoder are used to optimize the one-dimensional vectors of each image modality to capture the global context information of the image and generate the optimized features;

[0013] S6. In the difference learning module, calculate the difference features between modalities by element-wise subtraction of the features optimized by the Transformer encoder, and after convolution and normalization processing, fuse them with the difference features to obtain the difference fusion features;

[0014] S7. The difference fusion features are processed by the Transformer encoder and reshaped into the same shape as the high-level features to obtain the final vascular segmentation result.

[0015] The 3D original vascular image set includes CTA images, MRA images and DSA images; each local patch contains the local region information of the vascular image; each local patch corresponds to a small cut block of the same size marked with modal features, that is, the modal feature marked cut block; the feature marking of the modal feature marked cut block is continuously trained and optimized through a pre-trained model to adapt to the features of a specific image modality and capture the specific information of the image modality.

[0016] Preferably, the fusion and reconstruction of the local patches of each image with the corresponding modal feature marked cut blocks are carried out as follows:

[0017] Each original image in the 3D original vascular image set is cut into several local patches, and a feature marking related to the modality of the original image is introduced into each local patch. Each local patch is fused with the corresponding modal feature marked cut block through feature fusion. The formula is:

[0018] ;

[0019] where represents the local patch The fused patch; [ ] represents the feature splicing operation; Encoder(s) represents converting the modal feature token chunk s into a feature vector of the same dimension as the corresponding local patch Feature vector of the same dimension;

[0020] Then, through the decoder, all the fused patches are combined and reconstructed to obtain a reconstructed map embedded with modal features , and the formula is:

[0021] ;

[0022] where Decoder is the decoder, represents the Pth local patch, and P is the number of local patches.

[0023] Preferably, the S4 is specifically:

[0024] Each input low-level feature adds position information through the position encoder; wherein, the position encoder is generated by sine and cosine functions;

[0025] The feature with position encoding is input into the multi-head self-attention mechanism, and after linear transformation, the self-attention result of each head is output;

[0026] The self-attention results of all heads are concatenated and linearly transformed to obtain the final self-attention result;

[0027] The final self-attention result is input into a multi-layer perceptron after layer normalization to perform deep non-linear processing on the normalized feature, and then added to the input low-level feature through a residual connection to obtain the high-level feature of the image modality.

[0028] Preferably, the Transformer encoder includes a normalization layer, a multi-head self-attention mechanism and a multi-layer perceptron; using the Transformer encoder to optimize the one-dimensional vectors of each image modality, specifically:

[0029] The normalized feature is input into the multi-head self-attention mechanism, and each self-attention head learns different relationships in the input feature in parallel;

[0030] Then the output results of all heads are concatenated together and linearly transformed into the output space to obtain the multi-head self-attention result;

[0031] The multi-head self-attention result is input into a multi-layer perceptron after layer normalization processing to perform deep non-linear processing on the normalized feature, and then added to the input feature through a residual connection to generate the feature optimized by the Transformer.

[0032] Preferably, S6 is specifically as follows:

[0033] Reshape the features optimized by the Transformer into the same shape as the high-level features, and then calculate the differential features between each image modality by element-wise subtraction. ;

[0034] Send the differential features into a convolutional layer and a normalization layer in sequence to obtain the normalized differential features.

[0035] Perform residual connection on the normalized differential features and the differential features to obtain the fused differential features. ;

[0036] Reshape the fused differential features back to one-dimensional features , which is the differential fusion feature.

[0037] Preferably, it further includes continuously optimizing the accuracy of multi-modal vascular image segmentation by using a loss function. The loss function combines multi-class weighted cross-entropy loss and , and its expression is:

[0038] ;

[0039] ;

[0040] ;

[0041] Among them, represents the true value of the c-th class and the n-th vascular image pixel, , n N, C is the total number of marked categories; N is the total number of pixels in the vascular image; is the predicted value.

[0042] Preferably, it further includes continuously optimizing the vascular image segmentation model by using a loss function to improve the accuracy of the model in multi-modal vascular image segmentation. The loss function used combines multi-class weighted cross-entropy loss and , and the expression is:

[0043] ;

[0044] Among them, and respectively represent the modal feature marked chunks corresponding to the i-th local patch through the encoder and The encoded feature representation; Denotes the inner product operation, i.e., the dot product of two vectors; Denotes the L2 norm; Denotes the total number of local patches;

[0045] By minimizing the loss function , the difference between the modal feature-labeled chunks and is reduced to improve the effectiveness of the labels.

[0046] The present invention also provides a vascular image segmentation device based on a difference learning model, comprising:

[0047] An input unit for acquiring a set of 3D vascular original images of different modalities and inputting them into a vascular image segmentation model based on difference learning; wherein, the vascular image segmentation model includes a modal difference labeling module, a 3D UNet encoder, a position encoder, a difference learning module, and a Transformer encoder module;

[0048] A modal difference labeling unit for fusing and reconstructing the local patches obtained by slicing each image with the small chunks labeled with corresponding modal features in the modal difference labeling module to obtain a reconstructed map embedded with modal features;

[0049] A 3D UNet encoding unit for inputting the reconstructed map into the 3D UNet encoder to extract the low-level features of each image modality;

[0050] A position encoding unit for adding position information to the low-level features of each image modality through the position encoder and capturing the global information of the input features in combination with the self-attention mechanism to generate the high-level features corresponding to each image modality;

[0051] A Transformer optimization unit for expanding the high-level features into corresponding one-dimensional vectors through a linear layer and optimizing the one-dimensional vectors of each image modality using the normalization layer, multi-head self-attention mechanism, and multi-layer perceptron of the Transformer encoder to capture the global context information of the image and generate optimized features;

[0052] A difference learning unit for calculating the difference features between modalities in the difference learning module by element-wise subtraction of the features optimized by the Transformer encoder, and after convolution and normalization processing, fusing them with the difference features to obtain difference fusion features;

[0053] An output unit, configured to process the difference fusion feature through a Transformer encoder, reshape it into the same shape as the high-level feature, and obtain the final blood vessel segmentation result.

[0054] The present invention also provides a blood vessel image segmentation device based on a differential learning model, including a processor and a memory. A computer program is stored in the memory and can be executed by the processor to implement a blood vessel image segmentation method based on a differential learning model as described above.

[0055] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a blood vessel image segmentation method based on a differential learning model as described above is implemented.

[0056] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0057] The present invention obtains multi-modal blood vessel image segmentation by constructing a blood vessel image segmentation model based on a differential learning model. The model includes a modal difference marking module, a 3D UNet encoder, a position encoder, a differential learning module, and a Transformer encoder module. In the modal difference marking module, the obtained original image is reconstructed using the modal difference marking method to capture the specific information of the image modality, and a reconstructed image embedded with modal features is obtained to make up for the differences between different imaging modalities and enhance the understanding and processing ability of the multi-modal blood vessel image segmentation model for multi-modal images.

[0058] The present invention introduces a multi-modal fusion method based on differential learning. By calculating and fusing the difference features between different modalities, the complementary information between different modality images is fully exploited, thereby overcoming the limitations of traditional single-modal methods in the segmentation of small blood vessels, improving the accuracy and robustness of blood vessel segmentation, and particularly performing outstandingly in the segmentation task of small blood vessels. Compared with the existing methods, the method of the present invention can more accurately capture the morphological features of small blood vessels, reduce the influence of noise interference and differences between modalities, and obtain accurate and fine segmentation results.

[0059] In addition, the present invention uses different loss functions to optimize the modal difference marking module and the multi-modal blood vessel image segmentation model respectively. By minimizing the loss function, the modal difference marking is optimized, the differences between different modalities are reduced, and the accuracy of multi-modal blood vessel image segmentation is improved. Description of the Drawings

[0060] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0061] Figure 1 Schematic diagram of a vascular image segmentation method based on a differential learning model provided for Embodiment 1.

[0062] Figure 2 Structural schematic diagram of a vascular image segmentation method based on a differential learning model provided for Embodiment 1.

[0063] Figure 3 Schematic diagram of the structure of the modal difference marking module provided for Embodiment 1.

[0064] Figure 4 Visualization display diagram of the segmentation result provided for Embodiment 1.

[0065] Figure 5 Schematic diagram of a vascular image segmentation device based on a differential learning model provided for Embodiment 2.

[0066] The following further details the present invention in conjunction with the accompanying drawings and specific embodiments. Specific Embodiments

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0068] Embodiment 1

[0069] Embodiment 1 of the present invention provides a vascular image segmentation method based on a differential learning model, which can be implemented by a vascular image segmentation device based on a differential learning model (hereinafter referred to as the segmentation device), and particularly, is executed by one or more processors in the segmentation device.

[0070] In this embodiment, the segmentation device may be an electronic device equipped with a processor, and the processor has a computer program of the vascular image segmentation method based on the differential learning model and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0071] Existing vascular segmentation methods have significant deficiencies in multi-modal data fusion and differential processing. Traditional methods mainly rely on a single imaging modality (such as CTA or MRA), and fail to fully utilize the complementarity of multi-modal data, resulting in limited segmentation accuracy and robustness. Especially in the segmentation of complex and small blood vessels, single modality cannot comprehensively capture blood vessel features, affecting the accuracy of the segmentation results. In addition, existing methods fail to effectively handle the differences between different modalities (such as visual features, resolution, and noise differences), resulting in inconsistent cross-modal segmentation results and further reducing the overall performance of the model. Therefore, this embodiment aims to propose a multi-modal fusion method based on differential learning to effectively integrate the differences between modalities and improve the accuracy and robustness of vascular segmentation, especially in the segmentation task of complex and small blood vessels.

[0072] As Figures 1 - 2 shown, a vascular image segmentation method based on a differential learning model includes steps S1 to S7.

[0073] S1. Obtain a set of 3D original vascular images of different modalities and input them into a vascular image segmentation model based on differential learning; wherein, the vascular image segmentation model includes a modality difference marking module, a 3D UNet encoder, a position encoder, a differential learning module, and a Transformer encoder module.

[0074] Specifically, the set of 3D original vascular images of different modalities includes but is not limited to: CTA images, MRA images, DSA images, etc. Among them, CTA (Computed Tomography Angiography) images are obtained by injecting a contrast agent intravenously to increase the contrast agent concentration in the blood vessels, so that the blood vessel structure can be clearly shown on the CT image. MRA (Magnetic Resonance Angiography) images can show the blood vessel structure without using a contrast agent. It uses the signal difference generated by blood flow to distinguish blood vessels from other tissues. DSA (Digital Subtraction Angiography) is an imaging technique that, after injecting a contrast agent, uses computer processing to remove the images of bones and soft tissues and only retains the blood vessel images.

[0075] As Figures 2 - 3As shown in the figure, the multi-modal vascular image segmentation model based on differential learning includes a modal difference marking module, a 3D UNet encoder, a position encoder, a differential learning module, and a Transformer encoder module.

[0076] S2. In the modal difference marking module, using the modal difference marking method, the local patches obtained by slicing each image are fused and reconstructed with the small blocks marked with corresponding modal features to obtain a reconstructed image embedded with modal features.

[0077] As Figure 3 shown, taking CTA images and MRA images as examples, we reconstruct the acquired original images using the modal difference marking method to obtain a reconstructed image embedded with modal features, so as to make up for the differences between different imaging modalities (CTA and MRA), and enhance the model's understanding and processing ability of multi-modal images. In this embodiment, each local patch contains local region information of the vascular image. The feature markers corresponding to the modal feature marking blocks corresponding to each local patch are obtained through continuous training and optimization of the pre-trained model to adapt to the features of specific image modalities and capture the specific information of the image modalities.

[0078] In this embodiment, by introducing the fusion strategy of modal specificity markers and local patches, the unique features in different modalities are fully mined to improve the expressiveness and generalization ability of the model. The local patches of each image are fused and reconstructed with the corresponding feature markers. The specific steps are as follows:

[0079] Slice each original image in the 3D vascular original image set into several local patches , and introduce feature markers related to the modality of the original image for each local patch, that is, each small block (local patch) corresponds to a small block of the same size used to mark their modal features, that is, the corresponding modal feature marking block . For each local patch and its corresponding modal feature marking block s, adopt a feature fusion strategy (such as using the splicing method in this embodiment, or other methods such as weighted fusion according to the situation) to enhance the expression ability of the patch, and fuse each local patch with the corresponding modal feature marking block through feature fusion. The formula is:

[0080] ;

[0081] where represents the local patch after fusion; [ ] represents the feature splicing operation; Encoder(s) represents converting the modal feature marking block s into the corresponding local patch through the encoder Feature vectors of the same dimension; i = 1, 2, …, P; P is the number of local patches.

[0082] The fused patch Not only retains the spatial information of the local image, but also incorporates modality-specific features, thus being able to better represent cross-modal features.

[0083] After all local patches complete feature fusion, through the decoder, all the fused patches are combined and reconstructed to obtain a reconstructed map embedded with modality features , the formula is:

[0084] ;

[0085] Among them, Decoder is the decoder, represents the Pth local patch.

[0086] S3, input the reconstructed map into the 3D UNet encoder to extract the low-level features of each image modality.

[0087] In this embodiment, the 3D UNet encoder is a variant of the UNet architecture, specifically designed to process three-dimensional medical image data, mainly composed of multiple convolutional layers and pooling layers. These layers gradually extract and compress the features of the input three-dimensional medical image in a stacked manner.

[0088] Convolutional layer: Responsible for extracting local features in the image. In 3D UNet, the convolutional layer uses a three-dimensional convolutional kernel and can extract features in volumetric data. Usually, a non-linear activation function (such as ReLU) follows each convolutional layer to enhance the non-linear ability of the network.

[0089] Pooling layer: Used to gradually reduce the size of the feature map while increasing the receptive field size. This helps the network capture more global feature information. In 3D UNet, the pooling layer usually uses three-dimensional max pooling or average pooling.

[0090] S4, add position information to the low-level features of each image modality through the position encoder, and combine the self-attention mechanism to capture the global information of the input features, generating high-level features corresponding to each image modality.

[0091] Furthermore, in the position encoder part of the vascular image segmentation model, position information is added to each input low-level feature through the position encoder; among them, the position encoder is generated by sine and cosine functions, aiming to add position information to each input feature, thereby helping the model understand the relative positions of elements in the sequence. Through position encoding, the input features not only contain the original image information but also introduce the order relationship of the data, enabling the model to capture the internal structural information of the data.

[0092] The features with positional encoding are input into the multi-head self-attention mechanism of the Transformer encoder. The multi-head self-attention mechanism generates a query matrix (Q), a key matrix (K), and a value matrix (V) by performing a linear transformation on the input features. Then, the dot product similarity between the query matrix and the key matrix is calculated and normalized to obtain the attention weight matrix. The attention weight matrix is weighted and summed with the corresponding value matrix to output the results of each head of self-attention, and the outputs of all heads are concatenated. In this way, the self-attention mechanism can dynamically adjust the weights of each element according to the relationships between the input features, thereby effectively capturing global information.

[0093] After processing the self-attention, the final self-attention result is input into a multi-layer perceptron after layer normalization to perform deep non-linear processing on the normalized features, and then added to the input low-level features through a residual connection to obtain the high-level features of the image modality, ensuring the effective flow of information.

[0094] Normalization processing can ensure the stability of the mean and variance of the features, thereby helping to accelerate training and prevent gradient vanishing or explosion. The multi-layer perceptron in this embodiment consists of multiple fully connected layers, and the transformation between each layer is usually performed through a non-linear activation function (such as the rectified linear unit). The role of the multi-layer perceptron is to perform deep non-linear processing on the normalized features, thereby further enhancing the expressive power of the model.

[0095] Through the iterative processing of the above multiple steps, the output of the position encoder is a set of high-dimensional feature representations that fuse global context information and .

[0096] In this step, high-order features of the input low-level features are extracted through the position encoder and the self-attention mechanism. These high-level features can not only effectively capture the complex relationships in the input data but also provide accurate feature representations for downstream tasks (such as vessel segmentation).

[0097] S5. The high-level features are expanded into corresponding one-dimensional vectors through a linear layer, and the normalization layer, multi-head self-attention mechanism, and multi-layer perceptron of the Transformer encoder are used to optimize the one-dimensional vectors of each image modality to capture the global context information of the image and generate optimized features.

[0098] In this step, the role of the linear layer is to perform a linear transformation on the input high-level features and convert their dimensions to the input dimensions required by the Transformer encoder, that is, to expand them into one-dimensional vectors and 。Then, the one-dimensional vectors of each image modality are input into the Transformer encoder.

[0099] As Figure 2 shown, the Transformer encoder structure in this embodiment sequentially includes: a normalization layer, a multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron MLP. The normalized input features are input into the multi-head self-attention mechanism and divided into multiple subspaces. Each subspace is processed by an independent self-attention head, and each self-attention head learns different relationships in the input features in parallel. Finally, the output results of all heads are concatenated together and mapped to the final output space through a linear transformation to obtain the final self-attention result. In this way, the model can capture rich information in the input data from multiple perspectives and enhance the modeling ability for complex data.

[0100] The multi-head self-attention (Multi-Head Self-Attention) layer is the core part of the Transformer encoder, which allows the model to pay attention to all other elements when processing each input element, thereby learning the global dependencies in the input sequence. The multi-head self-attention mechanism calculates the results of multiple attention heads in parallel and concatenates them to enhance the model's representation ability.

[0101] Each Transformer encoder layer also includes residual connection (Residual Connection) and layer normalization (Layer Normalization) operations. Residual connection helps alleviate the vanishing gradient problem in deep neural networks, while layer normalization helps accelerate the training process and improve the model's stability.

[0102] The global dependencies between different positions (or different modalities) are captured through the multi-head self-attention mechanism, thereby generating a feature representation containing global context information. Through these operations, the Transformer encoder can gradually optimize the representation of the input vector to make it contain richer global context information.

[0103] After the output of the multi-head self-attention is processed by layer normalization, it is input into the multi-layer perceptron to perform deep non-linear processing on the normalized features, and then added to the input features through residual connection to generate optimized features.

[0104] The Transformer encoder outputs optimized feature representations and . These feature representations not only contain the high-level features of the original image but also the global dependencies between different positions (or different modalities).

[0105] In this embodiment, the Transformer encoder is an important part of the Transformer model. It is responsible for converting the input sequence information into a set of representation vectors for subsequent training and prediction by the feed-forward neural network or decoder.

[0106] S6. In the difference learning module, the features optimized by the Transformer are used to calculate the difference features between modalities by element-wise subtraction. After convolution and normalization, they are fused with the difference features to obtain the difference fusion features.

[0107] The difference learning module is used to solve the modality difference problem in multi-modal data processing. Through difference modeling and feature fusion, complementary information from different modalities (CTA and MRA) is accurately extracted to optimize the segmentation results. First, the two modality features and processed by the Transformer encoder are reshaped into the same shape as the original high-level features, obtaining and respectively:

[0108] ;

[0109] where reshape is the reshaping function.

[0110] This operation ensures the consistency of the shapes of the features of different modalities, providing a unified basis for subsequent difference calculation. Next, the model calculates the difference features between modalities by element-wise subtraction , and the expression is:

[0111] ;

[0112] The obtained difference features describe the difference information between the two image modalities in terms of visual features, resolution, noise, etc., providing a basis for subsequent feature fusion and optimization.

[0113] Then, the difference features are fed into the convolutional layer and normalized to obtain the normalized difference features, which helps to smooth the difference features, reduce the influence of noise, and enhance the model's ability to capture key features. Next, the normalized difference features are subjected to residual connection with the original difference features to obtain the fused difference features , and the expression is:

[0114] ;

[0115] The residual connection can retain the useful information of the original differential features, and at the same time deepen the feature expression ability of the model through reinforcement learning, thereby improving the recognition ability of complex and small blood vessel structures. Finally, the fused differential features are reshaped back into one-dimensional features , which is the differential fusion feature described above:

[0116] ;

[0117] Through this differential learning module, the model corresponding to the method of the present invention (referred to as the DL_Net model) can fully fuse the differential information between different modalities, thereby improving the accuracy and robustness of blood vessel segmentation.

[0118] S7, the differential fusion feature is processed by the Transformer encoder and reshaped into the same shape as the high-level feature to obtain the final blood vessel segmentation result.

[0119] In this embodiment, the differential fusion feature is reshaped through the Transformer encoder part to obtain the final blood vessel segmentation result.

[0120] In another preferred embodiment, the present invention further includes continuously optimizing the accuracy of the blood vessel image segmentation model based on the differential learning model by using a loss function. The loss function used combines multi-class weighted cross-entropy loss and , and its expression is:

[0121] ;

[0122] ;

[0123] ;

[0124] wherein, represents the true value of the c-th class and the n-th blood vessel image pixel, , n N, C is the total number of categories; N is the total number of pixels of the blood vessel image; is the predicted value.

[0125] In addition, when the feature markers corresponding to each local patch involved in the modality difference marking module are continuously trained and optimized by the pre-trained model. In order to optimize the modality difference marking and reduce the difference between different modalities, a loss function based on cosine similarity is used to optimize the representation of the feature markers corresponding to each local patch to reduce the difference between different image modalities, and the expression is:

[0126] ;

[0127] Among them, and respectively represent the feature representations obtained by encoding the modality feature token chunks and corresponding to the i-th local patch through the encoder; represents the inner product operation, that is, the dot product of two vectors; represents the L2 norm; represents the total number of local patches;

[0128] By minimizing the loss function , the difference between the modality feature token chunks and is reduced, the representations of the same feature in different modalities are more consistent, while different features maintain their differences, thereby improving the effectiveness of cross-modal feature fusion.

[0129] As Figure 4 shown, the segmentation results of the method model of the present invention (the vascular image segmentation model based on difference learning, i.e., Ours in the figure) are visually compared with those of six common medical image segmentation models (3DUnet, SegResNet, SegResNetVAE, UNETR, SwinUNETR, and CSNet3D). The red box area of the Ground Truth (true annotation or reference segmentation result) in the figure shows the small and easily overlooked blood vessels between different blood vessels. In the CT and MR image modalities, the model we proposed (i.e., Ours in the figure) successfully pays attention to these small blood vessels and performs correct segmentation. While other models either ignore these blood vessels or mis-segment them into other types. It can be seen that the model we proposed can pay more attention to the details in the image through difference learning, thereby achieving a more accurate segmentation effect.

[0130] Note: Figure 4 In, CT-case A represents a case of CTA image. MR-case A represents a case of MRA image. Ground Truth: Refers to the true annotation or reference segmentation result, which is used as a benchmark for evaluating the segmentation performance. Ours: The segmentation result obtained by using the method model of the present invention.

[0131] The following are several commonly used medical image segmentation models:

[0132] 3D U-Net: A 3D image segmentation model based on convolutional neural network (CNN), which is widely used in medical image segmentation tasks. Its basic principle is: input a 3D image into an encoder, and reduce its dimension through a series of convolutional layers. Then, upsample the output of the encoder through a series of transposed convolutional layers, and finally output a segmentation result with the same size as the original input.

[0133] SegResNet: A segmentation network based on the ResNet architecture, which uses deep residual learning to improve image segmentation performance. Its basic principle is: train a deeper network through residual learning, that is, learn the residual between the network input and output, rather than directly learning the mapping relationship from the original input to the output.

[0134] SegResNetVAE: A SegResNet model combined with variational autoencoder (VAE), which improves the robustness of segmentation by enhancing the generation ability. VAE is a generative model used to learn the data distribution and generate new data samples. It realizes this by mapping the input data to an encoded vector in a low-dimensional latent space, and then mapping the encoded vector back to the original data space through a decoder.

[0135] UNETR: A 3D image segmentation model that combines the Transformer architecture with U-Net, suitable for 3D medical image segmentation.

[0136] SwinUNETR: A 3D image segmentation model that fuses Swin Transformer and convolutional neural network (CNN). It captures long-range dependencies through Swin Transformer and uses CNN for local feature extraction at the same time.

[0137] CSNet3D: A 3D convolutional neural network (CNN) model specifically designed for medical image segmentation.

[0138] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0139] The multi-modal vascular image segmentation model based on differential learning of the present invention first reconstructs the original image into a reconstructed image embedded with modal features in the modal difference marking module; then uses the 3D U-Net encoder to extract the low-level features of each modality of the reconstructed image respectively, and combines the position encoding to generate the high-level features of the modality. Next, the high-level features of each modality are processed by the Transformer encoder to capture the global context information. Differential learning is used to model and integrate the differences between different modalities (such as visual features, resolution, and noise differences), generate differential features, and after processing, fuse them with the original differential features. Finally, the fused features are further encoded to generate the vascular segmentation result.

[0140] The present invention effectively integrates 3D image data of different modalities by introducing a differential learning mechanism to improve the accuracy of blood vessel segmentation, especially having obvious advantages in the segmentation tasks of small blood vessels and complex blood vessels.

[0141] Embodiment 2

[0142] As Figure 5 shown, the second embodiment of the present invention also provides a blood vessel image segmentation device based on a differential learning model, including:

[0143] An input unit for obtaining a set of 3D blood vessel original images of different modalities and inputting them into a blood vessel image segmentation model based on differential learning; wherein, the blood vessel image segmentation model includes a modality difference marking module, a 3D UNet encoder, a position encoder, a differential learning module, and a Transformer encoder module;

[0144] A modality difference marking unit for fusing and reconstructing local patches obtained by slicing each image with small blocks marked with corresponding modality features in the modality difference marking module to obtain a reconstructed image embedded with modality features;

[0145] A 3D UNet encoding unit for inputting the reconstructed image into the 3D UNet encoder to extract low-level features of each image modality;

[0146] A position encoding unit for adding position information to the low-level features of each image modality through the position encoder and capturing the global information of the input features in combination with the self-attention mechanism to generate high-level features corresponding to the image modality;

[0147] A Transformer optimization unit for expanding the high-level features into corresponding one-dimensional vectors through a linear layer and optimizing the one-dimensional vectors of each image modality by using the normalization layer, multi-head self-attention mechanism, and multi-layer perceptron of the Transformer encoder to capture the global context information of the image and generate optimized features;

[0148] A differential learning unit for calculating the modality difference features by element-wise subtraction of the features optimized by the Transformer encoder in the differential learning module, and after convolution and normalization processing, fusing them with the difference features to obtain difference fusion features;

[0149] An output unit for processing the difference fusion features through the Transformer encoder and reshaping them into the same shape as the high-level features to obtain the final blood vessel segmentation result.

[0150] Embodiment 3

[0151] The third embodiment of the present invention also provides a vascular image segmentation device based on a differential learning model, which includes a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement the vascular image segmentation method based on the differential learning model as described above.

[0152] Embodiment Four

[0153] The fourth embodiment of the present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the vascular image segmentation method based on the differential learning model as described above is implemented.

[0154] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0155] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0156] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.

[0157] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0158] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0159] Depending on the context, the word "if" as used herein can be interpreted as "when", "while", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".

[0160] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0161] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A blood vessel image segmentation method based on a difference learning model, characterized in that: include: S1, obtaining a set of 3D vascular original images of different modalities, and inputting them into a vascular image segmentation model based on difference learning; wherein the vascular image segmentation model includes a modality difference labeling module, a 3D UNet encoder, a position encoder, a difference learning module and a Transformer encoder module; S2, in the modality difference labeling module, using the modality difference labeling method, the local patches obtained by segmenting each image are fused and reconstructed with the small blocks labeled with the corresponding modality features to obtain a reconstructed image embedded with the modality features; S3, inputting the reconstructed image into a 3D UNet encoder to extract low-level features of each image modality; S4, adds position information to the low-level features of each image modality through the position encoder, and combines the self-attention mechanism to capture the global information of the input features to generate high-level features of the corresponding image modality; S5, the high-level features are expanded into corresponding one-dimensional vectors through a linear layer, and the one-dimensional vectors of each image modality are optimized using the normalization layer of the Transformer encoder, the multi-head self-attention mechanism and the multi-layer perceptron to capture the global context information of the image and generate optimized features; S6, in the difference learning module, the features optimized by the Transformer encoder are subtracted element by element to calculate the difference features between the modalities, and after convolution and normalization, they are fused with the difference features to obtain the difference fusion features; S7, the difference fusion feature is processed by the Transformer encoder and reshaped into the same shape as the high-level feature to obtain the final blood vessel segmentation result.

2. The vascular image segmentation method based on the difference learning model according to claim 1 is characterized in that The 3D vascular original image set includes CTA images, MRA images and DSA images; each of the local patches contains the local area information of the vascular image; each local patch corresponds to a small block of the same size marked with modal features, namely, the modal feature marked block; the feature labels of the modal feature marked block are obtained by continuous training and optimization of the pre-trained model to adapt to the characteristics of a specific image modality and capture the specific information of the image modality.

3. The vascular image segmentation method based on the difference learning model according to claim 2 is characterized in that ,The local patch of each image is fused and reconstructed with the corresponding modal feature marked block. The specific steps are as follows: Each original image of the 3D vascular original image set is divided into several local patches, each local patch introduces a feature tag related to the original image modality, and each local patch is fused with the corresponding modality feature tag block through feature fusion. The formula is: ; in, Indicates a local patch The fused patch; [ ] indicates the feature concatenation operation; Encoder(s) indicates the conversion of the modal feature marker block s into the corresponding local patch Feature vectors of the same dimension; ; P is the number of local patches; Then, through the decoder, all the fused patches are combined and reconstructed to obtain the reconstructed image embedded with the modal features. , the formula is: ; Among them, Decoder is a decoder, represents the Pth local patch.

4. The vascular image segmentation method based on the difference learning model according to claim 1 is characterized in that , the S4 is specifically: Each low-level feature inputted is added with position information through the position encoder; wherein the position encoder is generated by sine and cosine functions; The features with position encoding are input into the multi-head self-attention mechanism, and after linear transformation, the self-attention results of each head are output; Concatenate and linearly transform the self-attention results of all heads to obtain the final self-attention result; The final self-attention result is normalized by layers and then input into a multi-layer perceptron to perform deep nonlinear processing on the normalized features. It is then added to the input low-level features through residual connections to obtain high-level features of the image modality.

5. The vascular image segmentation method based on the difference learning model according to claim 1 is characterized in that , the Transformer encoder includes a normalization layer, a multi-head self-attention mechanism and a multi-layer perceptron; the Transformer encoder is used to optimize the one-dimensional vector of each image modality, specifically: The normalized features are input into the multi-head self-attention mechanism, and each self-attention head learns different relationships in the input features in parallel; Then the output results of all heads are concatenated together and mapped to the output space through a linear transformation to obtain the multi-head self-attention result; The multi-head self-attention results are normalized by layers and then input into a multi-layer perceptron to perform deep nonlinear processing on the normalized features. They are then added to the input features through residual connections to generate Transformer optimized features.

6. The vascular image segmentation method based on the difference learning model according to claim 1, characterized in that , the S6 is specifically: Reshape the Transformer optimized features into the same shape as the high-level features, and then calculate the difference features between the image modalities by element-by-element subtraction ; Sending the difference features to the convolution layer and the normalization layer in sequence to obtain normalized difference features; The normalized difference feature is compared with the difference feature Perform residual connection to obtain the fused difference features ; The fused difference features Reshape back to one-dimensional features , which is the difference fusion feature.

7. The vascular image segmentation method based on the difference learning model according to claim 1 is characterized in that , also includes, using the loss function to continuously optimize the vascular image segmentation model to improve the accuracy of the model in multimodal vascular image segmentation. The loss function used Combining multi-class weighted cross entropy loss and , whose expression is: ; ; ; in, represents the true value of the c-th and n-th blood vessel image pixel, , n N, C is the total number of labeled categories; N is the total number of pixels in the vascular image; is the predicted value.

8. The vascular image segmentation method based on the difference learning model according to claim 3 is characterized in that , also includes: when the feature labels of the modal feature label blocks are continuously trained and optimized through the pre-training model, a loss function based on cosine similarity is used The representation of the modality feature marker slice corresponding to each local patch is optimized to reduce the difference between different image modalities. The expression is: ; in, and Respectively represent the modal feature labeling blocks corresponding to the i-th local patch through the encoder and Feature representation after encoding; Represents the inner product operation, that is, the dot product of two vectors; represents the L2 norm; represents the total number of local patches; By minimizing the loss function , so that the modal feature markers are cut and The differences between them are reduced to improve the effectiveness of marking.

9. A blood vessel image segmentation device based on a difference learning model, characterized in that: include: An input unit, used to obtain a set of 3D vascular original images of different modalities and input them into a vascular image segmentation model based on difference learning; wherein the vascular image segmentation model includes a modality difference labeling module, a 3D UNet encoder, a position encoder, a difference learning module and a Transformer encoder module; A modality difference labeling unit, used in the modality difference labeling module, to fuse and reconstruct the local patches obtained by segmenting each image with the small blocks labeled with the corresponding modality features by using the modality difference labeling method, so as to obtain a reconstructed image embedded with the modality features; A 3D UNet encoding unit, used for inputting the reconstructed image into a 3D UNet encoder to extract low-level features of each image modality; The position encoding unit is used to add position information to the low-level features of each image modality through the position encoder, and combine the self-attention mechanism to capture the global information of the input features to generate high-level features of the corresponding image modality; A Transformer optimization unit is used to expand the high-level features into corresponding one-dimensional vectors through a linear layer, and optimize the one-dimensional vectors of each image modality using the normalization layer, multi-head self-attention mechanism and multi-layer perceptron of the Transformer encoder to capture the global context information of the image and generate optimized features; The difference learning unit is used to calculate the difference features between the modalities by element-by-element subtraction of the features optimized by the Transformer encoder in the difference learning module, and then fuse them with the difference features after convolution and normalization to obtain the difference fusion features; The output unit is used to process the difference fusion feature through the Transformer encoder and reshape it into the same shape as the high-level feature to obtain the final blood vessel segmentation result.

10. A blood vessel image segmentation device based on a difference learning model, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a blood vessel image segmentation method based on a difference learning model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal image segmentation model and method based on space attention mechanism

    CN116977346A

  • Microvascular decompression-oriented multi-modal multi-target segmentation method and system

    CN117671748A