An automatic segmentation method for the trigeminal nerve based on a deep network

Through the multimodal data fusion method based on deep network, fully automatic and accurate segmentation of trigeminal nerves is achieved, solving the problems of artificial dependence and operational non-repeatability in traditional methods, and improving segmentation accuracy and efficiency.

CN114972745BActive Publication Date: 2025-05-27ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210379996.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-05-27
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

The traditional trigeminal nerve segmentation method has artificial dependence and operational non-repeatability, resulting in low segmentation accuracy and long time-consuming, and existing algorithms are difficult to complete segmentation of trigeminal nerves and segmentation of other smaller cranial nerves.

Method used

The trigeminal nerve automatic segmentation method is adopted based on deep networks, and the deep information fusion of multimodal data is used to perform fully automatic and precise segmentation of trigeminal nerves using a pyramid deep convolution network.

Benefits of technology

The fully automatic and accurate segmentation of trigeminal nerves is realized, which greatly reduces the workload of expert label identification and provides stable, efficient and repeatable analysis methods for the segmentation of other cranial nerves.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972745B_ABST
    Figure CN114972745B_ABST
Patent Text Reader

Abstract

A method for automatic segmentation of the trigeminal nerve based on a deep network. The network performs modality fusion on the individual modalities of T1, DEC, and FOD images and consists of two independent analysis paths and a shared synthesis path to achieve multimodal fusion. The fusion of multimodal data can provide more information for model prediction, thereby improving the accuracy of the trigeminal nerve prediction results. The present invention effectively fuses the depth information of multimodal data, realizes the full-automatic and accurate segmentation of the trigeminal nerve, greatly reduces the workload of expert annotation and recognition, and at the same time provides a stable, efficient, and repeatable analysis method for the segmentation research of other cranial nerves besides the trigeminal nerve.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical imaging and neuroanatomy under computer graphics, and in particular to an automatic trigeminal nerve segmentation method based on a deep network. Background Art

[0002] Cranial nerves are twelve pairs of left and right nerves that emanate from the human brain, controlling important functions such as smell, vision, and eye movement. Among them, the trigeminal nerve is the fifth pair of cranial nerves among the twelve pairs of cranial nerves, and is the largest and most complex nerve among all cranial nerves. The segmentation of cranial nerves can well show the relationship between the lesion area and the cranial nerves in medical images, providing reliable theoretical support for doctors' preoperative planning.

[0003] In the traditional trigeminal nerve segmentation process, experts with many years of work experience are required to select the trigeminal nerve by placing a region of interest (ROI) before segmentation. This ROI is used as a seed point, and the sample is tracked using probabilistic and deterministic tracking algorithms. Finally, experts with anatomical knowledge are required to manually remove fibers that do not belong to the trigeminal nerve to obtain the final trigeminal nerve segmentation result. As can be seen from the above, there are two key problems with the traditional method, namely, manual dependence and non-repeatability of operations. Specifically, first, researchers are easily disturbed by subjective factors during the actual operation process, which affects the accuracy of the research results to a certain extent; second, the researchers' drawing methods are not fixed, the standards are not unified, and repeatability is difficult to achieve; in addition, the above manual operation process needs to be repeated for each brain data, which is time-consuming.

[0004] In order to classify a pixel, the traditional convolutional neural network-based segmentation method uses an image block around the pixel as the input of the CNN for training and prediction. The disadvantages of this type of method are: ① The storage overhead is large and the computational efficiency is low; ② The size of the pixel block limits the size of the perception area. Usually, the size of the pixel block is much smaller than the size of the entire image, and only some local features can be extracted, resulting in limited classification performance. In order to solve the above problems, a fully convolutional network was proposed, which can classify images at the pixel level. Later, the U-Net network was proposed to further improve the accuracy of image segmentation. U-Net is a deep learning network based on a fully convolutional network for semantic segmentation applied to biomedicine.

[0005] However, existing algorithms face the challenge of difficulty in segmenting the trigeminal nerve completely and other smaller cranial nerves cannot be segmented. The research methods of automatic and accurate segmentation of the trigeminal nerve still need to be improved. Summary of the invention

[0006] To overcome the deficiencies of the prior art and address the problems of cumbersome manual labeling and low segmentation accuracy in traditional segmentation methods, the present invention provides an automatic segmentation method for the trigeminal nerve based on a deep network. This method effectively integrates the depth information of multi-modal data, achieving fully automatic and precise segmentation of the trigeminal nerve. It not only greatly reduces the workload of expert annotation and recognition but also provides a stable, efficient, and repeatable analysis method for the segmentation research of other cranial nerves besides the trigeminal nerve. From the perspective of the network framework, the network proposed in the present invention performs modal fusion on the individual modalities of T1, DEC, and FOD images, consisting of two independent analysis paths and a shared synthesis path to achieve multi-modal fusion. The fusion of multi-modal data can improve the accuracy of the trigeminal nerve prediction results because it can provide more information for model prediction.

[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0008] An automatic segmentation method for the trigeminal nerve based on a deep network, the method comprising the following steps:

[0009] Step 1, dataset preparation

[0010] Use the MRI and DWI images of N cases of HCP data, select the T1 image in the MRI image and the direction-encoded color DEC image and fiber orientation distribution function FOD image generated from the DWI image. Among them, the DEC image shows fibers with different directions of travel through different colors, clearly showing the normal anatomy and its travel of the cerebral white matter fibers; the FOD image is a macroscopic manifestation of the water molecule diffusion direction, quantifying the proportion of fiber orientations in each direction within a voxel.

[0011] Step 2, labeled dataset preparation

[0012] First, draw a rough region MASK containing the path of the trigeminal nerve in the common space and register it to N individuals. Then, use the dual-tensor unscented Kalman filter (UKF-2) method to scatter seed points within the MASK for fiber tracking. After that, first automatically screen the tracked fibers using the region of interest (ROI) and region of avoidance (ROA), and then perform layer-by-layer screening at the expert level. The double screening ensures the accuracy of the results. Finally, map the screened fibers onto the voxels to generate the final more accurate labeled data GroundTruth;

[0013] Step 3, data preprocessing

[0014] Slice the data obtained in Step 1 into a size of 128×160×128, perform histogram equalization, gray histogram normalization, and image augmentation operations on the image data to complete the preparation of training data and test data;

[0015] Step 4: Network design and training

[0016] Construct a pyramid deep convolutional network for multimodal fusion, perform multimodal fusion of T1, DEC and FOD images, and then use the training samples generated in steps 1 and 2 above and the constructed network model with GroundTruth training;

[0017] Step 5: Predict segmentation

[0018] Use the network model trained in step 4 to predict the optic nerve area of ​​the test data, compare the predicted results with the marked results, and calculate the prediction accuracy.

[0019] Furthermore, in step 2, the preparation of the labeled data set includes the following steps:

[0020] 2.1 Draw the mask

[0021] Based on anatomical knowledge, an elliptical area MASK is drawn in the common space of the T1 data of N subjects. This MASK includes the area where the trigeminal nerve passes through, so that all trigeminal nerve fibers can be tracked. The MASK drawn in the common space is then registered to the individual as the MASK for individual fiber tracking.

[0022] 2.2 Fiber tracking

[0023] A set number of seed points was selected within the MASK, and then the dual tensor unscented Kalman filter UKF-2T method was used for fiber tracking. The hybrid model of the two tensors was fitted to the dMRI data at the same time, providing highly sensitive fiber tracking capabilities, especially in the presence of crossing fibers, to track the putative trigeminal spinal tract and the putative trigeminal mesencephalic tract;

[0024] 2.3 Automatic screening

[0025] First, the nerve fibers were automatically screened, using the Meckel chamber drawn on the b=0 image and the cisternal segment drawn on the direction-coded color map of the diffusion tensor imaging (DTI) as the region of interest (RO) to automatically screen and preliminarily select the fibers belonging to the trigeminal nerve;

[0026] 2.4 Manual screening

[0027] Because the results of the above operations contained a large number of false positive results, in order to further ensure the accuracy of the mask, after automatically screening the fibers using the region of interest, the experts used anatomical knowledge to further manually screen out the trigeminal nerve fiber bundles of all subjects;

[0028] 2.5 Mask Generation

[0029] After the above double screening, high-quality trigeminal nerves are generated. Then, we map the three-dimensional streamlines onto voxels to obtain binary nerve fiber masks; these binary masks are analyzed as connected regions, and only the largest connected region is retained, thus filtering out single voxels and small groups that are not connected to other voxels, so as to obtain the GroundTruth of the final trigeminal nerve.

[0030] Furthermore, in the fourth step described above, the network design and training process are as follows:

[0031] Build a multi-modal fusion model based on the U-Net pyramid deep convolutional network. Considering that the segmentation map gradually forms during the synthesis process, it may be beneficial to perform modal fusion starting from the synthesis stage after analyzing each individual modality. The built network consists of two independent analysis paths and a shared synthesis path. The first independent analysis path is as follows: Combine the T1 and DEC images as the input, perform convolutional operations and then enter the encoder module. The encoder module contains 4 convolutional layers and max-pooling layers, with 32, 64, 128, and 256 feature maps respectively; the decoder module contains 4 deconvolutional layers and convolutional layers, with 256, 128, 64, and 32 feature maps respectively. For all convolutional layers, the size of the convolutional kernel is 3×3×3; for all max-pooling layers, the pool size is 2×2×2 and the stride is 2; for all deconvolutional layers, the deconvolved feature maps are combined with the corresponding features in the encoder module. The other independent analysis path is: Use the FOD image as the input, and after convolution, perform the encoding and decoding process in the same way as the first independent analysis path. The difference is that the encoder-decoder of this path only has three layers. The shared synthesis path is formed by connecting the outputs of the independent analysis paths. When the synthesis progresses from one stage to another, the corresponding stages from each analysis path are also connected to the input. The synthesis path uses a Softmax classifier to generate voxel-level probability maps and predictions. After the network is built, use the training samples generated in the previous steps to train the built network model, further modify and adjust its parameters, and finally use the network model for experiments on new data.

[0032] Preferably, in order to measure the quality of the network model prediction, that is, the degree of difference between the predicted value and the true value of the model, the Dice coefficient is used as the Loss function of the network. The definition of the Dice coefficient is:

[0033]

[0034] where TP is the number of true positive voxels, FP is the number of false positive voxels, FN is the number of false negative voxels, and p x ∈P:Ω→{0,1} is the predicted binary segmentation content, and g x∈G:Ω→{0,1} is the binary content with true value.

[0035] The beneficial effects of the present invention are as follows: effectively integrating the depth information of multi-modal data, realizing the full-automatic and accurate segmentation of the trigeminal nerve, greatly reducing the workload of expert annotation and recognition, and at the same time providing a stable, efficient and repeatable analysis method for the segmentation research of other cranial nerves except the trigeminal nerve. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is the implementation step flowchart of this method.

[0037] Figure 2 is the schematic diagram of the encoding and decoding process of the deep network. DETAILED DESCRIPTION OF THE INVENTION

[0038] The present invention will be further described below.

[0039] Referring to Figure 1 and Figure 2 , an automatic segmentation method for the trigeminal nerve based on a deep network includes the following steps:

[0040] Step 1, dataset preparation

[0041] Using the MRI and DWI images of 100 cases of HCP data, select the T1 image in the MRI image and the directionally encoded color (DEC) image and fiber orientation distribution function (FOD) image generated from the DWI image. Among them, the DEC image shows fibers with different directions of travel through different colors, clearly showing the normal anatomy and its travel of the white matter fibers in the brain; the FOD image is the macroscopic manifestation of the diffusion direction of water molecules, quantifying the proportion of fiber orientations in each direction within a voxel. Mathematically, it is a probability distribution on a sphere because each point on the sphere corresponds to a unique direction;

[0042] Step 2, labeled dataset preparation

[0043] First, draw a rough region MASK containing the path of the trigeminal nerve in the common space and register it to 100 individuals. Then, use the dual-tensor traceless Kalman filter (UKF-2) method to scatter seed points within the MASK for fiber tracking. After that, the tracked fibers are first automatically screened using the region of interest (ROI) and the region of avoidance (ROA), and then layer-by-layer screening at the expert level is carried out. The double screening ensures the accuracy of the results. Finally, the screened fibers are mapped onto the voxels to generate the final more accurate labeled data GroundTruth;

[0044] In the said Step 2, the labeled dataset preparation includes the following steps:

[0045] 2.1 Draw the mask

[0046] Based on anatomical knowledge, an elliptical area MASK was drawn in the common space of the T1 data of 100 subjects. This MASK included the area through which the trigeminal nerve passed, making it possible to track all trigeminal nerve fibers. The MASK drawn in the common space was then registered to the individual as the MASK for individual fiber tracking.

[0047] 2.2 Fiber tracking

[0048] The set number of seed points was selected in the MASK, and then the dual tensor unscented Kalman filter UKF-2T method was used for fiber tracking. This UKF method simultaneously fits a mixed model of two tensors to the dMRI data when tracking fibers, providing highly sensitive fiber tracking capabilities, especially in the presence of crossing fibers, which is very important for tracking the brainstem portion of the trigeminal nerve (including the presumed trigeminal spinal tract and the presumed trigeminal mesencephalic tract).

[0049] 2.3 Automatic screening

[0050] First, the nerve fibers were automatically screened using the Meckel chamber drawn on the b=0 image and the cerebral cistern segment drawn on the direction-coded color map of diffusion tensor imaging (DTI) as the region of interest (ROI) for automatic screening, and the fibers belonging to the trigeminal nerve were preliminarily selected.

[0051] 2.4 Manual screening

[0052] Because the results of the above operations contained a large number of false positive results, in order to further ensure the accuracy of the mask, after automatically screening the fibers using the region of interest, the experts used anatomical knowledge to further manually screen out the trigeminal nerve fiber bundles of all subjects.

[0053] 2.5 Mask Generation

[0054] After the above double screening, a high-quality trigeminal nerve was generated. Then we mapped the 3D streamlines to the voxels to obtain binary nerve fiber masks. These binary masks were analyzed as connected areas, and only the largest connected areas were retained, thereby filtering out single voxels and small groups that were not connected to other voxels. Thus, the final GroundTruth of the trigeminal nerve was obtained.

[0055] Step 3: Data preprocessing

[0056] The data obtained in step 1 is sliced ​​into 128×160×128 size, and the image data is subjected to histogram equalization, grayscale histogram normalization, and image augmentation operations to complete the preparation of training data and test data;

[0057] Step 4, Network Design and Training

[0058] Construct a pyramid deep convolutional network for multimodal fusion, perform multimodal fusion on T1, DEC, and FOD images, and then use the training samples and GroundTruth generated in the above Step 1 and Step 2 to train the constructed network model.

[0059] In the above-mentioned Step 4, the network design and training process is as follows:

[0060] The present invention constructs a multimodal fusion model based on a U-Net pyramid deep convolutional network. Considering that the segmentation map gradually forms during the synthesis process, it may be beneficial to perform modal fusion starting from the synthesis stage after analyzing each individual modality. Therefore, the network constructed by the present invention consists of two independent analysis paths and a shared synthesis path. The first independent analysis path is as follows: Combine the T1 and DEC images as the input, perform convolution operations and then enter the encoder module. The encoder module contains 4 convolutional layers and max-pooling layers, with 32, 64, 128, and 256 feature maps respectively; the decoder module contains 4 deconvolutional layers and convolutional layers, with 256, 128, 64, and 32 feature maps respectively. For all convolutional layers, the size of the convolutional kernel is 3×3×3; for all max-pooling layers, the pool size is 2×2×2 and the stride is 2; for all deconvolutional layers, the deconvolved feature maps are combined with the corresponding features in the encoder module. The other independent analysis path is: Use the FOD image as the input, and after convolution, perform the encoding and decoding process in the same way as the first independent analysis path. The difference is that the encoder-decoder of this path only has three layers. The shared synthesis path is connected by the outputs of the independent analysis paths. When the synthesis progresses from one stage to another, the corresponding stages from each analysis path are also connected to the input. The synthesis path uses a Softmax classifier to generate voxel-level probability maps and predictions. After the network is constructed, use the training samples generated in the previous steps to train the constructed network model, further modify and adjust its parameters, and finally use the network model for experiments on new data.

[0061] To measure the quality of the network model prediction, that is, the degree to which the predicted value of the model is different from the true value, the Dice coefficient is used as the Loss function of the network. The following formula is the definition of the Dice coefficient:

[0062]

[0063] where TP is the number of true positive voxels, FP is the number of false positive voxels, FN is the number of false negative voxels, and p x ∈P:Ω→{0,1} is the predicted binary segmentation content, and gx ∈G:Ω→{0,1} is the binary content of the true value;

[0064] Step Five, prediction and segmentation

[0065] Use the network model trained in Step Four to predict the optic nerve region of the test data, compare the predicted result with the labeled result, and calculate the prediction accuracy.

[0066] The specific implementation described above is only one of the best implementation modes of the present invention, and is not used to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the spirit and principles of the present invention and the content of the drawings shall be included in the patent protection scope of the present invention.

Claims

1. A method for automatic segmentation of the trigeminal nerve based on a deep network, characterized in that: The method includes the following steps: Step 1, Dataset preparation Use the MRI and DWI images of N cases of HCP data, select the T1 image in the MRI image and the direction-encoded color DEC image and fiber orientation distribution function FOD image generated from the DWI image. Among them, the DEC image shows fibers with different directions of travel in different colors, clearly showing the normal anatomy and its travel of the white matter fibers in the brain; the FOD image is a macroscopic manifestation of the diffusion direction of water molecules, quantifying the proportion of fiber orientations in each direction within a voxel. Step 2, Preparation of the labeled dataset First, draw a rough area MASK containing the area where the trigeminal nerve passes in the common space and register it to N individuals. Then, use the dual-tensor unscented Kalman filter UKF-2T method to scatter seed points within the MASK for fiber tracking. After that, first automatically screen the tracked fibers using the region of interest ROI and the region of avoidance ROA, and then perform layer-by-layer screening at the expert level. The double screening ensures the accuracy of the results. Finally, map the screened fibers onto the voxels to generate the final more accurate labeled data GroundTruth. Step 3, Data preprocessing Slice the data obtained in Step 1 into a size of 128×160×128, perform histogram equalization, gray histogram normalization, and image augmentation operations on the image data to complete the preparation of the training data and test data. Step 4, Network design and training Construct a multi-modal fusion pyramid deep convolutional network to perform multi-modal fusion on the T1, DEC, and FOD images, and then use the training samples and GroundTruth generated in the above Step 1 and Step 2 to train the constructed network model. Step 5, Predictive segmentation Use the network model trained in Step 4 to predict the optic nerve region of the test data, compare the predicted result with the marked result, and calculate the prediction accuracy. In the above Step 4, in order to measure the quality of the network model prediction, that is, the degree of difference between the predicted value and the true value of the model, the Dice coefficient is used as the Loss function of the network. The definition of the Dice coefficient is: where TP is the number of true positive voxels, FP is the number of false positive voxels, FN is the number of false negative voxels, and p x ∈ P: Ω → {0, 1} is the predicted binary segmentation content, and g x ∈ G: Ω → {0, 1} is the ground truth binary content.

2. A method for automatic segmentation of the trigeminal nerve as described in claim 1, characterized in that, In the above Step 2, the preparation of the labeled dataset includes the following steps: 2.1 Drawing the MASK According to anatomical knowledge, draw an elliptical-like area MASK in the common space of the T1 data of N subjects. This MASK contains the area where the trigeminal nerve passes, enabling all trigeminal nerve fibers to be tracked. Then register the MASK drawn in the common space to the individual as the MASK for individual fiber tracking. 2.2 Fiber tracking A set number of seed points was selected within the MASK, and then the dual tensor unscented Kalman filter UKF-2T method was used for fiber tracking. The hybrid model of the two tensors was fitted to the dMRI data at the same time, providing highly sensitive fiber tracking capabilities, especially in the presence of crossing fibers, to track the putative trigeminal spinal tract and the putative trigeminal mesencephalic tract; 2.3 Automatic screening First, the nerve fibers were automatically screened, using the Meckel chamber drawn on the b=0 image and the cisternal segment drawn on the direction-coded color map of the diffusion tensor imaging (DTI) as the region of interest (RO) to automatically screen and preliminarily select the fibers belonging to the trigeminal nerve; 2.4 Manual screening Because the results of the above operations contained a large number of false positive results, in order to further ensure the accuracy of the mask, after automatically screening the fibers using the region of interest, the experts used anatomical knowledge to further manually screen out the trigeminal nerve fiber bundles of all subjects; 2.5 Mask Generation After the above double screening, a high-quality trigeminal nerve was generated. Then we mapped the three-dimensional streamlines onto voxels to obtain a binary nerve fiber mask. These binary masks were analyzed as connected regions and only the largest connected regions were kept, thus filtering out single voxels and small groups that were not connected to other voxels to obtain the final trigeminal GroundTruth.

3. A trigeminal nerve automatic segmentation method as claimed in claim 1 or 2, It is characterized in that In step 4, the network design and training process is as follows: A multimodal fusion model based on the U-Net pyramid deep convolutional network is built. Considering that the segmentation map is gradually formed during the synthesis process, it may be beneficial to start modal fusion from the synthesis stage after analyzing each individual modality. The built network consists of two independent analysis paths and a shared synthesis path. The first independent analysis path is as follows: the T1 and DEC images are combined as input, and after the convolution operation, they enter the encoder module. The encoder module contains 4 convolutional layers and maximum pooling layers, which contain 32, 64, 128, and 256 feature maps respectively; The decoder module contains 4 deconvolution layers and convolution layers, which contain 256, 128, 64, and 32 feature maps respectively. For all convolution layers, the size of the convolution kernel is 3×3×3; for all max-pooling layers, the pool size is 2×2×2 and the stride is 2; for all deconvolution layers, the deconvolved feature maps are combined with the corresponding features in the encoder module; another independent analysis path is as follows: taking the FOD image as the input, after convolution, it undergoes the encoding and decoding process in the same way as the first independent analysis path. The difference is that the encoder-decoder of this path has only three layers. The shared synthesis path is formed by connecting the outputs of the independent analysis paths. When the synthesis progresses from one stage to another, the corresponding stages from each analysis path are also connected to the input. The synthesis path uses a Softmax classifier to generate voxel-level probability maps and predictions. After the network is built, the constructed network model is trained using the training samples generated in the previous steps, and its parameters are further modified and adjusted. Finally, the network model is used for experiments with new data.

Citation Information

Patent Citations

  • Cranial nerve automatic imaging method based on deep network learning

    CN111710010A

  • Optic nerve automatic segmentation method based on deep network

    CN112489048A