Brain target localization method based on dual-mode MRI and recurrent convolutional neural network
By combining dual-mode MRI and recurrent convolutional neural networks, the problems of insufficient target positioning accuracy and efficiency in existing technologies have been solved, and high-precision and efficient target positioning has been achieved, which is suitable for deep brain stimulation treatment.
Patent Information
- Application Number
- CN202411601422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-11-11
AI Technical Summary
Existing deep brain stimulation target localization methods have deficiencies in accuracy and efficiency. In particular, methods based on convolutional neural networks cannot effectively utilize multiple MRI modes and image sequences, resulting in low positioning accuracy and low computational efficiency.
A method based on dual-mode MRI and recurrent convolutional neural networks was adopted. T1 and T2 mode MRI image sequences were processed simultaneously through a single neural network, and recurrent neural networks were combined for target classification and positioning, thereby improving computational efficiency and accuracy.
The accuracy and efficiency of target positioning are improved, approaching or exceeding the accuracy of DRTT analysis, significantly improving computational efficiency, and is suitable for the treatment of essential tremor and other similar diseases.
Smart Images

Figure CN119477867B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of medical image processing, and relates to a brain deep electric stimulation target positioning method, in particular to an improved brain deep electric stimulation target positioning method based on dual-mode magnetic resonance imaging (MRI) and a cyclic convolutional neural network. BACKGROUND
[0002] Essential tremor (ET) is one of the most common movement disorders. Its global incidence is about 0.4%-6%[5]. The first choice for ET is drug treatment, but only about half of the patients can significantly improve the symptoms after using drugs. For ET patients who do not respond to drug treatment, deep brain stimulation (DBS) is another possible treatment. The basic idea of DBS is to stimulate specific locations in the brain using electrodes (see Figure 1 ). These locations are called target points. The selection of target points plays an important role in the treatment effect. The classic target area for DBS treatment of ET is the ventral intermediate nucleus (Vim) of the thalamus, but the clinical effect is not the same, and the tremor improvement rate varies from 48% to 57%[7]. Some scholars recommend using the posterior subthalamic area (PSA) as the DBS target area for treating ET. Its symptom improvement rate can reach 81%[6]. A recent randomized double-blind clinical trial compared the treatment effects of Vim and PSA target points and found that PSA-DBS is at least as effective as Vim-DBS, and may be more effective than Vim-DBS[1]. However, in different studies, the specific stimulation sites in the PSA region vary greatly, the main reason being that the target point is invisible on magnetic resonance imaging (MRI) and is usually in a person-specified range. At present, several teams have proposed different target point calculation methods, but the best target point positions obtained are significantly different, and the clinical symptom improvement rate varies from 55% to 81%[6, 2, 8, 7]. Therefore, the best stimulation target point based on PSA is still unknown. In clinical practice, most neurosurgeons still determine the target point position according to MRI using visual inspection and experience, but the accuracy level varies.
[0003] How to determine the optimal target point position in the PSA region is a frontier problem in the medical field today, and is the basic problem faced by the present application. There are currently several methods to solve this problem.
[0004] 1. Visual inspection method;
[0005] Experienced doctors can determine the approximate position of the target point by visual inspection based on the MRI image.
[0006] 2. DRTT analysis;
[0007] Based on diffusion tensor imaging (DTI) studies, the dentato-rubro-thalamic tract (DRTT) in the PSA is considered to be the anatomical basis for the reduction of tremor by PSA-DBS [3, 4]. Different research teams found that the optimal stimulation target located by anatomical structure positioning of MRI is located near the center or edge of the DRTT fiber [7]. Therefore, a new high-precision target positioning method called DRTT analysis [9] has emerged, which is to first draw the DRTT structure and then determine the target position according to the DRTT.
[0008] 3. Target estimation based on convolutional neural network;
[0009] Another technology based on DRTT analysis has emerged, which is to estimate the target position based on magnetic resonance imaging data using a convolutional neural network. This method uses machine learning to summarize the results of DRTT analysis to obtain the rules of the target position. For details, refer to Chinese patent CN202111386383.7 "Brain deep electric stimulation target positioning method based on magnetic resonance imaging and neural network".
[0010] Disadvantages of the background art:
[0011] 1. Visual method;
[0012] Advantages: Does not require excessive reliance on computer analysis, and the results are interpretable.
[0013] Disadvantages: Relies on the experience and state of the doctor, lacks assurance of accuracy, and cannot be automatically implemented.
[0014] 2. DRTT analysis;
[0015] Advantages: High precision, results tend to be optimal.
[0016] Disadvantages:
[0017] 1) Technical requirements: In addition to MRI, DTI technology is also needed to generate DRTT structure. However, DTI technology is different from general MRI and does not belong to routine medical examination methods;
[0018] 2) Personnel requirements: Generating DRTT structure and determining target position involves several complex data processing steps, which requires professional personnel with medical imaging knowledge to manually operate;
[0019] 3) Time requirements: The entire process is slow, and it takes several hours or even more to process a case.
[0020] 3. Target estimation based on convolutional neural network;
[0021] Advantages:
[0022] 1) Higher positioning accuracy than visual inspection;
[0023] 2) Higher speed than DRTT analysis;
[0024] 3) More convenient to use than DRTT analysis.
[0025] Disadvantages:
[0026] 1) Lower positioning accuracy than DRTT analysis;
[0027] 2) Can only handle one type of magnetic resonance imaging mode (e.g. T1 or T2), cannot take advantage of data from multiple modes;
[0028] 2) The overall structure uses a convolutional neural network, which can only process single images sequentially, cannot directly process image sequences;
[0029] 3) Uses a two-step overall structure, requires the use of two neural networks (classification network and positioning network).
[0030] In summary, the evaluation of target positioning mainly considers accuracy and efficiency. Visual inspection lacks both accuracy and efficiency. DRTT analysis technology has high accuracy but low efficiency, and is still in the academic stage, with considerable distance from widespread clinical application. On the other hand, the method based on convolutional neural network achieves a compromise between accuracy and efficiency, but is still not perfect and has a lot of room for improvement.
[0031] References:
[0032] [1] M. T. Barbe, P. Reker, S. Hamacher, J. Franklin, D. Kraus, T. A. Dembek, J. Becker, J. K. Steffen, N. Allert, J. Wirths, H. S. Dafsari, J. Voges, G. R. Fink, V. Visser-Vandewalle, and L. Timmennann. DBS of the PSA and the VIM in essential tremor: A randomized, double-blind, crossover trial. Neurology, 91(6):e543-e550, 082018.
[0033] [2]P.Blomstedt,U.Sandvik,and S.Tisch.Deep brain stimulation in theposterior subthalamic area in the treatment of essential tremor.Mov Disord,25(10):1350–1356,Jul 2010.
[0034] [3]V.A.Coenen,N.Allert,and B.Madler.A role of diffusion tensorimaging fiber tracking in deep brain stimulation surgery:DBS of the dentato-rubro-thalamic tract(drt)for the treatment of therapy-refractory tremor.ActaNeurochir(Wien),153(8):1579–1585,Aug 2011.
[0035] [4]V.A.Coenen,N.Allert,S.Paus,M.Kronenburger,H.Urbach,andB.Madler.Modulation of the cerebello-thalamo-cortical network in thalamicdeep brain stimulation for tremor:a diffusion tensor imagingstudy.Neurosurgery,75(6):657–669,Dec 2014.
[0036] [5]E D Louis.The roles of age and aging in essential tremor:Anepidemiological perspective.Neuroepidemiology,52(1-2):111–118,2019.
[0037] [6] J. Murata, M. Kitagawa, H. Uesugi, H. Saito, Y. Iwasaki, S. Kikuchi, K. Tashiro, and Y. Sawamura. Electrical stimulation of the posterior subthalamic area for the treatment of intractable proximal tremor. J Neurosurg, 99(4):708-715, Oct 2003.
[0038] [7] A. Nowacki, I. Debove, F. Rossi, J. A. Schlaeppi, K. Petermann, R. Wiest, M. Schupbach, and C. Pollo. Targeting the posterior subthalamic area for essential tremor: proposal for MRI-based anatomical landmarks. J Neurosurg, 131(3):820-827, 102018.
[0039] [8] P. Plaha, S. Javed, D. Agombar, G. O’Farrell, S. Khan, A. Whone, and S. Gill. Bilateral caudal zona incerta nucleus stimulation for essential tremor: outcome and quality of life. J Neurol Neurosurg Psychiatry, 82(8):899-904, Aug 2011.
[0040] [9]Juergen R.Schlaier, Anton L.Beer, Rupert Faltermeier, Claudia Fellner, Kathrin Steib, Max Lange, Mark W. Greenlee, Alexander T. Brawanski, and Judith M. Anthofer. Probabilistic vs. deterministic fiber tracking and the influence of different seed regions to delineate cerebellar-thalamic fibers in deep brain stimulation. European Journal of Neuroscience, 45(12):1623–1633, 2017. Summary of the Invention
[0041] To address the shortcomings of the existing technology, the present invention provides a brain target localization method based on dual-modal MRI and recurrent convolutional neural networks. The goal of this invention is to design a machine learning algorithm based on neural networks and deep learning to estimate the optimal target location for deep brain stimulation (DBS) therapy based on two-modal (T1 and T2) magnetic resonance imaging (MRI) data, surpassing the accuracy of existing technologies (primarily based on convolutional neural networks). Figure 2 Demonstration diagram for target positioning.
[0042] The method of the present invention can process magnetic resonance imaging data from either a single mode or both T1 and T2 modes simultaneously. Its overall structure employs a recurrent convolutional neural network, enabling direct processing of image sequences. It uses a single neural network to combine the classification and localization steps, further improving computational efficiency. It utilizes a high-resolution network structure, enhancing computational accuracy. Furthermore, it employs more sophisticated data preprocessing techniques, increasing the robustness of the method. Through these innovations, the method of the present invention further improves accuracy, efficiency, and flexibility over previously proposed methods based on convolutional neural networks.
[0043] Complete technical solution:
[0044] The method of the present application estimates the optimal DBS target position according to the MRI data of the patient's head. The data of a single MRI scan consists of many image layers, but the target point can only exist in a few layers. In addition, the scan of a conventional MRI can usually be divided into two modes, T1 and T2, which show different emphases. The basic idea of the present application is to use the MRI image sequence of the two modes as input, judge whether each image layer can have a target point, and then select the target point estimate on the image layer that is most likely to contain the target point. This dual-mode processing method can further improve the accuracy and is very different from the single-mode processing method in the background art. This idea can be specifically expanded into the flowchart in Figure 3 .
[0045] The core of the present application is to use a recurrent neural network for image sequences, called "classification and positioning network". The input of this network is two (one when running in single mode) MRI image sequences of length N: one is T1 mode and the other is T2 mode, corresponding to two scans of the same patient, and it is assumed that the brain positions corresponding to any i-th layer of mode 1 and mode 2 are consistent. The role of the classification and positioning network is to judge whether each image layer in the input image sequence contains a target point, and at the same time estimate the optimal target position on each image layer in the input image sequence. Its output is an array sequence of length N, where each element is an array containing two parts (A, B): the first part (i.e. A) is called the classification score, which can be understood as the probability that the corresponding image layer contains a target point, and the larger the value, the more likely it is to contain a target point; the second part (i.e. B) is a two-dimensional coordinate pair, corresponding to the two-dimensional target coordinates of the left and right brain respectively. The two parts are represented separately, and the output sequence can be split into a score sequence and a target coordinate pair sequence. The index of the image layer that is most likely to contain the target point is combined with the two-dimensional coordinate output for that layer, which uniquely determines the three-dimensional coordinates of the target point. The use of a neural network to simultaneously complete the classification and positioning tasks improves the implementation efficiency, and is very different from the use of two neural networks in the background art.
[0046] The basic process of the present application is as follows:
[0047] Step 1: data preprocessing;
[0048] Step 2: image sequence classification and target positioning;
[0049] For the input single-patient dual-mode MRI image layer sequence T1 and T2, use the classification and positioning network to simultaneously calculate the likelihood of each image layer containing a target point and the target point coordinate estimate for each image layer; save the output of the classification and positioning network and split it into a score sequence and a target coordinate pair sequence.
[0050] Step 3: Select the top K layers that are most likely to contain the target point and output the corresponding target point coordinates.
[0051] Step 4: Calculate the final score for each image layer based on the score sequence, set the score below T to 0, select the top K layers with the highest final score in the image layer, and output the corresponding target point coordinates. If no layer is selected, the algorithm is aborted and manual analysis is performed.
[0052] Step 4: Verification and output;
[0053] According to the pre-set rules, verify whether the target point is standard. If all target points are not standard, the algorithm is aborted and manual analysis is performed. Output the target point coordinates and corresponding layer index that pass the verification;
[0054] Further, the specific method of step 1 is as follows:
[0055] The preprocessing of MRI images includes the following steps:
[0056] a) Image Resampling: Resample all images to a uniform voxel size.
[0057] b) Denoising: Remove noise from MRI data.
[0058] c) Bias Field Correction: Correct the uneven distribution of image brightness.
[0059] d) Image Registration: Register the image to the standard space (such as MNI space) for comparison with other images.
[0060] e) Normalization: Normalize the pixel intensity of the image to the same scale. Divide all pixel values by x% (e.g. 95%) of the maximum pixel value, so that all pixel values are limited between 0 and 1 (values exceeding 1 are taken as 1).
[0061] f) Skull Stripping: Remove non-brain tissue from the image.
[0062] Since the size range of the target point is relatively small (several millimeters) and only exists in a few layers of MRI data, most of the image layers of the head MRI are actually irrelevant to the target point positioning. Therefore, the layer indexes where the target point exists are counted and a range is artificially set to ignore image layers that exceed this range.
[0063] Crop all retained MRI image layers to the set standard size.
[0064] Further, the classification and localization network structure of step 2 is specifically as follows:
[0065] The classification and localization network generally adopts a "circular convolution" neural network structure. The network is composed of a circular neural network part and a dual-mode convolutional neural network part. Among them, the dual-mode convolutional neural network part is responsible for processing the dual-mode image sequence input, for feature extraction and outputting a feature sequence; the circular neural network part is responsible for comprehensively processing the output sequence of the dual-mode convolutional neural network and calculating the target position. The dual-mode convolutional neural network part includes two branches of the same structure (corresponding to T1 and T2 modes respectively) and a fusion block (concatenation). A single branch is based on a known high-resolution network (HRNet) structure, mainly composed of a convolutional layer and two high-resolution blocks (HR blocks). Each high-resolution block is based on a known residual network (ResNet) structure, composed of several residual blocks, which realize multi-scale information exchange (repeated n times) through the residual blocks; for the output data of the two branches, a fusion block (concatenation) is used for fusion to obtain a dual-mode feature vector. The circular neural network is mainly composed of multiple fully connected layers.
[0066] The last layer of the classification and localization network is a fully connected layer, which is used to process the output of the circular neural network. The output result is a 5-dimensional array, which can be split into a 1-dimensional array and a 4-dimensional array. Among them:
[0067] 1) The 1-dimensional array is defined as the classification score. Due to the action of the sigmoid function, the classification score is between 0 and 1, which is the probability that the input image contains a target point. According to the score, different MRI image layers can be sorted.
[0068] 2) The 4-dimensional array can be converted into a 2x2 form, that is, two two-dimensional coordinates, corresponding to the positions of the left and right target points. Due to the action of the sigmoid function, the coordinate value is between 0 and 1, which is the coordinate after normalization according to the image length and width. Assuming that the normalized coordinate is (x, y), the resolution of the image is HxW, and the length and width of the pixel are p and q respectively, then the actual coordinate is (xHp, yWq).
[0069] Further, supervised learning is used to train the classification and localization network, which requires a certain number of MRI images with known target point positions for training and verification. A stochastic gradient descent algorithm based on mini batch is used to optimize the network parameters using back propagation. The known target point positions of the MRI images are the optimal target points obtained by DRTT analysis technology and manually labeled;
[0070] Further, the DRTT analysis technique needs to use three kinds of magnetic resonance data, including T1-MRI, T2-MRI and diffusion tensor imaging (DTI) data.
[0071] T1-MRI and T2-MRI preprocessing procedure: the default MRI image original layer thickness is 0.7-1.0mm. The MRI is corrected according to the AC-PC line. The midpoint of the AC-PC line is set as the coordinate axis origin, and the re-layering is carried out.
[0072] Diffusion tensor imaging preprocessing procedure: first, deterministic fiber tracking is carried out. The generalized q sampling imaging method is used to reconstruct the magnetic resonance diffusion tensor imaging data to obtain the tensor in each voxel. The MRI and diffusion tensor imaging data are registered by rigid transformation. Three regions of interest are defined on the MRI: 1) the ipsilateral dentate nucleus manually segmented; 2) the ipsilateral PSA region manually segmented; 3) the ipsilateral precentral gyrus automatically extracted using the FreeSurfer DKT atlas. The seed point is set as the ipsilateral dentate nucleus. The minimum and maximum fiber length of the deterministic fiber tracking is set as 20 and 200mm respectively. The angle threshold is set as 55, the step length is set as 0.5, and the upper limit of seed tracking is set as 50000. The reconstructed DRT fiber is manually adjusted to remove the anatomically unreasonable fiber bundle.
[0073] Optimal target labeling: the MRI and the reconstructed fiber bundle are fused and exported as a new MRI_label image. The MRI_label image is opened in the FSLview software, the maximum diameter layer of the red nucleus and the upper and lower 5 layers (a total of 11 layers) are identified, the midpoint of the bilateral DRT fiber is identified, the coordinates are recorded, saved and exported.
[0074] The beneficial effects brought by the technical scheme of the present application are as follows:
[0075] The existing target positioning methods mainly have three kinds, namely, a traditional method (a clinician analyzes an MRI image by naked eyes and according to experience), a DRTT analysis method (fiber tractography is performed by using DTI), and a method of positioning a target by using a convolutional neural network. Compared with the background art, the present application can achieve a better balance between precision and efficiency. In terms of precision, the target positioning precision of the present application can exceed the traditional method, and is further closer to the target positioning precision of the DRTT analysis compared with the new method based on the convolutional neural network. In terms of efficiency, the analysis speed of the present application scheme is much higher than the traditional method, and is superior to or close to the latest convolutional neural network method; since the analysis of the entire magnetic resonance image sequence can be completed at one time, the convenience is greatly improved, and it is the most efficient method at present. The present application is mainly aimed at primary tremor (ET) disease, but is also applicable to other similar diseases, such as Parkinson's disease. BRIEF DESCRIPTION OF DRAWINGS
[0076] Figure 1 is a schematic diagram of DBS;
[0077] Figure 2 is a target positioning demonstration diagram, the white mark indicated by the arrow is a real target point, and the red mark is an estimated target point. The numbers are distance errors. The classification score (confidence) of the MRI image layer is displayed in the lower right corner;
[0078] Figure 3 is a flowchart of the technical scheme of the embodiment of the present application;
[0079] Figure 4 is a schematic diagram of MRI data representation;
[0080] Figure 5 is an example of the recurrent convolutional neural network structure of the embodiment of the present application;
[0081] Figure 6 is an example of the dual-mode convolutional neural network structure of the embodiment of the present application;
[0082] Figure 7 is an example of the recurrent neural network structure of the embodiment of the present application;
[0083] Figure 8 is a schematic diagram of the high-resolution block and residual block structure of the embodiment of the present application. DETAILED DESCRIPTION
[0084] The scheme details of the technical scheme of the present application which are not mentioned in the embodiments are described and supplemented in detail in combination with the drawings and the embodiments.
[0085] Representation of magnetic resonance imaging (MRI) data and target position
[0086] The basic working mode of MRI is to scan the human body in successive sections during its movement. Therefore, a complete MRI image is composed of a stack of gray-scale pictures. The interval between two adjacent layers depends on the device and is fixed and known, for example, 0.7 mm. A human head MRI usually contains hundreds of image layers. They can be regarded as a sequence of two-dimensional arrays or as a three-dimensional array. Relatively, the position of the target point can be represented as a combination of a two-dimensional coordinate and a layer index, or as a three-dimensional coordinate. The scheme of the present invention represents the position of the target point in the former way, that is, to estimate the two-dimensional coordinate of the target point in the layer under the premise of selecting the layer. Figure 4 For the schematic diagram of MRI data representation, the position of the target point can be determined by the coordinates (x, y) and the layer index z.
[0087] The brain target positioning method based on dual-mode MRI and cyclic convolutional neural network includes the following steps:
[0088] Step 1: Image preprocessing;
[0089] The preprocessing of MRI images includes the following steps:
[0090] a) Image resampling: resample all images to a uniform voxel size, for example, 0.7 mm.
[0091] b) Denoising: use methods such as Gaussian smoothing and non-local mean filtering (NLM) to remove noise in MRI data.
[0092] c) Bias field correction: use methods such as N4ITK and FASTMRI to correct the uneven distribution of image brightness.
[0093] d) Image registration: use algorithms such as linear registration and nonlinear registration to register the image to the standard space (such as MNI space) for comparison with other images.
[0094] e) Normalization: normalize the pixel intensity of the image so that it is on the same scale. Divide all pixel values by x% (for example, 95%) of the maximum pixel value, so that all pixel values are limited between 0 and 1 (values exceeding 1 are taken as 1). On this basis, in order to match the input end of the neural network, the data can be further standardized, for example, to make the mean close to zero and the variance close to one.
[0095] f) Skull Stripping: Tools such as BET (FSL) and 3dSkullStrip (AFNI) are used to remove skull, skin, and other non-brain tissues from the images.
[0096] These preprocessing steps can help improve the accuracy and stability of subsequent image analysis, making the features extracted from the MRI data more reliable.
[0097] Since the target size range is relatively small (several millimeters) and only exists in a few layers of MRI data, most of the image layers of the head MRI are actually irrelevant to the target positioning. Therefore, the layer index where the target exists can be counted, and a range can be set manually, and the image layers that exceed this range can be ignored. In this way, unnecessary calculations can be reduced.
[0098] On the other hand, an MRI image can contain a lot of useless black blank areas. In order to reduce data redundancy and match subsequent machine learning, a standard size needs to be specified. All MRI image layers need to be cropped to the standard size. For example, the standard image layer size used in an embodiment of the present application is about 220x280 (pixel spacing is 0.7 to 1 millimeter). A relatively simple way of cropping is fixed-position cropping. Since the patient's head is not necessarily in the middle of the image, this method may cut off a small part of the edge of the useful image (i.e. the non-blank part of the image). Another possible but more complex way of cropping is to first center the non-blank part of the image and then crop it.
[0099] Step 2: Neural network structure;
[0100] The classification and positioning network in the present application scheme adopts a "circular convolution" neural network structure. This type of network is mainly composed of a recurrent neural network part and a dual-mode convolutional neural network part. Figure 5 is a general schematic diagram of this type of network. Among them, the dual-mode convolutional neural network part is responsible for processing the dual-mode image sequence input, for feature extraction and outputting the feature sequence; the recurrent neural network part is responsible for comprehensively processing the output sequence of the dual-mode convolutional neural network and calculating the target position. Figure 6 is a specific example of a convolutional neural network part. The dual-mode convolutional neural network part includes two branches of the same structure (corresponding to T1 and T2 modes, respectively) and a fusion block (concatenation). A single branch is based on the known high-resolution network (HRNet) structure, mainly composed of a convolutional layer and two high-resolution blocks (HR block). Each high-resolution block is based on the known residual network (ResNet) structure, composed of several residual blocks (residual block), whose structure is determined by Figure 8(left) to achieve multi-scale information exchange (repeat n times); for the output data of the two branches, a fusion block (concatenation) is used to obtain a dual-mode feature vector. The residual block is mainly composed of convolutional layers, and its specific structure is determined by Figure 8 (right). Figure 7 is a specific example of a recurrent neural network, mainly composed of multiple fully connected layers.
[0101] In addition to convolutional layers and fully connected layers, some other common neural network operations are also used in the example diagram, including max pooling, ReLU, sigmoid, batch normalization, and concatenation. Their descriptions and some basic parameter settings are listed in Table 1. Since these are standard operations in the field of neural networks and deep learning, their specific mathematical definitions are described in related literature (e.g., 《Dive into Deep Learning》https: / / d2l.ai), and will not be described in detail in this document.
[0102] In practice, the classification and positioning network can also use different backbone structures. For example, a relatively simple way is to change the network structure by changing the number of high-resolution blocks, residual blocks, or fully connected layers in the network.
[0103] Table 1 describes the components of the neural network in the example.
[0104]
[0105] The last layer of the classification and positioning network is a fully connected layer, which is used to process the output of the recurrent neural network. The output is a 5-dimensional array, which can be split into a 1-dimensional array and a 4-dimensional array. Among them:
[0106] 1) The 1-dimensional array is defined as the classification score. Due to the effect of the sigmoid function, the classification score is between 0 and 1, which can be understood as the probability of the input image containing the target point. According to the score, different MRI image layers can be sorted.
[0107] 2) The 4-dimensional array can be converted into a 2x2 form, i.e., two two-dimensional coordinates, corresponding to the positions of the left and right target points. Due to the effect of the sigmoid function, the coordinate value is between 0 and 1, which can be understood as the normalized coordinate after the image length and width. Assuming that the normalized coordinate is (x, y), the resolution of the image is HxW, and the length and width of the pixel are p and q, respectively, then the actual coordinate is (xHp, yWq).
[0108] Step 3: Training of the classification and positioning network;
[0109] The neural network algorithm needs to be trained before use. The purpose of training is to determine the parameters of the network. The present scheme adopts supervised learning, and a certain number of MRI images with known target positions are needed for training and verification. Generally, the more training data, the better the verification effect. The embodiment in the present scheme uses about 150 sets of head MRI scan images in the experiment, of which 4 / 5 are randomly taken for training, and the rest are used for verification, and good results are obtained. The present scheme adopts a stochastic gradient descent algorithm based on a small batch, for example, 5 to 10 sets of training data are used each time, and the network parameters are optimized by back-propagation. These methods are standard operations in neural network calculation, and their specific mathematical definitions are explained in related literature, which will not be described in detail in this document. The known target position data of the MRI image is the optimal target point obtained by using the DRTT analysis technology and manually labeled;
[0110] Step 4: DRTT analysis;
[0111] One of the focuses of the present scheme is to obtain the optimal target point by using the DRTT analysis technology and manually label it. This technology needs to use three kinds of magnetic resonance data, including T1-MRI, T2-MRI and diffusion tensor imaging (DTI) data. In the case of missing one kind of MRI mode data, another kind of MRI mode can still be used to complete the analysis. The specific process is as follows.
[0112] T1-MRI and T2-MRI preprocessing process: The default original layer thickness of the MRI image is 0.7-1.0 mm. The MRI is corrected according to the AC-PC line by using special software (such as SPM 12.0 based on MATLAB). The midpoint of the AC-PC line is set as the coordinate axis origin, and the re-layering is performed.
[0113] Diffusion tensor imaging preprocessing procedure: deterministic fiber tracking was performed using a dedicated software (e.g. DSI studio). The diffusion tensor imaging data was reconstructed using generalized q-sampling imaging method to obtain the tensor in each voxel. The MRI and diffusion tensor imaging data were registered using rigid transformation. Three regions of interest were defined on the MRI: 1) the ipsilateral dentate nucleus which was manually segmented; 2) the ipsilateral PSA region which was manually segmented; 3) the ipsilateral precentral gyrus which was automatically extracted using FreeSurfer DKT atlas. The seed point was set as the ipsilateral dentate nucleus. The minimum and maximum fiber length for deterministic fiber tracking were set as 20 and 200 mm, respectively. The angle threshold was set as 55 and the step length was set as 0.5. The upper limit of seed tracking was set as 50000. The reconstructed DRT fibers were manually adjusted to remove anatomically unreasonable fiber bundles.
[0114] Optimal target labeling: the MRI and reconstructed fibers were fused and exported as a new MRI_label image. The MRI_label image was opened in FSLview software. The midplane of the maximum diameter of the red nucleus and 5 planes above and below it (11 planes in total) were identified to locate the midpoints of the bilateral DRT fibers. The coordinates were recorded, saved and exported.
[0115] Step 5: target localization by trained classification and localization network
[0116] For the inputted single patient's dual-mode MRI image sequence T1 and T2, the classification and localization network was used to calculate the possibility of each image layer containing the target point, as well as the target point coordinate estimation of each image layer; the output of the classification and localization network was saved, and different MRI image layers were sorted according to the classification score to obtain a score sequence; the actual coordinates of the target point were determined according to the obtained left and right target point positions to obtain a target point coordinate pair sequence.
[0117] According to the score sequence, the final score of each image layer was calculated, and the score below T was set to 0; the K image layers with the maximum final score were selected in the image layer, and the corresponding target point coordinates were output; K is a parameter (integer), which is usually less than or equal to the maximum number of layers that may contain the target point, for example, 11. If no layer is selected, the algorithm is aborted and manual analysis is performed.
[0118] Step 6: verification and output
[0119] According to the pre-set rules, the target points were verified for standardization; if all the target points were not standardized, the algorithm was aborted and manual analysis was performed; the verified target point coordinates and corresponding layer index were output.
[0120] Predefined rules: In order to reduce the generation of unreasonable target positioning results, some simple target verification rules need to be defined. These rules represent the impossible cases of the target. For example, the target of the left brain cannot appear in the right brain. The target verification rules are determined according to the corresponding specifications in the medical field.
[0121] In addition to the basic flow, some small probability events may also occur. For example, the classification score of the selected layer is too low (i.e. less than T). T is a normalized threshold parameter (floating point number), usually between 0 and 1. Another case is that the target position derived by the positioning network does not conform to common sense (i.e. some basic assumptions) and cannot pass the post-validation. When these small probability events occur, the results obtained by the algorithm will be considered invalid; if all the results are invalid, manual analysis is required.
[0122] The method of the present application can be used for fully automatic target positioning, or can be used for semi-automatic target positioning in cooperation with doctors. There are three specific use methods as follows:
[0123] 1. The appropriate layer is selected by the classification network, and the doctor performs target positioning on the selected layer.
[0124] 2. The appropriate layer is selected by the doctor, and the positioning network performs target positioning on the selected layer.
[0125] 3. The appropriate layer is selected by the classification network, and the positioning network performs target positioning on the selected layer.
[0126] It can be seen that the first two methods are part of the last method, and the last method is the complete scheme described in the foregoing. In addition, the method of the present application can be used for dual-mode data or single-mode data, and there are three specific methods as follows:
[0127] 1. The input data is T1 mode (T2 mode input is empty).
[0128] 2. The input data is T2 mode (T1 mode input is empty).
[0129] 3. The input data is T1, T2 dual mode.
[0130] The terms involved in the above analysis process are described in the relevant neurology and medical imaging literature, and will not be described in detail in this document.
[0131] The above description is further detailed in connection with specific / preferred embodiments of the present application, and should not be construed as limiting the specific implementation of the present application to these descriptions. For those skilled in the art to which the present application belongs, without departing from the concept of the present application, a number of substitutions or modifications can be made to the described embodiments, and these substitutions or modifications should be considered to fall within the protection scope of the present application.
[0132] The part of the present application not described in detail belongs to the technology known to those skilled in the art.
Claims
1. A brain target localization method based on dual-mode MRI and recurrent convolutional neural network, characterized in that: The steps are as follows: Step 1: Data preprocessing; Step 2: Image sequence classification and target location; For the input dual-mode MRI image layer sequences T1 and T2 of a single patient, a classification and localization network is used to simultaneously calculate the probability that each image layer contains the target and the target coordinate estimation of each image layer; the output of the classification and localization network is saved and split into a sequence of component values and a sequence of target coordinate pairs; The classification and localization network generally adopts a "recurrent convolutional" neural network structure; the network consists of two parts: a recurrent neural network part and a dual-mode convolutional neural network. The dual-mode convolutional neural network part is responsible for processing the dual-mode image sequence input for feature extraction and outputting a feature sequence; the recurrent neural network part is responsible for comprehensively processing the output sequence of the dual-mode convolutional neural network and calculating the target position. The dual-mode convolutional neural network part includes two branches with the same structure and a fusion block. A single branch is based on a known high-resolution network structure, mainly consisting of a convolution layer and two high-resolution blocks. Each high-resolution block is based on a known residual network structure and consists of several residual blocks. Multi-scale information exchange is achieved through the residual blocks. The output data of the two branches are fused through a fusion block to obtain a dual-mode feature vector. Step 3: Select the top K layers that are most likely to contain the target and output the corresponding target coordinates; Calculate the final score of each image layer based on the score sequence, and set the scores below T to 0; select the K layers with the largest final scores from the image layers and output the corresponding target coordinates; if no layer is selected, terminate the algorithm and switch to manual analysis; Step 4: Verification and output; Verify that the targets are standardized according to pre-set rules; if all targets are not standardized, terminate the algorithm and switch to manual analysis; output the verified target coordinates and corresponding layer index.
2. The brain target localization method based on dual-mode MRI and recurrent convolutional neural network according to claim 1, characterized in that: Step 1: MRI image preprocessing includes the following steps: a) Image resampling: resample all images to a uniform voxel size; b) Denoising: removing noise from MRI data; c) Bias field correction: corrects the uneven distribution of image brightness; d) Image registration: registering an image to a standard space for comparison with other images; e) Normalization: Normalize the pixel intensity of the image to make it on the same scale; Divide all pixel values by x% of the maximum pixel value, so that all pixel values are limited to between 0 and 1, and values exceeding 1 are set to 1; f) Remove non-brain tissue: remove non-brain tissue from the image; The layer indexes with target points are counted and a range is set artificially. Image layers beyond this range are ignored. All retained MRI image layers are cropped to the set standard size.
3. The brain target localization method based on dual-mode MRI and recurrent convolutional neural network according to claim 1, characterized in that: The recurrent neural network is composed of multiple fully connected layers; The last layer of the classification and positioning network is a fully connected layer, which is used to process the output of the recurrent neural network. The output result is a 5-dimensional array, which can be split into a 1-dimensional array and a 4-dimensional array; wherein: 1) A 1-dimensional array is defined as the classification score; due to the sigmoid function, the classification score ranges from 0 to 1 and is the probability that the input image contains the target. Different MRI image layers are sorted according to the score; 2) The 4-dimensional array can be converted into a 2x2-dimensional form, that is, two 2-dimensional coordinates, corresponding to the positions of the left and right targets respectively. Due to the action of the sigmoid function, the coordinate values are between 0 and 1, which are the coordinates after normalization based on the length and width of the image. Assuming the normalized coordinates are (x, y), the image resolution is H×W, and the length and width of the pixel are p and q respectively, the actual coordinates are (xHp, yWq).
4. The brain target localization method based on dual-mode MRI and recurrent convolutional neural network according to claim 3, characterized in that: Supervised learning is used to train the classification and localization network, which requires a certain number of MRI images with known target locations for training and verification. A mini-batch-based stochastic gradient descent algorithm is used to optimize the network parameters using back propagation. The known target locations in the MRI images are the optimal targets obtained and manually annotated using DRTT analysis technology.
5. The brain target localization method based on dual-mode MRI and recurrent convolutional neural network according to claim 4, characterized in that: The DRTT analysis technique requires three types of magnetic resonance imaging data, including T1-MRI, T2-MRI, and diffusion tensor imaging data; the specific process is as follows: T1-MRI and T2-MRI preprocessing process: The default MRI image original layer thickness is 0.7-1.0mm; the MRI is corrected according to the AC-PC line; the midpoint of the AC-PC line is set as the coordinate axis origin and re-slicing is performed; Diffusion tensor imaging preprocessing process: first, deterministic fiber tracking is performed; Generalized q-sampling was used to reconstruct MRI diffusion tensor imaging data to obtain the tensor within each voxel. Rigid transformation was used to register the MRI and diffusion tensor imaging data. Three regions of interest were defined on the MRI: 1) the manually segmented ipsilateral dentate nucleus; 2) the manually segmented ipsilateral PSA region; and 3) the ipsilateral precentral gyrus automatically extracted using the FreeSurfer DKT atlas. The seed point was set to the ipsilateral dentate nucleus. The minimum and maximum fiber lengths for deterministic fiber tracking were set to 20 and 200 mm, respectively. The angle threshold was set to 55, the step size was set to 0.5, and the seed tracking upper limit was set to 50,000. The reconstructed DRT fibers were manually adjusted to remove anatomically unreasonable fiber bundles. Optimal target annotation: MRI and reconstructed fiber bundles were fused and exported as a new MRI_label image. The MRI_label image was opened in FSLview software, and the midpoints of the bilateral DRT fibers were identified at the level with the largest red nucleus diameter and five levels above and below. The coordinates were recorded, saved, and exported.
Citation Information
Patent Citations
Deep brain stimulation target positioning method based on magnetic resonance imaging and neural network
CN114255209A
Personalized target selection method for non-invasive neuromodulation technology
WO2024193324A1