A fiber classification method based on the longest common subsequence

Through the fiber classification method based on the longest common subsequence, the problem of low fiber classification accuracy and false positive fibers in the prior art is solved, and fiber classification with higher accuracy and speed is achieved, which can better describe the nerve fiber structure.

CN114372973BActive Publication Date: 2025-06-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210025943.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-06-13
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

When using neural fiber classification algorithms, it is difficult to distinguish fibers with similar distances but different shapes, resulting in false positive fiber problems and low accuracy.

Method used

The fiber classification method based on the longest common subsequence is adopted to solve the longest common subsequence between feature description sequences through discretization of centroid fibers, transformation of multi-feature descriptors, and dynamic programming, and combined with deep learning training weights, high-precision classification of fibers is achieved.

Benefits of technology

The problem of false positive fibers in the fiber cluster is effectively avoided, and the accuracy and speed of fiber classification is greatly improved, and the neural fiber structure in three-dimensional space can be described more fully.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_5
    Figure SMS_5
  • Figure SMS_6
    Figure SMS_6
  • Figure SMS_7
    Figure SMS_7
Patent Text Reader

Abstract

A fiber classification method based on the longest common subsequence. First, estimate the fiber direction of the diffusion information of the original data and perform fiber tracking to obtain anatomically conforming brain fiber streamlines. Secondly, extract the centroid fibers after discretization and the multi-feature information of each node of the fibers to be classified, and obtain a multi-feature description sequence describing each fiber. Then, add control parameters ε and δ to constrain whether the corresponding points of the geometric feature description sequence match, and solve the longest common subsequence of the feature sequence based on the control parameters. Finally, by adding a weight parameter θ i Describe the proportion of the fiber structure information described by different fiber description multi-feature sequences; train the weights of the manually labeled fibers through deep learning to obtain a fiber classification model for the fiber classification clusters to be classified. The present invention can complete efficient and accurate fiber classification, and at the same time propose a new measurement standard for fiber similarity, providing an effective tool for identifying brain fibers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to medical image processing, and in particular to a fiber classification method based on the longest common subsequence. Background Art

[0002] Diffusion magnetic resonance imaging based on magnetic resonance imaging technology is currently the only non-invasive nerve fiber imaging method. Many algorithms based on diffusion magnetic resonance imaging have been developed and applied in the fields of neuroscience and psychiatry. In clinical research, the processing of single brain nuclear magnetic resonance data includes fiber direction estimation and fiber tracking. To ensure the integrity of nerve fibers, fiber tracking imaging algorithms usually output millions of fibers, which brings a lot of difficulties to subsequent fiber processing and analysis. At this time, researchers at home and abroad have proposed fiber bundle segmentation methods to classify whole-brain fibers into multiple fiber clusters with clinical anatomical significance, and then perform subsequent processing and analysis. At this time, a large number of three-dimensional streamlines obtained by fiber tracking algorithms will be classified into ordered fiber clusters according to their morphological characteristics, positions in the brain, and functional area connection conditions. Each cluster represents a nerve fiber bundle with corresponding functions in the brain, such as the corticospinal tract, arcuate fasciculus, medial longitudinal fasciculus, and so on.

[0003] Currently, the methods of fiber classification at home and abroad are mainly based on the ROI template of the region of interest, the method of streamline classification, and the method of automatically segmenting the corresponding nerve region by using deep learning. First of all, the method based on the ROI template is very dependent on the accuracy of image registration. At the same time, if there is a neurosurgical disease tumor, the two-dimensional image data of the patient cannot be well registered with the template, resulting in recognition failure. In particular, the ROI template still requires a lot of anatomical knowledge for users. The volume effect that cannot be avoided in medical two-dimensional original images is more difficult for users. The method of automatically segmenting nerve regions based on deep learning depends on a large number of samples for training, which is time-consuming and laborious for annotators. And when a lesion occurs, the correct region cannot be segmented either. The fiber streamline classification method is based on the three-dimensional streamline similarity between fibers for classification. The current methods generally use the minimum average distance between fibers as a metric, but this method cannot distinguish fibers with similar distances but different shapes, which causes a lot of problems in practical applications. For example, when a fiber cluster satisfies the connection of the corresponding functional area, but there are false-positive fibers with obvious structural errors, the traditional method cannot eliminate the false-positive fibers in the cluster. For these problems, the current research has not found a good solution. The essence of the problem is that after discretizing the fibers, some structural information is lost, and the new fiber structure cannot be completely described, resulting in low classification accuracy. Summary of the Invention

[0004] In order to overcome the problem of a large number of false positive fibers within clusters in the existing fiber classification algorithms with low accuracy, the present invention proposes a problem of fiber classification based on the longest common subsequence, effectively avoiding the problem of false positive fibers with similar distances but different shapes within the classified fiber clusters, greatly improving the accuracy and speed of fiber classification, and proposing a new fiber structure descriptor, thereby more completely describing the nerve fiber structure in three-dimensional space.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0006] A method for fiber classification based on the longest common subsequence, the method comprising the following steps:

[0007] Step 1: Solve the centroid fiber in the fiber cluster and discretize it, the process is as follows:

[0008] First, estimate the fiber direction of the diffusion information of the original data and perform fiber tracking to obtain anatomically conforming brain fiber streamlines, manually screen the fiber clusters to be classified, discretize all the fibers in all the clusters, and solve their centroid fibers;

[0009] Step 2: Use a multi-feature descriptor to convert the fiber streamline node sequence into a geometric feature description sequence, the process is as follows:

[0010] Extract the multi-feature information of each node of the discretized centroid fiber and the fibers to be classified, including fiber geometric morphology features and distance features, to obtain a multi-feature description sequence describing each fiber;

[0011] Step 3: Adopt a metric model based on dynamic programming to solve the longest common subsequence between feature description sequences to describe the geometric morphological similarity between fiber streamline models, the process is as follows:

[0012] Add control parameters ε and δ to constrain whether the corresponding points of the geometric feature description sequence match, and solve the longest common subsequence of the feature sequence based on the control parameters;

[0013] Step 4: Assign weights to the multi-feature descriptor for the similarity metric model respectively, and use deep learning to train its weights to obtain a fiber classification model based on the longest subsequence, the process is as follows:

[0014] Add a weight parameter θ i Describe the proportion of the multi-feature sequence of different fiber descriptions in describing the fiber structure information, and train its weights through deep learning on the manually labeled marked fibers to obtain a fiber classification model for the clusters to be classified.

[0015] Furthermore, in the above step 2, the geometric feature description sequence is:

[0016] Discretize each streamline in the fiber cluster into three-dimensional nodes, and obtain the centroid fiber by calculating the arithmetic mean of the three-dimensional coordinates of the corresponding points. Let a represent the centroid fiber and b represent the new fiber, then a i is represented as the i-th three-dimensional node on the streamline of a, and b i is represented as the i-th three-dimensional node on the streamline of b;

[0017] Calculate the spatial curvatures of the three-dimensional nodes on the a fiber and the b fiber respectively as and Then the spatial curvature characteristic description sequences of the a fiber and the b fiber are and

[0018] For the distance feature sequence, first use the maximum Euclidean distance constraint to avoid the problem of the storage order of three-dimensional nodes caused by inconsistent fiber tracking directions. Discretize the fibers into 21 three-dimensional nodes at equal distances respectively. d E (a i , b i ) represents the Euclidean distance between the corresponding points of the a and b fibers, as shown in Equation 1; according to Equation 2, d E (a, b) represents the maximum Euclidean distance between the a and b fibers; assume that a F and b F are respectively described as the flips of the a and b fibers, then d EF = d E (a, b F ) = d E (a F , b). If d E (a, b) > d E (a, b F ), then the fiber tracking directions of a and b are consistent, and the distance feature sequence is D = {d E (a 1 , b 1 ), d E (a 2 , b 2 ), …, d E (a 21 , b 21 )}; if d E (a, b) < d E (a, b F ), then flip the b fiber, and the distance feature sequence is D = {d E (a 1 , b 1 ), d E (a 2 , b 2 ), …, d E (a 21 , b 21)};

[0019]

[0020] d E (a, b) = max i∈21 (d E (a i , b i )) Formula 2

[0021] Brain fibers are responsible for connecting functional areas, and the positions reached by their head and tail ends are particularly important. Let the distance between the head and tail ends of two fibers a and b be d c1 and d c21 , corresponding to the first point and the last point in the maximum Euclidean distance sequence respectively. Then the endpoint distance feature sequence is {d c1 (a 1 , b 1 ), d c21 (a 21 , b 21 )}. At the same time, it is hypothesized that the shortest distances between the fiber bundle to be classified and the three brain functional areas P1, P2, and P3 representing its position are d p1 , d p2 and d p3 , then the cortical distance feature sequence of this fiber is {d p1 (a 1 , p 1 ), d p2 (a 1 , p 2 ), d p3 (a 1 , p 3 )}.

[0022] Furthermore, in step 3 described above, solving the longest common subsequence between feature description sequences based on dynamic programming includes the following steps:

[0023] 3.1: Set controllable parameters ε and δ

[0024] Due to the differences within the fiber cluster, the same type of fibers connecting the corresponding functional areas may have similar but not exactly the same fiber trajectories. Discretely, different nodes have different morphological feature information, not a strictly corresponding relationship of equal point values. Therefore, a controllable parameter ε is added to control the threshold of corresponding point matching, thereby relaxing the identified fiber interval;

[0025] Because the fiber lengths within the same cluster are not the same, after equidistant discretization, the same morphological features may be different in the positions of discrete points; therefore, a controllable parameter δ is added to describe the serial number difference of corresponding matching points;

[0026] 3.2: Solving the Longest Common Subsequence of Feature Sequences Based on Controllable Parameters ε and δ

[0027] Under the constraints of the controllable parameters ε and δ, the above-mentioned fiber curvature feature description sequences A c and B c are introduced into Formula 3 to obtain the solution of the longest common subsequence LCS based on dynamic programming, where n represents the number of discrete three-dimensional nodes of fiber a, and m represents the number of discrete three-dimensional nodes of fiber b;

[0028]

[0029] 3.3: Fiber Morphological Similarity Measurement Based on the Longest Common Subsequence

[0030] Under the constraints of the controllable parameters, the longer the longest common subsequence of the two fiber curvature feature sequences, the more similar the two feature sequences are; in three-dimensional space, the existence of a long longest common subsequence of the curvature feature sequences of two fibers indicates a strong geometric morphological consistency between them; as shown in Formula 4, a similarity index H is designed to describe the similarity measurement based on the longest common subsequence. The purpose of the control parameters ε and δ is to maximize the similarity index H, so as to achieve the goal of minimizing the difference within the cluster, where Ts represents all the fibers in the fiber cluster, and the number of fibers is k; fiber a represents the centroid fiber of the cluster Ts, and H i (a, b) represents the similarity measurement between fiber b and the centroid fiber a in Ts. As can be seen from Formula 4, H i (a, b) ∈ (0, 1), and H i (a, b) is closer to 1, indicating that the similarity between the two fibers is higher. When the two fibers completely overlap in three-dimensional space, that is, when the node positions, the number of nodes, and the morphological features of each point are all consistent, H i (a, b) = 1.

[0031]

[0032] 3.4: Merging Feature Sequences

[0033] There are a total of multiple description sequences in the above process, including the spatial curvature feature description sequence, the distance feature sequence, the end-point distance feature sequence, and the cortical distance feature sequence; this step describes the importance of these feature sequences to the fiber spatial structure by assigning different weights to them and superimposing them;

[0034] Weights θ 1 , θ 2 , θ 3 , and θ 4 are assigned to the spatial curvature feature description sequence, the distance feature sequence, the end-point distance feature sequence, and the cortical distance feature sequence respectively. They are superimposed into the fiber geometric feature description sequence L i, such as Formula 5:

[0035]

[0036] Furthermore, in the fourth step, the design of the classification model based on the longest common subsequence includes the following steps:

[0037] 4.1: Construct the feature space

[0038] Calculate the fiber geometric feature description sequence L of the fiber to be trained respectively i , and merge them into a two-dimensional feature space matrix L s , as shown in Formula 6, and mark whether the trained fiber I belongs to the classified cluster. If it belongs, mark I = 1; if not, mark I = 0;

[0039] L s =(L 1 , L 2 ,…, L i ) Formula 6

[0040] 4.2: Calculate the weight of the feature sequence based on deep learning

[0041] Use the training binary logistic regression classifier, the Stochastic Average Gradient (SAG) solver in Scikit-learn, where the default parameters are used, except that the number of iterations of the solver is increased to 1000 times to ensure convergence, and the parameter reduces the negative impact of residual class imbalance. In all cases, the residual class imbalance is set to 1:3, and finally the weight θ of the feature sequence in the fiber cluster classification model is obtained 1 , θ 2 , θ 3 , and θ 4 .

[0042] The beneficial effects of the present invention are manifested in: effectively avoiding the problem of false positive fibers with similar distances but different shapes in the fiber cluster of the classification result, greatly improving the accuracy and speed of fiber classification. At the same time, controllable parameters are added to adjust the classification threshold, and when a neurosurgical disease tumor occurs, the screening threshold can be appropriately enlarged artificially to capture the fibers of interest around the tumor. Detailed implementation method

[0043] The present invention will be further described below.

[0044] A method for fiber classification based on the longest common subsequence, the method includes the following steps:

[0045] Step 1: Solve the centroid fiber in the fiber cluster and discretize it, and the process is as follows:

[0046] Discretize each streamline in the fiber cluster into three-dimensional nodes, and obtain the centroid fiber by calculating the arithmetic mean of the three-dimensional coordinates of the corresponding points. Let a represent the centroid fiber, then a i represents the i-th three-dimensional node on the streamline of a. Similarly, for the new fiber b;

[0047] Step 2: Use a multi-feature descriptor to convert the fiber streamline node sequence into a geometric feature description sequence. The process is as follows:

[0048] Calculate the spatial curvature of the three-dimensional nodes on the a fiber and the b fiber respectively as and Then the spatial curvature feature description sequences of the a fiber and the b fiber are and

[0049] For the distance feature sequence, first use the maximum Euclidean distance constraint to avoid the problem of the storage order of three-dimensional nodes caused by inconsistent fiber tracking directions. Discretize the fibers into 21 three-dimensional nodes at equal distances respectively. d E (a i ,b i ) represents the Euclidean distance between the corresponding points of the a and b fibers, as shown in Equation 1; According to Equation 2, d E (a,b) represents the maximum Euclidean distance between the a and b fibers. Assume that a F and b F are described as the flips of the a and b fibers respectively, then d EF = d E (a,b F ) = d E (a F ,b). If d E (a,b) > d E (a,b F ), then the fiber tracking directions of a and b are consistent, and the distance feature sequence is D = {d E (a 1 ,b 1 ), d E (a 2 ,b 2 ), …, d E (a 21 ,b 21 )}; If d E (a,b) < d E (a,b F ), then flip the b fiber, and the distance feature sequence is D = {d E (a 1 ,b 1 ), d E (a 2 ,b 2 ), …, dE (a 21 ,b 21 )};

[0050]

[0051] d E (a,b)=max i∈21 (d E (a i ,b i )) Formula 2

[0052] Brain fibers are mainly responsible for the connection function area, and the positions reached by their head and tail ends are particularly important. Let the distance between the head and tail endpoints of two fibers a and b be d c1 and d c21 , corresponding to the first point and the last point in the maximum Euclidean distance sequence respectively. Then the endpoint distance feature sequence is {d c1 (a 1 ,b 1 ),d c2 (a 21 ,b 21 )}. At the same time, it is hypothesized that the shortest distances between the fiber bundle to be classified and the three brain functional areas P1, P2, and P3 that can represent its position are d p1 , d p2 and d p3 , then the cortical distance feature sequence of this fiber is {d p1 (a 1 ,p 1 ),d p2 (a 1 ,p 2 ),d p3 (a 1 ,p 3 )};

[0053] Step 3: Adopt a metric model that uses dynamic programming to solve the longest common subsequence between feature description sequences to describe the geometric morphological similarity between fiber streamline models, including the following steps:

[0054] 3.1 Set controllable parameters ε and δ

[0055] Due to the differences within the fiber cluster, the same type of fibers connecting the corresponding functional areas may have similar but not exactly the same fiber trajectories. Discretely, it is manifested that different nodes have different morphological feature information, not a strictly corresponding relationship with equal point values. Therefore, a controllable parameter ε is added to control the threshold for corresponding point matching, thereby relaxing the identified fiber interval;

[0056] Since the fiber lengths in the same cluster are inconsistent, after equidistant discretization, the same morphological features may show differences in the positions of the discrete points. Therefore, a controllable parameter δ is introduced to describe the serial number difference of the corresponding matching points;

[0057] 3.2 Solving the longest common subsequence of the feature sequences based on the controllable parameters ε and δ

[0058] Under the constraints of the controllable parameters ε and δ, the above-mentioned fiber curvature feature description sequences A c and B c are introduced into Formula 3 to obtain the solution of the longest common subsequence (LCS) based on dynamic programming, where n represents the number of three-dimensional nodes discretized from fiber a, and m represents the number of three-dimensional nodes discretized from fiber b.

[0059]

[0060] 3.3 Fiber morphological similarity measurement based on the longest common subsequence

[0061] Under the constraints of the controllable parameters, the longer the longest common subsequence of the two fiber curvature feature sequences, the more similar the two feature sequences are. In three-dimensional space, the existence of a long longest common subsequence of the curvature feature sequences of two fibers indicates a strong geometric morphological consistency between them. As shown in Formula 4, a similarity index H is designed to describe the similarity measurement based on the longest common subsequence. The purpose of the control parameters ε and δ is to maximize the similarity index H, so as to achieve the goal of minimizing the difference within the cluster. Among them, Ts represents all the fibers in the fiber cluster, and the number of fibers is k; fiber a represents the centroid fiber of the cluster Ts, and H i (a, b) represents the similarity measurement between fiber b and the centroid fiber a in Ts. As can be seen from Formula 4, H i (a, b) ∈ (0, 1), H i (a, b) the closer it is to 1, the higher the similarity between the two fibers. When the two fibers completely overlap in three-dimensional space, that is, when the node positions, the number of nodes, and the morphological features of each point are all consistent, H i (a, b) = 1;

[0062]

[0063] 3.4 Merging feature sequences

[0064] In the above process, there are a total of multiple description sequences, including the spatial curvature feature description sequence, the distance feature sequence, the endpoint distance feature sequence, and the cortical distance feature sequence. This step describes the importance of these feature sequences to the fiber spatial structure in the form of assigning different weights and superimposing them.

[0065] Weights θ are respectively assigned to the spatial curvature feature description sequence, the distance feature sequence, the endpoint distance feature sequence, and the cortical distance feature sequence 1 , θ 2 , θ 3 , and θ 4 . They are superimposed to form the fiber geometry feature description sequence L i , as shown in Equation 5;

[0066]

[0067] Step 4: Weights for the similarity measurement model are respectively assigned to the multi-feature descriptors, and their weights are trained using deep learning to obtain a fiber classification model based on the longest subsequence:

[0068] 4.1 Construct the feature space

[0069] The fiber geometry feature description sequence L of the fiber to be trained is respectively calculated i , and combined into a two-dimensional feature space matrix L s , as shown in Equation 6, and it is marked whether the fiber I to be trained belongs to the classified cluster. If it belongs, it is marked as I = 1; if not, it is marked as I = 0;

[0070] L s =(L 1 , L 2 , …, L i ) Equation 6

[0071] 4.2 Calculate the weights of the feature sequence based on deep learning

[0072] Using the training binary logistic regression classifier and the Stochastic Average Gradient (SAG) solver in Scikit-learn, with default parameters used, except that the number of iterations of the solver is increased to 1000 to ensure convergence, and the parameter to reduce the negative impact of residual class imbalance, and the residual class imbalance is set to 1:3 in all cases. Finally, the weights θ 1 , θ 2 , θ 3 , and θ 4 of the feature sequence in the fiber cluster classification model are obtained.

[0073] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept and is only for illustrative purposes. The protection scope of the present invention should not be regarded as limited to the specific forms stated in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be conceived by those of ordinary skill in the art based on the inventive concept of the present invention.

Claims

1. A fiber classification method based on the longest common subsequence, characterized in that, the method comprises the following steps: Step 1: Solve the centroid fiber in the fiber cluster and discretize it. The process is as follows: First, estimate the fiber direction of the diffusion information of the original data and perform fiber tracking to obtain anatomically conforming brain fiber streamlines. Manually screen the fiber clusters to be classified, discretize all the fibers in all the clusters, and solve their centroid fibers; Step 2: Use a multi-feature descriptor to convert the fiber streamline node sequence into a geometric feature description sequence. The process is as follows: Extract the multi-feature information of each node of the discretized centroid fiber and the fibers to be classified, including fiber geometric morphology features and distance features, to obtain a multi-feature description sequence describing each fiber; Step 3: Adopt a metric model based on dynamic programming to solve the longest common subsequence between feature description sequences and then describe the geometric morphological similarity between fiber streamline models. The process is as follows: Add control parameters ε and δ to constrain whether the corresponding points of the geometric feature description sequences match, and solve the longest common subsequence of the feature sequences based on the control parameters; Step 4: Assign weights to the multi-feature descriptors for the similarity metric model respectively, and use deep learning to train their weights to obtain a fiber classification model based on the longest subsequence. The process is as follows: Add the weight parameter θ i Describe the proportion of multi-feature sequences describing fiber structure information for different fibers, and train the weights of the manually labeled marked fibers through deep learning to obtain a fiber classification model for the quasi-classified clusters; In the above-mentioned Step 3, solving the longest common subsequence between feature description sequences based on dynamic programming includes the following steps: 3.1: Set controllable parameters ε and δ Due to the differences within the fiber cluster, the same type of fibers connecting the corresponding functional areas may have similar but not exactly the same fiber trajectories. Discretely, it is manifested that different nodes have different morphological feature information and do not have a strictly corresponding point numerical equality relationship. Therefore, a controllable parameter ε is added to control the threshold of corresponding point matching, thereby relaxing the identified fiber interval; Because the fiber lengths in the same cluster are not the same, after equidistant discretization, the same morphological features may be different in the positions of the discrete points; so a controllable parameter δ is added to describe the serial number difference of the corresponding matching points; 3.2: Solve the longest common subsequence of the feature sequences based on the controllable parameters ε and δ Under the constraints of controllable parameters ε and δ, the fiber curvature feature description sequence A is introduced into Equation 3 c and B c , and the longest common subsequence LCS solution based on dynamic programming is obtained, where n represents the number of three-dimensional nodes discretized by fiber a, and m represents the number of three-dimensional nodes discretized by fiber b; 3.3: Fiber morphological similarity metric based on the longest common subsequence Under the constraints of controllable parameters, the longer the longest common subsequence of the two fiber curvature feature sequences, the more similar the two feature sequences are; in three-dimensional space, the existence of a long longest common subsequence of the curvature feature sequences of two fibers indicates a strong geometric morphological consistency between the two; as shown in Equation 4, a similarity index H is designed to describe the similarity measure based on the longest common subsequence. The purpose of the control parameters ε and δ is to maximize the similarity index H, thereby achieving the goal of minimizing the intra-cluster difference, where Ts represents all the fibers in the fiber cluster, and the number of fibers is k; fiber a represents the centroid fiber of the cluster Ts, and H i (a, b) represents the similarity measure between fiber b and the centroid fiber a in Ts. As can be seen from Equation 4, H i (a, b) ∈ (0, 1), H i (a, b) The closer it is to 1, the higher the similarity between the two fibers. When the two fibers completely overlap in three-dimensional space, that is, at the node position, when the number of nodes and the morphological features of each point are all consistent, H i (a, b) = 1; 3.4: Merge the feature sequences There are a total of multiple description sequences, including a spatial curvature feature description sequence, a distance feature sequence, an endpoint distance feature sequence, and a cortical distance feature sequence; this step describes the importance of these feature sequences to the fiber spatial structure in the form of assigning different weights to them and superimposing them; Assign weights θ to the spatial curvature feature description sequence, distance feature sequence, end-point distance feature sequence, and cortical distance feature sequence respectively 1 , θ 2 , θ 3 and θ 4 , and superimpose them into the fiber geometry feature description sequence L i , as shown in Equation 5:

2. A fiber classification method based on the longest common subsequence according to claim 1, characterized in that, in the above-mentioned Step 2, the geometric feature description sequence is: Discretize each streamline in the fiber cluster into three-dimensional nodes, and obtain the centroid fiber by calculating the arithmetic mean of the three-dimensional coordinates of the corresponding points. Let a represent the centroid fiber and b represent the new fiber. Then a i is represented as the i-th three-dimensional node on the streamline of a, and b i is represented as the i-th three-dimensional node on the streamline of b; Calculate the spatial curvature of the three-dimensional nodes on the a-fiber and b-fiber respectively as and Then the spatial curvature characteristic description sequences of the a-fiber and b-fiber are and The distance feature sequence first uses the maximum Euclidean distance constraint to avoid the problem of the storage order of 3D nodes caused by inconsistent fiber tracking directions. The fibers are discretized into 21 3D nodes at equal distances, d E (a i ,b i ) represents the Euclidean distance between the corresponding points of two fibers a and b, as shown in Equation 1; according to Equation 2, d E (a,b) represents the maximum Euclidean distance between fibers a and b; assume that a F and b F are described as the flips of fibers a and b respectively, then d EF = d E (a,b F ) = d E (a F ,b); if d E (a,b) > d E (a,b F ), then the fiber tracking directions of a and b are consistent, and the distance feature sequence is D = {d E (a 1 ,b 1 ), d E (a 2 ,b 2 ), …, d E (a 21 ,b 21 )}; if d E (a,b) < d E (a,b F ), then flip fiber b, and the distance feature sequence is D = {d E (a 1 ,b 1 ), d E (a 2 ,b 2 ), …, d E (a 21 ,b 21 )}; d E (a, b) = max i∈21 (d E (a i , b i ) Formula 2 Brain fibers are responsible for connecting functional areas, and the positions reached by their head and tail ends are particularly important. Let the distance between the head and tail ends of two fibers a and b be d c1 and d c21 , corresponding to the first point and the last point in the maximum Euclidean distance sequence respectively. Then the end-point distance feature sequence is {d c1 (a 1 ,b 1 ),d c21 (a 21 ,b 21 )}. At the same time, it is hypothesized that the shortest distances between the fiber bundle to be classified and the three brain functional areas P1, P2, and P3 representing its position are d p1 ,d p2 and d p3 . Then the cortical distance feature sequence of this fiber is {d p1 (a 1 ,p 1 ),d p2 (a 1 ,p 2 ),d p3 (a 1 ,p 3 )}.

3. A fiber classification method based on the longest common subsequence according to claim 1 or 2, characterized in that, in the above-mentioned Step 4, the design of the classification model based on the longest common subsequence includes the following steps: 4.1: Build a feature space Calculate the fiber geometric feature description sequence L of the fibers to be trained respectively i and merge them into a two-dimensional feature space matrix L s as shown in Formula 6, and mark whether the trained fiber I belongs to the classified cluster. If it belongs, mark I = 1; if not, mark I = 0; L s = (L 1 , L 2 , …, L i ) Equation 6 4.2: Calculate the weights of the feature sequences based on deep learning Using the training binary logistic regression classifier, the Stochastic Average Gradient (SAG) solver in Scikit-learn, with default parameters used, except that the number of iterations of the solver is increased to 1000 to ensure convergence, and the parameter to reduce the negative impact of residual class imbalance, with the residual class imbalance set to 1:3 in all cases, finally obtaining the feature sequence weights θ in the fiber classification model of the quasi-classified cluster 1 , θ 2 , θ 3 and θ 4 .

Citation Information

Patent Citations

  • Method for establishing white matter fiber parameterized model based on healthy crowd

    CN104537711A

  • Cranial nerve automatic segmentation method based on large sample data driving

    CN110533664A