An Estimation Model and Method for Tissue Microstructure Based on Undersampled DMRI Data
Through the efficient diffusion wave vector space learning module and the three-dimensional physical coordinate space learning module, the problems of insufficient utilization of DMRI data and low learning efficiency of the HGT model are solved, and more efficient tissue microstructure imaging is achieved.
Patent Information
- Application Number
- CN202310723232.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-06-19
AI Technical Summary
The existing HGT model does not fully utilize DMRI data in tissue microstructure estimation, ignores three-dimensional spatial information, and the learning efficiency of diffuse wave vector space is not high.
The efficient diffused wave vector space learning module and the three-dimensional physical coordinate space learning module are adopted to use the SGC network and the Transformer model for feature learning to improve the utilization and learning efficiency of three-dimensional space information.
The accuracy and efficiency of tissue microstructure imaging are improved, and the three-dimensional spatial information of DMRI data is fully utilized.
Smart Images

Figure CN116824155B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer image processing, and relates to a 3D HGT tissue microstructure estimation model and method based on undersampled DMRI data. Background Art
[0002] With the continuous development of artificial intelligence technology, deep learning is increasingly used in medical image processing. Using intelligent image processing technology can achieve quantitative analysis of human organs, lesions, etc., and assist doctors in disease diagnosis. Diffusion tissue microstructure imaging has developed rapidly in recent years, and a series of imaging models have achieved significant results. However, most of these models require densely sampled Diffusion MRI (DMRI) data to complete calculations, so they are not satisfactory in clinical applications. In terms of tissue microstructure estimation using undersampled DMRI data, the Hybrid Graph Transformer (HGT) has been proposed and has become a representative method. It significantly improves the accuracy of tissue microstructure estimation by performing graph learning in the diffusion wave vector space (q-space) of undersampled DMRI data and Transformer-based spatial domain learning in the physical structure space (x-space).
[0003] However, the HGT model relies on training using two-dimensional slices, so there are the following defects: (1) The model does not make full use of DMRI data and ignores three-dimensional spatial information; (2) The learning of the model in the diffusion wave vector space is not efficient enough. Summary of the Invention
[0004] The technical problem to be solved by the present invention is:
[0005] In order to avoid the deficiencies of the prior art, the present invention provides a 3D HGT tissue microstructure estimation model and method based on undersampled DMRI data.
[0006] In order to solve the above technical problem, the technical solution adopted by the present invention is:
[0007] A tissue microstructure estimation model based on undersampled DMRI data, characterized by comprising an efficient diffusion wave vector space learning module and a three-dimensional physical coordinate space learning module. After the data is input, it first undergoes feature learning through the efficient diffusion wave vector space learning module, and then the three-dimensional spatial information of the data is learned through the three-dimensional physical structure space learning module.
[0008] A further technical solution of the present invention: The high-efficiency diffusion wave vector space learning module is specifically as follows: Based on the SGC network, first compare the angle between the sampling points in the diffusion wave vector space with a preset threshold. If the angle is less than the preset threshold, the geometric structure feature element is 1, otherwise it is 0, and a binary correlation matrix representing the graph network is generated; then substitute it into the SGC network simplification formula for calculation to achieve high-efficiency diffusion wave vector space learning.
[0009] A further technical solution of the present invention: The three-dimensional physical coordinate space learning module is specifically as follows: Based on the Transformer model, a U-shaped network composed of three parts: Encoder, Decoder, and Bottleneck constitutes a cascaded self-attention Transformer block in the local volume; the structures of each part are as follows:
[0010] The Encoder is composed of two Local Volume-based Multi-head Self-Attention units; before the data enters the network through the Encoder, it needs to be processed by the Embedding Layer into three matrices Q, K, and V, representing the query vector, key vector, and value vector respectively;
[0011] The Decoder is composed of two Wide Volume-based Multi-head Self-Attention units, and the result processed by the Decoder needs to be integrated with features through the Expanding Layer before output;
[0012] The Bottleneck includes two Local Volume-based Multi-head Self-attentions units and a Wide Volume-based Multi-head Self-attention unit.
[0013] A further technical solution of the present invention: The Embedding Layer includes two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer.
[0014] A method for estimating tissue microstructure, characterized in that the steps are as follows:
[0015] Step 1: Input the collected undersampled DMRI data;
[0016] Step 2: Use the diffusion wave vector space learning module to perform feature learning on the DMRI data to obtain a feature matrix with the same dimension as the input data matrix;
[0017] Step 3: Add each element of the obtained feature matrix to each element of the input data matrix, and use the resulting matrix as the input of the three-dimensional physical coordinate space learning module;
[0018] Step 4: Use the three-dimensional physical coordinate space learning module to further learn the input data and output a new feature matrix;
[0019] Step 5: Add each element of the new feature matrix obtained in Step 4 to each element of the result matrix obtained in Step 3;
[0020] Step 6: Perform operations on the matrix obtained in Step 5 through a convolutional layer and a fully connected layer to obtain the indices of the microstructure predicted by the model.
[0021] A further technical solution of the present invention: Step 2 is specifically as follows:
[0022] First, compare the angle θ between the sampling points v i and v j in the diffusion wave vector space with a preset threshold θ. If θ ij is less than θ, then a ij = 1; otherwise, a ij = 0. Extract the geometric structure features of the diffusion wave vector space in the DMRI data and save them as a binary adjacency matrix A = {a ij} to realize the graph representation of the DMRI data. Then, use the SGC network to remove the non-linear part in the graph convolutional network layer and collapse the function of K layers into a linear function, simplify the multi-layer network calculation to matrix multiplication, and improve the operation efficiency. The simplification formula is: ij where K represents the number of layers,
[0023]
[0024] is the adjacency matrix with self-loops added, represents the input feature matrix, and Θ represents the product result of the weight matrices of the K-layer graph network.
[0025] A further technical solution of the present invention: Step 4 includes the following steps:
[0026] Step 4-1: The data first enters the Embedding Layer in the Encoder. The Embedding Layer contains two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer. The convolution kernel and the convolution stride are both 1, and the data can be divided into three matrices Q, K, and V, which represent the query vector, the key vector, and the value vector respectively;
[0027]
[0027] Step 4-2: Input the obtained Q, K, and V matrices into the LSA for calculating the multi-head self-attention. The LSA is an offset window self-holding module, which defines the window as the Local 3D Volume Block. The self-attention calculation formula is as follows:
[0028]
[0029] where Q, K, and V are three matrices representing the query vector, key-value vector, and value vector respectively, B is the encoding matrix of the relative position, and D is the dimension of Q / K; the calculated result is denoted as L1. After decomposing L1 into Q, K, and V matrices, input them into the LSA again to obtain L2. Repeat the above operations to obtain L3 and L4;
[0030] Step 4-3: The obtained L3 and L4 will be input into the Skip Attention for operation; the Skip Attention is another attention calculation method. Use a single-layer neural network to decompose the output X of layer l l into a key-value matrix K l and a value matrix V l , and use the output X of another layer l' l′ as the query matrix Q l′ , and substitute it into the calculation formula:
[0031]
[0032] where B l′ is the encoding matrix of the relative position, D l′ is the dimension of Q l′ / K l ; L3 and L4 are used as X l and X l′ in the above description respectively, and substitute them into the formula to obtain S1;
[0033] Step 4-4: After decomposing S1 into Q, K, and V matrices, input them into the WSA. The WSA has a similar structure to the LSA and is obtained by multiplying the window size of the LSA by the constant 4; S1 is operated through the WSA to obtain W1;
[0034] Step 4-5: Input W1 and L2 into the Skip Attention to calculate S2;
[0035] Step 4-6: Input S2 into the WSA to calculate W2;
[0036] Step 4-7: Input W2 and L1 into the Skip Attention to calculate S3;
[0037] Step 4-8: Input S3 into the WSA to calculate W3;
[0038] Step 4-9: Finally, process W3 through the Expanding Layer to obtain the prediction matrix of the microstructure and output it.
[0039] A computer system, comprising: one or more processors, a computer-readable storage medium for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0040] A computer-readable storage medium, storing computer-executable instructions, which are used to implement the above method when executed.
[0041] The beneficial effects of the present invention are as follows:
[0042] An organization microstructure estimation model and method based on undersampled DMRI data provided by the present invention have the following advantages compared with the prior art:
[0043] 1. The present invention proposes a three-dimensional physical coordinate space learning module, which can utilize the three-dimensional space information of DMRI data for learning, and makes better use of the three-dimensional space information on the basis of HGT.
[0044] 2. The present invention proposes an efficient diffusion wave vector space learning module, which effectively improves the learning efficiency of the diffusion wave vector space and overall improves the accuracy of the model's microstructure imaging. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation to the present invention. Throughout the drawings, the same reference signs denote the same components.
[0046] Figure 1 Schematic diagram of the 3D-HGT structure;
[0047] Figure 2 Flowchart of generating microstructure images from DMRI data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0049] The present invention provides an efficient diffusion wave vector space learning module to improve the learning efficiency of the model in the diffusion wave vector space.
[0050] The efficient diffusion wave vector space learning module: Based on the Simplified Graph Convolutional Network (SGC), first, the q-space is characterized as a graph, and the angle θ i between point v j and v ij in the diffusion wave vector space is compared with a preset threshold θ. If θ ij is less than θ, then a ij = 1; otherwise, a ij = 0. In this way, the geometric structure features in the diffusion wave vector space of the DMRI data are extracted and saved as a binary matrix A = {a ij}. Then, using SGC, the non-linear part in the graph convolutional network layer is removed, and the function of K layers is collapsed into a linear function, which can simplify the multi-layer network calculation into matrix multiplication and improve the operation efficiency. The simplification formula is: where K represents the number of layers, is the adjacency matrix with self-loops added, represents the input feature matrix, and Θ represents the product result of the weight matrices of the K-layer graph network.
[0051] The present invention also provides a three-dimensional physical coordinate learning module, which increases the utilization of three-dimensional space information on the basis of HGT.
[0052] The three-dimensional physical coordinate space learning module: Based on the Transformer model, it is a U-shaped network composed of three parts: Encoder, Decoder, and Bottleneck, forming a cascaded self-attention Transformer block in the Local Volume. The structures of each part are as follows:
[0053] 1) Encoder is composed of two Local Volume-based Multi-head Self-Attention (LSA) units. Before the data enters the network through Encoder, it needs to be processed by the Embedding Layer into three matrices Q, K, and V, representing the query vector, key-value vector, and value vector respectively. The Embedding Layer contains two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer.
[0054] 2) Decoder is composed of two Wide Volume-based Multi-head Self-Attention (WSA) units. Before the result processed by Decoder is output, it needs to be integrated with features through the Expanding Layer.
[0055] 3) The Bottleneck includes two LSA units and one WSA unit. As a buffer block, Skip Attention integrates the shallow attention and deep attention between the Encoder and the Decoder here. Skip Attention is another attention calculation method, and its explanation is as follows.
[0056] The LSA, as an offset window self-holding module, defines the window as the Local 3D Volume Block, and the multi-head self-attention is calculated in this Local Volume. The formula is:
[0057]
[0058] where Q, K, and V are three matrices representing the query vector, key-value vector, and value vector respectively, B is the encoding matrix of the relative position, and D is the dimension of Q / K.
[0059] The WSA is similar to the LSA and is obtained by multiplying the window size of the LSA by the constant 4. It has a similar computational complexity to the LSA but greatly expands the attention perception field.
[0060] The Skip Attention is another attention calculation method. It decomposes the output X of layer l into a key-value matrix K l and a value matrix V l using a single-layer neural network, and uses the output X l of another layer l' as the query matrix Q l′ . The calculation formula of Skip Attention is: l′ where B
[0061]
[0062] is the encoding matrix of the relative position, and D l′ is the dimension of Q l′ / K l′ . l The dimension of
[0063] As Figure 1 shown, it is the structure diagram of each part of the 3D HGT, where:
[0064] a. It is the structure diagram of the 3D HGT model;
[0065] b. It is the structure diagram of the Transformer layer with the LSA / WSA as the self-attention module with window offset;
[0066] c. It is the structure diagram of the three-dimensional physical structure space learning module in the model.
[0067] AsFigure 2 As shown, it is a flowchart of the tissue microstructure estimation of the present invention. The process of generating a tissue microstructure index map from undersampled DMRI data is as follows:
[0068] Step 1: Input the undersampled DMRI data;
[0069] Step 2: Use the diffusion wave vector space learning module to perform feature learning on the DMRI data.
[0070] See Figure 1 a in Efficient q-Space Learning: First, the angle θ i between the sampling points v j in the diffusion wave vector space is compared with the preset threshold θ. If θ ij is less than θ, then a ij = 1; otherwise, a ij = 0. In this way, the geometric structure features in the diffusion wave vector space of the DMRI data can be extracted and saved as the binary adjacency matrix A = {a ij}, and then the DMRI data can be characterized as a graph. Then, use the SGC network to remove the non-linear part in the graph convolutional network layer and collapse the function of K layers into a linear function, so that the multi-layer network calculation can be simplified to matrix multiplication, improving the operation efficiency. The simplification formula is: ij where K represents the number of layers,
[0071]
[0072] is the adjacency matrix with self-loops added, represents the input feature matrix, and Θ represents the product result of the K-layer graph network weight matrix.
[0073] Step 3: The obtained feature matrix has the same dimension as the input data matrix. Add each element of the two matrices respectively, and the resulting matrix is used as the input of the three-dimensional physical structure space learning module.
[0074] Step 4: Use the three-dimensional physical structure space learning module to process the input data
[0075] The structure of this module can be seen in Figure 1 a in 3D x-Space Learning: It is a U-shaped network composed of an Encoder, a Decoder, and a Bottleneck.
[0076] The learning process of this module can be seen in Figure 1 c, specifically:
[0077] Step 4-1: The data first enters the Embedding Layer in the Encoder. The Embedding Layer contains two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer. The convolution kernel and the convolution stride are both 1, and the data can be divided into three matrices Q, K, and V, which represent the query vector, the key-value vector, and the value vector respectively.
[0078] Step 4-2: The three matrices Q, K, and V obtained are input into the LSA for calculating the multi-head self-attention. The self-attention calculation formula is:
[0079]
[0080] where Q, K, and V are three matrices representing the query vector, the key-value vector, and the value vector respectively, B is the encoding matrix of the relative position, and D is the dimension of Q / K. The calculation result is denoted as L1. After decomposing L1 into the Q, K, and V matrices, it is input into the LSA again to obtain L2. Repeat the above operations to obtain L3 and L4.
[0081] Step 4-3: The obtained L3 and L4 will be input into Skip Attention for operation. Skip Attention is another attention calculation method. A single-layer neural network decomposes the output X of layer l l into a key-value matrix K l and a value matrix V l , and takes the output X of another layer l' l′ as the query matrix Q l′ , and substitutes it into the calculation formula:
[0082]
[0083] where B l′ is the encoding matrix of the relative position, and D l′ is the dimension of Q / K. L3 and L4 are used as X l and X l′ described above and substituted into the formula to obtain S1.
[0084] Step 4-4: After decomposing S1 into the three matrices Q, K, and V, it is input into the WSA. S1 is operated by the WSA to obtain W1.
[0085] Step 4-5: Use Skip Attention to calculate S2 with W1 and L2.
[0086] Step 4-6: Input S2 into the WSA to calculate W2.
[0087] Step 4-7: Use Skip Attention to calculate S3 with W2 and L1.
[0088] Step 4-8: Input S3 into WSA for calculation to obtain W3.
[0089] Step 4-9: Finally, process W3 through the Expanding Layer to obtain the prediction matrix of the tissue microstructure and output it.
[0090] Step 5: Add the prediction matrix obtained in Step 4 to the matrix obtained in Step 3, element by element.
[0091] Step 6: Perform operations on the matrix obtained in Step 5 through the convolutional layer and the fully connected layer to obtain various indices of the tissue microstructure predicted by the model.
[0092] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. An organizational microstructure estimation model based on undersampled DMRI data, characterized in that It includes an efficient diffusion wave vector space learning module and a three-dimensional physical coordinate space learning module. After the data is input, it first undergoes feature learning through the efficient diffusion wave vector space learning module, and then the three-dimensional spatial information of the data is learned through the three-dimensional physical coordinate space learning module; The specific implementation of the efficient diffusion wave vector space learning module is as follows: Based on the SGC network, first compare the angles between sampling points in the diffusion wave vector space with a preset threshold. If the angle is less than the preset threshold, the geometric structure feature element is 1, otherwise it is 0, to generate a binary correlation matrix representing the graph network; then substitute it into the SGC network simplification formula for calculation to achieve efficient diffusion wave vector space learning; the SGC network simplification formula: SGC = K Among them, represents the number of layers, is the adjacency matrix with self-loops added, represents the input feature matrix, represents the product result of the weight matrices of the 𝐾-layer graph network; The specific implementation of the three-dimensional physical coordinate space learning module is as follows: Based on the Transformer model, a U-shaped network composed of three parts, namely Encoder, Decoder, and Bottleneck, constitutes a cascaded self-attention Transformer block in the local volume; the structures of each part are as follows: The Encoder is composed of two Local Volume-based Multi-head Self-Attention units; before the data enters the network through the Encoder, it needs to be processed by the Embedding Layer into and and three matrices, representing the query vector, the key-value vector, and the value vector respectively; The Decoder is composed of two Wide Volume-based Multi-head Self-Attention units, and the results processed by the Decoder need to be integrated with features through the Expanding Layer before output; The Bottleneck includes two Local Volume-based Multi-head Self-attentions units and a Wide Volume-based Multi-head Self-attention unit.
2. The tissue microstructure estimation model based on undersampled DMRI data according to claim 1, wherein The Embedding Layer contains two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer.
3. An organizational microstructure estimation method implemented based on the model described in claim 2, characterized in that The steps are as follows: Step 1: Input the collected undersampled DMRI data; Step 2: Use the diffusion wave vector space learning module to perform feature learning on the DMRI data to obtain a feature matrix with the same dimension as the input data matrix; Step 3: Add each element of the obtained feature matrix to each element of the input data matrix, and use the resulting matrix as the input of the three-dimensional physical coordinate space learning module; Step 4: Use the three-dimensional physical coordinate space learning module to further learn the input data and output a new feature matrix; Step 5: Add each element of the new feature matrix obtained in Step 4 to each element of the result matrix obtained in Step 3; Step 6: Perform operations on the matrix obtained in Step 5 through a convolutional layer and a fully connected layer to obtain the various indices of the microstructure predicted by the model.
4. The method for estimating tissue microstructure according to claim 3, wherein: The specific implementation of Step 2 is as follows: First, the angle between the sampling points in the diffusion wave vector space and is compared with a preset threshold . If it is less than , then = 1; otherwise = 0. The geometric structure features of the diffusion wave vector space in the DMRI data are extracted and saved as a binary adjacency matrix = to achieve the graph representation of the DMRI data; then, using the SGC network, the non-linear part in the graph convolutional network layer is removed and the function of the layer is collapsed into a linear function, reducing the multi-layer network calculation to matrix multiplication to improve the operation efficiency. The simplification formula is: SGC = K Among them, represents the number of layers, is the adjacency matrix with self-loops added, represents the input feature matrix, represents the product result of the weight matrices of the 𝐾-layer graph network.
5. The tissue microstructure estimation method according to claim 3, wherein: Step 4 includes the following steps: Step 4-1: The data first enters the Embedding Layer in the Encoder. The Embedding Layer contains two convolutional blocks composed of a convolutional layer, a GELU layer, and a normalization layer. The convolution kernel and the convolution stride are both 1, and the data can be divided into , , three matrices representing the query vector, the key-value vector, and the value vector respectively; Step 4-2: Input the obtained , , into the three matrices into LSA for calculating the multi-head self-attention. LSA is an offset window self-holding module, which defines the window as the Local 3D Volume Block. The self-attention calculation formula is: Attention , , =Softmax Among them, , , are three matrices representing the query vector, key vector, and value vector respectively. is the encoding matrix of relative positions. is / 's dimension; the calculated result is denoted as . Decompose into , , matrices and then input them into LSA again to obtain . Repeat the above operations to obtain , ; Step 4-3: The obtained is operated with the input Skip Attention. Using a single-layer neural network, the output of the layer is decomposed into a key-value matrix and a value matrix . Taking the output of another layer as the query matrix , substitute it into the calculation formula: Attention , , =Softmax Among them, is the coding matrix of the relative position, is / of which the dimension is; and are respectively used as and substituted into the formula to obtain ; Step 4-4: Decompose into , , and then input them into WSA. WSA has a similar structure to LSA and is obtained by multiplying the window size of LSA by the constant 4; Obtain through the operation of WSA; Step 4-5: Take and input into Skip Attention for calculation to obtain ; Step 4-6: Input into the WSA calculation to obtain ; Step 4-7: Combine with and input them into Skip Attention for calculation to obtain ; Step 4-8: Input into the WSA calculation to obtain ; Step 4-9: Finally, is processed by the Expanding Layer to obtain the prediction matrix of the microstructure and output it.
6. A computer system, characterized in that It includes: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 3-5.
7. A computer-readable storage medium, characterized in that Stores computer-executable instructions that, when executed, are used to implement the method according to any one of claims 3 to 5.
Citation Information
Patent Citations
Deep learning-based ocean wide swath distance ambiguity resolution method
CN114780911A
CRISPR / Cas9 single guide RNA targeting activity prediction method based on Transformer
CN114999576A