Three-dimensional brain network dynamic segmentation method based on deep learning
By combining multi-level cube partitioning and multi-scale sparse Transformer encoder with graph convolution and multi-head graph attention, the shortcomings of existing three-dimensional brain network segmentation methods in multimodal data fusion accuracy and dynamic segmentation continuity are solved, and more accurate and stable dynamic segmentation of brain areas is achieved.
Patent Information
- Application Number
- CN202510794726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-14
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-14
AI Technical Summary
Existing three-dimensional brain network segmentation methods have shortcomings in multimodal data fusion accuracy, spatial structure expression ability, cross-scale structural consistency modeling and dynamic segmentation continuity. It is difficult to effectively capture the evolution of the functional state of brain regions, especially when dealing with complex brain networks, and the segmentation results are poor.
A multi-level cube partitioning strategy and a multi-scale sparse Transformer encoder are adopted, combined with graph convolution and multi-head graph attention, to construct temporally and spatially consistent fused image volume data. Through the cross-scale node alignment mechanism and Transformer-GNN interactive update, the collaborative extraction of local fine-grained and global coarse-grained features and the residual update feedback of graph semantic information are achieved.
It significantly improves the coherence and medical interpretability of the segmentation map, can more accurately capture the dynamic changes of brain regions, enhances the ability to model the spatial structure and functional connection characteristics in cross-modal fusion images, and improves the accuracy and stability of the segmentation results.
Smart Images

Figure CN120689619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to a three-dimensional brain network dynamic segmentation method based on deep learning. Background Art
[0002] With the rapid development of neuroimaging technology, multimodal brain imaging methods such as functional magnetic resonance imaging, structural magnetic resonance imaging and diffusion tensor imaging have been widely used in the study of human brain cognitive mechanisms and the early diagnosis of brain diseases. In recent years, researchers have paid more and more attention to the dynamic evolution relationship between brain functional connectivity and brain area maps. Three-dimensional brain network segmentation, as a basic link in brain imaging analysis, has gradually developed from traditional static brain area division to a brain network segmentation task with temporal dynamic characteristics.
[0003] Currently, mainstream three-dimensional brain network segmentation methods are mostly based on convolutional neural networks to perform voxel-level semantic segmentation of single-modality imaging data. Although some progress has been made, many key issues still exist. On the one hand, traditional convolutional neural network models are difficult to fully capture the spatial relationships with strong non-Euclidean characteristics in brain imaging data, especially when dealing with cross-scale and unstructured brain area connection relationships. They lack global modeling capabilities; on the other hand, there are significant differences in spatial resolution, image registration, and signal-to-noise ratio between multimodal data. Traditional fusion strategies often simply adopt cascade or averaging operations, which cannot effectively extract deep collaborative features between modalities, limiting the further improvement of fusion segmentation accuracy.
[0004] In addition, although some existing brain region segmentation methods have introduced graph neural networks to model the topological structure between brain regions, the graph structure is usually constructed statically and cannot be dynamically adjusted for different scales or time periods. It is difficult to support the cross-scale evolution and information feedback of brain region connection maps in real neural activities. Especially when dealing with complex brain networks, there is a lack of effective coupling mechanism between graph modeling and segmentation models, resulting in poor performance of segmentation results in terms of boundary coherence, local details and dynamic consistency.
[0005] Furthermore, existing brain parcellation methods generally ignore the coherence modeling across time frames and are unable to maintain the consistency and stability of brain region identification at continuous time points, which is particularly critical for capturing the evolution of the functional state of brain regions. Existing segmentation methods usually process single-frame data during the training and inference stages and lack modeling capabilities in the time dimension, resulting in delays and ambiguity in the detection of dynamic brain activity, making it difficult to meet the actual needs of neuroscience and clinical diagnosis for high-temporal and spatial resolution segmentation technology.
[0006] In summary, existing three-dimensional brain network segmentation methods still have major technical defects in multimodal data fusion accuracy, spatial structure expression ability, cross-scale structural consistency modeling and dynamic segmentation continuity. There is an urgent need for a new segmentation strategy that can fully integrate multimodal features and combine spatial topological structure and temporal dynamic information to achieve more accurate, stable and temporally evolving three-dimensional brain area dynamic segmentation. Summary of the Invention
[0007] One purpose of the present invention is to propose a three-dimensional brain network dynamic segmentation method based on deep learning. The present invention realizes the collaborative extraction of local fine-grained and global coarse-grained features, which can not only effectively suppress redundant information, but also enhance the modeling ability of spatial structure and functional connection features in cross-modal fusion images.
[0008] According to an embodiment of the present invention, a three-dimensional brain network dynamic segmentation method based on deep learning includes the following steps:
[0009] S1. Construct temporally and spatially consistent fused image volume data;
[0010] S2. Divide the fused image volume data into cubes of different scales according to a multi-level cube partitioning strategy. Perform position encoding on the cubes of each scale and input them into a multi-scale sparse Transformer encoder to generate a multi-scale sparse Transformer encoded feature pyramid.
[0011] S3. Based on predefined medical partition templates and diffusion tensor imaging connectivity information, an initial brain region map node set is established at each scale for the multi-scale sparse Transformer encoding feature pyramid to obtain the brain region map structure at the same scale;
[0012] S4. Perform cross-scale node alignment on brain region map nodes with the same name at different scales, and use the cross-scale node alignment fusion graph structure as the initial output of the cross-scale node alignment fusion mechanism;
[0013] S5. Execute inter-node message passing within the cross-scale node alignment fusion graph structure through graph convolution and multi-head graph attention to obtain the first round of fused brain region graph node embedding features;
[0014] S6. Feed the first round of fused brain region map node embedding features as residual information to the next layer input of the multi-scale sparse Transformer encoder to obtain the updated multi-scale sparse Transformer encoding feature pyramid, and repeat steps S3 to S5 until the multi-scale sparse Transformer-GNN interactive update is completed;
[0015] S7. Perform skip connection fusion and hierarchical upsampling reconstruction on the final multi-scale sparse Transformer encoded feature pyramid to generate voxel-level segmentation probability volume data;
[0016] S8. Post-process the voxel-level segmentation probability volume data in the training phase and the inference phase respectively to obtain the three-dimensional brain region dynamic segmentation results.
[0017] Optionally, the S1 includes the following steps:
[0018] S11. Acquire structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data.
[0019] S12. performing initial resampling processing on the structural magnetic resonance imaging data, the functional magnetic resonance imaging data, and the diffusion tensor imaging data respectively;
[0020] S13. performing voxel-level rigid registration of the functional magnetic resonance imaging data and the diffusion tensor imaging data relative to the structural magnetic resonance imaging data;
[0021] S14. further performing non-rigid registration processing on the functional magnetic resonance imaging data and the diffusion tensor imaging data to establish a deformation field, where the deformation field describes the small displacement of each voxel position (x, y, z) in three spatial dimensions;
[0022] S15. Performing intensity normalization processing on the registered structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data to obtain normalized structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data;
[0023] S16. performing artifact removal processing on the normalized structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data;
[0024] S17. After registration, normalization, and artifact removal, the structural MRI data, functional MRI data, and diffusion tensor imaging data are fused at the channel level along the voxel dimension to form the final fused image volume data D fusion .
[0025] Optionally, the S2 includes the following steps:
[0026] S21. The fused image volume data D fusion Divide according to the multi-level cube partitioning strategy, using different scale partitioning factors s l , divide the original 3D space voxel grid into a set of non-overlapping cube blocks corresponding to the scale, and each scale level corresponds to a set of cube blocks at a resolution where N lrepresents the total number of cubes in layer l, l represents the scale level, Indicates the Nth at scale level l l A cube is a small cubic area that has been cut into three-dimensional space;
[0027] S22. For each cube at each scale level Perform position encoding to obtain the position encoding vector
[0028] S23. Encode the position vector at each scale level and the fused feature vector of the corresponding cube Splice to form a scale feature representation set F (l) ;
[0029] S24. For shallow scale level l high , the scale feature representation set Enter the local window self-attention module and set the window size to Perform dense self-attention mechanism within each window to output shallow local fine-grained representation
[0030] S25. For deep scale level l low , the scale feature representation set Input cross-window sparse self-attention module, only establish connections between the preset part of query and its related keys, and output deep global coarse-grained representation
[0031] S26. For each shallow scale level, the shallow local fine-grained representation is element-wise fused with the matching deep sparse global representation to form a fused encoding representation Z (l) ;
[0032] S27. Fusion encoding representation of all scale levels Unified organization is carried out to construct a multi-scale sparse Transformer encoding feature pyramid with cross-hierarchical structure perception capabilities.
[0033] Optionally, S3 includes the following steps:
[0034] S31. Construct a predefined medical partition template to form a brain region partition set R. The brain region partition set consists of multiple brain region template regions. In each scale level l, the fusion coding representation set Z (l) The spatial position of each cube block in the brain area template area is judged to coincide with the space, and the cube block is mapped to the graph node in the brain area graph structure through the mapping function to form the initial node mapping relationship;
[0035] S32. At scale level l, generate an initial brain region map node set V based on the node mapping relationship (l) Each graph node in the initial brain region graph node set is used to represent a set of cube blocks of a brain region template at the current scale. For each graph node, the average value of all cube block encoding representations contained in the current node is calculated to obtain the initial feature vector of the current graph node.
[0036] S33. Using the initial feature vector set of all graph nodes at scale level l, calculate the anatomical similarity between each pair of graph nodes and form an anatomical similarity matrix;
[0037] S34. Extract the number of white matter fiber tracts between each pair of brain template regions based on the normalized diffusion tensor imaging data, calculate the normalized connectivity based on the fiber distribution density within and outside the region, and organize the functional connectivity results between all regions into a functional similarity matrix;
[0038] S35. Construct a graph adjacency matrix at scale level l by linearly weighted fusion of the anatomical similarity matrix and the functional similarity matrix;
[0039] S36. Combine the brain region graph node set at scale level l with the graph adjacency matrix to form the brain region graph structure G at the same scale level at the current scale level. (l) .
[0040] Optionally, S3 includes the following steps:
[0041] S41. Set a cross-scale alignment threshold, select all scale levels with resolutions lower than the cross-scale alignment threshold to form a low-resolution set, and select all scale levels with resolutions higher than or equal to the cross-scale alignment threshold to form a high-resolution set;
[0042] S42. For each low-resolution scale level of the brain area map structure, extract the spatial center coordinates of each map node and the initial feature vector of the map node;
[0043] S43. At the high-resolution scale level, for all candidate graph nodes that have the same medical template name as the low-resolution graph node, calculate the spatial Euclidean distance between each candidate graph node and the low-resolution graph node, and select the candidate node with the smallest Euclidean distance as the high-resolution alignment target graph node;
[0044] S44. Match each pair of low-resolution graph nodes with their nearest neighbor high-resolution alignment target graph nodes to form a cross-scale alignment mapping set;
[0045] S45. For each pair of matching graph nodes in the cross-scale alignment mapping set, calculate feature similarity using the initial feature vectors of the two graph nodes;
[0046] S46. All cross-scale matching graph node pairs and corresponding alignment weights are combined to form a cross-scale edge set;
[0047] S47. Merge the low-resolution scale-level brain region map structure with the high-resolution scale-level brain region map structure and the cross-scale edge set to obtain the final cross-scale node-aligned fusion brain region map structure.
[0048] Optionally, the S5 includes the following steps:
[0049] S51. Based on the mapping relationship between graph nodes and cubes in the cross-scale node alignment fusion brain region graph structure, the cube encoding features at each scale level in the multi-scale sparse Transformer encoding feature pyramid are mapped to the initial attribute features of the corresponding graph nodes;
[0050] S52. Concatenate and integrate the initial attribute features of the graph nodes in all scale levels according to the node numbers to form an initial feature matrix of the cross-scale node set;
[0051] S53. Input the initial feature matrix into the graph neural network message passing module, perform the first round of graph convolution operation in combination with the adjacency matrix of the brain region graph structure of cross-scale node alignment fusion, and obtain the first round of node embedding representation H (1) ;
[0052] S54. Input the first round of graph convolution representation into the multi-head graph attention mechanism module, and integrate the output of each graph node obtained from all attention heads through the splicing operation to generate the first round of fused brain area graph node embedding features.
[0053] Optionally, the S6 includes the following steps:
[0054] S61. According to the mapping relationship between the graph nodes and the original cube blocks, the first round of fusion brain area graph nodes are embedded into the features Project and broadcast back to the original cube blocks at each scale level in the multi-scale sparse Transformer encoding feature pyramid to obtain the graph feedback residual information corresponding to each cube block;
[0055] S62. For each scale level l, the residual information R is fed back to the graph (1) The corresponding residual vector in The cubic block features output by the previous layer Transformer encoding Perform a joint update to form a new cubic input representation
[0056] S63. Input the new cube into the representation Input to the next layer of multi-scale sparse Transformer encoder, perform local window dense self-attention and cross-window sparse self-attention mechanism, extract multi-scale features of the new input, and output the updated multi-scale sparse Transformer encoding feature pyramid {
[0057]
[0058] S64. After each round of updating, repeat the process of constructing the brain region map structure at the same scale in S3 to obtain a new brain region map structure;
[0059] S65. Based on the new brain region map structure, repeat the cross-scale node alignment process in S4 to generate a new cross-scale node aligned fusion brain region map structure;
[0060] S66. Execute graph neural network message passing and multi-head graph attention mechanism on the new cross-scale node alignment fusion brain region graph structure to generate the second round of fusion brain region graph node embedding features And feed it back to the encoder for residual update;
[0061] S67. Repeat steps S61 to S66 to perform multiple rounds of Transformer-GNN dual-path interactive updates until the set number of iterations or the residual convergence threshold is reached, and finally a stable multi-scale sparse Transformer encoding feature pyramid and fused brain area map node embedding representation are obtained.
[0062] Optionally, the S7 includes the following steps:
[0063] S71. Encoding feature pyramids based on stable multi-scale sparse Transformer Encoded features at all scale levels A skip connection fusion strategy is used to integrate information, and channel splicing and convolution fusion are performed with the current level encoding features to generate a fused cross-scale voxel feature representation;
[0064] S72. Based on the mapping relationship between graph nodes and cube blocks, the embedded representation of the fused brain region graph nodes is projected back to the corresponding voxel block area via the nearest neighbor, and a graph feature field is established in the voxel space. The cross-scale voxel feature representation and the graph feature field are then concatenated in the channel dimension to form the final fused voxel-level segmentation feature.
[0065] S73. The final fused voxel-level segmentation features are input into the segmentation head network composed of convolutional layers. The segmentation head network performs a set of 1×1×1 convolution operations and outputs voxel-level multi-category probability predictions through the softmax function to generate voxel-level segmentation probability volume data.
[0066] Optionally, the S8 specifically includes introducing adjacent time frame probability consistency constraints and deformation field smoothing regularization to the voxel-level segmentation probability volume data of continuous time frames in the training stage, applying recursive Bayesian filtering and B-spline-based deformation compensation to the voxel-level segmentation probability volume data in the inference stage, performing conditional random field refinement, flight time mapping-based hole filling and confidence evaluation on the voxel-level segmentation probability volume data that evolves smoothly in the time domain, and obtaining three-dimensional brain area dynamic segmentation results.
[0067] The beneficial effects of the present invention are:
[0068] (1) The present invention constructs a multi-level cube partitioning strategy to divide the fused sMRI, fMRI and DTI image volumes into cubic blocks of different scales, and introduces a local window dense self-attention mechanism in the shallow scale and a sparse cross-window self-attention mechanism in the deep scale, thereby realizing the collaborative extraction of local fine-grained and global coarse-grained features. It can not only effectively suppress redundant information, but also enhance the modeling ability of spatial structure and functional connectivity features in cross-modal fusion images.
[0069] (2) Based on the predefined medical brain region template, the present invention constructs initial brain region map nodes at different scales and introduces a cross-scale node alignment mechanism. Through the spatial nearest neighbor principle and feature similarity weighting strategy, a one-to-one mapping between low-resolution and high-resolution map nodes is established and fused into a cross-scale map structure, which effectively bridges the problem of inconsistent semantics between maps of different resolutions, maintains spatial stability in overlapping brain regions, and maintains node semantic consistency in anatomically ambiguous areas, significantly improving the coherence and medical interpretability of the segmentation map.
[0070] (3) The present invention constructs a dual-path nested interaction mechanism of Transformer and GNN. By back-projecting the node embedding feature vector obtained in the graph neural network to the input channel of the original multi-scale Transformer, the residual update feedback of the graph semantic information is realized, and the graph structure and encoding features are reconstructed in each cycle, thereby achieving the coordinated evolution of the cross-scale graph structure and the encoding pyramid until the residual converges. When processing the dynamic changes of brain areas, it can more accurately capture the boundary evolution and regional expansion, and improve the model's perception of the dynamic changes of small areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0072] Figure 1 This is a flowchart of a three-dimensional brain network dynamic segmentation method based on deep learning proposed by the present invention. DETAILED DESCRIPTION
[0073] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0074] refer to Figure 1 , a three-dimensional brain network dynamic segmentation method based on deep learning, comprising the following steps:
[0075] S1. Construct temporally and spatially consistent fused image volume data;
[0076] S2. Divide the fused image volume data into cubes of different scales according to a multi-level cube partitioning strategy. Perform position encoding on the cubes of each scale and input them into a multi-scale sparse Transformer encoder. In the shallow layer, perform local window dense self-attention on the cubes of the same scale, and in the deep layer, perform cross-window sparse self-attention on the cubes of different scales to generate a multi-scale sparse Transformer encoding feature pyramid.
[0077] S3. Based on predefined medical partitioning templates and diffusion tensor imaging connectivity information, we establish an initial brain region map node set for the multi-scale sparse Transformer encoding feature pyramid at each scale. We then calculate the anatomical and functional similarities between nodes within the scale to obtain the brain region map structure at the same scale.
[0078] S4. Perform cross-scale node alignment on brain region map nodes with the same name at different scales. Establish an alignment mapping between brain region map nodes with resolutions below a threshold and those with resolutions above the threshold through nearest neighbor matching. Under the guidance of the alignment mapping, construct a cross-scale node alignment fusion graph structure. This cross-scale node alignment fusion graph structure is used as the initial output of the cross-scale node alignment fusion mechanism.
[0079] S5. Map the multi-scale sparse Transformer encoding feature pyramid to node attributes of a cross-scale node-aligned fusion graph structure. Graph convolution and multi-head graph attention are used to perform inter-node message passing within the cross-scale node-aligned fusion graph structure to obtain the first round of fused brain region graph node embedding features.
[0080] S6. Feed the first round of fused brain region map node embedding features as residual information to the next layer input of the multi-scale sparse Transformer encoder, and update them together with the previous layer features in the next layer of multi-scale sparse Transformer encoder to obtain the updated multi-scale sparse Transformer encoding feature pyramid. Repeat steps S3 to S5 until the multi-scale sparse Transformer-GNN interactive update is completed;
[0081] S7. Perform skip connection fusion and hierarchical upsampling reconstruction on the final multi-scale sparse Transformer encoded feature pyramid, project the final node embedding features of the cross-scale node alignment fusion graph structure back to the voxel space, and generate voxel-level segmentation probability volume data;
[0082] S8. Post-process the voxel-level segmentation probability volume data in the training phase and the inference phase respectively to obtain the three-dimensional brain region dynamic segmentation results.
[0083] In this embodiment, S1 includes the following steps:
[0084] S11. Acquire structural MRI data, functional MRI data, and diffusion tensor imaging data. Structural MRI data are used to provide a three-dimensional anatomical reference. Functional MRI data characterize fluctuations in blood oxygen levels caused by neural activity. Diffusion tensor imaging data are used to describe the orientation of white matter fiber bundles and connectivity information between brain regions.
[0085] S12. Perform initial resampling on the structural MRI data, functional MRI data, and diffusion tensor imaging data, respectively, and normalize the spacing of the three-dimensional voxels in the three spatial dimensions of x, y, and z to v. x 、v y 、v z , the normalized spacing is uniformly expressed as the voxel space spacing Δv;
[0086] S13. performing voxel-level rigid registration of the functional magnetic resonance imaging data and the diffusion tensor imaging data relative to the structural magnetic resonance imaging data;
[0087] S14. After performing rigid registration, further non-rigid registration is performed on the fMRI and DTI data to establish a deformation field. The deformation field describes the small displacement of each voxel position in the three spatial dimensions (x, y, z). This deformation field is used to locally fine-tune the registration results to eliminate subtle misalignments caused by differences in brain tissue morphology or individual differences.
[0088] S15. Performing intensity normalization processing on the registered structural MRI data, functional MRI data, and diffusion tensor imaging data, respectively, using an intra-modality intensity normalization method to normalize the voxel intensities of each image using the mean and standard deviation, respectively, to obtain normalized structural MRI data, functional MRI data, and diffusion tensor imaging data;
[0089] S16. Perform artifact removal on the normalized structural MRI data, functional MRI data, and diffusion tensor imaging data using spatial high-pass filtering and boundary enhancement based on morphological reconstruction to eliminate false structural responses caused by imaging artifacts, noise interference, or magnetic susceptibility effects, while preserving anatomical boundaries and true tissue information to the greatest extent possible.
[0090] S17. After registration, normalization, and artifact removal, the structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data are fused at the channel level along the voxel dimension to form the final fused image volume data.
[0091] In this embodiment, S2 includes the following steps:
[0092] S21. The fused image volume data D fusion Divide according to the multi-level cube partitioning strategy, using different scale partitioning factors s l , where l represents the scale level, the original three-dimensional space voxel grid is divided into a set of non-overlapping cubic blocks corresponding to the scale, and each scale level corresponds to a set of cubic blocks at a resolution where N l represents the total number of cubes in layer l, Indicates the Nth at scale level l l A cube is a small cubic area that has been cut into three-dimensional space;
[0093] The spatial structure of the fused image volume data itself is a regular three-dimensional voxel grid structure, which is regarded as the original three-dimensional spatial voxel grid of the input image.
[0094] S22. For each cube at each scale level Perform position encoding, i is greater than or equal to 1 and less than or equal to N l , define the position encoding vector Where (x i ,y i ,z i ) is the 3D spatial coordinate of the cube center in the fused image volume data, and PE(·) is a learnable position encoding function;
[0095] S23. Encode the position vector at each scale level and the fused feature vector of the corresponding cube Splice to form a scale feature representation set in Represents a splicing operation;
[0096] S24. For shallow scale level l high , the scale feature representation set Enter the local window self-attention module and set the window size to Perform dense self-attention mechanism within each window to output shallow local fine-grained representation
[0097] S25. For deep scale level l low , the scale feature representation set Input cross-window sparse self-attention module, only establish connections between the preset part of query and its related keys, and output deep global coarse-grained representation
[0098] S26. For each shallow scale level, the shallow local fine-grained representation is element-wise fused with the matching deep sparse global representation to form a fused encoding representation.
[0099] S27. Fusion encoding representation of all scale levels Unified organization is carried out to construct a multi-scale sparse Transformer encoding feature pyramid with cross-hierarchical structure perception capabilities.
[0100] In this embodiment, S3 includes the following steps:
[0101] S31. Construct a predefined medical partition template to form a brain region partition set R. The brain region partition set consists of multiple brain region template regions. In each scale level l, the fusion coding representation set Z (l) The spatial position of each cube block in the brain area template area is judged to coincide with the space, and the cube block is mapped to the graph node in the brain area graph structure through the mapping function to form the initial node mapping relationship;
[0102] S32. At scale level l, generate an initial brain region map node set V based on the node mapping relationship (l) Each graph node in the initial brain region graph node set is used to represent a set of cube blocks of a brain region template at the current scale. For each graph node, the average value of all cube block encoding representations contained in the current node is calculated to obtain the initial feature vector of the current graph node. The initial feature vector of the graph node is used to represent the representation information of the current brain region graph node at the current scale.
[0103] S33. Using the initial set of feature vectors for all graph nodes at scale level l, calculate the anatomical similarity between each pair of graph nodes. The current anatomical similarity is obtained by calculating the cosine similarity between the feature vectors of the two corresponding graph nodes. The anatomical similarity is used to measure the local structural homogeneity of the two brain region graph nodes in the encoding feature space, thereby forming an anatomical similarity matrix.
[0104] S34. Extract the number of white matter fiber tracts between each pair of brain template regions based on the normalized diffusion tensor imaging data. Calculate the normalized connectivity based on the fiber density within and outside the region to represent the strength of functional connectivity between the two regions. Organize the functional connectivity results between all regions into a functional similarity matrix, which is used to measure the coupling relationship between neural signal pathways between brain regions.
[0105] S35. Construct a graph adjacency matrix at scale level l by linearly weighted fusion of the anatomical similarity matrix and the functional similarity matrix. The fusion process uses a fusion weight coefficient, which is used to control the proportion of anatomical information and functional information that contribute to the graph structure.
[0106] S36. Combine the brain region graph node set at scale level l with the graph adjacency matrix to form the brain region graph structure G at the same scale level at the current scale level. (l) .
[0107] In this embodiment, S3 includes the following steps:
[0108] S41. Set a cross-scale alignment threshold to determine the resolution relationship between different scale levels, select all scale levels with resolutions lower than the cross-scale alignment threshold to form a low-resolution set, and select all scale levels with resolutions higher than or equal to the cross-scale alignment threshold to form a high-resolution set;
[0109] S42. For each low-resolution scale level of the brain area map structure, extract the spatial center coordinates of each map node and the initial feature vector of the map node;
[0110] S43. At the high-resolution scale level, for all candidate graph nodes that have the same medical template name as the low-resolution graph node, calculate the spatial Euclidean distance between each candidate graph node and the low-resolution graph node, and select the candidate node with the smallest Euclidean distance as the high-resolution alignment target graph node;
[0111] S44. Match each pair of low-resolution graph nodes with the nearest neighbor high-resolution alignment target graph node to form a matching pair, and form a cross-scale alignment mapping set, where each mapping in the cross-scale alignment mapping set is used to describe a mapping path from a low-resolution scale hierarchical graph node to a high-resolution scale hierarchical graph node;
[0112] S45. For each pair of matching graph nodes in the cross-scale alignment mapping set, calculate feature similarity using the initial feature vectors of the two graph nodes;
[0113] S46. All cross-scale matching graph node pairs and corresponding alignment weights are combined to form a cross-scale edge set, which is used to record the connection relationship and weight strength between the low-resolution graph nodes and the high-resolution graph nodes;
[0114] S47. Merge the low-resolution scale-level brain region graph structure with the high-resolution scale-level brain region graph structure and the cross-scale edge set to obtain the final cross-scale node-aligned fusion brain region graph structure. The cross-scale node-aligned fusion brain region graph structure is composed of the graph node sets of the two scale levels, their respective graph adjacency matrices and the cross-scale edge set, which serves as the initial brain region graph structure output of the cross-scale node alignment fusion mechanism.
[0115] In this embodiment, S5 includes the following steps:
[0116] S51. Based on the mapping relationship between graph nodes and cubes in the cross-scale node alignment fusion brain region graph structure, map the cube encoding features at each scale level in the multi-scale sparse Transformer encoding feature pyramid to the initial attribute features of the corresponding graph node. The initial attribute features of the graph node are obtained by averaging all the cube encoding features mapped to the graph node.
[0117] S52. Concatenate and integrate the initial attribute features of the graph nodes in all scale levels according to the node numbers to form an initial feature matrix of the cross-scale node set, where each row in the initial feature matrix corresponds to a graph node, and each column corresponds to a feature dimension;
[0118] S53. Input the initial feature matrix into the graph neural network message passing module, and perform the first round of graph convolution operation in combination with the adjacency matrix of the brain region graph structure that is fused with cross-scale node alignment. The first round of graph convolution operation uses the spectral domain graph convolution to add self-loops to the adjacency matrix and then perform degree matrix normalization. The convolution weight matrix is used to perform linear transformation on the features, and the activation function is applied to obtain the first round of node embedding representation H. (1) ;
[0119] S54. Input the first round of graph convolution representation into the multi-head graph attention mechanism module, set the number of attention heads to K, and each graph node performs weighted summation of the features of neighboring nodes on each attention head according to the attention weight of the adjacent nodes. The attention weight is determined by the correlation between node features, and the feature projection is provided by the learnable linear transformation matrix of each attention head. The output of each graph node obtained from all attention heads is integrated through the splicing operation to generate the first round of fused brain area graph node embedding features.
[0120] In this embodiment, S6 includes the following steps:
[0121] S61. According to the mapping relationship between the graph nodes and the original cube blocks, the first round of fusion brain area graph nodes are embedded into the features Project and broadcast back to the original cube blocks at each scale level in the multi-scale sparse Transformer encoding feature pyramid to obtain the graph feedback residual information corresponding to each cube block;
[0122] S62. For each scale level l, the residual information R is fed back to the graph (1) The corresponding residual vector in The cubic block features output by the previous layer Transformer encoding Perform a joint update to form a new cubic input representation
[0123]
[0124] Among them, γ is the graph residual control coefficient, Represents the cubic feature input after fusion of graph node residual information;
[0125] S63. Input the new cube into the representation Input to the next layer of multi-scale sparse Transformer encoder, perform local window dense self-attention and cross-window sparse self-attention mechanism, extract multi-scale features of the new input, and output the updated multi-scale sparse Transformer encoding feature pyramid
[0126] S64. After each round of updating, repeat the process of constructing the brain region map structure at the same scale in S3 to obtain a new brain region map structure;
[0127] S65. Based on the new brain region map structure, repeat the cross-scale node alignment process in S4 to generate a new cross-scale node aligned fusion brain region map structure;
[0128] S66. Execute graph neural network message passing and multi-head graph attention mechanism on the new cross-scale node alignment fusion brain region graph structure to generate the second round of fusion brain region graph node embedding features And feed it back to the encoder for residual update;
[0129] S67. Repeat steps S61 to S66 to perform multiple rounds of Transformer-GNN dual-path interactive updates until the set number of iterations or the residual convergence threshold is reached, and finally a stable multi-scale sparse Transformer encoding feature pyramid and fused brain area map node embedding representation are obtained.
[0130] In this embodiment, S7 includes the following steps:
[0131] S71. Encoding feature pyramids based on stable multi-scale sparse Transformer Encoded features at all scale levels A skip connection fusion strategy is used for information integration. The skip connection fusion strategy is defined as upsampling from the highest scale level l = L to the lowest scale level l = 1 step by step, and performing channel concatenation and convolution fusion with the current level encoding features to generate a fused cross-scale voxel feature representation.
[0132] S72. Based on the mapping relationship between graph nodes and cube blocks, the fused brain region graph node embedding representation is projected back to its corresponding voxel block area via the nearest neighbor, and a graph feature field is established in the voxel space. The cross-scale voxel feature representation and the graph feature field are then concatenated in the channel dimension to form the final fused voxel-level segmentation feature, which represents the comprehensive expression of each voxel position in the local space and cross-node semantic dual pathways.
[0133] S73. The final fused voxel-level segmentation features are input into the segmentation head network composed of convolutional layers. The segmentation head network performs a set of 1×1×1 convolution operations and outputs voxel-level multi-category probability predictions through the softmax function to generate voxel-level segmentation probability volume data P voxel .
[0134] In this embodiment, S8 specifically includes introducing adjacent time frame probability consistency constraints and deformation field smoothing regularization to the voxel-level segmentation probability volume data of continuous time frames in the training stage, applying recursive Bayesian filtering and B-spline-based deformation compensation to the voxel-level segmentation probability volume data in the inference stage to achieve a segmentation boundary that evolves smoothly in the time domain, performing conditional random field refinement, flight time mapping-based hole filling and confidence evaluation on the voxel-level segmentation probability volume data that evolves smoothly in the time domain, and obtaining a spatially refined and temporally smooth three-dimensional brain area dynamic segmentation result.
[0135] Example 1: A patient suspected of temporal lobe epilepsy was admitted to the neurosurgery department of Hospital A. The doctor needed to obtain high-resolution, dynamic, continuous, and spatially realistic brain region segmentation results before surgery to support the formulation of precise surgical plans and intraoperative navigation. The patient was a 26-year-old male, right-handed, and had no serious comorbidities. The hospital's imaging center collected multimodal whole-brain images, including T1 structural magnetic resonance imaging (spatial resolution 1 mm 3 , size 256×256×176), dynamic fMRI images (TR=0.8s, time series length 240 frames, spatial resolution 2.5mm 3, 96×96×60), and diffusion tensor imaging (64-dimensional, high-resolution, with the same spatial dimensions as T1). Prior annotations of medical brain regions based on the standard AAL template and Desikan-Killiany partitioning were also collected. To simulate the temporal microdeformation of brain tissue caused by intraoperative cisternal filtration and traction, real-time dynamic image sequences of the brain surface and deep sulci during surgical manipulation were also collected to verify the stability of dynamic scene segmentation.
[0136] Using the preprocessing platform of the hospital imaging center, T1 structural MRI, fMRI and DTI data were subjected to head motion artifact removal, intensity normalization, and voxel space resampling (unified to 1.0 mm 3 ), then all modalities are aligned to the T1 structural space based on a mutual information-optimized registration algorithm, and the cross-modality registration is further refined using non-rigid B-spline deformation compensation technology. Finally, the fusion generates temporally and spatially consistent multi-channel image volume data, with each frame of data having a spatial size of 256×256×176, which remains consistent with the original fMRI on the time axis.
[0137] In the multi-scale encoding stage, the fused image volume data is divided into multi-level patches (8×8×8, 16×16×16, 32×32×32), forming a high, medium, and low-level cubic block set. The system performs learnable position encoding on each layer of cubic blocks and inputs them into the self-developed multi-scale sparse Transformer encoder. In the shallow layer, dense self-attention is performed within the local window (16×16×16) to capture the microstructure, while in the deep layer, sparse global attention is established across windows only for important key points, effectively reducing the video memory and computing pressure of the original Transformer under large-scale voxels. A single GPU card can stably support 512 3 Real-time inference on volume data.
[0138] Using AAL template partitioning and patient-specific DTI fiber tracking data, the system automatically establishes an initial brain region map structure at multiple scales. Node adjacency relationships are generated based on cosine similarity of structural features and white matter fiber connectivity. For nodes in the same brain region at different scales, the system automatically establishes cross-scale alignment mappings using spatial nearest neighbor and feature similarity, generating a cross-scale fused heterogeneous graph structure encompassing all scales and brain regions.
[0139] After the multi-scale sparse Transformer encoding features are transferred to the node attributes of the fusion graph structure through the graph node mapping function, multi-layer spectral domain graph convolution and multi-head attention mechanism are used to perform inter-node message passing. The first round of node embedding features obtained are fed back to the next layer of Transformer to realize the sequence global-topological structure dual-pathway cyclic interaction of segmentation features. Three rounds of iterations are set in the entire interaction process, and the residual convergence threshold is 0.005, and finally stable segmentation candidate features are obtained.
[0140] During the segmentation output phase, U-shaped skip connections are used to fuse features from each layer, and hierarchical upsampling is used to reconstruct low-resolution features to the original voxel resolution. Combined with the latest round of graph node embedding features, these features are added to the voxel space through interpolation and projection, and then spliced together as the final segmentation head input. Finally, a voxel-level segmentation probability map is output for each frame. KL consistency constraints and Kalman filter-based smoothing compensation are introduced for adjacent frames in the temporal sequence, significantly improving the smoothness of the segmentation over time and the authenticity of physiological evolution. Segmentation results are output in DICOM and NIfTI formats, which can be directly accessed by physicians through preoperative planning and intraoperative real-time navigation systems.
[0141] The research team used patient data and retrieved 25 epilepsy / tumor / bleeding cases (25 subjects aged 14-62 years) with similar segmentation requirements from the hospital data platform between 2019 and 2024. They compared the proposed method with three mainstream segmentation methods on an equivalent GPU server (NVIDIA RTX4090, 128GB memory): traditional 3D U-Net, multi-stage convolutional segmentation based on SegResNet, and a segmentation algorithm using only 2.5DSwin-Transformer. The specific comparison results are as follows:
[0142] Table 1 Comparison results between the method of the present invention and three mainstream segmentation methods
[0143]
[0144] In the test data, the training sample size of the segmentation candidate is 18,000 frames, and the test sample size is 4,100 frames. The training adopts the AdamW optimizer, the initial learning rate is 0.0006, the batch size is 2, and the training is 150 rounds. All methods are carried out under the same conditions of artifacts, deformation and signal drift. In terms of quantitative indicators, the method of the present invention is superior to traditional convolution and pure Transformer methods in the core indicators of whole-brain segmentation Dice score (0.913), boundary continuity (Hausdorff distance 1.82mm), dynamic frame segmentation flicker rate (0.9%), and deep groove / vascular area breakage rate (2.2%). In the dynamic segmentation scenario, the fused graph neural structure can quickly adapt to the changes in functional connectivity and effectively suppress the drift and breakage problems of temporal segmentation. In terms of reasoning efficiency, the present invention can achieve real-time reasoning of 1.48 seconds / frame under single-frame segmentation, which is significantly faster than most similar Transformer models, and the video memory consumption is lower than the standard 2.5D Swin-Transformer solution.
[0145] Based on predefined medical brain region templates, the present invention constructs initial brain region map nodes at different scales and introduces a cross-scale node alignment mechanism. Through the spatial nearest neighbor principle and feature similarity weighting strategy, a one-to-one mapping between low-resolution and high-resolution map nodes is established and fused into a cross-scale map structure, effectively bridging the problem of semantic inconsistency between maps of different resolutions. It can maintain spatial stability in overlapping brain region areas and maintain node semantic consistency in anatomically ambiguous areas, significantly improving the coherence and medical interpretability of the segmentation map.
[0146] The present invention constructs a dual-path nested interaction mechanism of Transformer and GNN. By back-projecting the node embedding feature vector obtained in the graph neural network to the input channel of the original multi-scale Transformer, the residual update feedback of the graph semantic information is realized, and the graph structure and encoding features are reconstructed in each cycle, thereby achieving the coordinated evolution of the cross-scale graph structure and the encoding pyramid until the residual converges. When processing dynamic changes in brain areas, it can more accurately capture boundary evolution and regional expansion, thereby improving the model's perception of dynamic changes in small areas.
[0147] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A three-dimensional brain network dynamic segmentation method based on deep learning, characterized by: The steps include: S1. Construct temporally and spatially consistent fused image volume data; S2. Divide the fused image volume data into cubes of different scales according to a multi-level cube partitioning strategy. Perform position encoding on the cubes of each scale and input them into a multi-scale sparse Transformer encoder to generate a multi-scale sparse Transformer encoded feature pyramid. S3. Based on predefined medical partition templates and diffusion tensor imaging connectivity information, an initial brain region map node set is established at each scale for the multi-scale sparse Transformer encoding feature pyramid to obtain the brain region map structure at the same scale; S4. Perform cross-scale node alignment on brain region map nodes with the same name at different scales, and use the cross-scale node alignment fusion graph structure as the initial output of the cross-scale node alignment fusion mechanism; S5. Execute inter-node message passing within the cross-scale node alignment fusion graph structure through graph convolution and multi-head graph attention to obtain the first round of fused brain region graph node embedding features; S6. Feed the first round of fused brain region map node embedding features as residual information to the next layer input of the multi-scale sparse Transformer encoder to obtain the updated multi-scale sparse Transformer encoding feature pyramid, and repeat steps S3 to S5 until the multi-scale sparse Transformer-GNN interactive update is completed; S7. Perform skip connection fusion and hierarchical upsampling reconstruction on the final multi-scale sparse Transformer encoded feature pyramid to generate voxel-level segmentation probability volume data; S8. Post-process the voxel-level segmentation probability volume data in the training phase and the inference phase respectively to obtain the three-dimensional brain region dynamic segmentation results.
2. A three-dimensional brain network dynamic segmentation method based on deep learning according to claim 1, characterized in that: Said S1 comprises the following steps: S11. Acquire structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data. S12. performing initial resampling processing on the structural magnetic resonance imaging data, the functional magnetic resonance imaging data, and the diffusion tensor imaging data respectively; S13. performing voxel-level rigid registration of the functional magnetic resonance imaging data and the diffusion tensor imaging data relative to the structural magnetic resonance imaging data; S14. further performing non-rigid registration processing on the functional magnetic resonance imaging data and the diffusion tensor imaging data to establish a deformation field, where the deformation field describes the small displacement of each voxel position (x, y, z) in three spatial dimensions; S15. Performing intensity normalization processing on the registered structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data to obtain normalized structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data; S16. performing artifact removal processing on the normalized structural magnetic resonance imaging data, functional magnetic resonance imaging data, and diffusion tensor imaging data; S17. After registration, normalization, and artifact removal, the structural MRI data, functional MRI data, and diffusion tensor imaging data are fused at the channel level along the voxel dimension to form the final fused image volume data D fusion .
3. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 2, characterized in that: The S2 comprises the following steps: S21. The fused image volume data D fusion Divide according to the multi-level cube partitioning strategy, using different scale partitioning factors s l , divide the original 3D space voxel grid into a set of non-overlapping cube blocks corresponding to the scale, and each scale level corresponds to a set of cube blocks at a resolution where N l represents the total number of cubes in layer l, l represents the scale level, Indicates the Nth at scale level l l A cube is a small cubic area that has been cut into three-dimensional space; S22. For each cube at each scale level Perform position encoding to obtain the position encoding vector S23. Encode the position vector at each scale level and the fused feature vector of the corresponding cube Splice to form a scale feature representation set F (l) ; S24. For shallow scale level l high , the scale feature representation set Enter the local window self-attention module and set the window size to Perform dense self-attention mechanism within each window to output shallow local fine-grained representation S25. For deep scale level l low , the scale feature representation set Input cross-window sparse self-attention module, only establish connections between the preset part of query and its related keys, and output deep global coarse-grained representation S26. For each shallow scale level, the shallow local fine-grained representation is element-wise fused with the matching deep sparse global representation to form a fused encoding representation Z (l) ; S27. Fusion encoding representation of all scale levels Unified organization is carried out to construct a multi-scale sparse Transformer encoding feature pyramid with cross-hierarchical structure perception capabilities.
4. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 3, characterized in that: The S3 includes the following steps: S31. Construct a predefined medical partition template to form a brain region partition set R. The brain region partition set consists of multiple brain region template regions. In each scale level l, the fusion coding representation set Z (l) The spatial position of each cube block in the brain area template area is judged to coincide with the space, and the cube block is mapped to the graph node in the brain area graph structure through the mapping function to form the initial node mapping relationship; S32. At scale level l, generate an initial brain region map node set V based on the node mapping relationship (l) Each graph node in the initial brain region graph node set is used to represent a set of cube blocks of a brain region template at the current scale. For each graph node, the average value of all cube block encoding representations contained in the current node is calculated to obtain the initial feature vector of the current graph node. S33. Using the initial feature vector set of all graph nodes at scale level l, calculate the anatomical similarity between each pair of graph nodes and form an anatomical similarity matrix; S34. Extract the number of white matter fiber tracts between each pair of brain template regions based on the normalized diffusion tensor imaging data, calculate the normalized connectivity based on the fiber distribution density within and outside the region, and organize the functional connectivity results between all regions into a functional similarity matrix; S35. Construct a graph adjacency matrix at scale level l by linearly weighted fusion of the anatomical similarity matrix and the functional similarity matrix; S36. Combine the brain region graph node set at scale level l with the graph adjacency matrix to form the brain region graph structure G at the same scale level at the current scale level. (l) .
5. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 4, characterized in that: The S3 includes the following steps: S41. Set a cross-scale alignment threshold, select all scale levels with resolutions lower than the cross-scale alignment threshold to form a low-resolution set, and select all scale levels with resolutions higher than or equal to the cross-scale alignment threshold to form a high-resolution set; S42. For each low-resolution scale level of the brain area map structure, extract the spatial center coordinates of each map node and the initial feature vector of the map node; S43. At the high-resolution scale level, for all candidate graph nodes that have the same medical template name as the low-resolution graph node, calculate the spatial Euclidean distance between each candidate graph node and the low-resolution graph node, and select the candidate node with the smallest Euclidean distance as the high-resolution alignment target graph node; S44. Match each pair of low-resolution graph nodes with their nearest neighbor high-resolution alignment target graph nodes to form a cross-scale alignment mapping set; S45. For each pair of matching graph nodes in the cross-scale alignment mapping set, calculate feature similarity using the initial feature vectors of the two graph nodes; S46. All cross-scale matching graph node pairs and corresponding alignment weights are combined to form a cross-scale edge set; S47. Merge the low-resolution scale-level brain region map structure with the high-resolution scale-level brain region map structure and the cross-scale edge set to obtain the final cross-scale node-aligned fusion brain region map structure.
6. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 5, characterized in that: The S5 comprises the following steps: S51. Based on the mapping relationship between graph nodes and cubes in the cross-scale node alignment fusion brain region graph structure, the cube encoding features at each scale level in the multi-scale sparse Transformer encoding feature pyramid are mapped to the initial attribute features of the corresponding graph nodes; S52. Concatenate and integrate the initial attribute features of the graph nodes in all scale levels according to the node numbers to form an initial feature matrix of the cross-scale node set; S53. Input the initial feature matrix into the graph neural network message passing module, perform the first round of graph convolution operation in combination with the adjacency matrix of the brain region graph structure of cross-scale node alignment fusion, and obtain the first round of node embedding representation H (1) ; S54. Input the first round of graph convolution representation into the multi-head graph attention mechanism module, and integrate the output of each graph node obtained from all attention heads through the splicing operation to generate the first round of fused brain area graph node embedding feature h i (1) .
7. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 1, characterized in that: The S6 comprises the following steps: S61. According to the mapping relationship between the graph nodes and the original cube blocks, the first round of fusion brain area graph nodes are embedded into the features Project and broadcast back to the original cube blocks at each scale level in the multi-scale sparse Transformer encoding feature pyramid to obtain the graph feedback residual information corresponding to each cube block; S62. For each scale level l, the residual information R is fed back to the graph (1) The corresponding residual vector in The cubic block features output by the previous layer Transformer encoding Perform a joint update to form a new cubic input representation S63. Input the new cube into the representation Input to the next layer of multi-scale sparse Transformer encoder, perform local window dense self-attention and cross-window sparse self-attention mechanism, extract multi-scale features of the new input, and output the updated multi-scale sparse Transformer encoding feature pyramid { S64. After each round of updating, repeat the process of constructing the brain region map structure at the same scale in S3 to obtain a new brain region map structure; S65. Based on the new brain region map structure, repeat the cross-scale node alignment process in S4 to generate a new cross-scale node aligned fusion brain region map structure; S66. Execute graph neural network message passing and multi-head graph attention mechanism on the new cross-scale node alignment fusion brain region graph structure to generate the second round of fusion brain region graph node embedding features And feed it back to the encoder for residual update; S67. Repeat steps S61 to S66 to perform multiple rounds of Transformer-GNN dual-path interactive updates until the set number of iterations or the residual convergence threshold is reached, and finally a stable multi-scale sparse Transformer encoding feature pyramid and fused brain area map node embedding representation are obtained.
8. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 7, characterized in that: The S7 comprises the following steps: S71. Encoding feature pyramids based on stable multi-scale sparse Transformer Encoded features at all scale levels A skip connection fusion strategy is used to integrate information, and channel splicing and convolution fusion are performed with the current level encoding features to generate a fused cross-scale voxel feature representation; S72. Based on the mapping relationship between graph nodes and cube blocks, the embedded representation of the fused brain region graph nodes is projected back to the corresponding voxel block area via the nearest neighbor, and a graph feature field is established in the voxel space. The cross-scale voxel feature representation and the graph feature field are then concatenated in the channel dimension to form the final fused voxel-level segmentation feature. S73. The final fused voxel-level segmentation features are input into the segmentation head network composed of convolutional layers. The segmentation head network performs a set of 1×1×1 convolution operations and outputs voxel-level multi-category probability predictions through the softmax function to generate voxel-level segmentation probability volume data.
9. The method for dynamic segmentation of three-dimensional brain networks based on deep learning according to claim 8, characterized in that: The S8 specifically includes introducing adjacent time frame probability consistency constraints and deformation field smoothing regularization to the voxel-level segmentation probability volume data of continuous time frames in the training stage, applying recursive Bayesian filtering and B-spline-based deformation compensation to the voxel-level segmentation probability volume data in the inference stage, performing conditional random field refinement, flight time mapping-based hole filling and confidence evaluation on the voxel-level segmentation probability volume data that evolves smoothly in the time domain, and obtaining three-dimensional brain area dynamic segmentation results.
Citation Information
Patent Citations
Three-dimensional brain tumor segmentation model based on deformable feature aggregation
CN118967712A
Single-channel epileptic seizure early warning method based on deep double-cross metric learning
CN120093219A
Unified representation calculation method and apparatus for brain network, and electronic device and storage medium
WO2024119337A1
Automatic trajectory prediction method based on graph spatial-temporal pyramid
WO2024193334A1
Cited By
Controllable three-dimensional graphic content generation method and system based on multi-image fusion
CN121095470A
Image semantic segmentation method based on graph convolutional network
CN121305087A