Brain image processing method and apparatus

By extracting and fusion of fMRI and sMRI data, combined with lightweight models, the problems of high computational complexity and low accuracy in the prior art are solved, and efficient brain image processing and mental illness diagnosis are achieved.

WO2025148608A1PCT designated stage expired Publication Date: 2025-07-17BEIJING JINGDONG TUOXIAN TECH CO LTD

Patent Information

Application Number
PCT/CN2024/139069
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-12-13
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The existing brain image processing model is difficult to effectively apply due to the limited data set size, high computational complexity and calculation cost, and low image processing accuracy and accuracy, especially in the diagnosis of bipolar disorder.

Method used

The spatiotemporal feature extraction unit and specific convolutional layer are used to extract the feature of fMRI and sMRI data, combined with lightweight models for feature fusion and classification, and the double pyramid structure and grouped convolution are used to reduce the computational complexity and improve the feature extraction efficiency.

Benefits of technology

It improves the accuracy and accuracy of brain image processing, reduces calculation costs, and achieves the clinical level of image recognition and detection of brain mental illnesses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139069_17072025_PF_FP_ABST
    Figure CN2024139069_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a brain image processing method and apparatus. One specific implementation of the method comprises: using a spatiotemporal feature extraction unit to perform feature extraction on a brain magnetic resonance image of a first modality to obtain spatiotemporal features of a brain region; using a specific convolutional layer to perform feature extraction on a brain magnetic resonance image of a second modality to obtain structural features of the brain region, the specific convolutional layer having a convolution kernel of a set size and a convolution kernel stride; and performing feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification on the basis of the fused features to obtain an image processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Brain image processing method and device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese invention patent application No. 202410032783.5 filed on January 9, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the fields of medical image analysis and computer vision technology, and in particular to a method and device for processing brain images. Background Art

[0004] In recent years, magnetic resonance imaging (MRI) has become crucial for slowing the progression of symptoms of psychiatric illnesses and improving patients' quality of life. Functional Magnetic Resonance Imaging (fMRI) and structural Magnetic Resonance Imaging (sMRI) are commonly used medical imaging modalities for diagnosing psychiatric brain disorders and have demonstrated an important role in the latest diagnosis of brain diseases. Taking the diagnosis of bipolar disorder (BD) as an example, the application of fMRI and sMRI technology can help us gain a deeper understanding of the neurobiological mechanisms of BD, provide more objective and accurate BD diagnostic methods, and offer new ideas for the treatment and prevention of BD.

[0005] In the process of implementing the present disclosure, the inventors found that the related art generally performs brain image processing by processing fMRI data or sMRI data. However, due to the limited size of BD diagnostic data sets, the computational complexity and computational cost of existing brain image processing models are high, and the image processing precision and accuracy are low. Summary of the Invention

[0006] In view of this, embodiments of the present disclosure provide a method and apparatus for processing brain images.

[0007] To achieve the above objectives, according to one aspect of an embodiment of the present disclosure, a method for processing a brain image is provided, comprising:

[0008] Using a spatiotemporal feature extraction unit to extract features from the brain magnetic resonance image of the first modality to obtain spatiotemporal features of the brain region;

[0009] Performing feature extraction on the brain magnetic resonance image of the second modality through a specific convolutional layer to obtain structural features of the brain region, wherein the specific convolutional layer has a convolution kernel of a set size and a convolution kernel movement step size;

[0010] The spatiotemporal features of the brain region and the structural features of the brain region are subjected to feature fusion to obtain fusion features, and feature classification is performed based on the fusion features to obtain an image processing result.

[0011] Optionally, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging; and the brain magnetic resonance image of the second modality is T1-weighted structural magnetic resonance imaging.

[0012] Optionally, the spatiotemporal feature extraction unit includes a spatial feature extraction module and a temporal feature extraction module connected in series, the spatial feature extraction module includes a fully connected layer and a batch normalization layer, the receptive field of the fully connected layer is the same size as the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporal feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, wherein the receptive field is the range of perception of neurons in a neural network to an image.

[0013] Optionally, the spatiotemporal features of the brain region are obtained by using a spatiotemporal feature extraction unit to perform feature extraction on the brain magnetic resonance image of the first modality, including: using the spatial feature extraction module to perform feature extraction on the brain magnetic resonance image of the first modality to obtain spatial features; using the temporal feature extraction module to perform feature extraction on the spatial features to obtain temporal features; and obtaining the spatiotemporal features of the brain region based on the spatial features and the temporal features.

[0014] Optionally, the time feature extraction module is used to perform feature extraction on the spatial feature to obtain the time feature, including: setting the convolution kernel moving step of the convolution layer to 2, and using the time feature extraction module to perform time feature extraction on the spatial feature to obtain the time feature.

[0015] Optionally, the spatiotemporal feature extraction unit further includes an original feature extraction module, and the spatiotemporal feature extraction unit utilizes the residual structure of a neural network and combines the original feature extraction module, the spatial feature extraction module and the temporal feature extraction module to perform feature extraction.

[0016] Optionally, the specific convolutional layer includes at least one group of convolutional channels; and feature extraction is performed on the brain magnetic resonance image of the second modality through the specific convolutional layer to obtain structural features of the brain region, including: for the brain magnetic resonance image of the second modality, feature extraction is performed using at least one group of convolutional channels included in the specific convolutional layer according to the set dimension of each group of convolutional channels to obtain structural features of the brain region.

[0017] Optionally, the spatiotemporal features and the structural features of the brain region are subjected to feature fusion to obtain fused features, including: performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of specified dimensions; and using an array connection function to concatenate the intermediate structural features and the spatiotemporal features of the brain region to perform feature fusion to obtain fused features.

[0018] Optionally, performing feature classification based on the fused features includes: performing feature classification on the fused features using a linear classifier.

[0019] According to another aspect of an embodiment of the present disclosure, a brain image processing apparatus is provided, comprising:

[0020] a spatiotemporal feature extraction module, configured to extract features from the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit to obtain spatiotemporal features of the brain region;

[0021] a structural feature extraction module, configured to extract features from the second modality brain magnetic resonance image through a specific convolutional layer to obtain structural features of the brain region, wherein the specific convolutional layer has a convolution kernel of a set size and a convolution kernel movement step size;

[0022] The feature fusion classification module is used to perform feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fusion features, and perform feature classification based on the fusion features to obtain image processing results.

[0023] Optionally, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging; and the brain magnetic resonance image of the second modality is T1-weighted structural magnetic resonance imaging.

[0024] Optionally, the spatiotemporal feature extraction unit includes a spatial feature extraction module and a temporal feature extraction module connected in series, the spatial feature extraction module includes a fully connected layer and a batch normalization layer, the receptive field of the fully connected layer is the same size as the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporal feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, wherein the receptive field is the range of perception of neurons in a neural network to an image.

[0025] Optionally, the spatiotemporal feature extraction module is further used to: use the spatial feature extraction module to perform feature extraction on the brain magnetic resonance image of the first modality to obtain spatial features; use the temporal feature extraction module to perform feature extraction on the spatial features to obtain temporal features; and obtain the spatiotemporal features of the brain region based on the spatial features and the temporal features.

[0026] Optionally, when the spatiotemporal feature extraction module uses the time feature extraction module to extract the spatial features to obtain the time features, the spatiotemporal feature extraction module is also used to: set the convolution kernel moving step of the convolution layer to 2, and use the time feature extraction module to extract the time features of the spatial features to obtain the time features.

[0027] Optionally, the spatiotemporal feature extraction unit further includes an original feature extraction module, and the spatiotemporal feature extraction unit utilizes the residual structure of a neural network and combines the original feature extraction module, the spatial feature extraction module and the temporal feature extraction module to perform feature extraction.

[0028] Optionally, the specific convolutional layer includes at least one set of convolutional channels; the structural feature extraction module is also used to: for the brain magnetic resonance image of the second modality, use at least one set of convolutional channels included in the specific convolutional layer to perform feature extraction according to the set dimensions of each set of convolutional channels to obtain structural features of the brain area.

[0029] Optionally, the feature fusion classification module is also used to: perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of specified dimensions; use array connection functions to splice the intermediate structural features and the spatiotemporal features of the brain region to perform feature fusion to obtain fused features.

[0030] Optionally, the feature fusion classification module is further used to: use a linear classifier to perform feature classification on the fused features.

[0031] According to yet another aspect of the embodiments of the present disclosure, an electronic device is provided.

[0032] An electronic device comprises: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the brain image processing method provided by the embodiment of the present disclosure.

[0033] According to yet another aspect of the embodiments of the present disclosure, a computer-readable medium is provided.

[0034] A computer-readable medium stores a computer program, which, when executed by a processor, implements the brain image processing method provided by an embodiment of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The accompanying drawings are used to better understand the present disclosure and do not constitute an improper limitation of the present disclosure.

[0036] FIG1 is a schematic diagram of the main steps of a method for processing a brain image according to an embodiment of the present disclosure;

[0037] FIG2 is a network structure diagram of a spatiotemporal feature extraction unit according to an embodiment of the present disclosure;

[0038] FIG3 is a network structure diagram of a structural feature extraction unit according to an embodiment of the present disclosure;

[0039] FIG4 is a schematic diagram of a brain image processing system according to an embodiment of the present disclosure;

[0040] FIG5 is a schematic diagram of main modules of a brain image processing apparatus according to an embodiment of the present disclosure;

[0041] FIG6 is a diagram of an exemplary system architecture in which embodiments of the present disclosure may be applied;

[0042] FIG7 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0044] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments.

[0045] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0046] In order to solve the problems existing in the related art, the present disclosure provides a method for processing brain images, which combines computer vision technology and deep learning technology to process data of fMRI and sMRI modalities respectively, and uses deep learning methods to fully extract the features of brain magnetic resonance images, and fuse the feature data of the two modalities to improve the precision of image processing results and image processing accuracy. In addition, the present disclosure also proposes a lightweight model based on convolution, which improves the performance of the model while reducing the computational cost and complexity of the model. In the application scenario of diagnosing brain mental illness, the image recognition detection accuracy of brain mental illness can reach clinical levels.

[0047] In the introduction to the embodiments of the present disclosure, the professional terms and their definitions are as follows:

[0048] Flatten: In deep learning, flatten is used to reduce multi-dimensional data to one dimension, usually used for conversion between fully connected layers and convolutional layers;

[0049] concat: used to connect two or more arrays along a specified axis to form a new array;

[0050] Stride: In convolutional neural networks, stride refers to the length of the steps the filter moves across the input data. For example, a stride of 2 means the filter moves two units at a time.

[0051] groups: In the convolution operation, the groups parameter is used to specify the connection mode between the input and output, which means dividing the channels into several groups for convolution operation. One is the depth-wise convolution operation (the number of groups is equal to the number of input channels), and the other is the conventional convolution operation (the number of groups is equal to 1);

[0052] BN: short for Batch Normalization, a technique commonly used in deep learning networks to accelerate training and prevent overfitting. It can keep the input of each layer in the network in the same distribution;

[0053] ReLU: This is a commonly used activation function in deep learning. It resets negative values ​​to 0 and leaves positive values ​​unchanged. This increases the nonlinearity and sparsity of the model, helping to alleviate the vanishing gradient problem.

[0054] Transformer: Transformer is a model framework widely used in natural language processing. It is based on the self-attention mechanism and effectively handles long-distance dependency problems.

[0055] Receptive Field: This refers to the area of ​​the input image where pixels on the feature map output by each layer of a convolutional neural network are mapped back to the input image. In simpler terms, this means that the area of ​​the input image that a convolutional neural network feature can see relative to the original image is the size of a point on the feature map.

[0056] The problem of image recognition and diagnosis of brain psychiatric disorders based on sMRI and fMRI can be viewed as an image classification problem. Deep learning methods for image classification tasks can also be used for image recognition and diagnosis of brain psychiatric disorders. Convolutional neural networks are typically composed of convolutional layers, pooling layers, and batch normalization layers. The convolutional layer uses discrete convolution functions to calculate local features in the image based on the image's translation invariance and extracts rich high-level semantic information from the image; the batch normalization layer normalizes the features in a small batch, which can accelerate the convergence of the model and make the training process more stable. It plays a significant role in overcoming the problems of gradient explosion and gradient vanishing; the pooling layer aggregates the features within the window through sampling, which can reduce the number of model parameters, increase the receptive field, and alleviate overfitting.

[0057] Existing Transformer-based methods for brain image processing have high computational complexity, resulting in prohibitive computational and time costs when processing sMRI data. Transformer-based methods leverage their large receptive field to learn adaptive global information, demonstrating superior performance in many downstream tasks in computer vision. However, their large number of parameters requires extensive data for optimization, and collecting sMRI and fMRI data is difficult. Therefore, this paper proposes a convolution-based approach to construct a model for brain image processing.

[0058] Furthermore, related art generally performs brain image processing based solely on sMRI or fMRI data. To more fully utilize the features of brain magnetic resonance images, the present disclosure processes data from both fMRI and sMRI modalities separately, and fuses the feature data from both modalities to improve the precision and accuracy of the image processing results. Therefore, the present disclosure aims to address the issues of insufficient intra-modal feature extraction and incomplete inter-modal feature fusion in existing methods for multimodal data. In an embodiment of the present disclosure, a dual-pyramid structure (i.e., a model learning model using two pyramid-shaped structures; the pyramid structure is considered an effective model structure) can be used to construct an encoder. A hierarchical strategy is used to gradually expand the receptive field and extract features at different scales. For sMRI data, the present disclosure proposes a window pyramid network that divides the sMRI data into windows and performs feature learning on each window to extract the structural features of the brain regions included in the sMRI data. For fMRI data, a spatiotemporal pyramid structure is proposed that simultaneously learns both the temporal and spatial features of the fMRI data. The extracted spatial and temporal features are then fused to obtain the spatiotemporal features of the fMRI data. Finally, the structural features and spatiotemporal features extracted from the data of the two modalities are fused, and the fused features are output through the classifier to obtain the image classification results to obtain the brain image processing results.

[0059] FIG1 is a schematic diagram of the main steps of the brain image processing method according to an embodiment of the present disclosure. As shown in FIG1 , the brain image processing method according to the embodiment of the present disclosure mainly includes the following steps S101 to S103.

[0060] Step S101: extracting features from the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit to obtain spatiotemporal features of the brain region.

[0061] In an embodiment of the present disclosure, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging, namely, rs-fMRI data.

[0062] Because fMRI (functional magnetic resonance imaging) data is a continuous sMRI (structural magnetic resonance imaging) data stream with a large data volume, directly processing the features of fMRI data requires a significant computational cost. Related techniques typically extract a certain number of brain regions from fMRI data and then derive the functional connectivity matrix between these regions. By analyzing the functional connectivity matrix, brain region images can be processed and analyzed. When applied to brain disease diagnosis scenarios, a diagnosis of the patient's corresponding disease can be made. To efficiently process fMRI data, the present disclosure does not directly process the fMRI data, but instead analyzes the features of the brain regions corresponding to the fMRI data. Therefore, the present disclosure does not involve preprocessing the functional connectivity matrix, but directly extracts brain region features from the fMRI data. To better utilize the temporal dimension information contained in the time series, the present disclosure uses a serially connected spatial feature extraction module and a temporal feature extraction module as a spatiotemporal feature extraction unit for rs-fMRI, extracting spatial features and temporal features, respectively.

[0063] According to one embodiment of the present disclosure, the spatiotemporal feature extraction unit includes a spatial feature extraction module and a temporal feature extraction module connected in series. The spatial feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same size as the brain magnetic resonance image of the first modality, and the receptive field in the temporal dimension is 1. The temporal feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1. The receptive field is the range of receptive perception of neurons in a neural network to an image. The batch normalization layer is used to avoid errors caused by data scale and other factors and improve learning efficiency. Typically, the batch normalization layer and the convolutional layer are integrated as a whole and displayed only as the convolutional layer. When the receptive field in a certain dimension is 1, it means that features in the corresponding dimension are not extracted. The receptive field of the fully connected layer is the same size as the input data (the brain magnetic resonance image of the first modality), which can fully learn the global features of the fMRI data. For features in the temporal dimension, the convolutional layer is directly used to extract temporal features.

[0064] According to one embodiment of the present disclosure, using a spatiotemporal feature extraction unit to extract features from a first modality brain magnetic resonance image to obtain spatiotemporal features of a brain region may specifically include: using the spatial feature extraction module to extract features from the first modality brain magnetic resonance image to obtain spatial features; using the temporal feature extraction module to extract features from the spatial features to obtain temporal features; and obtaining spatiotemporal features of the brain region based on the spatial features and the temporal features. When using the spatiotemporal feature extraction unit to extract spatiotemporal features from the first modality brain magnetic resonance image, the spatial feature extraction module and the temporal feature extraction module are used sequentially to perform spatial feature extraction and temporal feature extraction, thereby obtaining spatiotemporal features.

[0065] According to another embodiment of the present disclosure, the temporal feature extraction module is used to extract the spatial features to obtain the temporal features, including: setting the convolution kernel movement step of the convolution layer to 2, and using the temporal feature extraction module to extract the temporal features of the spatial features to obtain the temporal features. The convolution kernel movement step refers to the step size of each movement of the convolution kernel in the convolution operation. By setting the convolution kernel movement step of the second convolution layer in the temporal feature extraction module to 2, the temporal features can be downsampled while performing temporal feature extraction, thereby gradually increasing the receptive field of the model in the temporal dimension.

[0066] According to another embodiment of the present disclosure, the spatiotemporal feature extraction unit further includes an original feature extraction module, and the spatiotemporal feature extraction unit utilizes the residual structure of a neural network and combines the original feature extraction module, the spatial feature extraction module, and the temporal feature extraction module to perform feature extraction. The original feature extraction module, for example, is extracted by performing convolution processing on the brain magnetic resonance image of the first modality. The original feature extraction module, for example, is a 2D convolution layer, and the convolution kernel is, for example, 1×1.

[0067] Figure 2 is a network structure diagram of a spatiotemporal feature extraction unit of an embodiment of the present disclosure. As shown in Figure 2, in the embodiment of the present disclosure, the original feature extraction module is omitted. The first box shows the spatial feature extraction module, in which "F×1, 1, Stride=(1, 1)" is shown. "F×1" indicates that its convolution kernel size is F×1, where F is the number of dimensions of the spatial feature; "1" represents a feature point, and "Stride=(1, 1)" indicates that the step size in both the spatial dimension and the temporal dimension is 1. The second box shows the temporal feature extraction module, in which "1×3, 1, Stride=(1, 2)" is shown. "1×3" is its convolution kernel size, "1" represents a feature point, and "Stride=(1, 2)" indicates that the step size in the spatial dimension is 1 and the step size in the temporal dimension is 2. In the network structure of the spatiotemporal feature extraction unit, a residual structure is added to further improve the accuracy of the model.

[0068] Assume that the preprocessed fMRI dataset P = {(x i1 ,…,x iN )|i=1,…,M}, where M represents the number of fMRI data frames, N represents the number of brain regions, and x refers to the features of each point. Therefore, the embodiments of the present disclosure differ from traditional approaches in their network structure design. Instead of gradually reducing image size and increasing feature dimensionality for feature extraction, they maintain the number of brain regions N constant as the network depth increases. This reduces the number of model parameters, lowers the computational complexity of the model, and avoids redundant features.

[0069] Step S102: The brain magnetic resonance image of the second modality is subjected to feature extraction through a specific convolution layer to obtain the structural features of the brain region, wherein the specific convolution layer has a convolution kernel of a set size and a convolution kernel moving step size. In the embodiment of the present disclosure, the brain magnetic resonance image of the second modality is a T1-weighted structural magnetic resonance imaging (abbreviated as T1w). The principle of T1-weighted imaging is based on the difference in the magnetic resonance relaxation time T1 of human tissue. T1 refers to the time it takes for the magnetic resonance signal to change from a high energy state to a low energy state, that is, the time it takes for the magnetic resonance signal to change from an excited state to a ground state.

[0070] sMRI data consists of voxels, resulting in a relatively large data size. To address the high computational complexity of feature extraction modules in existing methods, this disclosure proposes a novel Patch Pyramid Feature Extraction Module (P2FEM). P2FEM is built based on convolution operations. Existing convolutional neural networks consist of convolutional layers, pooling layers, and batch normalization layers. The pooling layer is used to reduce the number of model parameters, increase the receptive field, and alleviate overfitting. However, performing convolution followed by pooling introduces certain redundancy. To improve the computational efficiency of the model, this disclosure removes the pooling layer and adjusts the convolution layer. First, a large convolution kernel (for example, a 7×7×7 three-dimensional convolution kernel in the embodiments of this disclosure) is used to expand the model's receptive field and enhance its learning ability. Second, to reduce the model's computational complexity and alleviate overfitting, the convolution kernel's step size is set. This step size reduces the number of convolution calculations, completing the downsampling process while reducing computational complexity. In the embodiments of the present disclosure, by using large convolution kernels and setting step sizes, the convolution layer has the same function as the pooling layer. In addition, to further reduce the computational complexity of the model and alleviate overfitting, in the embodiments of the present disclosure, grouped convolution is used to replace the traditional convolution layer.

[0071] According to one embodiment of the present disclosure, the specific convolution layer includes no less than one group of convolution channels; the structural features of the brain region are obtained by extracting features from the brain magnetic resonance image of the second modality through the specific convolution layer, including: for the brain magnetic resonance image of the second modality, according to the dimension of each set group of convolution channels, respectively using no less than one group of convolution channels included in the specific convolution layer to extract features, so as to obtain the structural features of the brain region. Specifically, assuming that a data contains features of multiple dimensions, grouped convolution refers to grouping the feature dimensions, and then using the corresponding convolution channels to extract features for each group of features. For example, after an 8-dimensional data is divided into two groups, the dimension of each group of features is 4, and then the corresponding convolution channels are used to perform convolution respectively to implement grouped convolution, thereby reducing the redundancy of features and reducing the computational complexity of convolution.

[0072] FIG3 is a network structure diagram of a structural feature extraction unit of an embodiment of the present disclosure. As shown in FIG3 , in an embodiment of the present disclosure, for T1-weighted structural magnetic resonance imaging (abbreviated as T1w), it is described as (H, W, D, C), where W, H, and D represent the width, height, and depth of the data, respectively, and C is the dimension or number of channels of the feature. Among them, "7×7×7, C2, Stride=2, Groups=4" is a specific convolutional layer, "7×7×7" represents the convolution kernel, "C2" refers to the output dimension value of C2, "Stride=2" means that the stride in the three dimensions of width, height, and depth of the data is 2, and "Groups=4" means that there are 4 groups of convolution channels in total. Accordingly, the meaning of the specific convolutional layer "5×5×5, C3, Stride=2, Groups=4" can be understood.

[0073] Compared to traditional convolutional neural networks, the structural feature extraction unit of the disclosed embodiment has higher computational efficiency and fewer parameters, and the extracted features have stronger generalization capabilities. Due to the large size of sMRI data, each additional convolutional layer significantly increases the computational complexity of the model. To minimize the computational cost of the model, only one convolutional layer is used for feature extraction at each level.

[0074] Step S103: Fusing the spatiotemporal features and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result. Before performing feature classification, the features may also be subjected to dimensionality reduction, feature fusion, and other processing to facilitate more accurate feature classification.

[0075] According to one embodiment of the present disclosure, the spatiotemporal features of the brain region and the structural features of the brain region are subjected to feature fusion to obtain fused features, which may specifically include: performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; and using an array concatenation function to perform feature fusion processing on the intermediate structural features and the spatiotemporal features of the brain region to obtain fused features. According to an embodiment of the present disclosure, feature classification based on the fused features may specifically include: using a linear classifier to perform feature classification on the fused features.

[0076] When fusing features from multimodal data, the flatten function (used for data dimensionality reduction) and a fully connected neural network can first be used to process the structural features into intermediate structural features of the target size. Then, the concat function is used to concatenate the intermediate structural features with the spatiotemporal features to obtain the fused features corresponding to the multimodal data. Finally, the fused features are input into a linear classifier for feature classification, resulting in a predicted probability value. In machine learning, the goal of classification is to cluster objects with similar features. A linear classifier achieves this by making classification decisions based on linear combinations of features.

[0077] FIG4 is a schematic diagram of the principle architecture of a brain image processing system according to an embodiment of the present disclosure. As shown in FIG4 , in an embodiment of the present disclosure, the 4-dimensional data of the first modality brain magnetic resonance image rs-fMRI and the 3-dimensional data of the second modality brain magnetic resonance image T1-weighted sMRI (abbreviated as T1w) are respectively input into two encoder branches for feature extraction. For rs-fMRI data, a plurality of spatiotemporal feature extraction units composed of a spatial feature extraction module and a temporal feature extraction module are used to perform spatiotemporal feature extraction to obtain spatiotemporal features of multiple brain regions. Among them, the spatiotemporal feature extraction unit mainly includes a residual structure and a 2D convolution batch normalization layer (the spatial feature extraction module and the temporal feature extraction module are both 2D convolution batch normalization layers), and by setting the step size corresponding to the convolution layer in the temporal feature extraction module, the temporal features can be downsampled while extracting the temporal features, thereby gradually expanding the receptive field of the model in the temporal dimension.

[0078] For T1w data, a specialized convolutional layer consisting of multiple 3D convolutional batch normalization layers is used to perform 3D convolution to extract structural features from multiple brain regions. By using grouped convolution, this structural feature extraction unit achieves higher computational efficiency and fewer parameters, and the extracted features have stronger generalization capabilities.

[0079] After extracting the spatiotemporal and structural features of each brain region, these features can be fused and classified. During feature fusion, the flatten function (used for data dimensionality reduction) and a fully connected neural network can be used to process the structural features into the target size as intermediate structural features. The concat function is then used to concatenate the intermediate structural features with the spatiotemporal features of the brain region to obtain fused features corresponding to the multimodal data. Finally, the fused features are input into a linear classifier for feature classification, resulting in a predicted probability value.

[0080] FIG5 is a schematic diagram of the main modules of the brain image processing apparatus according to an embodiment of the present disclosure. As shown in FIG5 , the brain image processing apparatus 500 according to the present disclosure mainly includes a spatiotemporal feature extraction module 501 , a structural feature extraction module 502 , and a feature fusion classification module 503 .

[0081] The spatiotemporal feature extraction module 501 is configured to extract the spatiotemporal features of the brain region from the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit;

[0082] a structural feature extraction module 502 for performing feature extraction on the brain magnetic resonance image of the second modality through a specific convolutional layer to obtain structural features of the brain region, wherein the specific convolutional layer has a convolution kernel of a set size and a convolution kernel shift step size;

[0083] The feature fusion classification module 503 is used to perform feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fused features, and perform feature classification based on the fused features to obtain image processing results.

[0084] According to one embodiment of the present disclosure, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging; and the brain magnetic resonance image of the second modality is T1-weighted structural magnetic resonance imaging.

[0085] According to another embodiment of the present disclosure, the spatiotemporal feature extraction unit includes a spatial feature extraction module and a temporal feature extraction module connected in series, the spatial feature extraction module includes a fully connected layer and a batch normalization layer, the receptive field of the fully connected layer is the same size as the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporal feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, wherein the receptive field is the range of perception of neurons in a neural network to an image.

[0086] According to another embodiment of the present disclosure, the spatiotemporal feature extraction module 501 can also be used to: use the spatial feature extraction module to perform feature extraction on the brain magnetic resonance image of the first modality to obtain spatial features; use the temporal feature extraction module to perform feature extraction on the spatial features to obtain temporal features; and obtain the spatiotemporal features of the brain region based on the spatial features and the temporal features.

[0087] According to another embodiment of the present disclosure, when the spatiotemporal feature extraction module 501 uses the time feature extraction module to extract the spatial features to obtain the time features, it can also be used to: set the convolution kernel moving step of the convolution layer to 2, and use the time feature extraction module to extract the time features of the spatial features to obtain the time features.

[0088] According to another embodiment of the present disclosure, the spatiotemporal feature extraction unit also includes an original feature extraction module, and the spatiotemporal feature extraction unit utilizes the residual structure of the neural network and combines the original feature extraction module, the spatial feature extraction module and the temporal feature extraction module to perform feature extraction.

[0089] According to another embodiment of the present disclosure, the specific convolutional layer includes at least one group of convolutional channels; the structural feature extraction module 502 can also be used to: for the brain magnetic resonance image of the second modality, use the at least one group of convolutional channels included in the specific convolutional layer to perform feature extraction according to the set dimensions of each group of convolutional channels, so as to obtain the structural features of the brain area.

[0090] According to another embodiment of the present disclosure, the feature fusion classification module 503 can also be used to: perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of specified dimensions; use array connection functions to splice the intermediate structural features and the spatiotemporal features of the brain region to perform feature fusion to obtain fused features.

[0091] According to yet another embodiment of the present disclosure, the feature fusion classification module 503 may also be configured to: perform feature classification on the fused features using a linear classifier.

[0092] According to the technical solution of the embodiment of the present disclosure, the spatiotemporal features of the brain region are obtained by extracting features from the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit; the structural features of the brain region are obtained by extracting features from the brain magnetic resonance image of the second modality through a specific convolution layer, and the specific convolution layer has a convolution kernel of a set size and a convolution kernel moving step size; the spatiotemporal features and the structural features of the brain region are fused to obtain fused features, and feature classification is performed based on the fused features to obtain the image processing results. The technical solution can adopt a dual-branch processing method to process the brain magnetic resonance image data of two different modalities respectively, and use the fused features for feature classification. It can make full use of the features of the brain magnetic resonance image data for feature classification and image processing, thereby improving the accuracy of the image processing results and the accuracy of image processing. In addition, the present disclosure extracts structural features through a lightweight model based on convolution, which improves the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of diagnosing brain mental illness, the image recognition detection accuracy of brain mental illness can reach the clinical level.

[0093] FIG6 shows an exemplary system architecture 600 to which the brain image processing method or brain image processing apparatus according to the embodiments of the present disclosure may be applied.

[0094] As shown in Figure 6, system architecture 600 may include terminal devices 601, 602, and 603, a network 604, and a server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, and 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0095] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0096] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0097] The server 605 may be a server that provides various services, such as a background management server that supports websites browsed by users using the terminal devices 601, 602, and 603 (for example only). The background management server may perform feature extraction on the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit to obtain spatiotemporal features of the brain region; perform feature extraction on the brain magnetic resonance image of the second modality using a specific convolution layer to obtain structural features of the brain region, wherein the specific convolution layer has a convolution kernel of a set size and a convolution kernel movement step; perform feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fused features, perform feature classification based on the fused features to obtain image processing results, and feed back the processing results (such as feature classification results, image processing results - for example only) to the terminal device.

[0098] It should be noted that the brain image processing method provided in the embodiment of the present disclosure is generally executed by the server 605 , and accordingly, the brain image processing device is generally provided in the server 605 .

[0099] It should be understood that the number of terminal devices, networks and servers in Figure 6 is merely illustrative and any number of terminal devices, networks and servers may be provided as required.

[0100] 7, which shows a schematic diagram of a computer system 700 suitable for implementing a terminal device or server according to an embodiment of the present disclosure. The terminal device or server shown in FIG7 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0101] As shown in FIG7 , a computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the system 700 are also stored in the RAM 703. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0102] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read therefrom can be installed into the storage section 708 as needed.

[0103] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from a removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-mentioned functions defined in the system of the present disclosure are performed.

[0104] It should be noted that the computer-readable medium described in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0106] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units or modules may also be provided in a processor. For example, they may be described as: a processor including a spatiotemporal feature extraction module, a structural feature extraction module, and a feature fusion classification module. In some cases, the names of these units or modules do not constitute a limitation on the units or modules themselves. For example, the feature fusion classification module may also be described as "a module for performing feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fusion features, and performing feature classification based on the fusion features to obtain image processing results."

[0107] As another aspect, the present disclosure further provides a computer-readable medium, which may be included in the device described in the above embodiment; or may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device includes: using a spatiotemporal feature extraction unit to extract features from a brain magnetic resonance image of a first modality to obtain spatiotemporal features of a brain region; performing feature extraction on a brain magnetic resonance image of a second modality through a specific convolutional layer to obtain structural features of the brain region, wherein the specific convolutional layer has a convolution kernel of a set size and a convolution kernel movement step; performing feature fusion on the spatiotemporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain image processing results.

[0108] According to the technical solution of the embodiment of the present disclosure, the spatiotemporal features of the brain region are obtained by extracting features from the brain magnetic resonance image of the first modality using a spatiotemporal feature extraction unit; the structural features of the brain region are obtained by extracting features from the brain magnetic resonance image of the second modality through a specific convolution layer, and the specific convolution layer has a convolution kernel of a set size and a convolution kernel moving step size; the spatiotemporal features and the structural features of the brain region are fused to obtain fused features, and feature classification is performed based on the fused features to obtain the technical solution of the image processing result. A dual-branch processing method can be used to process the brain magnetic resonance image data of two different modalities respectively, and the fused features are used for feature classification. The features of the brain magnetic resonance image data can be fully utilized for feature classification and image processing, thereby improving the accuracy of the image processing results and the accuracy of image processing. In addition, the present disclosure extracts structural features through a lightweight model based on convolution, which improves the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of diagnosing brain mental illness, the recognition and detection accuracy of brain mental illness images can reach clinical levels.

[0109] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for processing brain images, comprising: Performing feature extraction on a brain magnetic resonance image of the first modality using a spatio-temporal feature extraction unit to obtain spatio-temporal features of a brain region; Performing feature extraction on a brain magnetic resonance image of the second modality through a specific convolutional layer to obtain structural features of the brain region, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; Performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result.

2. The method according to claim 1, wherein The brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging; the brain magnetic resonance image of the second modality is structural magnetic resonance imaging after T1 weighting.

3. The method according to claim 1, wherein, The spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1; where the receptive field is the range of image perception by neurons in a neural network.

4. The method according to claim 3, wherein, Performing feature extraction on a brain magnetic resonance image of the first modality using a spatio-temporal feature extraction unit to obtain spatio-temporal features of a brain region, including: Performing feature extraction on the brain magnetic resonance image of the first modality using the spatially connected feature extraction module to obtain spatial features; Performing feature extraction on the spatial features using the temporally connected feature extraction module to obtain temporal features; Obtaining spatio-temporal features of the brain region according to the spatial features and the temporal features.

5. The method according to claim 4, wherein, Performing feature extraction on the spatial features using the temporally connected feature extraction module to obtain temporal features, including: By setting the convolutional kernel moving step of the convolutional layer to 2, performing temporal feature extraction on the spatial features using the temporally connected feature extraction module to obtain the temporal features.

6. The method according to claim 3, wherein, The spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of a neural network to combine the original feature extraction module, the spatially connected feature extraction module, and the temporally connected feature extraction module for feature extraction.

7. The method according to claim 1, wherein The specific convolutional layer includes not less than one set of convolutional channels; Performing feature extraction on a brain magnetic resonance image of the second modality through a specific convolutional layer to obtain structural features of the brain region, including: For the brain magnetic resonance image of the second modality, respectively performing feature extraction on the basis of the dimensions of each set of convolutional channels set using not less than one set of convolutional channels included in the specific convolutional layer to obtain structural features of the brain region.

8. The method according to claim 1, wherein Performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, including: Performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; Using an array connection function to splice and process the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fused features.

9. The method according to claim 1 or 8, wherein, Performing feature classification based on the fused features, including: Performing feature classification on the fused features using a linear classifier.

10. A processing device for brain images, comprising: A spatio-temporal feature extraction module, configured to use a spatio-temporal feature extraction unit to extract features from a brain magnetic resonance image of a first modality to obtain spatio-temporal features of a brain region; A structural feature extraction module, configured to extract structural features of a brain region from a brain magnetic resonance image of a second modality through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; A feature fusion and classification module, configured to perform feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and perform feature classification based on the fused features to obtain an image processing result.

11. The apparatus according to claim 10, wherein, The brain magnetic resonance image of the first modality is a resting-state functional magnetic resonance imaging; the brain magnetic resonance image of the second modality is a T1-weighted structural magnetic resonance imaging.

12. The device according to claim 10, wherein, The spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, where the receptive field is the range of perception of neurons in a neural network for an image.

13. The apparatus according to claim 12, wherein, The spatio-temporal feature extraction module is further configured to: Use the spatially connected feature extraction module to extract spatial features from the brain magnetic resonance image of the first modality; Use the temporally connected feature extraction module to extract temporal features from the spatial features; Obtain spatio-temporal features of a brain region according to the spatial features and the temporal features.

14. The device according to claim 13, wherein, When the spatio-temporal feature extraction module uses the temporally connected feature extraction module to extract temporal features from the spatial features, it is further configured to: By setting the convolutional kernel moving step of the convolutional layer to 2, use the temporally connected feature extraction module to perform temporal feature extraction on the spatial features to obtain the temporal features.

15. The apparatus according to claim 12, wherein, The spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of a neural network to combine the original feature extraction module, the spatially connected feature extraction module, and the temporally connected feature extraction module for feature extraction.

16. The device according to claim 10, wherein The specific convolutional layer includes not less than one set of convolutional channels; The structural feature extraction module is further configured to: For the brain magnetic resonance image of the second modality, respectively use not less than one set of convolutional channels included in the specific convolutional layer according to the dimensions of each set of convolutional channels set to extract features, so as to obtain structural features of a brain region.

17. The apparatus according to claim 10, wherein The feature fusion and classification module is further configured to: Perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; Use an array connection function to splice and process the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fused features.

18. The apparatus according to claim 10 or 17, wherein, The feature fusion and classification module is further configured to: Use a linear classifier to perform feature classification on the fused features.

19. An electronic device, comprising: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, such that the one or more processors implement the method according to any one of claims 1-9.

20. A computer-readable medium having stored thereon a computer program, which when executed by a processor implements the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • SMRI image classification method and device based on multi-input convolutional neural network

    CN110992351A

  • Brain map classification method based on deep multi-modal graph convolution

    CN113592836A

  • Cranial nerve image feature extraction method based on scale equalization coupling convolution architecture

    CN115409843A

  • Brain network classification method based on space-time diagram convolution

    CN115496953A

  • Multi-modal brain network calculation method and device, equipment and storage medium

    CN115775626A

Cited By

  • Space-time fusion neural network construction method and system

    CN121031680A

  • Magnetic resonance image classification method and system based on double-branch network

    CN121033552A