Brain image processing method and device

By performing dual-branch processing of the FMR and structural magnetic resonance imaging data, extracting space-time and structural features and fusion classification, the problems of high computational complexity and low accuracy in the prior art are solved, and efficient brain mental illness image recognition is achieved.

CN120298736APending Publication Date: 2025-07-11BEIJING JINGDONG TUOXIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410032783.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The brain image processing model has high computational complexity and low image processing accuracy and accuracy in the prior art. Especially in the diagnosis of bipolar disorder disease, the limited data set size leads to high computational cost and insufficient accuracy of the model.

Method used

The data of fMRI and structural magnetic resonance imaging are respectively extracted by the dual-branch processing method. The spatiotemporal and structural magnetic resonance imaging are extracted using spatiotemporal and spatial feature extraction units and specific convolutional layers, and the spatiotemporal and structural features of the brain region are classified through feature fusion, and the calculation complexity is reduced using a lightweight convolution model.

Benefits of technology

It improves the accuracy and accuracy of brain image processing, reduces calculation costs, and achieves the clinical level accuracy of brain mental illness image recognition detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298736A_ABST
    Figure CN120298736A_ABST
Patent Text Reader

Abstract

The invention discloses a brain image processing method and device, and relates to the technical field of medical image analysis and computer vision. A specific embodiment of the method comprises the following steps: performing feature extraction on a brain magnetic resonance image of a first mode by using a spatial-temporal feature extraction unit to obtain spatial-temporal features of a brain region; performing feature extraction on the brain magnetic resonance image of the second modal through a specific convolutional layer to obtain structural features of a brain region, wherein the specific convolutional layer has a convolution kernel with a set size and a convolution kernel moving step length; and performing feature fusion on the spatial-temporal features of the brain region and the structural features of the brain region to obtain fusion features, and performing feature classification based on the fusion features to obtain an image processing result. According to the embodiment, the features of the brain magnetic resonance image data can be fully utilized to perform feature classification and image processing, the precision of the image processing result and the image processing accuracy are improved, the structural features are extracted through the convolution-based lightweight model, and the calculation cost and calculation complexity of the model are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of medical image analysis and computer vision, and particularly to a method and apparatus for processing brain images. Background Art

[0002] In recent years, magnetic resonance imaging has been crucial for slowing down the deterioration of symptoms of mental illnesses and improving the quality of life of patients. Functional magnetic resonance imaging (fMRI) and structural magnetic resonance imaging (sMRI) are common medical imaging modalities for diagnosing brain mental illnesses and have shown important roles in the latest diagnosis of brain diseases. Taking the disease diagnosis of Bipolar Disorder (BD) as an example, the application of fMRI and sMRI technologies can help us deeply understand the neurobiological mechanisms of BD, provide more objective and accurate BD diagnosis methods, and offer new ideas for the treatment and prevention of BD.

[0003] In the process of implementing the present invention, the inventors found that in the prior art, brain image processing is generally performed by processing fMRI data or sMRI data. However, due to the limited scale of the BD diagnosis dataset, the computational complexity and computational cost of existing brain image processing models are relatively high, and the image processing accuracy and accuracy rate are low. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and apparatus for processing brain images, which can make full use of the characteristics of brain magnetic resonance image data for feature classification and image processing, improving the accuracy of the image processing result and the image processing accuracy rate. Additionally, the present invention uses a lightweight model based on convolution to extract structural features, improving the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of diagnosing brain mental illnesses, the image recognition and detection accuracy of brain mental illnesses can reach the clinical level.

[0005] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for processing brain images is provided, including:

[0006] Performing feature extraction on a brain magnetic resonance image of a first modality using a spatio-temporal feature extraction unit to obtain spatio-temporal features of a brain region;

[0007] Performing feature extraction on a brain magnetic resonance image of a second modality through a specific convolutional layer to obtain structural features of a brain region, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step;

[0008] Fuse the spatio-temporal features and the structural features of the brain region to obtain fused features, and perform feature classification based on the fused features to obtain an image processing result.

[0009] Optionally, the brain magnetic resonance image of the first modality is a resting-state functional magnetic resonance imaging; the brain magnetic resonance image of the second modality is a structural magnetic resonance imaging after T1 weighting.

[0010] Optionally, the spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, where the receptive field is the range of image perception of neurons in the neural network.

[0011] Optionally, using the spatio-temporal feature extraction unit to extract features from the brain magnetic resonance image of the first modality to obtain the spatio-temporal features of the brain region includes: using the spatially connected feature extraction module to extract spatial features from the brain magnetic resonance image of the first modality; using the temporally connected feature extraction module to extract temporal features from the spatial features; and obtaining the spatio-temporal features of the brain region according to the spatial features and the temporal features.

[0012] Optionally, using the temporally connected feature extraction module to extract temporal features from the spatial features includes: setting the convolution kernel moving step of the convolutional layer to 2, and using the temporally connected feature extraction module to extract temporal features from the spatial features to obtain the temporal features.

[0013] Optionally, the spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network to combine the original feature extraction module, the spatially connected feature extraction module, and the temporally connected feature extraction module for feature extraction.

[0014] Optionally, the specific convolutional layer includes no less than one set of convolutional channels; extracting the structural features of the brain region from the brain magnetic resonance image of the second modality through the specific convolutional layer includes: for the brain magnetic resonance image of the second modality, using no less than one set of convolutional channels included in the specific convolutional layer to extract features according to the dimension of each set of convolutional channels set, so as to obtain the structural features of the brain region.

[0015] Optionally, feature fusion is performed on the spatio-temporal features and the structural features of the brain region to obtain fusion features, including: performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; using an array concatenation function to splice the intermediate structural features and the spatio-temporal features of the brain region to perform feature fusion and obtain fusion features.

[0016] Optionally, feature classification is performed based on the fusion features, including: using a linear classifier to perform feature classification on the fusion features.

[0017] According to another aspect of the embodiments of the present invention, there is provided a processing device for brain images, including:

[0018] A spatio-temporal feature extraction module, configured to use a spatio-temporal feature extraction unit to extract features from a first-modal magnetic resonance image of the brain to obtain spatio-temporal features of a brain region;

[0019] A structural feature extraction module, configured to extract structural features of a brain region from a second-modal magnetic resonance image of the brain through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step;

[0020] A feature fusion and classification module, configured to perform feature fusion on the spatio-temporal features and the structural features of the brain region to obtain fusion features, and perform feature classification based on the fusion features to obtain an image processing result.

[0021] Optionally, the first-modal magnetic resonance image of the brain is a resting-state functional magnetic resonance imaging; the second-modal magnetic resonance image of the brain is a T1-weighted structural magnetic resonance imaging.

[0022] Optionally, the spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the first-modal magnetic resonance image of the brain, and the receptive field in the time dimension is 1. The temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, where the receptive field is the range of image perception by neurons in a neural network.

[0023] Optionally, the spatio-temporal feature extraction module is further configured to: use the spatially connected feature extraction module to extract spatial features from the first-modal magnetic resonance image of the brain; use the temporally connected feature extraction module to extract temporal features from the spatial features; and obtain spatio-temporal features of a brain region according to the spatial features and the temporal features.

[0024] Optionally, when the spatio-temporal feature extraction module uses the temporal feature extraction module to extract features from the spatial features to obtain temporal features, it is further configured to: extract temporal features from the spatial features by setting the convolution kernel stride of the convolutional layer to 2 and using the temporal feature extraction module to obtain the temporal features.

[0025] Optionally, the spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network to combine the original feature extraction module, the spatial feature extraction module, and the temporal feature extraction module for feature extraction.

[0026] Optionally, the specific convolutional layer includes not less than one set of convolutional channels; the structural feature extraction module is further configured to: perform feature extraction on the brain magnetic resonance image of the second modality using not less than one set of convolutional channels included in the specific convolutional layer according to the dimension of each set of convolutional channels set, so as to obtain the structural features of the brain region.

[0027] Optionally, the feature fusion and classification module is further configured to: perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; use an array concatenation function to concatenate the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fusion features.

[0028] Optionally, the feature fusion and classification module is further configured to: perform feature classification on the fusion features using a linear classifier.

[0029] According to another aspect of the embodiments of the present invention, an electronic device is provided.

[0030] An electronic device includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method for processing brain images provided by the embodiments of the present invention.

[0031] According to still another aspect of the embodiments of the present invention, a computer-readable medium is provided.

[0032] A computer-readable medium has a computer program stored thereon, and when the program is executed by a processor, it implements the method for processing brain images provided by the embodiments of the present invention.

[0033] One embodiment of the above invention has the following advantages or beneficial effects: extracting the spatio-temporal features of brain regions by using a spatio-temporal feature extraction unit for the brain magnetic resonance images of the first modality; extracting the structural features of brain regions from the brain magnetic resonance images of the second modality through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; performing feature fusion on the spatio-temporal features of brain regions and the structural features of brain regions to obtain fused features, and performing feature classification based on the fused features to obtain the technical solution of the image processing result. The dual-branch processing method can be adopted to process the brain magnetic resonance image data of two different modalities respectively, and use the fused features for feature classification, which can make full use of the features of the brain magnetic resonance image data for feature classification and image processing, improving the accuracy of the image processing result and the image processing accuracy rate. In addition, the present invention extracts structural features through a lightweight model based on convolution, improving the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of brain mental disease diagnosis, the image recognition and detection accuracy of brain mental diseases can reach the clinical level.

[0034] The further effects of the above non-conventional optional methods will be described below in combination with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:

[0036] Figure 1 is a schematic diagram of the main steps of the method for processing brain images according to an embodiment of the present invention;

[0037] Figure 2 is a network structure diagram of the spatio-temporal feature extraction unit according to an embodiment of the present invention;

[0038] Figure 3 is a network structure diagram of the structural feature extraction unit according to an embodiment of the present invention;

[0039] Figure 4 is a principle architecture diagram of the processing system of brain images according to an embodiment of the present invention;

[0040] Figure 5 is a schematic diagram of the main modules of the processing device of brain images according to an embodiment of the present invention;

[0041] Figure 6 is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0042] Figure 7 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The exemplary embodiments of the present invention will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0044] It should be noted that in the technical solutions disclosed in the present invention, in terms of the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of the user's personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data and to safeguard the security of the user's personal information, network security, and national security.

[0045] To solve the problems existing in the prior art, the present invention provides a method for processing brain images. By jointly using computer vision technology and deep learning technology, data of two modalities, fMRI and sMRI, are processed respectively. The deep learning method is used to fully extract the features of brain magnetic resonance images, and the feature data of the two modalities are fused to improve the accuracy of the image processing result and the image processing accuracy rate. In addition, the present invention also proposes a lightweight model based on convolution, which improves the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of diagnosing brain mental diseases, the image recognition detection accuracy of brain mental diseases can reach the clinical level.

[0046] In the introduction of the embodiments of the present invention, the professional terms involved and their interpretations are as follows:

[0047] flatten: In deep learning, flatten is used to reduce multi-dimensional data to one dimension, usually for conversion between fully connected layers and convolutional layers;

[0048] concat: Used to concatenate two or more arrays along a specified axis to form a new array;

[0049] stride: In a convolutional neural network, stride refers to the step size at which the filter slides over the input data. For example, if the stride is 2, it means the filter moves two units each time;

[0050] groups: In the convolution operation, the groups parameter is used to specify the connection mode between the input and output, which means dividing the channels into several groups for convolution operation. One is the depth-wise convolution operation (the number of groups is equal to the number of input channels), and the other is the conventional convolution operation (the number of groups is equal to 1).

[0051] BN: Batch Normalization is a technique commonly used in deep learning networks to accelerate training and prevent overfitting. It can keep the input of each layer in the network in the same distribution.

[0052] ReLU: It is a commonly used activation function in deep learning. It directly returns the negative part to 0 and keeps the positive part unchanged. This can increase the nonlinearity and sparsity of the model and help alleviate the gradient vanishing problem.

[0053] Transformer: Transformer is a model framework widely used in natural language processing. It is based on the self-attention mechanism and effectively handles long-distance dependency problems.

[0054] Receptive Field: Receptive Field refers to the area size of the pixel points on the feature map (feature map) output by each layer of the convolutional neural network mapped back to the input image. A popular explanation is that the size of a point on the feature map relative to the original image is also the area of ​​the input image that the convolutional neural network feature can see.

[0055] The problem of image recognition and diagnosis of brain mental illness based on sMRI and fMRI can be regarded as an image classification problem. Deep learning methods for image classification tasks can also be used for image recognition and diagnosis of brain mental illness. Convolutional neural networks usually consist of convolutional layers, pooling layers, and batch normalization layers. Among them, the convolutional layer uses discrete convolution functions to calculate local features in the image based on the translation invariance of the image, and extracts rich high-level semantic information from the image; the batch normalization layer normalizes the features in the small batch, which can speed up the convergence of the model and make the training process more stable, and plays a significant role in overcoming the problems of gradient explosion and gradient disappearance; the pooling layer aggregates the features in the window through sampling, which can reduce the number of model parameters, increase the receptive field, and alleviate overfitting.

[0056] Existing Transformer-based methods for brain image processing have a high computational complexity and unaffordable computational and time costs when processing sMRI data. Transformer-based methods utilize their large receptive fields to learn adaptive global information and have good performance in many downstream tasks in the field of computer vision. However, their large number of parameters requires a large amount of data for optimization, and the collection of sMRI and fMRI data is very difficult. Therefore, the present invention constructs a model for brain image processing based on a convolutional method.

[0057] In addition, in the prior art, brain image processing is generally performed only based on sMRI data or fMRI data. In the present invention, in order to make more full use of the features of brain magnetic resonance images, data of two modalities, namely fMRI and sMRI, can be processed separately, and the feature data of the two modalities can be fused to improve the accuracy of the image processing result and the image processing accuracy rate. Then, the present invention needs to solve the problems of insufficient intra-modal feature extraction and incomplete inter-modal feature fusion for multi-modal data. In an embodiment of the present invention, a dual pyramid structure (i.e., using two pyramid-like structures to build a model to learn features, and the pyramid structure is considered an effective model structure) can be used to construct an encoder. A hierarchical strategy is used to gradually expand the receptive field and extract features at different scales. For sMRI data, the present invention proposes a window pyramid network, which divides the sMRI data into windows one by one, and performs feature learning for each window to extract the structural features of the brain regions included in the sMRI data. For fMRI data, a spatio-temporal pyramid structure is proposed, and the temporal features and spatial features of the fMRI data are learned simultaneously. Then, the extracted spatial features and temporal features are fused to obtain the spatio-temporal features of the fMRI data. Finally, the structural features and spatio-temporal features extracted from the data of the two modalities are fused, and the fused features are output through a classifier to obtain an image classification result, so as to obtain a brain image processing result.

[0058] Figure 1 It is a schematic diagram of the main steps of the method for processing brain images according to an embodiment of the present invention. As Figure 1 shown, the method for processing brain images according to an embodiment of the present invention mainly includes the following steps S101 to S103.

[0059] Step S101: Use a spatio-temporal feature extraction unit to extract features from the brain magnetic resonance image of the first modality to obtain the spatio-temporal features of the brain region.

[0060] In an embodiment of the present invention, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging, that is, rs-fMRI data.

[0061] Since fMRI (functional magnetic resonance imaging) data is a continuous sMRI (structural magnetic resonance imaging) data stream with a large amount of data, directly processing the features of fMRI data requires a large amount of computational cost. In the prior art, a certain number of brain regions are usually extracted from fMRI data, and then the functional connection matrix between these regions is exported. By analyzing the functional connection matrix, the brain region images can be processed and analyzed. When applied to the brain disease diagnosis scenario, a judgment can be made on the diagnosis of the corresponding disease of the patient. In order to efficiently process the image of fMRI data, the present invention will not directly process the fMRI data, but analyze the features of the brain regions corresponding to the fMRI data. Therefore, the present invention does not involve the preprocessed functional connection matrix, but directly extracts the brain region features from the fMRI data. In order to better utilize the time dimension information contained in the time series, the present invention uses a series-connected spatial feature extraction module and a time feature extraction module as a spatio-temporal feature extraction unit of rs-fMRI to extract spatial features and time features respectively.

[0062] According to an embodiment of the present invention, the spatio-temporal feature extraction unit includes a series-connected spatial feature extraction module and a time feature extraction module. The spatial feature extraction module includes a fully-connected layer and a batch normalization layer. The receptive field of the fully-connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1. The time feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1. Here, the receptive field is the range of the image that the neurons in the neural network can sense. Among them, the batch normalization layer is used to avoid errors caused by data scale, etc., and improve the learning efficiency. Usually, the batch normalization layer and the convolutional layer are regarded as a whole, and only the convolutional layer is shown. When the receptive field in a certain dimension is 1, it means that the features in the corresponding dimension are not extracted. The receptive field of the fully-connected layer is the same as the size of the input data (the brain magnetic resonance image of the first modality), which can fully learn the global features of the fMRI data. For the features in the time dimension, the convolutional layer is directly used to extract the time features.

[0063] According to an embodiment of the present invention, the spatio-temporal feature extraction unit is used to extract the spatio-temporal features of the brain regions from the brain magnetic resonance images of the first modality, which may specifically include: using the spatial feature extraction module to extract spatial features from the brain magnetic resonance images of the first modality; using the temporal feature extraction module to extract temporal features from the spatial features; and obtaining the spatio-temporal features of the brain regions according to the spatial features and the temporal features. When using the spatio-temporal feature extraction unit to extract spatio-temporal features from the brain magnetic resonance images of the first modality, the spatial feature extraction module and the temporal feature extraction module are sequentially used to perform spatial feature extraction and temporal feature extraction, so as to obtain the spatio-temporal features.

[0064] According to another embodiment of the present invention, using the temporal feature extraction module to extract temporal features from the spatial features includes: by setting the convolution kernel moving step of the convolutional layer to 2, using the temporal feature extraction module to extract temporal features from the spatial features, and obtaining the temporal features. Wherein, the convolution kernel moving step refers to the step by which the convolution kernel moves each time in the convolution operation. By setting the convolution kernel moving step of the second convolutional layer in the temporal feature extraction module to 2, downsampling processing can be performed on the temporal features while performing temporal feature extraction, gradually increasing the receptive field of the model in the temporal dimension.

[0065] According to another embodiment of the present invention, the spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network to combine the original feature extraction module, the spatial feature extraction module and the temporal feature extraction module to perform feature extraction. Among them, the original feature extraction module is, for example, extracted by performing convolution processing on the brain magnetic resonance images of the first modality. The original feature extraction module is, for example, a two-dimensional convolutional layer, and the convolution kernel is, for example, 1×1.

[0066] Figure 2 is the network structure diagram of the spatio-temporal feature extraction unit according to an embodiment of the present invention. As Figure 2As shown, in the embodiment of the present invention, the original feature extraction module is omitted. The spatial feature extraction module is shown in the first box, where "F×1,1, Stride=(1,1)" is shown. "F×1" indicates that the size of its convolution kernel is F×1, where F is the number of dimensions of the spatial feature; "1" represents a feature point, and "Stride=(1,1)" means that the stride in both the spatial dimension and the temporal dimension is 1. The temporal feature extraction module is shown in the second box, where "1×3,1, Stride=(1,2)" is shown. "1×3" is the size of its convolution kernel, "1" represents a feature point, and "Stride=(1,2)" means that the stride in the spatial dimension is 1 and the stride in the temporal dimension is 2. In the network structure of this spatio-temporal feature extraction unit, a residual structure is added to further improve the accuracy of the model.

[0067] Suppose the preprocessed fMRI dataset P = {(x i1 ,…,x iN )│i = 1,…,M}, where M represents the number of frames of fMRI data, N represents the number of brain regions, and x refers to the features of each point. Then, in the embodiment of the present invention, in terms of network structure design, different from the traditional scheme, the idea of gradually reducing the image size and increasing the feature dimension is not adopted for feature extraction. Instead, as the network depth increases, the number of brain regions N remains unchanged. This can reduce the number of model parameters, lower the computational complexity of the model, and avoid redundant features.

[0068] Step S102: Feature extraction is performed on the brain magnetic resonance image of the second modality through a specific convolutional layer to obtain the structural features of the brain region. The specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving stride. In the embodiment of the present invention, the brain magnetic resonance image of the second modality is T1-weighted structural magnetic resonance imaging (abbreviation: T1w). The principle of T1-weighted imaging is realized based on the different magnetic resonance relaxation times T1 of human tissues. T1 refers to the time for the magnetic resonance signal to go from the high-energy state to the low-energy state, that is, the time for the magnetic resonance signal to go from the excited state to the ground state.

[0069] The sMRI data is composed of voxels and has a large data size. To address the problem of high computational complexity in the feature extraction module of existing methods, the present invention proposes a novel Patch Pyramid Feature Extraction Module (P2FEM). P2FEM is constructed based on convolutional operations. Existing convolutional neural networks consist of convolutional layers, pooling layers, and batch normalization layers. Among them, the role of the pooling layer is to reduce the number of model parameters, increase the receptive field, and alleviate the overfitting phenomenon, etc. However, there is a certain redundancy in the operation of performing convolution first and then pooling. To improve the computational efficiency of the model, the present invention removes the pooling layer and adjusts the convolutional layer. First, a large convolutional kernel (in the embodiments of the present invention, for example, a three-dimensional convolutional kernel of 7×7×7) is used to expand the receptive field of the model and enhance the learning ability of the model. Second, to reduce the computational complexity of the model and alleviate the overfitting phenomenon, the stride of the convolutional kernel of the convolution is set. By setting the stride, the number of convolution calculations can be reduced, and the downsampling process can be completed while reducing the computational complexity. In the embodiments of the present invention, by using a large convolutional kernel and setting the stride, the convolutional layer has the same function as the pooling layer. In addition, to further reduce the computational complexity of the model and alleviate the overfitting phenomenon, in the embodiments of the present invention, grouped convolution is used to replace the traditional convolutional layer.

[0070] According to an embodiment of the present invention, the specific convolutional layer includes no less than one group of convolutional channels; the structural features of the brain region are obtained by performing feature extraction on the brain magnetic resonance image of the second modality through the specific convolutional layer, including: for the brain magnetic resonance image of the second modality, feature extraction is performed using no less than one group of convolutional channels included in the specific convolutional layer according to the dimensions of each set of convolutional channels set, so as to obtain the structural features of the brain region. Specifically, assuming that a data contains features of multiple dimensions, grouped convolution means grouping the feature dimensions, and then using the corresponding convolutional channels to perform feature extraction on each group of features. For example, for an 8-dimensional data, after dividing its dimensions into two groups, the dimension of each group of features is 4, and then the corresponding convolutional channels are used for convolution respectively to achieve grouped convolution, thereby reducing the redundancy of features and the computational complexity of convolution.

[0071] Figure 3 It is the network structure diagram of the structural feature extraction unit of an embodiment of the present invention. As Figure 3As shown, in an embodiment of the present invention, for the T1-weighted structural magnetic resonance imaging (abbreviated as T1w), it is described as (H, W, D, C), where W, H, and D respectively represent the width, height, and depth of the data, and C is the dimension or number of channels of the features. Among them, "7×7×7, C2, Stride = 2, Groups = 4" is a specific convolutional layer. "7×7×7" represents the convolutional kernel, "C2" refers to the output dimension value being C2, "Stride = 2" means that the stride in the three dimensions of the width, height, and depth of the data is 2, and "Groups = 4" means that there are a total of 4 groups of convolutional channels. Correspondingly, the meaning represented by the specific convolutional layer "5×5×5, C3, Stride = 2, Groups = 4" can be understood.

[0072] Compared with the traditional convolutional neural network, the structural feature extraction unit of the embodiment of the present invention has higher computational efficiency and fewer parameters, and the features extracted by it have stronger generalization ability. Since the data size of sMRI is large, adding each convolutional layer will greatly increase the computational complexity of the model. In order to minimize the computational cost of the model, only one convolutional layer can be used for feature extraction at each level.

[0073] Step S103: Perform feature fusion on the spatio-temporal features and the structural features of the brain region to obtain fused features, and perform feature classification based on the fused features to obtain the image processing result. Before performing feature classification, the features can also be processed such as dimensionality reduction and feature fusion to facilitate more accurate feature classification.

[0074] According to an embodiment of the present invention, performing feature fusion on the spatio-temporal features and the structural features of the brain region to obtain fused features may specifically include: performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; using an array connection function to splice the intermediate structural features and the spatio-temporal features of the brain region to perform feature fusion to obtain fused features. According to the embodiment of the present invention, performing feature classification based on the fused features may specifically include: using a linear classifier to perform feature classification on the fused features.

[0075] When performing feature fusion of multimodal data, first, the flatten function (used for data dimensionality reduction) and a fully connected neural network can be used to process the structural features into the target size as intermediate structural features; then, the concat function, an array connection function, is used to splice the intermediate structural features with the spatio-temporal features to obtain the fusion features corresponding to the multimodal data. Finally, the fusion features are input into a linear classifier for feature classification to obtain the predicted probability values. In the field of machine learning, the goal of classification is to group objects with similar features. A linear classifier makes classification decisions through a linear combination of features to achieve this goal.

[0076] Figure 4 It is the principle architecture diagram of the brain image processing system according to an embodiment of the present invention. As Figure 4 shown, in the embodiment of the present invention, the 4D data of the rs-fMRI, the first-modal brain magnetic resonance image, and the 3D data of the sMRI (abbreviated as T1w) of the second-modal brain magnetic resonance image after T1 weighting are respectively input into two encoder branches for feature extraction. For the rs-fMRI data, multiple spatio-temporal feature extraction units composed of a spatial feature extraction module and a temporal feature extraction module are used for spatio-temporal feature extraction to obtain the spatio-temporal features of multiple brain regions. Among them, the spatio-temporal feature extraction unit mainly includes a residual structure and a 2D convolutional batch normalization layer (both the spatial feature extraction module and the temporal feature extraction module are 2D convolutional batch normalization layers), and by setting the stride corresponding to the convolutional layer in the temporal feature extraction module, the temporal features can be downsampled while extracting the temporal features, thereby gradually expanding the receptive field of the model in the temporal dimension.

[0077] For the T1w data, multiple specific convolutional layers composed of 3D convolutional batch normalization layers are used for 3D convolution to obtain the structural features of multiple brain regions. By using grouped convolution, the structural feature extraction unit has higher computational efficiency and fewer parameters, and the features it extracts have stronger generalization ability.

[0078] After extracting the spatio-temporal features and structural features of each brain region, the fusion and classification of the spatio-temporal features and structural features of the brain region can be performed. When performing feature fusion, the flatten function (used for data dimensionality reduction) and a fully connected neural network can be used to process the structural features into the target size as intermediate structural features; then, the concat function, an array connection function, is used to splice the intermediate structural features with the spatio-temporal features of the brain region to obtain the fusion features corresponding to the multimodal data. Finally, the fusion features are input into a linear classifier for feature classification to obtain the predicted probability values.

[0079] Figure 5It is a schematic diagram of the main modules of a brain image processing device according to an embodiment of the present invention. As Figure 5 shown, the brain image processing device 500 according to the embodiment of the present invention mainly includes a spatio-temporal feature extraction module 501, a structural feature extraction module 502, and a feature fusion and classification module 503.

[0080] The spatio-temporal feature extraction module 501 is used to extract the spatio-temporal features of brain regions by using a spatio-temporal feature extraction unit for the brain magnetic resonance image of the first modality.

[0081] The structural feature extraction module 502 is used to extract the structural features of brain regions from the brain magnetic resonance image of the second modality through a specific convolutional layer, and the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step size.

[0082] The feature fusion and classification module 503 is used to perform feature fusion on the spatio-temporal features and the structural features of the brain regions to obtain fusion features, and perform feature classification based on the fusion features to obtain an image processing result.

[0083] According to an embodiment of the present invention, the brain magnetic resonance image of the first modality is resting-state functional magnetic resonance imaging; the brain magnetic resonance image of the second modality is structural magnetic resonance imaging after T1 weighting.

[0084] According to another embodiment of the present invention, the spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, where the receptive field is the range of image perception by neurons in a neural network.

[0085] According to still another embodiment of the present invention, the spatio-temporal feature extraction module 501 can also be used to: extract spatial features from the brain magnetic resonance image of the first modality by using the spatially connected feature extraction module; extract temporal features from the spatial features by using the temporally connected feature extraction module; and obtain the spatio-temporal features of brain regions according to the spatial features and the temporal features.

[0086] According to still another embodiment of the present invention, when the spatio-temporal feature extraction module 501 extracts temporal features from the spatial features by using the temporally connected feature extraction module, it can also be used to: set the convolutional kernel moving step size of the convolutional layer to 2, and use the temporally connected feature extraction module to perform temporal feature extraction on the spatial features to obtain the temporal features.

[0087] According to another embodiment of the present invention, the spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network to perform feature extraction in combination with the original feature extraction module, the spatial feature extraction module, and the temporal feature extraction module.

[0088] According to another embodiment of the present invention, the specific convolutional layer includes not less than one set of convolutional channels; the structural feature extraction module 502 can also be used to: for the brain magnetic resonance image of the second modality, use the not less than one set of convolutional channels included in the specific convolutional layer to perform feature extraction according to the dimensions of each set of convolutional channels set, so as to obtain the structural features of the brain region.

[0089] According to another embodiment of the present invention, the feature fusion and classification module 503 can also be used to: perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; use an array connection function to splice the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fusion features.

[0090] According to another embodiment of the present invention, the feature fusion and classification module 503 can also be used to: perform feature classification on the fusion features using a linear classifier.

[0091] According to the technical solution of the embodiment of the present invention, the spatio-temporal features of the brain region are obtained by using the spatio-temporal feature extraction unit to perform feature extraction on the brain magnetic resonance image of the first modality; the structural features of the brain region are obtained by performing feature extraction on the brain magnetic resonance image of the second modality through a specific convolutional layer, and the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; the spatio-temporal features of the brain region and the structural features of the brain region are fused to obtain fusion features, and feature classification is performed based on the fusion features to obtain the technical solution of the image processing result. A dual-branch processing method can be adopted to process the brain magnetic resonance image data of two different modalities respectively, and feature classification is performed using the fusion features, which can make full use of the features of the brain magnetic resonance image data for feature classification and image processing, and improve the accuracy of the image processing result and the image processing accuracy rate. In addition, the present invention extracts structural features through a lightweight model based on convolution, which improves the model performance while reducing the calculation cost and calculation complexity of the model. In the application scenario of brain mental disease diagnosis, the image recognition and detection accuracy of brain mental diseases can reach the clinical level.

[0092] Figure 6 An exemplary system architecture 600 of a processing method for brain images or a processing device for brain images to which the embodiments of the present invention can be applied is shown.

[0093] As Figure 6As shown, the system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. The network 604 serves as a medium for providing communication links between the terminal devices 601, 602, 603 and the server 605. The network 604 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0094] Users can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0095] The terminal devices 601, 602, 603 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers, etc.

[0096] The server 605 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the terminal devices 601, 602, 603 (only for example). The background management server can perform feature extraction on the received brain image processing requests and other data to obtain the spatio-temporal features of the brain region using a spatio-temporal feature extraction unit for the first-modal brain magnetic resonance image; perform feature extraction on the second-modal brain magnetic resonance image through a specific convolutional layer to obtain the structural features of the brain region, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step size; perform feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fusion features, and perform feature classification based on the fusion features to obtain processing results such as image processing results, etc., and feedback the processing results (such as feature classification results, image processing results - only for example) to the terminal device.

[0097] It should be noted that the method for processing brain images provided by the embodiments of the present invention is generally executed by the server 605. Correspondingly, the device for processing brain images is generally provided in the server 605.

[0098] It should be understood that Figure 6 the numbers of terminal devices, networks, and servers in

[0099] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 7 The following refers to Figure 7The terminal device or server shown is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present invention.

[0100] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0101] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as required. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as required so that a computer program read from it can be installed into the storage section 708 as required.

[0102] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present invention are executed.

[0103] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0105] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a spatio-temporal feature extraction module, a structural feature extraction module, and a feature fusion and classification module. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the feature fusion and classification module can also be described as "a module for performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result".

[0106] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: using a spatio-temporal feature extraction unit to extract features from a first-modal brain magnetic resonance image to obtain spatio-temporal features of a brain region; extracting structural features of the brain region from a second-modal brain magnetic resonance image through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result.

[0107] According to the technical solution of the embodiments of the present invention, by using a spatio-temporal feature extraction unit to extract features from a first-modal brain magnetic resonance image to obtain spatio-temporal features of a brain region; extracting structural features of the brain region from a second-modal brain magnetic resonance image through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step; performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result, a dual-branch processing method can be adopted to separately process brain magnetic resonance image data of two different modalities, and use the fused features for feature classification, which can make full use of the features of the brain magnetic resonance image data for feature classification and image processing, improving the accuracy of the image processing result and the image processing accuracy rate. In addition, the present invention uses a lightweight model based on convolution to extract structural features, improving the model performance while reducing the computational cost and computational complexity of the model. In the application scenario of brain mental disease diagnosis, the recognition and detection accuracy of brain mental disease images can reach the clinical level.

[0108] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for processing brain images, characterized in that, Comprising: Performing feature extraction on the brain magnetic resonance image of the first modality using a spatio-temporal feature extraction unit to obtain the spatio-temporal features of the brain region; Performing feature extraction on the brain magnetic resonance image of the second modality through a specific convolutional layer to obtain the structural features of the brain region, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving step size; Performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, and performing feature classification based on the fused features to obtain an image processing result.

2. The method according to claim 1, characterized in that, The brain magnetic resonance image of the first modality is a resting-state functional magnetic resonance imaging; the brain magnetic resonance image of the second modality is a structural magnetic resonance imaging after T1 weighting.

3. The method according to claim 1, wherein The spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the brain magnetic resonance image of the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1; where the receptive field is the range of the image that a neuron in the neural network can sense.

4. The method according to claim 3, characterized in that, Performing feature extraction on the brain magnetic resonance image of the first modality using a spatio-temporal feature extraction unit to obtain the spatio-temporal features of the brain region, including: Performing feature extraction on the brain magnetic resonance image of the first modality using the spatially connected feature extraction module to obtain spatial features; Performing feature extraction on the spatial features using the temporally connected feature extraction module to obtain temporal features; Obtaining the spatio-temporal features of the brain region according to the spatial features and the temporal features.

5. The method according to claim 4, characterized in that, Performing feature extraction on the spatial features using the temporally connected feature extraction module to obtain temporal features, including: By setting the convolutional kernel moving step size of the convolutional layer to 2, using the temporally connected feature extraction module to perform temporal feature extraction on the spatial features to obtain the temporal features.

6. The method according to claim 3, wherein The spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network and combines the original feature extraction module, the spatially connected feature extraction module, and the temporally connected feature extraction module to perform feature extraction.

7. The method according to claim 1, wherein The specific convolutional layer includes not less than one set of convolutional channels; Performing feature extraction on the brain magnetic resonance image of the second modality through a specific convolutional layer to obtain the structural features of the brain region, including: For the brain magnetic resonance image of the second modality, performing feature extraction using not less than one set of convolutional channels included in the specific convolutional layer according to the dimensions of each set of convolutional channels set to obtain the structural features of the brain region.

8. The method according to claim 1, wherein Performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fused features, including: Performing dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; Using an array connection function to splice and process the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fused features.

9. The method according to claim 1 or 8, characterized in that Performing feature classification based on the fused features, including: Performing feature classification on the fused features using a linear classifier.

10. A processing device for brain images, characterized in that, Comprising: A spatio-temporal feature extraction module for extracting spatio-temporal features of brain regions from magnetic resonance images of the brain in the first modality using a spatio-temporal feature extraction unit; A structural feature extraction module for extracting structural features of brain regions from magnetic resonance images of the brain in the second modality through a specific convolutional layer, where the specific convolutional layer has a convolutional kernel of a set size and a convolutional kernel moving stride; A feature fusion and classification module for performing feature fusion on the spatio-temporal features of the brain region and the structural features of the brain region to obtain fusion features, and performing feature classification based on the fusion features to obtain an image processing result.

11. The device according to claim 10, characterized in that, The magnetic resonance images of the brain in the first modality are resting-state functional magnetic resonance imaging; the magnetic resonance images of the brain in the second modality are structural magnetic resonance imaging after T1 weighting.

12. The device according to claim 10, wherein The spatio-temporal feature extraction unit includes a spatially connected feature extraction module and a temporally connected feature extraction module connected in series. The spatially connected feature extraction module includes a fully connected layer and a batch normalization layer. The receptive field of the fully connected layer is the same as the size of the magnetic resonance images of the brain in the first modality, and the receptive field in the time dimension is 1; the temporally connected feature extraction module includes a convolutional layer, and the receptive field in the spatial dimension is 1, where the receptive field is the range of image perception by neurons in a neural network.

13. The device according to claim 12, characterized in that, The spatio-temporal feature extraction module is further configured to: Use the spatially connected feature extraction module to extract spatial features from the magnetic resonance images of the brain in the first modality; Use the temporally connected feature extraction module to extract temporal features from the spatial features; Obtain spatio-temporal features of brain regions based on the spatial features and the temporal features.

14. The device according to claim 13, wherein When the spatio-temporal feature extraction module uses the temporally connected feature extraction module to extract temporal features from the spatial features, it is further configured to: By setting the convolutional kernel moving stride of the convolutional layer to 2, use the temporally connected feature extraction module to perform temporal feature extraction on the spatial features to obtain the temporal features.

15. The device according to claim 12, characterized in that, The spatio-temporal feature extraction unit further includes an original feature extraction module, and the spatio-temporal feature extraction unit uses the residual structure of the neural network to combine the original feature extraction module, the spatially connected feature extraction module, and the temporally connected feature extraction module for feature extraction.

16. The device according to claim 10, wherein The specific convolutional layer includes not less than one set of convolutional channels; The structural feature extraction module is further configured to: For magnetic resonance images of the brain in the second modality, use not less than one set of convolutional channels included in the specific convolutional layer to extract features according to the dimensions of each set of convolutional channels set, so as to obtain structural features of brain regions.

17. The device according to claim 10, characterized in that The feature fusion and classification module is further configured to: Perform dimensionality reduction processing on the structural features of the brain region to obtain intermediate structural features of a specified dimension; Use an array connection function to splice the intermediate structural features and the spatio-temporal features of the brain region for feature fusion to obtain fusion features.

18. The device according to claim 10 or 17, characterized in that, The feature fusion and classification module is further configured to: Use a linear classifier to perform feature classification on the fusion features.

19. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-9.

20. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1-9 is implemented.