Image Feature Selection Method for the Spark Model of Schizophrenia Medical Records

Through multi-scale hierarchical segmentation and adaptive noise reduction processing methods, combined with functional connection calculation and Fisher score evaluation, the instability and low reliability problems of traditional feature selection methods in processing fMRI data is solved, achieving more accurate and reliable feature selection, supporting more in-depth research and diagnosis of schizophrenia.

CN119649158BActive Publication Date: 2025-06-13AFFILIATED HOSPITAL OF JIANGNAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510173924.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-13
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The traditional image feature selection method used in the Spark model of schizophrenia medical records is difficult to accurately select the most biologically significant features, and there is a low signal-to-noise ratio and complex brain region structure when processing fMRI data, resulting in unstable analysis results and low reliability.

Method used

By obtaining brain magnetic resonance image data and resting state fMRI data of multiple subjects, a preset hierarchical brain area template is used for multi-scale hierarchical segmentation, adaptive hierarchical noise reduction processing is carried out, a hierarchical progressive functional connection calculation framework is constructed, and feature importance evaluation and feature screening are performed based on Fisher scores.

Benefits of technology

High-quality brain region localization and acquisition of functional network feature data is achieved, the accuracy and reliability of feature selection are improved, and the understanding and diagnosis of schizophrenia pathological mechanisms are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649158B_ABST
    Figure CN119649158B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of medical information extraction, and particularly to an image feature selection method for a schizophrenia medical record Spark model. The method includes the following steps: obtaining brain magnetic resonance image data and resting-state functional magnetic resonance imaging data of multiple subjects; using a preset hierarchical brain region template to perform multi-scale hierarchical segmentation on the brain magnetic resonance image data of each subject's brain image to obtain brain region localization data including three levels of local, regional, and global; performing adaptive hierarchical noise reduction processing on the resting-state functional magnetic resonance imaging data in a Spark distributed environment to obtain denoised brain functional time series data. The present invention uses a preset hierarchical brain region template to accurately divide the brain magnetic resonance image data into three levels of local, regional, and global, thereby taking into account both the local structural details and global network characteristics of the brain regions and realizing the effective integration of multi-scale information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical information extraction, and in particular to an image feature selection method for a Spark model of schizophrenia medical records. Background Art

[0002] Schizophrenia is a complex and severe mental disorder with unknown etiology, mainly manifested by disorders in multiple aspects such as perception, thinking, emotion, and behavior, as well as the incoordination of mental activities. Spark MLlib provides a variety of feature selection methods, which are suitable for the scenario of image feature selection. For example, the Filter method selects features based on the statistical relationship between features and target variables, and the Wrapper method evaluates the importance of features by training a model. In the Spark model for schizophrenia medical records, image feature selection is a key step, which can help improve the accuracy and efficiency of the model. In the study of schizophrenia, medical images (such as CT, MRI, etc.) provide important brain morphological and functional information. By selecting features related to schizophrenia, the pathological mechanism of the disease can be better understood, and the diagnosis and treatment effects can be improved.

[0003] However, the traditional image feature selection methods for the Spark model of schizophrenia medical records often have the following problems: The existing feature selection methods of Spark MLlib (such as chi-square test, information gain) are relatively limited in processing medical images and are difficult to accurately select the most biologically significant features. The inherently low signal-to-noise ratio and multi-level and complex brain region structure of fMRI data make it difficult for traditional noise reduction methods and single-scale analysis to simultaneously take into account the accurate extraction of local details and global networks, resulting in unstable analysis results and low reliability. Summary of the Invention

[0004] Based on this, it is necessary for the present invention to provide an image feature selection method for a Spark model of schizophrenia medical records to solve at least one of the above technical problems.

[0005] To achieve the above object, an image feature selection method for a Spark model of schizophrenia medical records includes the following steps:

[0006] Step S1: Obtain the brain magnetic resonance image data and resting-state functional magnetic resonance imaging data of multiple subjects;

[0007] Step S2: Use a preset hierarchical brain region template to perform multi-scale hierarchical segmentation on the brain magnetic resonance image data of each subject's brain image, obtaining brain region localization data including three levels: local, regional, and global. Among them, the multi-scale hierarchical segmentation specifically includes that the local level is subcortical structure segmentation, the regional level is functional sub-region segmentation based on anatomical connections by calculating the structural connection probability and spatial adjacency relationship between brain regions, and the global level is large-scale brain network segmentation based on the hierarchical organization analysis of the resting-state network;

[0008] Step S3: Perform adaptive hierarchical noise reduction processing on the resting-state functional magnetic resonance imaging data in the Spark distributed environment to obtain denoised brain functional time series data;

[0009] Step S4: Construct a hierarchical progressive functional connection calculation framework based on the brain region localization data and the brain functional time series data, calculate the functional connection strength within and between each level respectively, and perform dynamic adjustment of the connection weights between levels based on the principle of synaptic plasticity to obtain brain functional network feature data;

[0010] Step S5: Perform hierarchical feature importance evaluation on the brain functional network feature data based on Fisher scores, and set hierarchical adaptive thresholds according to the pre-acquired prior knowledge of schizophrenia for feature screening to obtain a schizophrenia feature set.

[0011] Preferably, Step S1 includes the following steps:

[0012] Step S11: Use a 3T magnetic resonance imaging device to perform T1-weighted magnetic resonance scanning on the brains of multiple subjects to obtain brain magnetic resonance image data, where the imaging contrast of the brain magnetic resonance image data is specifically determined by the T1 relaxation time;

[0013] Step S12: Use a 7T superconducting magnetic resonance imaging device to collect blood oxygenation level-dependent contrast based on a gradient echo sequence for the brains of multiple subjects in the resting state to obtain blood oxygenation level-dependent contrast data;

[0014] Step S13: Perform analysis of local blood flow changes caused by neuronal activity on the blood oxygenation level-dependent contrast data to obtain resting-state functional magnetic resonance imaging data.

[0015] Preferably, Step S2 includes the following steps:

[0016] Step S21: Perform image resolution adaptive normalization processing on the brain magnetic resonance image data, and perform voxel size resampling based on B-spline interpolation to obtain normalized image data with a spatial resolution of 1 mm isotropic voxels;

[0017] Step S22: Segment the subcortical structures from the standardized image data according to a preset hierarchical brain region template, and perform brain region localization processing, so as to obtain subcortical structure localization data at the local level;

[0018] Step S23: Perform functional partition clustering based on anatomical connections by calculating the structural connection probability and spatial adjacency relationship between brain regions according to the subcortical structure localization data at the local level, so as to obtain functional partition data at the regional level;

[0019] Step S24: Perform hierarchical organization analysis of the cerebral cortex based on the resting-state network according to the functional partition data and the resting-state functional magnetic resonance imaging data, so as to obtain large-scale brain network data at the global level, where the hierarchical organization analysis specifically calculates the time series correlation and spatial continuity between functional partitions;

[0020] Step S25: Perform hierarchical integration on the subcortical structure localization data, the functional partition data, and the large-scale brain network data, and establish a correspondence matrix between levels, so as to obtain brain region localization data.

[0021] Through a series of refined processes, the present invention realizes the standardization of brain magnetic resonance image data and multi-level brain region localization, providing high-quality basic data for subsequent analysis. First, the resolution of the image data is standardized to a 1-mm isotropic voxel format, eliminating the differences caused by equipment and scanning parameters, and providing a unified framework for subsequent analysis. Then, a hierarchical brain region template is used for subcortical structure segmentation and localization, accurately identifying important structures such as the thalamus and amygdala, and providing local-level information for understanding disease changes. Based on this, functional partition clustering is performed by calculating the structural connection probability and spatial adjacency relationship to obtain functional partition data at the regional level, reflecting the structural-functional relationship of the brain. Further combined with the resting-state functional magnetic resonance imaging data, the time series correlation and spatial continuity between functional partitions are analyzed to obtain large-scale brain network data at the global level, revealing the overall architecture of the brain functional network. Finally, the data at the local, regional, and global levels are integrated, and a correspondence matrix between levels is established to form a multi-level and multi-dimensional brain region localization system, providing comprehensive and high-quality input data for subsequent functional connectivity analysis and feature screening, and ensuring the scientificity and effectiveness of the research method.

[0022] Preferably, step S22 includes the following steps:

[0023] Step S221: Perform edge enhancement processing on the standardized image data by calculating the spatial gradient amplitude and direction of the three-dimensional image, so as to obtain boundary enhancement data of the subcortical structure;

[0024] Step S222: Analyze the gray - level distribution characteristics of the local area based on the boundary - enhanced data, set a dynamic threshold, and perform segmentation to obtain the preliminary segmentation mask data of the subcortical structure;

[0025] Step S223: Use morphological operations to optimize the preliminary segmentation mask data to obtain a continuous and complete subcortical structure mask data, where the morphological optimization specifically involves processing holes and discontinuous boundaries by designing a three - dimensional morphological operator sequence;

[0026] Step S224: Perform connected - component analysis on the subcortical structure mask data and label different subcortical structures to obtain the fine subcortical structure localization data at the local level.

[0027] Preferably, step S23 includes the following steps:

[0028] Step S231: Perform tractography of the connecting fiber bundles on the subcortical structure localization data at the local level and analyze the white - matter fiber connections between subcortical structures to obtain a structural connection probability map;

[0029] Step S232: Calculate the spatial adjacency relationship between brain regions based on the spatial distance metric according to the structural connection probability map to obtain brain region spatial relationship data;

[0030] Step S233: Perform weighted fusion of the structural connection probability map and the brain region spatial relationship data based on the structural connection relationship and the spatial adjacency relationship to obtain a comprehensive relationship matrix;

[0031] Step S234: Use the spectral clustering algorithm to partition the comprehensive relationship matrix to obtain the functional partition data at the regional level.

[0032] Preferably, step S25 includes the following steps:

[0033] Step S251: Extract the spatial position, volume, and shape features of each structure from the subcortical structure localization data and construct feature vectors to obtain a local - level feature matrix;

[0034] Step S252: Analyze the functional connection patterns and spatial distribution characteristics of each partition according to the functional partition data to obtain a regional - level feature matrix;

[0035] Step S253: Extract network features including topological attributes and dynamic features from the large - scale brain network data to obtain a global - level feature matrix;

[0036] Step S254: Process the local - level feature matrix, the regional - level feature matrix, and the global - level feature matrix to obtain brain region localization data.

[0037] Preferably, step S3 includes the following steps:

[0038] Step S31: Partition the resting-state functional magnetic resonance imaging data in the Spark distributed environment, and divide the data into multiple time windows based on the time series, so as to obtain time window sequence data;

[0039] Step S32: Perform multi-scale decomposition based on wavelet transform on the time window sequence data respectively, and decompose the noise signal into different frequency bands, so as to obtain wavelet band coefficient data;

[0040] Step S33: Evaluate the signal-to-noise ratio of each frequency band according to the wavelet band coefficient data, and determine the adaptive noise threshold, so as to obtain the frequency band noise threshold data;

[0041] Step S34: Perform soft threshold shrinkage processing on the wavelet band coefficient data by using the frequency band noise threshold data to suppress the noise component, so as to obtain the wavelet coefficient data after denoising;

[0042] Step S35: Perform wavelet inverse transform reconstruction on the wavelet coefficient data, and perform parallel computing in the Spark environment, so as to obtain the preliminary denoised sequence data;

[0043] Step S36: Perform physiological noise removal processing including heartbeat and respiration on the preliminary denoised sequence data, so as to obtain the noise-corrected sequence data;

[0044] Step S37: Perform head motion parameter correction and spatial smoothing processing on the noise-corrected sequence data, so as to obtain the brain functional time series data.

[0045] Preferably, step S4 includes the following steps:

[0046] Step S41: Partition the brain functional time series data in the Spark distributed environment according to the brain region localization data, and perform standardization processing on the time series within each level, so as to obtain the standardized time series data;

[0047] Step S42: Calculate the functional connection strength between subcortical structures within the local level according to the standardized time series data, and quantify it through Pearson correlation coefficient analysis, so as to obtain the local level functional connection matrix;

[0048] Step S43: Perform functional connection calculation on the functional partition of the regional level according to the local level functional connection matrix, and perform weighted integration based on the time series correlation of the internal nodes of the region, so as to obtain the regional level functional connection matrix;

[0049] Step S44: Evaluate the node importance and connection significance through the network theory method based on the regional hierarchical functional connection matrix, and determine the functional connection pattern of the global hierarchical large-scale brain network, so as to obtain the global hierarchical functional connection matrix;

[0050] Step S45: Dynamically adjust the weights between levels for the local hierarchical functional connection matrix, regional hierarchical functional connection matrix, and global hierarchical functional connection matrix through the principle of synaptic plasticity, and establish the mapping relationship of hierarchical connections, so as to obtain the brain functional network feature data.

[0051] Preferably, step S45 includes the following steps:

[0052] Step S451: Normalize the local hierarchical functional connection matrix, regional hierarchical functional connection matrix, and global hierarchical functional connection matrix, and perform data integration in the Spark distributed environment, so as to construct the hierarchical connection mapping matrix data;

[0053] Step S452: Construct a weight update mathematical model based on the principle of synaptic plasticity according to the hierarchical connection mapping matrix data, and initialize the weights of each cross-level connection, so as to obtain the initial cross-level connection weight matrix data;

[0054] Step S453: Perform synchrony analysis on the time series activation of cross-level nodes according to the initial cross-level connection weight matrix data, so as to obtain the preliminary dynamic connection weight data, where the synchrony analysis includes calculating the update factors of positive reinforcement and negative inhibition;

[0055] Step S454: Perform iterative dynamic weight adjustment processing on the preliminary dynamic connection weight data in the Spark distributed environment, and correct the connection weights according to the preset global convergence condition, so as to obtain the cross-level connection weight matrix data;

[0056] Step S455: Establish the mapping relationship of hierarchical connections according to the cross-level connection weight matrix data, so as to obtain the brain functional network feature data.

[0057] Preferably, step S5 includes the following steps:

[0058] Step S51: Group the brain functional network feature data according to the local, regional, and global three levels, and calculate the Fisher discriminant scores for the functional connection features of each level respectively, so as to obtain the hierarchical Fisher score data;

[0059] Step S52: Establish a knowledge base of brain functional connection patterns related to schizophrenia based on the pre-collected clinical data of schizophrenia patients and healthy controls, so as to obtain the prior knowledge of schizophrenia;

[0060] Step S53: Set differential importance weights for the functional connection features at different levels based on the prior knowledge of schizophrenia, and perform weighted processing on the hierarchical Fisher score data to obtain weighted feature importance data;

[0061] Step S54: Determine the feature selection thresholds based on the distribution characteristics and quantities of the features at different levels according to the weighted feature importance data, so as to obtain hierarchical threshold data;

[0062] Step S55: Use the hierarchical threshold data to screen the weighted feature importance data to obtain a candidate feature set;

[0063] Step S56: Perform redundancy analysis on the candidate feature set, calculate the correlation between features, and remove highly correlated redundant features according to a preset threshold to obtain an optimized feature set;

[0064] Step S57: Match the optimized feature set with the prior knowledge of schizophrenia to generate the biological significance and clinical interpretability of the optimized features, so as to obtain a schizophrenia feature set.

[0065] The present invention provides comprehensive and high-quality basic data for subsequent analysis by acquiring the brain magnetic resonance image data and resting-state functional magnetic resonance imaging data of the subject. The brain magnetic resonance image data can clearly show the structural characteristics of the brain, while the resting-state functional magnetic resonance imaging data reflects the functional activity pattern of the brain in the resting state. This data acquisition method combining structure and function provides multi-dimensional information support for the in-depth study of the neural mechanism of schizophrenia, ensuring the comprehensiveness and accuracy of subsequent analysis. Using a preset hierarchical brain region template to perform multi-scale hierarchical segmentation on the brain image can accurately locate the brain regions at the local, regional, and global levels. This hierarchical method not only covers multiple levels from micro to macro, but also enables targeted analysis according to the characteristics of different levels. The division of the subcortical structure at the local level, the functional partition at the regional level, and the large-scale brain network at the global level allows researchers to comprehensively analyze the structural and functional characteristics of the brain from multiple perspectives, providing an accurate positioning basis for subsequent functional connectivity calculation and feature screening. This multi-scale hierarchical segmentation method can effectively avoid the limitations of single-level analysis, making the research on the pathological mechanism of schizophrenia more comprehensive and in-depth. Performing adaptive hierarchical noise reduction processing on the resting-state functional magnetic resonance imaging data can effectively remove noise interference and improve data quality. By performing noise reduction processing in the Spark distributed environment, not only can the advantages of distributed computing be fully utilized to improve processing efficiency, but also the integrity and accuracy of the data can be ensured. This noise reduction method can accurately suppress the noise components while retaining important physiological signals, thereby providing high-quality data input for subsequent functional connectivity calculation. In addition, steps such as physiological noise removal processing and head motion parameter correction further improve the reliability and usability of the data, ensuring the accuracy of subsequent analysis. Constructing a hierarchical progressive functional connectivity calculation framework based on the brain region localization data and the denoised brain functional time series data can accurately calculate the functional connectivity strength within and between each level. By introducing the principle of synaptic plasticity to dynamically adjust the connection weights between levels, the dynamic changes of the brain functional network can be more realistically reflected. This dynamic adjustment mechanism enables the functional connectivity network to better adapt to the information transmission and interaction between different levels, thereby more accurately revealing the abnormal patterns of the brain functional network in schizophrenia patients. The finally obtained brain functional network feature data provides high-quality input for subsequent feature evaluation and screening, ensuring the scientificity and accuracy of feature selection. Adopting a hierarchical feature importance evaluation method based on Fisher score and setting a hierarchical adaptive threshold in combination with the pre-acquired prior knowledge of schizophrenia can accurately screen out the feature set related to schizophrenia. This method can not only effectively remove redundant features, reduce the feature dimension, but also improve the biological significance and clinical interpretability of the features.Through matching and verification with prior knowledge, the accuracy and reliability of the selected feature set are further ensured, providing strong support for the diagnosis, treatment, and research of schizophrenia. This feature selection method based on the combination of prior knowledge and data-driven can effectively improve the pertinence and effectiveness of feature selection, providing an important reference basis for subsequent clinical applications and research. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Other features, objectives, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0067] Figure 1 It is a schematic flowchart of the steps of the image feature selection method for the Spark model of schizophrenia medical records of the present invention;

[0068] Figure 2 is Figure 1 a detailed schematic flowchart of step S1 in

[0069] Figure 3 is Figure 1 a detailed schematic flowchart of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] To achieve the above object, please refer to Figures 1 to 3 The present invention provides an image feature selection method for the Spark model of schizophrenia medical records, and the method includes the following steps:

[0071] Step S1: Obtain brain magnetic resonance image data and resting-state functional magnetic resonance imaging data of multiple subjects;

[0072] In this embodiment, first, 50 subjects are selected, and their T1-weighted structural image data and resting-state functional magnetic resonance imaging (rs-fMRI) data are obtained using a 3.0T magnetic resonance imaging (MRI) device. During the scanning process, the subjects are required to keep their eyes closed but stay awake, and the acquisition time is 10 minutes with a time resolution of 2 seconds, that is, each subject obtains 300 time points of data. The spatial resolution of the T1-weighted image is set to 1mm×1mm×1mm, and the spatial resolution of the rs-fMRI data is 3mm×3mm×3mm. All data are stored in NIfTI format and undergo basic preprocessing, including time slice correction, head motion correction, spatial registration to the MNI standard space, linear drift removal, and band-pass filtering (0.01 - 0.08Hz), to reduce noise interference and improve the accuracy of subsequent analysis.

[0073] Step S2: Use the preset hierarchical brain region template to perform multi-scale hierarchical segmentation on the brain magnetic resonance image data of each subject's brain image, obtaining brain region localization data including three levels: local, regional, and global. The multi-scale hierarchical segmentation specifically includes subcortical structure segmentation at the local level, functional partition segmentation based on anatomical connections by calculating the structural connection probability and spatial adjacency relationship between brain regions at the regional level, and large-scale brain network segmentation based on the hierarchical organization analysis of the resting-state network at the global level;

[0074] In this embodiment, the Harvard-Oxford brain region template is used as the hierarchical brain region template. The local level includes 68 subcortical structures, the regional level divides 116 functional partitions based on the AAL (Automated Anatomical Labeling) template, and the global level divides 7 functional networks (such as the default mode network, central executive network, etc.) according to the large-scale brain network model. During the multi-scale segmentation process, first, brain tissue segmentation is performed based on the T1-weighted structural image to remove non-brain tissues such as the skull. Then, a graph cut algorithm based on probability maps (Graph Cut) is used to perform fine segmentation of the subcortical structures, and spatial matching of each brain region is performed based on the standard brain atlas to achieve the division of functional partitions at the regional level. Finally, each functional partition is further aggregated into a global functional network through a network module analysis method (such as k-means clustering) to complete the multi-scale hierarchical segmentation and obtain the three-level brain region localization data of each subject.

[0075] Step S3: Perform adaptive hierarchical noise reduction processing on the resting-state functional magnetic resonance imaging data in the Spark distributed environment to obtain denoised brain functional time series data;

[0076] In this embodiment, under the Spark distributed computing framework, adaptive hierarchical noise reduction processing is performed on the acquired rs-fMRI data. First, based on the time series interpolation method, the signal loss data caused by head movement is filled, and physiological noise (such as respiration and heartbeat signals) is removed through independent component analysis (ICA). Secondly, at the local level, a denoising method based on wavelet transform is used to decompose the time series into multiple frequency components and remove the high-frequency noise components; at the regional level, a method based on dynamic low-rank representation (DLR) is used to remove head movement artifacts and maintain the functional relationship between brain regions; at the global level, a sparse Bayesian regression model is used for global noise estimation, and the empirical mode decomposition (EMD) method is used to remove non-physiological artifacts. Finally, all the denoised signal data is re-integrated to obtain optimized brain functional time series data.

[0077] Step S4: Construct a hierarchical and progressive functional connectivity calculation framework based on the brain region localization data and the brain functional time series data. Calculate the functional connectivity strengths within and between each layer, and perform dynamic adjustment of the connection weights between layers based on the principle of synaptic plasticity to obtain the brain functional network feature data;

[0078] In this embodiment, a hierarchical and progressive functional connectivity calculation framework is constructed based on the brain region localization data and the denoised brain functional time series data. First, at the local layer, the dynamic time warping (DTW) method is used to calculate the functional connectivity strength between subcortical brain regions. Secondly, at the regional layer, the sliding window Pearson correlation method is used to calculate the dynamic connectivity strength between functional sub-regions, and redundant connections are eliminated through sparse regularization constraints. Then, at the global layer, the method based on Granger causality analysis (GCA) is used to calculate the causal influence weights between large-scale brain networks. At the same time, in order to enhance the dynamic adaptability between layers, based on the principle of synaptic plasticity, the cross-layer connection weights are adaptively adjusted. Specifically, for the connection from the local to the regional layer, a weight update strategy based on STDP (Spike-Timing-Dependent Plasticity) is adopted to strengthen the connections of synchronously activated brain regions; for the connection from the regional to the global layer, a non-linear steady-state distribution optimization method is used to keep the overall connectivity of the brain network in dynamic balance. Finally, the brain functional network feature data is output, including the local, regional, and global functional connection matrices.

[0079] Step S5: Perform hierarchical feature importance evaluation on the brain functional network feature data based on Fisher scores, and set hierarchical adaptive thresholds for feature screening according to the pre-acquired prior knowledge of schizophrenia to obtain the schizophrenia feature set.

[0080] In this embodiment, hierarchical feature importance evaluation based on Fisher scores is performed on the obtained brain functional network feature data, and feature screening is performed using prior knowledge related to schizophrenia. Specifically, first, the Fisher scores of the functional connection features of each layer are calculated to measure the discrimination ability of the features between the schizophrenia group and the healthy control group. At the local layer, for the connection features of the subcortical structure, the top 10% features with the highest Fisher scores are screened; at the regional layer, the recursive feature elimination (RFE) method is adopted, and continuous iteration and optimization are performed during the feature selection process to ensure that the selected features remain stable in multiple cross-validations; at the global layer, an adaptive threshold is set based on the abnormal brain network patterns of schizophrenia (such as overactivation of the default mode network), and low-weight connections unrelated to schizophrenia are removed. Finally, the most diagnostically valuable feature set is selected and used for subsequent disease classification model training and prediction.

[0081] Preferably, step S1 includes the following steps:

[0082] Step S11: Use a 3T magnetic resonance imaging device to perform T1-weighted magnetic resonance scanning on the brains of multiple subjects, thereby obtaining brain magnetic resonance image data, where the imaging contrast of the brain magnetic resonance image data is specifically determined by the T1 relaxation time;

[0083] In this embodiment, 40 healthy subjects were selected, and a 3T magnetic resonance imaging (MRI) device (such as Siemens Prisma 3T) was used to perform T1-weighted magnetic resonance scanning on their brains. During the imaging process, the subjects needed to remain stationary and wear noise-canceling earplugs to reduce motion artifacts and external noise interference. The T1-weighted imaging used a three-dimensional MPRAGE (Magnetization Prepared Rapid Acquisition Gradient Echo) sequence, and the scanning parameters were set as echo time (TE) 2.98 ms, repetition time (TR) 2300 ms, flip angle 9°, field of view (FOV) 256 mm×256 mm, matrix resolution 256×256, slice thickness 1 mm, and finally T1-weighted brain magnetic resonance image data with an isotropic voxel resolution of 1 mm³ was obtained. The imaging contrast is determined by the T1 relaxation time, that is, different tissues (such as gray matter, white matter, cerebrospinal fluid) show different signal intensities under the MPRAGE sequence due to the difference in T1 relaxation time to ensure clear and distinguishable structures. After the acquisition, the image data was stored in the NIfTI format and spatially registered to the MNI standard brain template using SPM software for subsequent analysis.

[0084] Step S12: Use a 7T superconducting magnetic resonance imaging device to collect blood oxygenation level-dependent contrast based on a gradient echo sequence for the brains of multiple subjects at rest, thereby obtaining blood oxygenation level-dependent contrast data;

[0085] In this embodiment, the same subjects as in step S11 are selected, and a 7T superconducting magnetic resonance imaging device (such as Siemens Magnetom Terra 7T) is used to collect blood oxygenation level-dependent (BOLD) contrast based on the gradient echo (GRE) sequence for the brains of the subjects at rest. During the scanning process, the subjects need to close their eyes but stay awake and avoid any thinking or movement. The BOLD signal is collected using a single-shot gradient echo EPI (Echo Planar Imaging) sequence with the parameter settings of TE 25 ms, TR 1000 ms, flip angle 60°, FOV 220 mm×220 mm, matrix resolution 128×128, slice thickness 2 mm, no gap, and acquisition time of 10 minutes. Finally, 600 time points of BOLD image data are obtained for each subject. The BOLD signal reflects the level of neuronal activity based on the local susceptibility differences caused by the change in the oxygenation state of hemoglobin. After the imaging is completed, the data is stored in the NIfTI format and undergoes head motion correction, time slice correction, and spatial normalization to improve the signal stability and comparability.

[0086] Step S13: Analyze the local blood flow changes caused by neuronal activity in the blood oxygenation level-dependent contrast data to obtain the resting-state functional magnetic resonance imaging data.

[0087] In this embodiment, based on the obtained BOLD signal data, the local blood flow changes caused by neuronal activity are analyzed to extract the resting-state functional magnetic resonance imaging (rs-fMRI) data. First, the BOLD signal is preprocessed based on low-pass filtering (0.01 - 0.08 Hz) to remove high-frequency physiological noise (such as heartbeat and respiration artifacts). Then, a method based on the steady-state oxygenated blood flow coupling model is used to infer the spontaneous neuronal activity of brain regions by calculating the amplitude change of the local blood oxygenation level-dependent signal. Next, the ALFF (Amplitude of Low-Frequency Fluctuations) method is used to calculate the low-frequency amplitude of each brain region to quantify the intensity of neural activity, and the ReHo (Regional Homogeneity) method is used to evaluate the time series synchrony between local brain regions to identify the local characteristics of the resting-state brain functional network. Finally, the functional connection strength between all brain regions in the whole brain is calculated through Pearson correlation analysis, and a functional connection matrix is constructed to finally obtain the resting-state functional magnetic resonance imaging data for subsequent analysis.

[0088] Preferably, step S2 includes the following steps:

[0089] Step S21: Perform image resolution adaptive normalization processing on the brain magnetic resonance image data. Through voxel size resampling based on B-spline interpolation, standardized image data with a spatial resolution of 1 mm isotropic voxels is obtained;

[0090] In this embodiment, based on the acquired brain magnetic resonance image data, the B-spline interpolation method is used for voxel size resampling to standardize the spatial resolution. First, read the brain magnetic resonance image in NIfTI format and extract its voxel size information to ensure that the anisotropic resolution of the input image can be uniformly converted. Then, use the third-order B-spline interpolation method to perform three-dimensional resampling on the image, calculate the gray values of the new voxel points, and make it match the original image signal as smoothly as possible. B-spline interpolation has better spatial smoothness than bilinear interpolation and avoids the edge artifacts that may be brought by cubic Lagrange interpolation. During the resampling process, ANTs (Advanced Normalization Tools) is used for spatial transformation to maintain the morphological integrity of the brain structure. Finally, the spatial resolution of all image data is standardized to 1 mm isotropic voxels and re-stored in NIfTI format to ensure the accuracy consistency of subsequent segmentation and analysis.

[0091] Step S22: Perform subcortical structure segmentation on the standardized image data according to the preset hierarchical brain region template and perform brain region localization processing to obtain local-level subcortical structure localization data;

[0092] In this embodiment, a method based on the hierarchical brain region template is used to perform subcortical structure segmentation and brain region localization on the standardized brain magnetic resonance image data. First, use the FIRST tool in FSL (FMRIB Software Library) to automatically segment the subcortical structures of the T1-weighted image. The target structures include subcortical regions such as the thalamus, caudate nucleus, putamen, amygdala, and hippocampus. During the segmentation process, through the Bayesian probability map matching method, the brain regions of different subjects are morphologically corrected, and the segmentation boundary is optimized based on the maximum a posteriori estimation. Then, align the segmentation result with the preset hierarchical brain region template, and use FLIRT (FMRIB’s Linear Image Registration Tool) for affine transformation to ensure that the individual brain region coordinates of the subject are consistent with the standard MNI coordinates. Finally, label and store each subcortical structure to generate local-level subcortical structure localization data, providing an accurate anatomical reference for subsequent functional partition clustering.

[0093] Step S23: Perform functional partition clustering based on anatomical connections by calculating the structural connection probability and spatial adjacency relationship between brain regions according to the local-level subcortical structure localization data to obtain regional-level functional partition data;

[0094] Based on the subcortical structure localization data at the local level, this embodiment calculates the structural connection probability and spatial adjacency relationship between brain regions, and performs functional partition clustering based on anatomical connections. First, fiber tracking is performed using DTI (Diffusion Tensor Imaging) data. The iFOD2 method of MRtrix3 is adopted to estimate the distribution probability of white matter fiber bundles based on the spherical deconvolution model, and the anatomical connection probability between subcortical structures is calculated. Then, a spatial adjacency matrix of subcortical brain regions is constructed to define pairs of brain regions that are directly connected anatomically. Next, a method based on Fuzzy C-Means (FCM) is used for functional partition clustering, and the number of partitions is optimized to conform to the functional distribution characteristics of brain regions, so that brain regions with higher functional similarity are grouped into functional partitions at the same regional level. Finally, the clustering results are mapped to the standard brain space and stored as functional partition data, providing the brain functional organizational structure at the regional level.

[0095] Step S24: Perform hierarchical organization analysis of the cerebral cortex based on the resting-state functional magnetic resonance imaging data according to the functional partition data, so as to obtain large-scale brain network data at the global level, where the hierarchical organization analysis specifically calculates the time series correlation and spatial continuity between functional partitions;

[0096] Based on the functional partition data and the resting-state functional magnetic resonance imaging data, this embodiment performs hierarchical organization analysis based on the resting-state network to construct a large-scale brain network at the global level. First, the BOLD time series data of each functional partition is extracted and band-pass filtered (0.01 - 0.08 Hz) to remove high-frequency noise. Then, the Pearson correlation coefficient is used to calculate the time series correlation between functional partitions to obtain a functional connection matrix, and a robust network connection structure is extracted based on the sparse regularization method (L1 regularization). Next, the modularity optimization method (such as the Louvain algorithm) is used for functional partition clustering to identify brain functional modules with strong functional consistency. Further, the betweenness centrality and global efficiency of each functional partition in the network are calculated to measure its role in the large-scale brain network. Finally, the time series correlation and spatial continuity between functional partitions are combined to optimize the network hierarchical organizational structure and generate large-scale brain network data at the global level.

[0097] Step S25: Perform hierarchical integration of the subcortical structure localization data, the functional partition data, and the large-scale brain network data, and establish a correspondence matrix between the hierarchies to obtain brain region localization data.

[0098] Based on the subcortical structure localization data, functional partition data, and large-scale brain network data, this embodiment performs hierarchical integration to establish a correspondence matrix between levels, thereby obtaining the final brain region localization data. First, define the hierarchical mapping relationships among subcortical structures, functional partitions, and large-scale networks, calculate the contribution ratio of each subcortical brain region to the regional hierarchical functional partition, and measure the distribution similarity through the Kullback-Leibler divergence to optimize the correspondence. Then, construct a multi-level brain network based on graph theory methods, and use a weighted adjacency matrix to represent the connection strength between local, regional, and global levels. Next, adopt a connection weight dynamic adjustment mechanism based on the principle of synaptic plasticity to optimize the information transmission method between levels, so that the brain region localization data at different levels are consistent in anatomical structure and functional characteristics. Finally, generate complete brain region localization data to provide a standardized input for subsequent brain functional network analysis.

[0099] Preferably, step S22 includes the following steps:

[0100] Step S221: Perform edge enhancement processing on the subcortical structure by calculating the spatial gradient magnitude and direction of the three-dimensional image based on the standardized image data, thereby obtaining the boundary enhancement data of the subcortical structure;

[0101] In this embodiment, first, utilize the standardized image data. By calculating the gray-scale change within the neighborhood of each voxel in the three-dimensional image, adopt the gradient operation method to calculate the spatial gradient magnitude and direction of the image. Among them, the gradient magnitude reflects the speed of gray-scale change, and the gradient direction indicates the main direction of gray-scale change. During specific operations, use the central difference method to obtain the first-order derivatives of the image in the X, Y, and Z directions respectively, and obtain the gradient magnitude by taking the square root of the sum of squares. Then, use the arctangent function to calculate the angle distribution of the gradient direction. Finally, combine to obtain an edge-enhanced image, which has a high response value at the boundary of the subcortical structure, thereby effectively distinguishing the boundaries between tissues. In terms of operation parameters, set the window size to 3×3×3 voxels to balance noise suppression and edge detection accuracy, and at the same time normalize the calculation results to ensure that the dynamic range of the enhanced data is suitable for subsequent processing. The finally generated boundary enhancement data of the subcortical structure can significantly highlight the boundary information between brain structures.

[0102] Step S222: Analyze the gray-scale distribution characteristics of the local region based on the boundary enhancement data, set a dynamic threshold, and perform segmentation to obtain the preliminary segmentation mask data of the subcortical structure;

[0103] Based on the obtained boundary-enhanced data, this embodiment conducts a detailed analysis of the gray-scale distribution characteristics within a local region. By statistically calculating the mean, variance, and histogram distribution of pixel gray-scale values in each region, a dynamic threshold segmentation method is used to achieve preliminary segmentation. Specifically, first, within a predefined region of interest, the average gray-scale level within the region is determined through a local mean filtering method. Then, a dynamic threshold is determined based on the local standard deviation. The meaning is that pixels with gray-scale values greater than this dynamic threshold are considered to belong to the subcortical structure edge. The threshold is automatically adjusted according to the peak and valley relationships of the local gray-scale histogram to ensure adaptation to the gray-scale distribution differences in different brain regions. Subsequently, the boundary-enhanced data is binarized, and finally, preliminary subcortical structure segmentation mask data is generated. In this data, high-gray-scale regions are marked as candidate regions of the subcortical structure, while low-gray-scale regions are regarded as the background or noise.

[0104] Step S223: Use morphological operations to optimize the preliminary segmentation mask data to obtain continuous and complete subcortical structure mask data. Specifically, the morphological optimization is achieved by designing a three-dimensional morphological operator sequence to process holes and discontinuous boundaries.

[0105] This embodiment uses three-dimensional morphological operations to optimize the preliminary segmentation mask data, aiming to eliminate problems of holes and discontinuous boundaries caused by noise or segmentation errors. In terms of specific operations, first, the preliminary mask is processed using three-dimensional dilation. The structuring element is designed as a 3×3×3 cube to fill small holes. Subsequently, three-dimensional erosion is performed to eliminate the redundant noise introduced by dilation while maintaining the continuity of the original structure. Then, a combination of opening operation (erosion first and then dilation) and closing operation (dilation first and then erosion) is used to further correct regions with uneven or broken edges, ensuring the continuity and smoothness of the mask edge. The entire morphological operator sequence is iteratively optimized multiple times. In terms of parameter settings, the number of iterations for each operation is usually between 1 and 3 times, and is adjusted according to the actual image quality. Finally, continuous and complete subcortical structure mask data is obtained.

[0106] Step S224: Conduct connected component analysis based on the subcortical structure mask data and label different subcortical structures to obtain local-level fine subcortical structure localization data.

[0107] Based on the subcortical structure mask data with optimized morphology obtained in this embodiment, the connected component analysis method is used to label the independent connected regions in the image to achieve fine segmentation of different subcortical structures. During specific operations, the three-dimensional connected component labeling algorithm (such as using the 26-neighborhood method) is used to scan the entire mask data, and all connected pixel blocks are grouped. Each connected component is regarded as an independent structure candidate region. Subsequently, by comparing the volume, shape, and position characteristics of each connected component with the preset subcortical structure database, accurate label assignment is performed for each connected component. In addition, for regions with weak edges or segmentation errors, the region growing algorithm is combined for fine-tuning to ensure that each labeled subcortical structure is a continuous, complete, and anatomically characteristic region. Finally, the fine subcortical structure localization data at the local level is output, which contains the spatial position, boundary shape, and label information of each subcortical structure, providing a high-precision anatomical basis for subsequent multi-level brain region analysis.

[0108] Preferably, step S23 includes the following steps:

[0109] Step S231: Perform tractography of the connecting fiber bundles on the subcortical structure localization data at the local level, and perform white matter fiber connection analysis between subcortical structures to obtain a structural connection probability map;

[0110] Based on the subcortical structure localization data at the local level in this embodiment, the extended fiber tracking algorithm is used to track the white matter fibers between each subcortical structure. First, the preprocessed diffusion-weighted imaging data is used to reconstruct the local fiber orientation distribution through a spherical model. Then, the probabilistic fiber tracking algorithm (such as using a random walk model based on the Monte Carlo method) is used to generate a large number of fiber bundles for the seed points of each subcortical structure. It is set to generate 5000 fibers for each seed point, and the maximum step size is set to 0.5 mm and the curvature threshold does not exceed 45 degrees during the fiber tracking process to ensure that the tracked fibers conform to anatomical constraints. Subsequently, the starting and ending points of each fiber are matched in the predefined subcortical structure segmentation map. By counting the number of all fibers between the two regions, the structural connection probability between the two regions is calculated, and the connection strength is represented by a probability value. Finally, a structural connection probability map is generated, which expresses the relative possibility and strength of the connection between each subcortical structure in matrix form, providing a quantitative basis for subsequent analysis.

[0111] Step S232: Calculate the spatial adjacency relationship between brain regions based on the spatial distance metric according to the structural connection probability map to obtain brain region spatial relationship data;

[0112] Based on the previously generated structural connection probability map, this embodiment constructs spatial adjacency relationship data by calculating the spatial distance between brain regions. First, the three-dimensional coordinates of each subcortical structure in the standardized image are extracted, and the centroid coordinates of each structure are used as representatives. Then, by calculating the Euclidean distance between any two centroids, a distance matrix is obtained, and a preset spatial distance threshold (for example, set to 20 mm) is used to determine whether there is a direct spatial adjacency relationship between two regions. If the distance is less than the threshold, it is considered that the two brain regions are adjacent in space, and the closer the distance, the higher the weight. By performing reciprocal normalization on the distance, the closer brain regions correspond to higher spatial correlation weights, and finally, the brain region spatial relationship data is constructed. This data presents the spatial connection strength between regions in the form of a weighted adjacency matrix.

[0113] Step S233: Perform weighted fusion on the structural connection probability map and the brain region spatial relationship data based on the structural connection relationship and the spatial adjacency relationship to obtain a comprehensive relationship matrix.

[0114] This embodiment performs weighted fusion on the structural connection probability map and the brain region spatial relationship data to obtain a comprehensive relationship matrix that comprehensively reflects the anatomical connection and spatial adjacency relationship between brain regions. First, the two types of data are respectively normalized, mapping the structural connection probability values and spatial adjacency weights to the range of 0 to 1. Subsequently, a linear weighted fusion strategy is adopted, and the normalized structural connection probability and the normalized spatial relationship data are weighted and averaged according to a preset weight ratio (for example, setting the structural connection weight to 0.7 and the spatial adjacency weight to 0.3). The weighted result of each corresponding element in the matrix during the fusion process is the comprehensive connection value of this pair of brain regions. This value not only reflects the possibility of fiber connection between the two regions but also considers their proximity in physical space. The finally generated comprehensive relationship matrix is a symmetric matrix, which can comprehensively display the anatomical and spatial connection information between brain regions and provide a sufficient basis for spectral clustering analysis.

[0115] Step S234: Use the spectral clustering algorithm to partition the comprehensive relationship matrix to obtain functional partition data at the regional level.

[0116] In this embodiment, the spectral clustering algorithm is used to partition the foregoing comprehensive relationship matrix to obtain functional partition data at the regional level. First, a normalized Laplacian matrix of the comprehensive relationship matrix is constructed, that is, the degree matrix is constructed using the comprehensive relationship matrix and then the normalized graph Laplacian operator is calculated. The eigenvectors corresponding to the first k smallest non-zero eigenvalues are extracted through eigenvalue decomposition, where the value of k is determined according to the pre-determined number of partitions (for example, set to 10 regions) or automatically determined by the Eigengap method. Then, the representation of each brain region in these eigenvector spaces is used as the clustering feature, and the K-means clustering algorithm is used to divide all brain regions into several clusters, and each cluster represents a functional partition. During the clustering process, the number of iterations is set to 100 times, and the Euclidean distance is used as the distance metric to ensure the stability and accuracy of the clustering result. Finally, each clustering cluster is mapped back to the original brain region coordinate space to generate functional partition data at the regional level. This data not only reflects the anatomical connection advantages between brain regions but also integrates spatial adjacency information, providing a more refined and physiologically relevant partition basis for subsequent brain functional network analysis.

[0117] Preferably, step S25 includes the following steps:

[0118] Step S251: Extract the spatial position, volume, and shape features of each structure from the subcortical structure localization data, and construct eigenvectors to obtain a local-level feature matrix.

[0119] In this embodiment of the present invention, for the subcortical structure localization data, a local-level feature matrix is constructed by extracting the three-dimensional spatial position, volume, and shape features of each subcortical structure. The specific operation first uses the preprocessed standardized T1-weighted MRI image and the corresponding subcortical structure mask data, and uses the three-dimensional region growing algorithm to calculate the center coordinates of each structure as its spatial position. Subsequently, the volume is estimated by counting the number of voxels of each structure in the image, and at the same time, the shape features including the compactness of the structure boundary, the surface area to volume ratio, and the irregularity parameter are extracted. These shape indexes are obtained through boundary extraction and contour analysis. On this basis, after standardizing each feature, the principal component analysis (PCA) method is used to reduce the feature dimension, and eigenvectors composed of the eigenvalues of each feature of each structure are constructed. Finally, the eigenvectors of all subcortical structures are arranged into a matrix to form a local-level feature matrix. During the whole process, the noise in the region is filtered, and the minimum threshold of the region size is set to 50 voxels to ensure the reliability and consistency of the extracted data.

[0120] Step S252: Analyze the functional connection pattern and spatial distribution features of each partition according to the functional partition data to obtain a regional-level feature matrix.

[0121] In the embodiment of the present invention, the functional connection mode and spatial distribution characteristics of each region are analyzed by using the functional partition data, so as to construct a regional hierarchical feature matrix. The specific operation is as follows: First, the preprocessed resting-state fMRI signals are extracted from each functional partition, and the functional connection strength between each brain region in the partition is calculated by using the time series correlation analysis method to obtain the connection mode within the partition; At the same time, based on the spatial coordinates of the partition in the standardized brain map, the central position, spatial range and shape distribution characteristics of the partition are calculated, such as the eccentricity and spatial distribution dispersion of the partition, and these characteristics are obtained by statistically analyzing the position coordinates of each pixel or voxel in the region. Then, the functional connection mode characteristics and spatial distribution characteristics of each functional partition are normalized, and the functional consistency of the partition is verified by clustering analysis. Finally, the comprehensive characteristics of each partition (including the functional connection strength index and the spatial morphology index) are combined into a column of feature vectors and arranged in matrix form to form a regional hierarchical feature matrix. In actual operation, the time window is set to 60 seconds and the step size is set to 30 seconds to ensure the stability and representativeness of the time series data.

[0122] Step S253: Extract network features including topological attributes and dynamic features from the large-scale brain network data to obtain a global hierarchical feature matrix;

[0123] In the embodiment of the present invention, based on the large-scale brain network data, a network analysis method is used to extract its topological attributes and dynamic features to construct a global hierarchical feature matrix. The specific steps are as follows: First, the functional connection matrix between each network node is calculated by using the preprocessed resting-state fMRI data and the large-scale network division result, and then graph theory indicators such as node degree, clustering coefficient, betweenness centrality, global efficiency, etc. are used to measure the network topological structure, and each indicator is obtained by statistically analyzing the connection situation of the nodes in the network; At the same time, the sliding window method is applied to the time series data for dynamic network analysis to extract the time-varying features of the network connection, such as the change amplitude and stability coefficient of the dynamic connectivity. Specifically, the parameter settings of the window length of 30 seconds and the overlap rate of 50% are adopted. After obtaining the topological indicators and dynamic features, the indicator data of all nodes are standardized, and the feature fusion is carried out by using the multivariate statistical method to construct the feature vector of each large-scale network. Finally, the feature vectors of each large-scale network are integrated and arranged into a matrix to form a global hierarchical feature matrix. This matrix not only reflects the static network structure but also captures the dynamic change characteristics, providing a quantitative basis for the overall brain network function evaluation.

[0124] Step S254: Process the local hierarchical feature matrix, the regional hierarchical feature matrix, and the global hierarchical feature matrix to obtain brain region localization data.

[0125] In an embodiment of the present invention, the foregoing local-level, regional-level, and global-level feature matrices are integrated to construct the final brain region localization data. The specific operation is as follows: First, the feature matrices of the three levels are aligned in features. The corresponding relationships between the matrices are determined by matching the shared samples, and a multi-level feature fusion method (such as feature concatenation and weighted average) is used to achieve feature fusion. Different weights are set according to the signal-to-noise ratio and statistical stability of the data at each level during fusion (for example, the local-level weight is 0.4, the regional-level weight is 0.35, and the global-level weight is 0.25) to ensure that the comprehensive data reflects the advantages of the features at each level. Then, dimensionality reduction processing is performed on the fused feature data. The linear discriminant analysis (LDA) method is used to extract the most discriminative feature vectors, and the dimensionality-reduced data is mapped back to the original brain region space. Through spatial interpolation and registration techniques, a brain region localization map with anatomical and functional significance is reconstructed, where the position, volume, and morphological information of each brain region correspond to the multi-level features. The finally obtained brain region localization data not only contains the fine localization information of the anatomical structure but also reflects the functional connectivity and network dynamic characteristics, providing reliable multi-dimensional information support for subsequent disease diagnosis and analysis.

[0126] Preferably, step S3 includes the following steps:

[0127] Step S31: Partition the resting-state functional magnetic resonance imaging data in a Spark distributed environment, and divide the data into multiple time windows based on the time series, thereby obtaining time window sequence data;

[0128] In this embodiment, first, the resting-state functional magnetic resonance imaging data is loaded into the Spark distributed computing platform in NIfTI format, and the data is segmented using the built-in data partitioning function of Spark. The entire dataset is divided into multiple consecutive time windows according to the order of the time series. The length of each window is set to 30 seconds according to the actual data acquisition parameters (for example, if the TR is 2 seconds, each window contains 15 time points). During the partitioning process, a custom function is used to ensure the integrity of the time order of the data within each partition. The finally obtained time window sequence data is stored in the form of an RDD or DataFrame, providing an efficient data access and computing basis for subsequent parallel processing.

[0129] Step S32: Perform multi-scale decomposition based on wavelet transform on the time window sequence data respectively, decompose the noise signal into different frequency bands, thereby obtaining wavelet band coefficient data;

[0130] In this embodiment, the multi-scale decomposition method based on wavelet transform is adopted for each time window sequence data, and the signal is decomposed at different scales to distinguish noise and real signals. Specifically, in terms of operation, the Daubechies wavelet basis (such as db4) is first selected, and the discrete wavelet transform is performed on the time series data in each window. The decomposition level is set to 5 layers. The original signal is decomposed into low-frequency approximation coefficients and high-frequency detail coefficients through a recursive filter bank. Among them, the high-frequency part mainly contains noise components, while the low-frequency part contains useful neural activity signals. The decomposition results are stored in the form of a wavelet frequency band coefficient data matrix. The coefficient data of each layer corresponds to the information of different frequency bands, providing a basis for subsequent noise processing.

[0131] Step S33: Evaluate the signal-to-noise ratio of each frequency band according to the wavelet frequency band coefficient data, and determine the adaptive noise threshold, so as to obtain the frequency band noise threshold data;

[0132] In this embodiment, based on the obtained wavelet frequency band coefficient data, the signal-to-noise ratio (Signal-to-Noise Ratio, SNR) of each frequency band is evaluated. The specific method is to statistically calculate the mean and standard deviation of the coefficients within each decomposition layer, and determine the SNR of each frequency band by comparing the amplitude differences between the signals and noises in different frequency bands. Then, an adaptive threshold algorithm (such as the method based on local maxima and median absolute deviation) is adopted to determine the noise threshold. The determination process of the noise threshold data for each frequency band takes into account the discrete degree of the coefficient distribution in this frequency band. The threshold is set to the statistic multiplied by a preset factor (such as 1.5 times). The finally generated frequency band noise threshold data is recorded in the form of an array, and each element corresponds to the threshold of a decomposition layer, which is used to guide the subsequent soft threshold processing.

[0133] Step S34: Perform soft threshold shrinkage processing on the wavelet frequency band coefficient data by using the frequency band noise threshold data to suppress the noise components, so as to obtain the denoised wavelet coefficient data;

[0134] In this embodiment, the soft threshold shrinkage processing is performed on each wavelet frequency band coefficient data by using the frequency band noise threshold data determined in step S33. Specifically, in terms of operation, for each wavelet coefficient, if its absolute value is less than the noise threshold of the corresponding frequency band, the coefficient is set to zero; if it is greater than the threshold, the coefficient is shrunk according to the threshold, that is, its amplitude is subtracted by the threshold and the original sign is retained, so as to achieve smooth transition and reduce the artifacts and discontinuities that may be introduced by the hard threshold processing. The coefficient data after the soft threshold processing retains the main features of the signal and effectively suppresses the noise components at the same time. The finally generated denoised wavelet coefficient data provides a clean basis for the reconstructed signals decomposed at each layer for each time window.

[0135] Step S35: Perform inverse wavelet transform reconstruction on the wavelet coefficient data and conduct parallel computing in the Spark environment to obtain preliminary noise-reduced sequence data;

[0136] In this embodiment, inverse wavelet transform reconstruction is performed on the wavelet coefficient data after soft threshold processing to restore the noise-reduced time series data. When operating specifically, the same wavelet basis (such as db4) and decomposition level as the forward wavelet transform are selected. Through inverse transform, the coefficients in each frequency band within each time window are combined and reconstructed into a complete time series signal. To improve the computing efficiency, this inverse transform process is carried out in the Spark distributed environment through the MapReduce paradigm for parallel computing, ensuring that the data within each partition is independently processed and the results are seamlessly integrated. The finally output preliminary noise-reduced sequence data has a higher signal-to-noise ratio and less high-frequency noise interference in the time domain.

[0137] Step S36: Perform physiological noise removal processing including heartbeat and respiration on the preliminary noise-reduced sequence data to obtain noise-corrected sequence data;

[0138] Based on the preliminary noise-reduced sequence data, this embodiment further removes the influence of physiological noises such as heartbeat and respiration on the signal. First, a method based on spectral analysis is used to detect the typical heartbeat frequency (usually in the range of 0.8 to 1.2 Hz) and respiration frequency (usually in the range of 0.2 to 0.4 Hz) in the signal. Subsequently, a corresponding filter is designed using a band-stop filter to suppress the noises in these frequency bands. The bandwidth and attenuation parameters of the filter are adaptively adjusted according to the specific physiological parameters of the subject (such as the pre-measured heart rate and respiration frequency). In addition, independent component analysis (ICA) can be combined to further separate and remove the physiological noise components. After these processes, the obtained noise-corrected sequence data significantly reduces the energy of the heartbeat and respiration noises in the frequency domain, making the neural signal components more prominent.

[0139] Step S37: Perform head movement parameter correction and spatial smoothing processing on the noise-corrected sequence data to obtain brain functional time series data.

[0140] In this embodiment, the noise-corrected sequence data is further subjected to head motion parameter correction and spatial smoothing processing to obtain the final brain functional time series data. Specifically, during the operation, first, a head motion correction algorithm (such as an algorithm based on six-parameter rigid registration) is used to correct the small movements generated by the subject during the scanning process. The displacement and rotation errors are estimated and corrected by matching the images at each time point with the reference image. The correction process is accelerated by parallel computing in the Spark environment. Then, a Gaussian smoothing kernel is used for spatial smoothing processing, and the standard deviation of the smoothing kernel is set to 6 mm to adapt to the spatial resolution of the fMRI data. The smoothing processing helps to improve the signal-to-noise ratio and the consistency of the signals within the region. After these preprocessing steps, the finally output brain functional time series data has high spatio-temporal consistency and low noise interference, which is suitable for subsequent functional connectivity and network analysis research.

[0141] Preferably, step S4 includes the following steps:

[0142] Step S41: Partition the brain functional time series data according to the brain region localization data in the Spark distributed environment, and perform standardization processing on the time series within each level, so as to obtain the standardized time series data;

[0143] In this embodiment, the preprocessed brain functional time series data is partitioned in the Spark distributed environment. First, the data is labeled according to the previously generated brain region localization data, and a mapping relationship is established between each time series and its corresponding brain region or level (such as subcortical structures, functional partitions, or large-scale networks). Then, the entire data set is segmented according to these mapping labels by using the partition function of Spark, ensuring that each partition only contains the time series data within the same brain region or the same level. Next, a standardization processing method is adopted for the time series data within each partition, that is, first calculate the mean and standard deviation of each time series, and then perform the transformation of subtracting the mean and dividing by the standard deviation on the value at each time point (commonly called Z-score standardization) to eliminate the influence of the signal amplitude differences in different brain regions. The standardized data is saved in the DataFrame of Spark, and each record contains the standardized time series, the corresponding brain region label, and the time index. The partitioning process uses a custom partitioner to ensure the balanced distribution of the data. Usually, each time window length is set to 30 seconds (for example, when the TR is 2 seconds, it contains 15 data points) to obtain sufficient statistical stability, and at the same time, linear interpolation is used to compensate for the data missing situation. The finally obtained standardized time series data provides a unified, smooth, and parallel-processable basic data for subsequent functional connectivity calculations.

[0144] Step S42: Calculate the functional connectivity strength between subcortical structures within the local level based on the standardized time series data, and quantify it through Pearson correlation coefficient analysis, so as to obtain the local-level functional connectivity matrix;

[0145] In this embodiment, the standardized time series data obtained in step S41 is used to quantify the functional connectivity strength for the connections between subcortical structures at the local level. First, for each pair of subcortical structures, signal sequences are extracted from their corresponding standardized time series, and then the Pearson correlation coefficient analysis method is used for calculation, that is, the covariance between each pair of sequences is calculated and divided by the product of their respective standard deviations, so as to obtain a correlation coefficient between -1 and 1. The closer the value is to 1, the higher the signal synchronization between the two regions, and the closer it is to -1, the more it represents reverse synchronization. During the calculation process, the data of each subject is analyzed independently, and the window length is usually the same as the time window used in S41 to ensure temporal matching; all the calculated correlation coefficients are arranged into a matrix according to the labels of the subcortical structures. Each row and each column represent a subcortical structure. This matrix is the local-level functional connectivity matrix, and the matrix is stored in a distributed file system for further processing. At the same time, in order to improve the calculation accuracy, robust correction is performed on extreme values to ensure that the functional connectivity strength reflected in the matrix has high biological significance and statistical reliability.

[0146] Step S43: Calculate the functional connectivity for the functional partitions at the regional level based on the local-level functional connectivity matrix, and perform weighted integration based on the time series correlation of the internal nodes of the region, so as to obtain the regional-level functional connectivity matrix;

[0147] In this embodiment, based on the local-level functional connectivity matrix, the functional partitions at the regional level are integrated and weighted. First, according to the preset functional partition data, the time series data of the subcortical structures belonging to the same functional partition at the local level are aggregated, and the inter-region connection index is constructed by calculating the weighted average correlation coefficient between all nodes within the partition. The weighting coefficient can be assigned according to the signal-to-noise ratio or signal stability of each node, ensuring that nodes with stronger or more stable signals within the functional partition play a greater role in the overall weighting; subsequently, the correlation calculation between regions is performed on the aggregated data, similar to the Pearson correlation coefficient method used in S42, so as to construct a new matrix, and each element represents the comprehensive functional connectivity strength between two functional partitions. The sliding window analysis method is used throughout the process to capture time-varying features. The window length is usually set to 60 seconds with 50% overlap to balance the time domain resolution and statistical stability. This regional-level functional connectivity matrix not only reflects the functional integration within the region but also reflects the interaction between different regions. The matrix data is stored in a distributed environment after normalization for subsequent network analysis.

[0148] Step S44: Evaluate the node importance and connection significance through the network theory method based on the regional hierarchical functional connection matrix, and determine the functional connection pattern of the global hierarchical large-scale brain network, so as to obtain the global hierarchical functional connection matrix;

[0149] In this embodiment, after obtaining the regional hierarchical functional connection matrix, the importance of each functional partition (node) is evaluated and the connection significance is detected through the network theory method to determine the functional connection pattern of the global hierarchical large-scale brain network. First, graph theory indexes are calculated for each node in the matrix, such as node degree, betweenness centrality, and clustering coefficient. The node degree reflects the number of regions connected to other regions, the betweenness centrality quantifies the bridging role of the node in the whole network, and the clustering coefficient is used to evaluate the tightness of the connections within the neighborhood of the node; these indexes are obtained through statistical analysis methods, and the Z-score normalization is used to unify the scales of different indexes. Then, a non-parametric statistical test method (such as permutation test) is used to judge the significance level of each connection, and the connections with significance lower than the preset p-value threshold (such as 0.05) are removed; in addition, the topological structure of the remaining connections is optimized, and the modular algorithm (such as Louvain algorithm) is used to identify the functional modules in the large-scale brain network. The finally generated global hierarchical functional connection matrix not only reflects the connection strength between regions, but also contains the relative importance of each node in the network and the module membership information, providing a detailed structural basis for the analysis of the overall network functional pattern. All processing parameters can be adjusted according to experimental data and prior knowledge to ensure the physiological relevance and statistical robustness of the network model.

[0150] Step S45: Dynamically adjust the weights between levels of the local hierarchical functional connection matrix, regional hierarchical functional connection matrix, and global hierarchical functional connection matrix through the principle of synaptic plasticity, and establish the mapping relationship of hierarchical connections, so as to obtain the brain functional network feature data.

[0151] In this embodiment, based on the principle of synaptic plasticity, dynamic weight adjustment is performed on the functional connection matrices at the local level, regional level, and global level to establish the connection mapping relationship between levels. First, the three-level connection matrices are normalized and then preliminarily fused. Then, based on the principle of the synaptic plasticity model (such as Spike-Timing-Dependent Plasticity, STDP), dynamic update rules are assigned to each cross-level connection. The specific operation is to detect the time series synchronization between cross-level nodes. If synchronous activation is detected, the connection weight is correspondingly increased, otherwise it is inhibited. This process is performed in parallel in the Spark environment through an iterative algorithm. During the iterative process, the mutual influence between the local connection matrix, regional connection matrix, and global connection matrix is considered for each update. The textual description of the weight update formula is: the current connection weight is equal to the weight at the previous moment plus an update factor proportional to the degree of synchronization, and at the same time subtract a penalty term related to the asynchronous component. The iteration continues until the global convergence condition is met (such as the update factor is less than the preset threshold of 0.001). The finally generated brain functional network feature data is stored in the form of a comprehensive matrix, which not only reflects the original strength of the functional connections at each level but also embodies the hierarchical mapping relationship after dynamic adjustment, providing multi-level functional connection features with high dynamic range and biological significance for subsequent pathological feature analysis and clinical decision-making.

[0152] Preferably, step S45 includes the following steps:

[0153] Step S451: Normalize the functional connection matrix at the local level, regional level, and global level, and perform data integration in the Spark distributed environment to construct the hierarchical connection mapping matrix data;

[0154] In this embodiment, first, the functional connection matrices at the local level, regional level, and global level are respectively normalized. The method based on min-max normalization is adopted, that is, the value of each element in each matrix is subtracted by the minimum value in the matrix and then divided by the difference between the maximum value and the minimum value, so that the numerical range of all matrices is standardized to between 0 and 1, thereby eliminating the bias caused by different data scales at each level. The normalized matrices are then integrated in the Spark distributed environment through the way of Join (union dataset). Using the DataFrame API of Spark, the normalized local, regional, and global matrices are aligned and merged according to the brain region correspondence relationship. Finally, a hierarchical connection mapping matrix data is constructed. This mapping matrix stores the comprehensive value of the cross-level connections of each brain region in the form of a multi-dimensional array, providing a unified data basis for subsequent dynamic weight update. At the same time, the data partition size is set to no more than 5000 data units per partition during the data integration process to ensure the efficiency and stability of distributed computing.

[0155] Step S452: Construct a weight update mathematical model based on the principle of synaptic plasticity according to the hierarchical connection mapping matrix data, and perform initialization processing on each cross-hierarchical connection weight, so as to obtain the initial cross-hierarchical connection weight matrix data;

[0156] Based on the hierarchical connection mapping matrix data constructed in step S451, this embodiment constructs a weight update mathematical model according to the principle of synaptic plasticity, and performs initialization processing on the cross-hierarchical connection weight. In the specific operation, a dynamic weight update framework is first designed. The core idea is to regard each cross-hierarchical connection weight as the product of the initial connection strength and the subsequent activity regulation factor. Then, a preset initialization function is applied to each element in the mapping matrix. This function is calculated according to the normalized value in the mapping matrix and a constant initialization factor (for example, setting the initial factor as a random value or a fixed value between 0.5 and 1.0), so as to set the initial value of each connection weight within a reasonable range; The initialization process is implemented by using the Map operation in the Spark environment. After the weights in each data partition are independently initialized, they are globally merged. The finally obtained initial cross-hierarchical connection weight matrix data provides a starting condition for subsequent iterative updates, and the initialization parameters can be fine-tuned according to the actual data distribution to ensure that the model has good convergence and biological rationality during subsequent weight adjustment.

[0157] Step S453: Perform synchronization analysis on the time series activation of cross-hierarchical nodes according to the initial cross-hierarchical connection weight matrix data, so as to obtain preliminary dynamic connection weight data, where the synchronization analysis includes calculating the update factors of positive reinforcement and negative inhibition;

[0158] Based on the initial cross-hierarchical connection weight matrix data, this embodiment performs synchronization analysis on the time series activation of cross-hierarchical nodes to obtain preliminary dynamic connection weight data. In the specific operation, the standardized time series data of the relevant brain regions of each cross-hierarchical connection is first extracted, and then the cross-correlation analysis method of time series is used to evaluate the synchronization of activations between different brain regions. By comparing the phase and amplitude of the activation signals, an update factor reflecting positive reinforcement and negative inhibition effects is calculated. The determination process of this update factor includes counting the peak synchronization events within the signal window and evaluating the inhibition in the low-correlation period. The positive update factor reflects that the connection weight should be appropriately increased in the case of high synchronization, while the negative update factor promotes the reduction of the connection weight in the case of low synchronization or anti-phase activity; The synchronization analysis uses the rolling window technique (for example, the window length is set to 30 seconds and the step size is 15 seconds) for local statistics, and the update factors of each connection are calculated in parallel on the Spark distributed platform. Finally, these factors are combined with the initial weights to form preliminary dynamic connection weight data, which provides a feedback basis and preliminary dynamic response index for subsequent iterative adjustment.

[0159] Step S454: Perform iterative dynamic weight adjustment processing on the preliminary dynamic connection weight data in the Spark distributed environment, and correct the connection weights according to the preset global convergence condition, so as to obtain the cross-level connection weight matrix data;

[0160] In this embodiment, iterative dynamic weight adjustment processing is performed on the preliminary dynamic connection weight data in the Spark distributed environment to achieve global optimization of the connection weights. In the specific operation, an iterative algorithm is adopted. In each iteration, the weights are updated according to the difference between the current cross-level connection weight data and the previous iteration result. The update process follows the preset global convergence condition (for example, it is considered converged when the change amplitude of all connection weights is less than 0.001). In each iteration, the aforementioned synchronization update factor is used to perform weighted adjustment on each cross-level connection, and an adaptive learning rate parameter is introduced (for example, the initial learning rate is set to 0.05 and gradually decays according to the number of iterations) to ensure stable convergence of the update process; on the Spark platform, the iterative process triggers Map and Reduce operations through loops to achieve local update of the weights within each data partition and then global integration until the convergence condition is met. The finally output cross-level connection weight matrix data not only has higher stability but also can more realistically reflect the biological characteristics of the dynamic interaction between brain regions.

[0161] Step S455: Establish a mapping relationship of hierarchical connections according to the cross-level connection weight matrix data, so as to obtain the brain functional network feature data.

[0162] In this embodiment, based on the finally obtained cross-level connection weight matrix data, a mapping relationship of hierarchical connections is established using these data. In the specific operation, first, the updated weight matrix is rearranged according to brain regions and levels, and the connection weights between each pair of brain regions are statistically summarized to form a mapping table that comprehensively describes the connection situation between levels. This mapping table is represented in the form of a multi-dimensional matrix or graph structure, where each node represents a brain region, and the edge weight between nodes is the final weight value of the cross-level connection; subsequently, community detection algorithms in graph theory (such as the Louvain algorithm) and visualization tools are used to analyze the mapping relationship to verify whether the connection pattern between levels conforms to the expected synaptic plasticity principle and physiological laws; finally, the mapping relationship data is saved in a standardized format (such as CSV or Parquet format) and stored in a distributed file system as brain functional network feature data for subsequent clinical decision-making and machine learning model training. In specific application scenarios, this data can help researchers discover abnormal connection patterns in the brain network and provide quantitative support for the diagnosis of diseases such as schizophrenia.

[0163] Preferably, step S5 includes the following steps:

[0164] Step S51: Group the brain functional network feature data according to three levels: local, regional, and global. Calculate the Fisher discriminant scores for the functional connectivity features at each level respectively, so as to obtain the hierarchical Fisher score data;

[0165] In this embodiment, after grouping the obtained brain functional network feature data according to three levels: local, regional, and global, the Fisher discriminant analysis method is used to evaluate each functional connectivity feature within each level, that is, calculate the mean difference and the ratio of within-group variance of the schizophrenia patient group and the healthy control group on this feature respectively. In specific operations, first calculate the mean and standard deviation of each feature in the two groups of data respectively, and then calculate the discriminant score. The literal expression of this score is the inter-group mean difference divided by the sum of within-group standard deviations. The higher the value, the stronger the ability of this feature to distinguish the two groups. The Fisher score data of all features are stored in a distributed database according to levels. Generally, an unbiased estimation method is used when calculating the standard deviation, and parallel computing is performed on the Spark platform to improve efficiency, so as to obtain the Fisher discriminant score data of each level of local, regional, and global.

[0166] Step S52: Establish a knowledge base of brain functional connectivity patterns related to schizophrenia based on the pre-collected clinical data of schizophrenia patients and healthy control groups, so as to obtain prior knowledge of schizophrenia;

[0167] Based on the pre-collected detailed clinical data of schizophrenia patients and healthy control groups in this embodiment, including symptom assessment, cognitive tests, and neuroimaging data, a knowledge base of brain functional connectivity patterns related to schizophrenia is constructed. Specifically, by integrating multi-modal data and literature reports from multiple centers, data mining and machine learning methods (such as clustering and pattern recognition algorithms) are used to identify abnormal functional connectivity patterns that repeatedly appear in schizophrenia, and combined with expert annotations to form a prior knowledge base. This knowledge base is stored in the form of a relational database, and descriptive statistics and clinical relevance indicators are attached. In the application scenario, a data set with a sample size of not less than 100 cases is often used to ensure statistical robustness, and finally a set of prior knowledge of schizophrenia for guiding feature screening is obtained.

[0168] Step S53: Set different importance weights for the functional connectivity features at different levels based on the prior knowledge of schizophrenia, and perform weighted processing on the hierarchical Fisher score data, so as to obtain weighted feature importance data;

[0169] In this embodiment, according to the prior knowledge of schizophrenia obtained in step S52, different levels of functional connection features are assigned different importance weights. The specific operation is as follows: First, preliminary weights are set for each level of features based on the key brain networks and connection patterns recorded in the prior knowledge. For example, higher weights are assigned to the connections in the default mode network and the central executive network, while lower weights are given to the less relevant networks. Then, the Fisher score data of each level is multiplied by the corresponding weight and weighted average processing is performed, so that the features with clinical indication significance are more prominent in the weighted feature importance data. This process is parallelly calculated through the Map operation on the Spark platform, and the weight factor used is usually dynamically adjusted between 0.5 and 2.0. The finally output weighted feature importance data can more accurately reflect the actual contributions of each feature in schizophrenia diagnosis.

[0170] Step S54: Determine the feature selection threshold based on the distribution characteristics and quantity of features at different levels according to the weighted feature importance data, so as to obtain the hierarchical threshold data;

[0171] In this embodiment, based on the weighted feature importance data obtained in step S53, a detailed analysis is carried out on the statistical characteristics (such as mean, standard deviation, and quantile) of feature distributions at different levels. An adaptive threshold determination method is adopted. In the specific operation, first, the histogram and box plot of the feature importance data are drawn to identify the outliers and central tendency in the data. Then, according to the preset distribution characteristics (such as selecting the 75th percentile or using the mean plus one standard deviation as the threshold), the feature selection threshold for each level is determined. In terms of parameter settings, in the application scenario, the threshold can be set at the upper-middle level of the weighted importance data, such as the 80% cumulative distribution position, to ensure that only the features with higher discrimination are selected. Finally, the threshold data corresponding to each level is saved as an independent threshold set.

[0172] Step S55: Screen the weighted feature importance data by using the hierarchical threshold data to obtain the candidate feature set;

[0173] In this embodiment, the hierarchical threshold data obtained in step S54 is used to screen the weighted feature importance data. The specific operation adopts a layer-by-layer comparison method. The weighted score of each functional connection feature is compared with the threshold of the corresponding level. If the feature score is higher than the threshold, it is retained; otherwise, it is excluded. At the same time, in the Spark distributed environment, the Filter operation is used to quickly screen the data set to ensure balanced data distribution and feature representativeness. The screened candidate feature set is saved in matrix form, where each column represents a screened functional connection feature. This candidate feature set provides preliminary input for further optimization and redundancy analysis. Usually, it is required that the number of features included in the candidate feature set accounts for 10% to 30% of the original feature set to balance information richness and computational burden.

[0174] Step S56: Perform redundancy analysis on the candidate feature set, calculate the correlation between features, and remove highly correlated redundant features according to a preset threshold, so as to obtain an optimized feature set;

[0175] In this embodiment, redundancy analysis is performed on the candidate feature set screened in step S55. Specifically, in the operation, the Pearson correlation coefficient matrix between each feature is calculated by using the correlation analysis method. All feature pairs in the matrix are traversed to identify those feature pairs with a correlation coefficient higher than the preset threshold (such as 0.9). These features are considered to have high redundancy. Then, the stepwise elimination method is adopted, and the features that perform more prominently in Fisher score and weighted importance are preferentially retained. At the same time, principal component analysis (PCA) is used to assist in verifying the contribution rate and independence of the features. The optimized feature set formed after removing redundant features has lower multicollinearity and higher information uniqueness. The whole process is carried out by using distributed matrix operations in the Spark environment to ensure the efficiency and accuracy when processing large-scale data.

[0176] Step S57: Match and verify the optimized feature set with the prior knowledge of schizophrenia to generate the biological significance and clinical interpretability of the optimized features, so as to obtain the schizophrenia feature set.

[0177] In this embodiment, the optimized feature set obtained in step S56 is matched and verified with the prior knowledge of schizophrenia established in step S52. Specifically, in the operation, by comparing the physiological and clinical descriptions corresponding to each optimized feature in the prior knowledge base, the biological significance of the features is verified by using text matching and rule-based methods. For example, it is checked whether the feature is involved in the known abnormal brain networks in schizophrenia (such as the default mode network or the central executive network) and whether its connection pattern conforms to the literature reports. At the same time, neuroscience experts are invited to conduct qualitative evaluations to ensure that each optimized feature has clinical interpretability and biological significance. Finally, the features that pass the matching verification are integrated to form the schizophrenia feature set. This feature set is output in the form of a multi-dimensional array and an annotation table for subsequent machine learning model training and clinical decision support system use. The parameters and matching rules of the whole process are dynamically adjusted according to specific clinical data and expert feedback to achieve the best matching effect.

Claims

1. An image feature selection method for the Spark model of schizophrenia medical records, characterized in that: The following steps are involved: Step S1: acquiring brain magnetic resonance image data and resting-state functional magnetic resonance imaging data of multiple subjects; Step S2: Using a preset hierarchical brain region template to perform multi-scale hierarchical segmentation on the brain magnetic resonance image data of each subject, and obtain brain region positioning data including three levels: local, regional and global. The multi-scale hierarchical segmentation specifically includes subcortical structure segmentation at the local level, functional partition segmentation based on anatomical connection by calculating the structural connection probability and spatial adjacency between brain regions at the regional level, and large-scale brain network segmentation based on hierarchical organization analysis of resting-state networks at the global level. Step S2 includes: Step S21: performing image resolution adaptive normalization processing on the brain magnetic resonance image data, and obtaining normalized image data with a spatial resolution of 1 mm isotropic voxel by resampling the voxel size based on B-spline interpolation; Step S22: performing subcortical structure segmentation on the standardized image data according to a preset hierarchical brain region template, and performing brain region positioning processing, thereby obtaining subcortical structure positioning data at a local level; Step S23: performing functional partitioning clustering based on anatomical connection by calculating the structural connection probability and spatial adjacency relationship between brain regions according to the local-level subcortical structure positioning data, thereby obtaining functional partitioning data at the regional level; Step S24: performing a hierarchical organization analysis of the cerebral cortex based on the resting-state network according to the functional partitioning data and the resting-state functional magnetic resonance imaging data, thereby obtaining large-scale brain network data at the global level, wherein the hierarchical organization analysis specifically calculates the time series correlation and spatial continuity between the functional partitions; Step S25: hierarchically integrating the subcortical structure positioning data, the functional partitioning data, and the large-scale brain network data, and establishing a corresponding relationship matrix between the hierarchies, thereby obtaining brain region positioning data; Step S3: performing adaptive hierarchical denoising on the resting-state functional magnetic resonance imaging data in a Spark distributed environment to obtain denoised brain function time series data; Step S4: construct a hierarchical functional connection calculation framework based on the brain area positioning data and the brain function time series data, calculate the functional connection strength within each level and between levels, and dynamically adjust the connection weights between levels based on the principle of synaptic plasticity to obtain brain function network feature data. Step S4 includes: Step S41: partitioning the brain function time series data according to the brain region positioning data in the Spark distributed environment, and standardizing the time series in each level to obtain standardized time series data; Step S42: Calculate the functional connectivity strength between subcortical structures in the local level according to the standardized time series data, and quantify it through Pearson correlation coefficient analysis to obtain a local level functional connectivity matrix; Step S43: Calculate the functional connectivity of the functional partitions at the regional level according to the local level functional connectivity matrix, and perform weighted integration based on the time series correlation of the nodes inside the region, so as to obtain the regional level functional connectivity matrix; Step S44: evaluating node importance and connection significance through network theory methods according to the regional level functional connection matrix, and determining the functional connection pattern of the global level large-scale brain network, thereby obtaining the global level functional connection matrix; Step S45: dynamically adjusting the inter-level weights of the local hierarchical functional connection matrix, the regional hierarchical functional connection matrix, and the global hierarchical functional connection matrix through the principle of synaptic plasticity, establishing a mapping relationship of hierarchical connections, and thereby obtaining brain functional network feature data; Step S5: Performing a hierarchical feature importance evaluation based on Fisher score on the brain functional network feature data, and setting a hierarchical adaptive threshold for feature screening according to the pre-acquired schizophrenia prior knowledge to obtain a schizophrenia feature set.

2. The image feature selection method for the schizophrenia medical record Spark model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: performing T1-weighted magnetic resonance scanning on the brains of multiple subjects using a 3T magnetic resonance imaging device, thereby obtaining brain magnetic resonance image data, wherein the imaging contrast of the brain magnetic resonance image data is specifically determined by the T1 relaxation time; Step S12: using a 7T superconducting magnetic resonance imaging device to perform blood oxygen level-dependent contrast acquisition based on a gradient echo sequence on the brains of multiple subjects in a resting state, thereby obtaining blood oxygen level-dependent contrast data; Step S13: Analyze the local blood flow changes caused by neuronal activities on the blood oxygen level dependence contrast data, so as to obtain resting state functional magnetic resonance imaging data.

3. The image feature selection method for the schizophrenia medical record Spark model according to claim 1, characterized in that: Step S22 includes the following steps: Step S221: performing edge enhancement processing by calculating the spatial gradient amplitude and direction of the three-dimensional image according to the standardized image data, thereby obtaining boundary enhancement data of the subcortical structure; Step S222: Analyze the grayscale distribution characteristics of the local area according to the boundary enhancement data, set a dynamic threshold, and perform segmentation, thereby obtaining preliminary segmentation mask data of the subcortical structure; Step S223: morphologically optimizing the preliminary segmentation mask data by using morphological operations, so as to obtain continuous and complete subcortical structure mask data, wherein the morphological optimization specifically processes holes and boundary discontinuities by designing a three-dimensional morphological operator sequence; Step S224: Perform connected domain analysis based on the subcortical structure mask data and mark different subcortical structures, so as to obtain fine subcortical structure positioning data at the local level.

4. The image feature selection method for the schizophrenia medical record Spark model according to claim 3, characterized in that: Step S23 includes the following steps: Step S231: performing connection fiber bundle tracing on the local level subcortical structure positioning data, and performing white matter fiber connection analysis between subcortical structures, thereby obtaining a structural connection probability map; Step S232: Calculating the spatial adjacency relationship between brain regions based on the spatial distance metric according to the structural connection probability map, thereby obtaining brain region spatial relationship data; Step S233: performing weighted fusion on the structural connection probability map and the brain region spatial relationship data based on the structural connection relationship and the spatial adjacency relationship, thereby obtaining a comprehensive relationship matrix; Step S234: partition the comprehensive relationship matrix using a spectral clustering algorithm, thereby obtaining functional partition data at the regional level.

5. The image feature selection method for the schizophrenia medical record Spark model according to claim 4, characterized in that: Step S25 includes the following steps: Step S251: extracting the spatial position, volume and shape features of each structure from the subcortical structure positioning data, and constructing a feature vector to obtain a local hierarchical feature matrix; Step S252: Analyze the functional connection pattern and spatial distribution characteristics of each partition according to the functional partition data, so as to obtain a regional level feature matrix; Step S253: extracting network features including topological attributes and dynamic features from the large-scale brain network data, thereby obtaining a global hierarchical feature matrix; Step S254: local level feature matrix, regional level feature matrix and global level feature matrix are processed to obtain brain region localization data.

6. The image feature selection method for the schizophrenia medical record Spark model according to claim 5, characterized in that: Step S3 includes the following steps: Step S31: partitioning the resting state functional magnetic resonance imaging data in the Spark distributed environment, dividing the data into multiple time windows based on the time series, thereby obtaining time window series data; Step S32: performing multi-scale decomposition based on wavelet transform on the time window sequence data, decomposing the noise signal into different frequency bands, thereby obtaining wavelet frequency band coefficient data; Step S33: evaluating the signal-to-noise ratio of each frequency band according to the wavelet frequency band coefficient data, and performing adaptive noise threshold determination, thereby obtaining frequency band noise threshold data; Step S34: using the frequency band noise threshold data to perform soft threshold shrinkage processing on the wavelet frequency band coefficient data to suppress the noise component, thereby obtaining the denoised wavelet coefficient data; Step S35: reconstruct the wavelet coefficient data by inverse wavelet transform, and perform parallel calculation in the Spark environment to obtain preliminary denoised sequence data; Step S36: performing physiological noise removal processing including heartbeat and breathing on the preliminary noise reduction sequence data, thereby obtaining noise correction sequence data; Step S37: performing head motion parameter correction and spatial smoothing processing on the noise-corrected sequence data to obtain brain function time series data.

7. The image feature selection method for the schizophrenia medical record Spark model according to claim 6, characterized in that: Step S45 includes the following steps: Step S451: normalizing the local level functional connectivity matrix, the regional level functional connectivity matrix, and the global level functional connectivity matrix, and integrating the data in the Spark distributed environment to construct the hierarchical connectivity mapping matrix data; Step S452: constructing a weight update mathematical model based on the synaptic plasticity principle according to the hierarchical connection mapping matrix data, initializing each cross-level connection weight, thereby obtaining initial cross-level connection weight matrix data; Step S453: performing synchronization analysis on the time series activation of cross-level nodes according to the initial cross-level connection weight matrix data, thereby obtaining preliminary dynamic connection weight data, wherein the synchronization analysis includes calculating the update factors of positive reinforcement and reverse inhibition; Step S454: performing iterative dynamic weight adjustment processing on the preliminary dynamic connection weight data in the Spark distributed environment, and correcting the connection weight according to the preset global convergence condition, so as to obtain cross-level connection weight matrix data; Step S455: Establishing a hierarchical connection mapping relationship according to the cross-hierarchical connection weight matrix data, thereby obtaining brain function network feature data.

8. The image feature selection method for the schizophrenia medical record Spark model according to claim 7, characterized in that: Step S5 includes the following steps: Step S51: grouping the brain function network feature data into three levels: local, regional and global, and calculating the Fisher discriminant score for the functional connection features of each level, thereby obtaining hierarchical Fisher score data; Step S52: establishing a knowledge base of brain functional connection patterns related to schizophrenia based on the pre-collected clinical data of schizophrenia patients and healthy controls, thereby obtaining prior knowledge of schizophrenia; Step S53: setting differentiated importance weights for functional connectivity features at different levels based on prior knowledge of schizophrenia, and weighting the hierarchical Fisher score data to obtain weighted feature importance data; Step S54: determining feature selection thresholds at different levels based on the distribution characteristics and quantity of features according to the weighted feature importance data, thereby obtaining hierarchical threshold data; Step S55: using the hierarchical threshold data to screen the weighted feature importance data, thereby obtaining a candidate feature set; Step S56: performing redundancy analysis on the candidate feature set, calculating the correlation between features, and removing highly correlated redundant features according to a preset threshold, thereby obtaining an optimized feature set; Step S57: Matching and testing the optimized feature set with the prior knowledge of schizophrenia to generate the biological significance and clinical interpretability of the optimized features, thereby obtaining the schizophrenia feature set.

Citation Information

Patent Citations

  • Brain toughness evaluation method and system based on magnetic resonance image and electronic equipment

    CN116823813A

  • Depression detection method and device based on pulse mixed supervision graph attention network

    CN118000751A

  • Individual function brain region subdivision method and device based on improved density clustering

    CN119067997A

  • Parkinson's disease early warning system and method based on graph neural network

    CN119339943A