A method for associating brain images with brain tissue genes
Patent Information
- Application Number
- CN202410556048.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-07
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-05-07
AI Technical Summary
[0008]本发明提供了一种脑影像与脑组织基因关联方法,以解决现有技术无法精确的分析基因组与脑部区域之间的相关性的技术问题
[0019] The advantage of this invention lies in its method for associating brain imaging with brain tissue genes, which can accurately locate brain regions and analyze the correlation between the genome and brain regions. It automatically locates targeted brain regions while simultaneously performing genomic analysis. This algorithm will improve the efficiency and reliability of brain research results, helping us to better understand the pathogenesis of brain diseases and develop more effective treatments.
Smart Images

Figure CN118314966B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method for linking brain images with brain tissue genes. Background Technology
[0002] Brain diseases pose a significant challenge to health worldwide, profoundly impacting the lives of patients and their families. These diseases encompass neurodegenerative disorders, mental disorders, and neurological illnesses such as Alzheimer's, Parkinson's, and depression. The causes of these diseases are complex and diverse, and our understanding of their mechanisms and treatments remains limited.
[0003] Currently, decoding brain genome information provides a powerful solution for elucidating the pathogenesis of various brain diseases and finding effective treatments. Clinically, obtaining brain gene data largely relies on manual sampling by neurologists. This method is not only invasive but also difficult to obtain data from some inoperable functional brain regions (such as the brainstem and thalamus). To better address this limitation, scientists have adopted a comprehensive research approach that combines brain imaging and genomic data analysis. Existing brain imaging techniques, such as magnetic resonance imaging (MRI) and positron emission tomography (PET), are the main means of quantitatively assessing abnormalities in brain structure and function in the context of brain diseases. Therefore, how to correlate brain imaging diagnosis with the analysis of genomic data has become a link connecting macroscopic to microscopic research on brain diseases.
[0004] However, the analysis of brain imaging and genomic data presents several challenges. First, processing brain imaging data is time-consuming and labor-intensive, and is often affected by various factors such as head movement and signal noise. Second, processing massive amounts of genomic data requires complex data transformation and statistical analysis, and the interpretation and validation of results require specialized knowledge and experience.
[0005] Matching and associating brain imaging and genomic data is a complex task that requires consideration of both spatial and genomic factors. To address these challenges, a fully automated brain imaging-genomic association algorithm is needed to quickly and accurately locate brain regions and analyze the correlation between the genome and these regions. Such an algorithm would provide a powerful tool for the research and treatment of brain diseases, accelerating diagnosis and treatment processes and improving treatment outcomes and quality of life.
[0006] In recent years, machine learning and artificial intelligence have made significant progress in the medical field. These technologies are capable of processing massive amounts of data, extracting useful information, and predicting future trends. Machine learning and artificial intelligence techniques are also widely used in brain imaging and genome-wide association studies. However, existing algorithms still face challenges, such as the accuracy of brain region localization, algorithm reliability, and efficiency. Therefore, further improvements and development of new algorithms are needed to meet the needs of brain disease research.
[0007] Traditional methods of studying targeted brain regions using anatomical sections involve dissecting the brain, preparing thin slices, staining the slices, and then observing and recording the location and structural features of the brain regions of interest under a microscope. However, this method has several significant drawbacks. First, traditional brain anatomical sectioning is a complex and tedious process. Obtaining slices from animal models or cadaveric brains requires time, specialized skills, and highly precise manipulation. Problems such as physical cutting leading to slice breakage or unstable staining techniques can occur during preparation, both of which can affect the accuracy of the research results. Second, brain anatomical sectioning has limitations in spatial resolution. While this method can provide highly detailed information on brain region structure, the slice thickness is typically in the range of tens to hundreds of micrometers, meaning that only information about local areas can be captured, rather than a comprehensive understanding of the characteristics and properties of the entire brain region. Furthermore, brain anatomical sectioning limits the number of research samples. Due to the need for cadaveric brains or animal models for dissection, the sample size for this method is usually small and limited. This makes it difficult to obtain comprehensive statistical data, limiting in-depth research on brain region variation, individual differences, and population characteristics. Summary of the Invention
[0008] This invention provides a method for linking brain images with brain tissue genes, thereby solving the technical problem that existing technologies cannot accurately analyze the correlation between genomes and brain regions.
[0009] To address the aforementioned problems, this invention provides a method for associating brain images with brain tissue genes. This method specifically includes the following steps: a data acquisition step, acquiring at least one frame of MRI images of the brain from multiple targets, the MRI image data including functional imaging data, structural imaging data, and diffusion tensor imaging data; a data preprocessing step, preprocessing the MRI images output from the data acquisition step to ensure comparability between different brain MRI images; a brain connectivity network construction step, constructing a brain connectivity network using the brain MRI image data output from the data preprocessing step, calculating the connectivity strength between different brain regions by locating various brain regions, and quantifying the structure and function of each brain region; and a grouping test. The steps include: grouping MRI images of the brains of healthy individuals with MRI images of the brains of patients; analyzing the brain connectivity networks output from the brain connectivity network construction step to examine whether there are significant differences between the two groups in different brain regions; and in the brain tissue gene association step, reading a standardized database, which is a standardized brain tissue gene expression profile; and acquiring multiple genomic data; for each brain region that shows significant differences in the grouping and testing step, associating it with the genomic data to obtain the brain tissue gene expression of the brain regions that show significant differences in the grouping and testing step.
[0010] Furthermore, the data preprocessing step specifically includes the following steps: a noise reduction step, which uses a filtering method to remove low-frequency signals below a critical value from the MRI image data output by the data acquisition step, thereby capturing and removing slow scan drift in the MRI image data; and a standardization step, which uses a standardized template to resample the MRI image data output by the noise reduction step using a linear or nonlinear algorithm, so that different brain MRI image data are comparable.
[0011] Furthermore, in the data preprocessing step, before the noise reduction step, a non-brain tissue removal step is included, which removes non-brain tissue from the MRI image data output by the data acquisition step. The non-brain tissue includes the skull and neck. The processed MRI image data is then used as the input to the noise reduction step.
[0012] Furthermore, in the data preprocessing step, after the non-brain tissue removal step and before the noise reduction step, a motion correction step is also included. If any target brain exhibits motion during the data acquisition step, any frame of the target brain is taken as a reference image, and motion correction is performed on each frame of the target brain's MRI image with the reference image to align each frame of the target brain's MRI image.
[0013] Furthermore, the brain connectivity network construction step specifically includes the following steps: a brain image segmentation step, which divides the MRI image data of the brain output from the data preprocessing step into different brain regions, each region being assigned a unique label, so that the brain is divided into different regions with specific functions or anatomical features; and a connectivity matrix construction step, which creates a connectivity matrix, where the rows and columns of the connectivity matrix correspond to different brain regions, and the values of the elements of the connectivity matrix represent the connectivity strength of the corresponding brain regions.
[0014] Further, the grouping test step specifically includes the following steps: a grouping step, grouping the MRI image data of the brains of healthy individuals with the MRI image data of the brains of patients; a hypothesis-establishing step, establishing a null hypothesis and an alternative hypothesis, wherein the null hypothesis is that the mean difference between the MRI image data of the brains of healthy individuals and the MRI image data of the brains of patients is zero, and the alternative hypothesis is that there is a significant difference between the mean of the MRI image data of the brains of healthy individuals and the MRI image data of the brains of patients; a sample difference calculation step, calculating the mean difference and standard error between each pair of paired samples; a t-statistic calculation step, using the mean difference and standard error output in the sample difference calculation step, calculating the t-statistic for each pair of paired samples; the t-statistic for paired samples represents the magnitude of the difference between the mean difference and zero relative to the standard error; and a p-value calculation step, calculating the p-value of the t-test for paired samples based on the t-statistic of paired samples and the degrees of freedom of paired samples. The p-value is used to determine whether the null hypothesis should be rejected, thereby determining whether there is a significant difference between the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population.
[0015] Furthermore, after the grouping test step and before the brain tissue gene association step, the method further includes: a multiple correction step, which reduces the false detection rate of the grouping test step by adjusting the threshold of the p-value; and by sorting all the p-values in ascending order and determining a critical value such that p-values above the critical value are considered to have significant differences.
[0016] Further, the brain tissue gene association step specifically includes the following steps: a genome data reading step, reading the standardized database to obtain multiple genome data; an association step, extracting the genome structure IDs and corresponding MNI spatial coordinates from the standardized database, calculating the brain regions corresponding to the MNI spatial coordinates; associating the set of genome structure IDs corresponding to the brain regions with significant differences in the grouping test step; a matrix calculation step, obtaining the set of genome structure IDs obtained in the association step, and taking their maximum common subset; obtaining the set of probe IDs from the standardized database, and obtaining the maximum common subset of the probe ID sets; calculating the gene expression matrix based on the maximum common subset of the genome structure ID sets and the maximum common subset of the probe ID sets, thereby calculating the gene expression of the brain regions with significant differences in the grouping test step.
[0017] Furthermore, in the matrix calculation step, after calculating the gene expression matrix, the following steps are also included: calculating the average value of the results of the gene expression value matrix; summing the results of each row in the gene expression value matrix that has the same gene ID.
[0018] The present invention also includes a data processing device, the data processing device including a memory for storing executable program code. The data processing device further includes a processor for reading the executable program code to run a computer program corresponding to the executable program code, to perform at least one step in the above-described method for associating brain images with brain tissue genes.
[0019] The advantage of this invention lies in its method for associating brain imaging with brain tissue genes, which can accurately locate brain regions and analyze the correlation between the genome and brain regions. It automatically locates targeted brain regions while simultaneously performing genomic analysis. This algorithm will improve the efficiency and reliability of brain research results, helping us to better understand the pathogenesis of brain diseases and develop more effective treatments. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method for associating brain images with brain tissue genes in an embodiment of the present invention; Figure 2 This is a flowchart of the data preprocessing steps in an embodiment of the present invention; Figure 3 This is a flowchart of the brain connectivity network construction steps in an embodiment of the present invention; Figure 4 This is a flowchart of the grouping test steps in an embodiment of the present invention; Figure 5 This is a flowchart of the brain tissue gene association step in an embodiment of the present invention. Specific Implementation
[0021] The following description, with reference to the accompanying drawings, illustrates preferred embodiments of the present invention to demonstrate its implementation. These embodiments fully explain the technical content of the invention to those skilled in the art, making the technical content clearer and easier to understand. However, the present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0022] like Figure 1 As shown, this embodiment provides a method for associating brain images with brain tissue genes, which specifically includes steps S1 to S6.
[0023] Step S1: Data acquisition step, acquiring at least one frame of MRI image data of the brain from multiple targets, including functional imaging data, structural imaging data, and diffusion tensor imaging data.
[0024] Magnetic resonance imaging (MRI) data uses magnetic fields and harmless radio waves to acquire detailed images of internal tissues. Functional imaging data is an application of MRI technology used to measure changes in blood oxygen levels during brain activity. Based on oxygen-dependent signals, functional imaging data infers brain function by monitoring changes in blood oxygenation in active brain regions. Structural imaging data is an imaging modality in MRI used to acquire detailed structural information about human tissues. Weighted imaging of structural imaging data uses short-time repetition and short-echo time settings, making cerebrospinal fluid appear black, while gray matter and white matter are displayed with different gray values. Structural imaging data can be used to locate and segment brain structures, and to study the anatomical structure of tissues and organs. Diffusion tensor imaging (DTI) data is an MRI technique used to assess the structure and connectivity of white matter fiber tracts in the brain. DTI measures the diffusion behavior of water molecules in tissues, providing information about white matter fiber tracts. By calculating the diffusion tensor, indices such as anisotropy and diffusion rate can be obtained to assess the orientation, connectivity, and integrity of fiber tracts, thus revealing characteristics of brain structure and connectivity.
[0025] These different types of MRI data provide different information for studying the structure and function of the brain. Structural imaging data is used to study the brain's anatomical structure; functional imaging data is used to study the brain's functional activities and connectivity patterns; and diffusion tensor imaging data is used to study the structure and connectivity of white matter fiber tracts. By comprehensively analyzing these different types of MRI data, we can gain a more complete understanding of the brain's structure and function, as well as their changes in various diseases and cognitive processes.
[0026] Step S2: Data preprocessing step, preprocessing the MRI images output from the data acquisition step to make different brain MRI image data comparable.
[0027] like Figure 2 As shown, the data preprocessing steps specifically include steps S21 to S24.
[0028] Step S21: Non-brain tissue removal step. The Bet command in the fsl tool is used to remove non-brain tissue from the raw data, thereby reducing interference. Non-brain tissue includes the skull and neck, and other common human body parts. The processed MRI image data is used as input for this noise reduction step.
[0029] Step S22: Motion correction step. If any target brain exhibits motion during the data acquisition step, any frame of the target brain's MRI image is taken as a reference image, and motion correction is performed on each frame of the target brain's MRI image with the reference image to align each frame of the target brain's MRI image.
[0030] This section primarily focuses on functional imaging data and diffusion tensor imaging data. Considering the potential head movement of subjects during scanning, motion correction is necessary for each time-point slice in the functional imaging data to ensure the accuracy of subsequent analysis. Motion correction is typically performed by aligning each time-point slice with a reference slice. For diffusion tensor imaging data, this is achieved by correcting each gradient direction image to a reference image.
[0031] Step S23: Noise reduction step, using filtering methods to remove low-frequency signals below the critical value in the nuclear magnetic resonance image data output from the data acquisition step, thereby capturing and removing slow scan drift in the nuclear magnetic resonance image data.
[0032] This step primarily targets functional imaging data, improving the signal-to-noise ratio by filtering out irrelevant frequency bands. In this embodiment, high-pass filtering is specifically utilized to capture and remove slow-scan drift by removing low-frequency signals below a threshold. The noise reduction step primarily targets functional imaging data, improving the signal-to-noise ratio by filtering out irrelevant frequency bands.
[0033] Step S24: Standardization step, using a standardized template to resample the MRI image data output from the noise reduction step through a linear or nonlinear algorithm, so that different brain MRI image data are comparable.
[0034] Because of individual differences in brain size and shape, and variations in their position within the scanner, the obtained brain data can be mismatched. To minimize these differences and ensure comparability of individual datasets, we need to resample the denoised individual data to a universal or standard template, such as the Montreal Neurological Institute (MNI) or Talairach space, using linear and nonlinear algorithms, thereby achieving comparability between different subjects.
[0035] Step S3: Brain connectivity network construction step. A brain connectivity network is constructed using the MRI image data of the brain output from the data preprocessing step. By locating each region of the target brain, the connection strength between each region of the brain is calculated, and the structure and function of each region of the brain are represented digitally.
[0036] like Figure 3 As shown, the brain connectivity network construction steps specifically include steps S31 to S32.
[0037] Step S31: Brain image segmentation step, dividing the MRI image data of the brain output from the data preprocessing step into different brain regions, each region is assigned a unique label, so that the brain is divided into different regions with specific functions or anatomical features. Step S32: Connection matrix construction step. Create a connection matrix. The rows and columns of the connection matrix correspond to different brain regions. The values of the elements of the connection matrix represent the connection strength of the corresponding brain regions.
[0038] Here are some commonly used methods for calculating the connectivity strength of brain regions: Correlation methods assess the strength of connections between brain regions by calculating correlation coefficients. Commonly used correlation coefficients include the Pearson correlation coefficient and cross-correlation coefficients. This method is applicable to functional imaging data, estimating functional associations between brain regions through changes in blood oxygenation level-dependent signals.
[0039] Diffusion methods are applicable to diffusion tensor imaging data, assessing the strength of connections between brain regions by measuring the diffusion behavior of water molecules in brain tissue. Metrics for measuring diffusion include the diffusion coefficient and the connectivity of fiber bundles.
[0040] Path length is a network theory-based method used to assess the strength of connections between brain regions. It measures connection strength by calculating the shortest path length from one brain region to another in a brain network. Shorter path lengths indicate stronger connections.
[0041] Functional connectivity models are statistical models that estimate the strength of connections between brain regions based on their temporal activity patterns. These models can establish connectivity relationships between brain regions using methods such as linear regression and time-lag correlation.
[0042] In practical applications, multiple methods are often combined to comprehensively assess the connection strength between brain regions in order to obtain more comprehensive and accurate results.
[0043] Step S4: Grouping and testing step. The MRI image data of the brains of healthy individuals and patients are grouped together, and the brain connectivity networks output from the brain connectivity network construction step are tested and analyzed to check whether there are significant differences in various brain regions between the two groups.
[0044] like Figure 4 As shown, the brain connectivity network construction steps specifically include steps S41 to S44.
[0045] Step S41: Grouping step, grouping the MRI images of the brains of healthy individuals with the MRI images of the brains of patients.
[0046] Step S42: Hypothesis establishment step, establishing null hypothesis and alternative hypothesis. The null hypothesis is that the mean difference between the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population is zero. The alternative hypothesis is that there is a significant difference between the mean of the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population.
[0047] Step S43: Sample difference calculation step, for each pair of paired samples, calculate the mean difference and standard error between them; t-statistic calculation step, using the mean difference and standard error output in the sample difference calculation step, calculate the t-statistic for each pair of paired samples; the t-statistic for the paired samples represents the magnitude of the difference between the mean difference and zero relative to the standard error.
[0048] Step S44: p-value calculation step. Based on the t-statistic of the paired samples and the degrees of freedom of the paired samples, calculate the p-value of the paired samples t-test. This p-value is used to determine whether the null hypothesis should be rejected, thereby determining whether there is a significant difference between the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population.
[0049] Step S5: Multiple correction step, which reduces the false discovery rate of the grouping test step by adjusting the p-value threshold; by sorting all p-values in ascending order, a corrected p-value threshold is calculated based on a preset FDR level (e.g., 0.05). Commonly used FDR adjustment methods include the Benjamini-Hochberg method and the Benjamini-Yekutieli method. Starting from the sorted p-value list, the position of the first p-value less than or equal to the FDR threshold is found and used as the threshold for the rejection region, such that p-values above this threshold are considered to have significant differences.
[0050] The FDR (Fault-Deviation Detection) multiple correction method can control the false discovery rate in multiple hypothesis tests, ensuring that the proportion of errors in rejected hypotheses does not exceed the pre-set FDR level. Compared to the traditional Bonferroni correction method, the FDR method has higher statistical power and can better balance the problems of false discovery and false omission.
[0051] Step S6: Brain tissue gene association step, read the standardized brain tissue gene expression profile and obtain multiple genomic data; for each brain region that shows significant differences in the grouping test step, use a standardized database to associate it with the genomic data, thereby obtaining the brain tissue gene expression of the brain regions that show significant differences in the grouping test step.
[0052] like Figure 5 As shown, the brain connectivity network construction steps specifically include steps S61 to S63.
[0053] Step S61: Genome data reading step, reading the standardized brain tissue gene expression profile and obtaining multiple genome data; the genome data includes the genome structure ID, the MNI spatial coordinates corresponding to the genome structure ID, probe ID, gene ID, gene name and gene expression value matrix.
[0054] A genome structure ID is a unique identifier for a specific brain region or anatomical structure. These structure IDs are typically encoded based on specific brain maps or anatomical reference models (such as AAL, Brodmann, etc.). The MNI (Montreal Neurological Institute) spatial coordinate system is a common standardized coordinate system used for aligning and comparing brain maps across different studies. The MNI spatial coordinates corresponding to a genome structure ID represent the structure's location in three-dimensional space. A probe ID is an identifier used in gene expression microarrays or genome sequencing technologies to uniquely identify each probe or probe object. A gene ID is an identifier used to uniquely identify a specific gene. Gene names, Ensembl IDs, Entrez Gene IDs, etc., are commonly used as gene IDs. A gene name is a human-readable gene identifier used to describe the naming of a specific gene. A gene expression matrix is a two-dimensional data matrix where each row represents a gene and each column represents a sample or experimental condition. Each element in the matrix represents the expression level of the corresponding gene in the corresponding sample, usually a numerical value (such as expression intensity, FPKM, TPM, etc.).
[0055] Genomic data provides quantitative information about the relationship between genes and brain structure, as well as gene expression. The location and identifier of specific brain regions can be determined using the genome's structural ID and MNI spatial coordinates. Probe IDs, gene IDs, and gene names are used to uniquely identify and describe specific genes.
[0056] In this embodiment, the proposed method obtains standardized microarray data from multiple subjects from the publicly available Allen Brain Atlas database. Each subject's data includes three core tables: the first table is MicroarrayExpression.csv, where the first column is the probe ID, and subsequent columns are gene expression value matrices. Each row of this matrix corresponds to a probe ID, and each column corresponds to a genome structure ID. The second table is Probes.csv, which contains probe IDs, gene IDs, and gene names. The third table is SampleAnnot.csv, which contains genome structure IDs and MNI spatial coordinates.
[0057] Step S62: The association step begins by using the genomic structure IDs and corresponding MNI spatial coordinates provided in SampleAnnot.csv to calculate the brain region corresponding to each MNI spatial coordinate using the label4MRI library in R. The index of the structure ID and the corresponding brain region name are then saved in the table BAname_i.csv. Next, for multiple abnormal brain regions obtained from brain imaging analysis (based on Bordmann Atlas partitioning, the Bordmann partition name for each brain region is obtained), we consider them as target brain regions. This leads to the finding of the set of structure IDs for each subject's genome corresponding to the target brain regions.
[0058] Step S63: Matrix calculation step. Since the set of genomic structure IDs of the target brain region differs among different subjects, it is necessary to find their maximum common subset.
[0059] First, the corresponding structure ID value is found in the SampleAnnot.csv table by using the index set of structure IDs in the structure_index_i.csv table, and duplicate values are removed. Then, the set of genome structure IDs obtained in the association step is obtained, and then the intersection function in the set type of Python language is used to get their maximum common subset.
[0060] Then, obtain the probe ID set from the Probes.csv table, and then use the intersection function in the set type of Python language to obtain the maximum common subset of the probe ID set.
[0061] Next, the gene expression matrix is calculated based on the greatest common subset of the genome structure ID set and the greatest common subset of the probe ID set. The gene expression value matrix is extracted from the MicroarrayExpression.csv table using the greatest common probe ID set in the common_subset_probe_id.csv table (used as the index of the expression value matrix rows). Since the gene expression values in the original MicroarrayExpression.csv table are the result after official log processing, the extracted expression value matrix is squared and restored. Then, the structure ID index in SampleAnnot.csv is obtained from the structure ID set in common_subset_structure_id.csv, and used as the column index to extract the gene expression value matrix. Columns with the same structure ID in the gene expression matrix are summed.
[0062] To obtain the statistical average, we averaged the gene expression value matrices for all subjects. Furthermore, considering that each gene probe targets a specific gene fragment, different probes may correspond to the same gene. Therefore, we summed the rows with the same gene ID in the gene expression value matrix based on the gene ID and probe ID in the Probes.csv table. Additionally, to obtain the average result for the entire targeted brain region, we calculated the average of the columns of the gene expression value matrix (each column representing a structural ID, corresponding to a portion of the brain region). Finally, the resulting gene expression value matrix was log-processed and saved to a table. The first column of this table records the gene ID, the third column records the corresponding gene name, and the second column records the gene expression value for the entire targeted brain region.
[0063] The advantage of this invention lies in its method for associating brain imaging with brain tissue genes, which can accurately locate brain regions and analyze the correlation between the genome and brain regions. It automatically locates targeted brain regions while simultaneously performing genomic analysis. This algorithm will improve the efficiency and reliability of brain research results, helping us to better understand the pathogenesis of brain diseases and develop more effective treatments.
[0064] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for linking brain images with brain tissue genes, characterized in that, Specifically, the steps include the following: The data acquisition step involves acquiring at least one frame of MRI image data of the brain from multiple targets. The MRI image data includes functional imaging data, structural imaging data, and diffusion tensor imaging data. The data preprocessing step preprocesses the MRI images output from the data acquisition step to make different brain MRI image data comparable. The brain connectivity network construction step involves using the MRI image data of the brain output from the data preprocessing step to construct a brain connectivity network. This is achieved by locating various regions of the target brain, calculating the connection strength between these regions, and quantifying the structure and function of each brain region. The grouping test step involves grouping the MRI image data of the brains of healthy individuals with the MRI image data of the brains of patients, and then testing and analyzing the brain connectivity networks output from the brain connectivity network construction step to examine whether there are significant differences between the two groups in various brain regions. The brain tissue gene association step involves reading a standardized database, which is a standardized brain tissue gene expression profile; and acquiring multiple genomic data. For each brain region that shows significant differences in the grouping test step, it is associated with the genomic data to obtain the brain tissue gene expression of the brain regions that show significant differences in the grouping test step. The brain tissue gene association step specifically includes the following steps: a genome data reading step, reading the standardized database to obtain multiple genome data; an association step, extracting the genome structure IDs and corresponding MNI spatial coordinates from the standardized database, and calculating the brain regions corresponding to the MNI spatial coordinates; associating the set of genome structure IDs corresponding to brain regions that show significant differences in the grouping test step; and a matrix calculation step, obtaining the set of genome structure IDs obtained in the association step and taking their maximum common subset. Obtain the probe ID set from the standardized database, and obtain the maximum common subset of the probe ID set; calculate the gene expression matrix based on the maximum common subset of the genome structure ID set and the maximum common subset of the probe ID set, thereby calculating the gene expression of brain regions that show significant differences in the grouping test step.
2. The method for associating brain images with brain tissue genes as described in claim 1, characterized in that, The data preprocessing steps specifically include the following steps: The noise reduction step uses a filtering method to remove low-frequency signals below a critical value from the nuclear magnetic resonance image data output by the data acquisition step, thereby capturing and removing slow scan drift in the nuclear magnetic resonance image data; The standardization step involves resampling the MRI image data output from the noise reduction step using a standardized template through a linear or nonlinear algorithm, thereby making different brain MRI image data comparable.
3. The method for associating brain images with brain tissue genes as described in claim 2, characterized in that, In the data preprocessing step, prior to the noise reduction step, the following is also included: The non-brain tissue removal step removes non-brain tissue from the MRI image data output by the data acquisition step. The non-brain tissue includes the skull and neck. The processed MRI image data is then used as the input to the noise reduction step.
4. The method for associating brain images with brain tissue genes as described in claim 3, characterized in that, The data preprocessing step, after the non-brain tissue removal step and before the noise reduction step, further includes: In the motion correction step, if any target's brain exhibits motion during the data acquisition step, any frame of the target's brain's MRI image is taken as a reference image, and motion correction is performed on each frame of the target's brain's MRI image with the reference image to align each frame of the target's brain's MRI image.
5. The method for associating brain images with brain tissue genes as described in claim 1, characterized in that, The brain connectivity network construction steps specifically include the following steps: The brain image segmentation step divides the MRI image data of the brain output from the data preprocessing step into different brain regions. Each region is assigned a unique label, so that the brain is divided into different regions with specific functions or anatomical features. The connection matrix construction steps involve creating a connection matrix where the rows and columns correspond to different brain regions, and the values of the elements in the connection matrix represent the connection strength of the corresponding brain region.
6. The method for associating brain images with brain tissue genes as described in claim 1, characterized in that, The grouping test step specifically includes the following steps: The grouping step involves grouping the MRI images of the brains of healthy individuals with those of patients. The hypothesis-establishing steps involve establishing a null hypothesis and an alternative hypothesis. The null hypothesis states that the mean difference between the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population is zero. The alternative hypothesis states that there is a significant difference between the mean values of the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population. The sample difference calculation steps are as follows: for each pair of paired samples, calculate the average difference and standard error between them; The t-statistic calculation step uses the average difference and standard error output in the sample difference calculation step to calculate the t-statistic for each pair of paired samples. The t-statistic for the paired samples represents the magnitude of the difference between the mean difference and zero relative to the standard error; The p-value calculation steps involve calculating the p-value of the paired sample t-test based on the t-statistic and degrees of freedom of the paired samples. This p-value is used to determine whether the null hypothesis should be rejected, thereby determining whether there is a significant difference between the MRI images of the brains of the healthy population and the MRI images of the brains of the patient population.
7. The method for associating brain images with brain tissue genes as described in claim 6, characterized in that, Following the grouping test step and before the brain tissue gene association step, the following is also included: The multiple correction step reduces the false detection rate of the grouping test step by adjusting the threshold of the p-value; by sorting all the p-values in ascending order and determining a critical value such that p-values above the critical value are considered to be significantly different.
8. The method for associating brain images with brain tissue genes as described in claim 1, characterized in that, In the matrix calculation step, after calculating the gene expression matrix, the following steps are also included: The average value of the gene expression value matrix is calculated; the summation is performed for each row in the gene expression value matrix that has the same gene ID.
9. A data processing device, characterized in that, include: Memory, used to store executable program code; as well as A processor for reading the executable program code to run a computer program corresponding to the executable program code, in order to perform the steps in the brain image and brain tissue gene association method according to any one of claims 1-8.
Citation Information
Patent Citations
Methodologies linking patterns from multi-modality datasets
CN101068498A
Classification model acquisition method and device, expression category determination method and device, equipment and medium
CN115457361A