Single-cell analysis method for brain function based on multimodal whole-brain omics data association

By using a multimodal whole-brain omics data association method, combined with partial least squares regression and spatial permutation test, the problem of analyzing brain function in single cells in existing technologies has been solved. This method enables accurate analysis of the association between cells and spatial brain phenotypes at the single-cell level, improving the accuracy and robustness of the analysis.

CN122493956APending Publication Date: 2026-07-31ZHEJIANG UNIV OF CHINESE MEDICINE JINHUA RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV OF CHINESE MEDICINE JINHUA RES INST
Filing Date
2026-05-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately analyze the correlation between brain function and cell type at the single-cell level; gene enrichment analysis has a high false positive rate; and deconvolution methods cannot pinpoint precise cell types.

Method used

By using a multimodal whole-brain omics data association method, preprocessing image transcriptomics and single-cell transcriptomics data, and using partial least squares regression and spatial permutation tests, we calculated the variable projection importance score of genes and the cell specificity score, constructed gene sets and gene specificity score matrices, eliminated false positives, and located significantly related cells.

Benefits of technology

It enables accurate analysis of the association between cells and spatial brain phenotypes at the single-cell level, improves the robustness of gene expression data, prevents batch effects and technical noise interference, and accurately analyzes cells that are significantly associated with the target spatial brain phenotype.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493956A_ABST
    Figure CN122493956A_ABST
Patent Text Reader

Abstract

This invention discloses a method for analyzing single-cell brain function based on multimodal whole-brain omics data association, belonging to the field of bioinformatics technology. The steps are as follows: obtaining the target gene expression matrix and the target spatial brain phenotype vector; performing partial least squares regression iteration and calculating the variable projection importance score of the genes; examining the correlation between the variable projection importance score of the genes and the spatial autocorrelation pattern using a zero distribution, constructing a gene set, and assigning gene weights to each gene in the gene set; obtaining a low-dimensional embedding, calculating the gene specificity score of the cells, and constructing a gene specificity score matrix; constructing a spatial permutation test gene set based on the spatial permutation test, calculating the significance probability of the association between a single cell and the spatial brain phenotype, and identifying cells significantly correlated with the target spatial brain phenotype. This invention solves the problem of accurately analyzing single-cell brain function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bioinformatics technology, and in particular relates to a single-cell analysis method for brain function based on multimodal whole-brain omics data association. Background Technology

[0002] As a complex system, the brain's functions emerge from a microscopic to a macroscopic level. At the macroscopic level, the brain is composed of complex circuits and tissues whose activity patterns are directly related to cognitive functions and behavior. The structural and functional specialization of different brain regions also determines their roles in behavior. At the microscopic level, the brain is composed of different types of neurons, glial cells, etc., whose characteristics and gene expression patterns collectively influence brain function. Neuropsychiatric disorders are often associated with abnormalities in specific brain circuits, structural changes in brain tissue, and abnormalities in specific cell types and genetic variations. Therefore, integrating macroscopic and microscopic brain characteristics is crucial for a comprehensive understanding of the brain's working mechanisms, predicting brain disease risk, identifying disease-related gene targets, and elucidating the dynamic relationship between cell types and macroscopic brain function.

[0003] In the practice of linking macroscopic and microscopic brain data, imaging transcriptomics has successfully correlated different imaging data, such as macroscopic brain function characterized by functional magnetic resonance imaging (fMRI), nuclear magnetic resonance imaging (MRI), and diffusion tensor imaging (DTI), with the high spatial resolution spatial transcriptomics database provided by the Allen Brain Atlas. However, since this database itself is not at single-cell resolution, it cannot reveal the different contributions of different cells to a particular macroscopic feature. Previous studies have attempted to further correlate cellular-level information, such as using gene enrichment analysis to examine whether gene sets of interest are enriched in certain cell types by identifying marker genes in different cells, and using the CIBERSORTx software to deconvolve spatial transcriptomics data with a single-cell transcriptomics database of the cortex as a reference, successfully correlating macroscopic brain features with the abundance of different cell types. However, these methods still have significant limitations. Gene enrichment analysis is considered to have a high false positive rate in the application of imaging transcriptomics, and deconvolution-based methods, considering the accuracy of deconvolution results, cannot pinpoint the final results to the fine cell type, let alone examine the cell-feature correlation level at the single-cell level. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, this invention provides a method for analyzing single-cell brain functions based on multimodal whole-brain omics data association, which solves the problem of accurately analyzing single-cell brain functions.

[0005] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: This invention provides a method for analyzing single-cell brain function based on multimodal whole-brain omics data association, comprising the following steps: S1. Preprocess probes and transcriptome samples from the spatial transcriptomics database, standardize and filter each gene, and map the spatial brain phenotype of interest onto the standard brain atlas to obtain the target gene expression matrix and the target spatial brain phenotype vector. S2. Perform partial least squares regression iteration based on the target gene expression matrix and the target space brain phenotype vector, and calculate the variable projection importance score of the gene. S3. Examine the correlation between the variable projection importance score of genes and the spatial autocorrelation pattern through zero distribution, sort and screen the genes, construct a gene set, and assign gene weights to each gene in the gene set. S4. Preprocess the single-cell transcriptomics dataset to obtain low-dimensional embeddings, and calculate the gene-specific score in each cell based on the neighboring cells of each cell and the ranking of each gene in the cell, and construct a gene-specific score matrix. S5. Based on the gene set and gene-specific score matrix, construct a spatial permutation test gene set based on the spatial permutation test, calculate the significance probability of association between a single cell and the spatial brain phenotype, and find cells that are significantly associated with the target spatial brain phenotype through multiple comparison correction.

[0006] The beneficial effects of this invention are as follows: This invention provides a single-cell analysis method for brain function based on multimodal whole-brain omics data association. Using genes as a link, it first calculates the gene set associated with the target spatial brain phenotype using traditional image transcriptomics procedures, and assigns a gene weight to each gene in the gene set. Then, based on the gene set and the preprocessed single-cell transcriptomics dataset, it calculates the enrichment score at the single-cell level. This score reflects the degree of association between each cell and the spatial brain phenotype. Statistical tests are used to eliminate false positives caused by spatial autocorrelation, identifying cells highly associated with the spatial brain phenotype. To improve the robustness of gene expression data in the single-cell transcriptomics dataset, this scheme replaces standardized gene expression with a gene-specific score for each cell, preventing interference from batch effects and technical noise. This invention can locate the cell analysis results at the single-cell level and realize the degree of association between cells and the spatial brain phenotype at the single-cell level, thereby accurately identifying cells significantly associated with the target spatial brain phenotype.

[0007] Other advantages of the present invention will be analyzed in more detail in the following embodiments. Attached Figure Description

[0008] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating the steps of a single-cell analysis method for brain function based on multimodal whole-brain omics data association in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0011] like Figure 1 As shown, in one embodiment of the present invention, the present invention provides a method for single-cell analysis of brain function based on multimodal whole-brain omics data association, comprising the following steps: S1. Preprocess probes and transcriptome samples from the spatial transcriptomics database, standardize and filter each gene, and map the spatial brain phenotype of interest onto the standard brain atlas to obtain the target gene expression matrix and the target spatial brain phenotype vector. S1 includes the following steps: S11. Use the Allen Brain Atlas as a spatial transcriptomics database at the whole-brain scale. This approach selects the Allen Brain Atlas database, which has the highest whole-brain coverage and spatial resolution, as the whole-brain scale spatial transcriptomics database from open-source databases. Spatial resolution refers to the probability that a transcriptome sample can be accurately matched with the corresponding three-dimensional spatial coordinates. S12. Preprocess probes and transcriptome samples from spatial transcriptomics databases; In this embodiment, the preprocessing operations for probes and transcriptome samples in the spatial transcriptomics database include labeling probe-genes from the Allen Human Brain Atlas library, filtering probes and transcriptome samples, and selecting probes. The use of the Abagen toolkit in this scheme to preprocess probes and transcriptome samples in the spatial transcriptomics database can effectively improve robustness.

[0012] S13. Based on the scaling-enhanced sigmoid function, the expression of each gene is standardized according to the interval of each gene expression to obtain standardized gene expression. The calculation expression for the standardized gene expression is as follows: , , in, This represents standardized gene expression. This represents the gene expression after preliminary standardization. This represents the minimum value of gene expression after initial standardization. This represents the maximum gene expression value after initial standardization. Indicates It is an exponential function of the basis constants. This indicates gene expression that has undergone standardized processing. Indicates the median gene expression. This represents the interquartile range of gene expression after standardization within the same tissue; in this embodiment, tissue refers to the cortex, subcortical structures, etc.; this scheme eliminates batch effects between different brain data sources in spatial transcriptomics databases by scaling and enhancing the sigmoid function to standardize gene expression.

[0013] S14. Generate a gene expression matrix based on standardized gene expression. ,in, Indicates belonging to, Represents the set of real numbers. Indicates the number of transcriptome samples. This indicates the number of genes after standardization. In this scheme, based on standardized gene expression, the gene expression matrix can be output using the Abagen toolkit. S15, according to Transcriptome samples from spatial transcriptomics databases are assigned to standard brain atlases to generate gene expression matrices under the standard brain atlases. ,in, This indicates standard brain atlas labels; in this scheme, standard brain atlases include the Schaefer 400 atlas, HCP-MMP1.0 atlas, Desikan-Killiany atlas, etc. The Abagen toolkit can automatically allocate transcriptome samples, thereby generating... ; S16, calculate separately The stability of variation in each gene; The expression for calculating the variation stability of genes in S16 is as follows: , in, Indicates gene The stability of the variation, This indicates the number of sources of brain data. Indicates gene Sources of brain data The spatial representation vector in the text, Indicates gene Sources of brain data The spatial representation vector in the text, express and The Pearson correlation coefficient between them; In this approach, by calculating the variation stability of each gene, genes with stable expression patterns and less interference from sequencing noise can be evaluated and selected from different brain data sources. The expression for calculating gene variation stability shows that the greater the variation stability of a gene, the more stable its expression pattern is across different brain data sources, and the less it is affected by sequencing noise and individual differences. S17, to The variation stability of each gene in the matrix is ​​sorted from largest to smallest, and the top 50% of genes are filtered out to obtain the target gene expression matrix. ; S18. Map the spatial brain phenotype of interest onto a standard brain atlas to obtain the target spatial brain phenotype vector. .

[0014] S2. Perform partial least squares regression iteration based on the target gene expression matrix and the target space brain phenotype vector, and calculate the variable projection importance score of the gene. S2 includes the following steps: S21, will Let it be denoted as the independent variable matrix and will Denote as spatial brain phenotype vector ,in, This indicates the number of brain regions in a standard brain atlas. This indicates the number of genes remaining after filtering. In this scheme, when the number of components in the partial least squares regression iteration increases, initially, the increased components can explain more latent components of the independent variable matrix, making the model better at predicting the univariate target vector. However, when the number of components increases beyond the optimal number, the increased components explain more irrelevant variance, which can lead to overfitting of the model. To prevent this from happening, this scheme uses the 10-fold cross-validation method to explore the optimal number of components.

[0015] S22. Based on the 10-fold cross-validation method, and Exploration yields the target optimal number of components ; S22 includes the following steps: S221. According to the 10-fold cross-validation method, based on random numbers... and Divided into training set and test set The training set is divided into an inner training set. and internal test set In this plan, and Used to explore the current optimal number of components ; S222, Let the number of components And for each ,according to The partial least squares regression model is trained to obtain the trained partial least squares regression model corresponding to each component score, where... The preset minimum number of components, The maximum number of components is preset. It is an integer; In this embodiment, it is preferred =1, It is 15; S223. Based on the partial least squares regression model trained according to each component score, predict... The prediction results for the internal test set were obtained. ; S224, based on and The current optimal number of components is calculated. ; The The calculation expression is as follows: , in, This represents the value of the independent variable when it reaches its minimum value. Represents the mean square error function; S225, with The number of components in the partial least squares regression model, and based on predict The prediction results for the test set are obtained. ; S226, according to and The score of the current optimal component number is calculated. ; The The calculation expression is as follows: , in, This represents the Pearson correlation coefficient; S227. Based on different random numbers, repeat the first preset number of iterations from S221 to S226 to obtain the current optimal component number for each iteration. The score of the current optimal component ; In this embodiment, the first preset number of cycles is 30 to 100. S228, Accumulate the data during each repetition. of The sum is used to obtain the total score of the optimal number of components for each repeated process; S229. Take the current optimal component number corresponding to the highest total score of the optimal component number as the target optimal component number. .

[0016] In this scheme, if the obtained equal Then it may be true Compared to the current Larger, requiring gradual enlargement The preferred interval for each increase is 5.

[0017] In this scheme, after obtaining the target optimal component score, to test the performance of the partial least squares regression model, S221~S222 are repeatedly executed based on different random numbers, with different effects of random numbers on the splitting of the training and test sets. Each trained partial least squares regression model is then tested, and its performance score is calculated. The median of the performance scores of all trained partial least squares regression models is taken as the performance score of the partial least squares regression model. The calculation expression for the performance score of each trained partial least squares regression model is as follows: , in, Indicates the first The performance score of a trained partial least squares regression model. The true label representing the brain phenotypic vector in the target space. Indicates the first The prediction results of a trained partial least squares regression model regarding the target space brain phenotypic vector.

[0018] S23, based on and The number of components is The partial least squares regression iterations are used to obtain the parameters of the partial least squares regression model, which include gene weights, loadings, and regression coefficients. In this solution, Python is used. sklearn.cross_decomposition under the module PLSRegression The function implements partial least squares regression iteration; S23 includes the following steps: S231, Let the independent variable matrix of the first round be... equal The univariate target vector in the first round equal The number of components equals ; S232, For each component According to the partial least squares regression algorithm, the first step is to execute the second step in sequence. Wheel S233~S235, up to the first After the first iteration ends, proceed to S236; S233, Obtain the first covariance objective function of the wheel Maximize gene weights And calculate the first The partial least squares regression score of the round, where, Represents the covariance function. Indicates the first The independent variable matrix of the wheel, Indicates the first The single-variable target vector of the wheel; The first The formula for calculating the partial least squares regression score of the round is as follows: , in, Indicates the first The partial least squares regression score of the wheel; In this scheme, the core idea of ​​the partial least squares regression algorithm is to find a projection direction that maximizes the covariance between the projected independent variable matrix and the univariate target vector. S234, Searching for the first The load of the wheel's genes So that closest ,in, for transpose, Indicates transpose; for the result obtained in S223 and Then calculate and To what extent can it explain and Variations; S235, based on , , and Update # The independent variable matrix of the wheel and univariate target vector And fit the linear relationship between the independent variable matrix and the single variable target vector; The and The calculation expression is as follows: , ; S236. Based on the linear relationship between the independent variable matrix and the univariate target vector, the regression coefficients of genes are obtained when predicting spatial brain phenotypes by extracting the coef attribute based on the partial least squares regression model. ,in, .

[0019] S24. Calculate the variable projection importance score of the gene based on the parameters of the partial least squares regression model. .

[0020] In this solution, the The calculation expression is as follows: , , in, The scale control constant represents the projection importance of the variable. Indicates the first The variable projection mutation of the wheel, Indicates the first The square of the gene weight of the wheel, express The transpose of .

[0021] In this scheme, to further measure whether genes are more focused on the positive or negative side of a given spatial brain phenotype, The sign of the value corresponds to the regression coefficient of the gene, where the positive side refers to the larger value and the negative side refers to the smaller value.

[0022] S3. Examine the correlation between the variable projection importance score of genes and the spatial autocorrelation pattern through zero distribution, sort and screen the genes, construct a gene set, and assign gene weights to each gene in the gene set. In this solution, The association between the variable projection importance score of genes and spatial autocorrelation patterns is examined by zero distribution to exclude false positive genes that are more related to spatial autocorrelation patterns. Whether it's the spatial gene expression patterns in the Allen Brain Atlas or the spatial patterns of spatial brain phenotypes, whole-brain-scale data generally exhibit spatial autocorrelation. This means that, in different directions, these statistics tend to increase or decrease with varying gradients, similar to terrain. Some spatial autocorrelation can exaggerate the correlation between gene expression and spatial brain phenotypes. To explain the impact of spatial autocorrelation, the best strategy is to generate multiple null models under spatial autocorrelation constraints based on the provided spatial brain phenotypes. These null models preserve the spatial autocorrelation characteristics of the spatial brain phenotypes but alter their spatial patterns, enabling spatial permutation tests, estimation of partial least squares regression models, and empirical distributions of subsequent statistics based on these null models. In this solution, the Python package is selected. neuromaps of nulls Nonparametric space null model rotation test in the module, i.e. Alexander-Bloch spin-test ; S3 includes the following steps: S31. Treat the human cerebral cortex as a sphere and rotate the spatial brain phenotype on the surface of the sphere to obtain the zero distribution. ,in, Indicates the number of space permutation checks; S32. Use the zero distribution to check the probability that the partial least squares regression model predicts spatial brain phenotype vectors better than spatial brain phenotypes, and calculate the significance probability of the spatial permutation test. In this solution, step S32 includes the following steps: S321. Initialize the first check count. =0; S322, for the first Wheel space permutation test, take the first Zero distribution corresponding to the wheel and in the replace In this case, the training and test sets are divided according to the S221 method, and the partial least squares regression model is calculated. The performance scores below, of which, Integer and ; The partial least squares regression model in S322 is... The expression for calculating the performance score is as follows: , in, Indicates The performance score of the partial least squares regression model obtained from training. The true label representing the brain phenotypic vector in the target space. Indicates The prediction results of the partial least squares regression model obtained from training regarding the target space brain phenotype vector; S323, if Greater than the final performance score of the partial least squares regression model Then update The value is ; S324, according to and The probability that the partial least squares regression model predicts spatial autocorrelation patterns better than spatial brain phenotypes is calculated and used as the significance probability of the permutation test. The expression for calculating the significance probability of the permutation test is as follows: , in, This represents the significance probability of the permutation test; In this scheme, if the significance probability of the permutation test is less than the preset significance probability threshold, the partial least squares regression model has a significantly better predictive effect on the target spatial brain phenotype than the predictive effect of the irrelevant spatial autocorrelation pattern. In this embodiment, the significance probability threshold is set to 0.05. S33. Use the null distribution to examine the significance probability between the variable projection importance score and the spatial autocorrelation pattern of each gene in the partial least squares regression model; In this solution, step S33 includes the following steps: S331. For each gene, initialize the second check count for that gene. =0; S332, for the first Wheel space permutation test, take the first Zero distribution corresponding to the wheel And directly fit the partial least squares regression model corresponding to the target optimal component number. and The variable projection importance scores of genes under this partial least squares regression model were extracted. ; In this scheme, S332 is executed according to the methods in S23~S24; S333, if Greater than Then update The value is ; S334, if No more than For a single gene, according to and the gene corresponding to The significance probability of the gene's variable projection importance score for the target spatial brain phenotype under the partial least squares regression model compared to the irrelevant spatial autocorrelation pattern was calculated and used as the significance probability at the gene level. The expression for calculating the significance probability at the gene level is as follows: , in, Represents the significance probability at the gene level; S34. Sort all genes in ascending order of their significance probability at the gene level. If genes have equal significance probabilities, sort them in descending order of their variable projection importance scores. This yields the gene ranking results. ; After the sequencing was completed, this scheme employed cross-validation to further determine which genes had a greater impact on the spatial brain phenotypes of interest, and which had little impact, or were even false positives or noise. What is the optimal number of the first few genes in the dataset? S35. Selecting through cross-validation Center front Construct a gene set from individual genes; In this solution, step S35 includes the following steps: S351, Defining the Number of Genes ; S352, respectively for Pick Center front Each gene forms a corresponding subset of independent variables. ; S353, based on and Construct a partial least squares regression model, and in replace In the case of S22, the corresponding optimal number of potential components is obtained through exploration, and with predict The prediction results of the subset of independent variables are obtained. ; S354, based on Calculate the corresponding Cross-validation scores of the partial least squares regression model under these conditions; The formula for calculating the cross-validation score is as follows: , in, express The cross-validation score of the partial least squares regression model. This indicates calculating the average value; In this plan, The larger the value, the more the currently selected gene can explain the major variations in the target space brain phenotype; S355, with The x and y coordinate values ​​and Construct a curve for the ordinate values ​​and select... Not following When the curve increases and continues to increase, it corresponds to the point where the curve inflects. Constructing a gene set; S36. To study genes more closely associated with the positive side of the spatial brain phenotype, additional screening is required. Genes with a value greater than 0 are added to the gene set, and after screening, the resulting gene set is denoted as . ; S37. If it is necessary to study genes on the negative side of the spatial brain phenotype, additional screening is required. Genes with a value less than 0 are selected, and the resulting gene set is denoted as [genes whose value is less than 0]. ; S38, based on and , respectively Each gene in the gene is assigned a gene weight; The formula for calculating the gene weight is as follows: in, Indicates gene Gene weights.

[0023] S4. Preprocess the single-cell transcriptomics dataset to obtain low-dimensional embeddings, and calculate the gene-specific score in each cell based on the neighboring cells of each cell and the ranking of each gene in the cell, and construct a gene-specific score matrix. Single-cell analysis of brain function requires examining the gene set associated with spatial brain phenotypes. Whether it is enriched in certain single-cell transcriptomics cells. However, the original single-cell transcriptomics dataset has the following problems and must undergo appropriate preprocessing: (1) poor sequencing quality of some cells, including too few expressed genes, too few total expressed genes, and too many expressed mitochondrial genes; (2) a sequencing tag barcode may capture the expression profiles of two or more cells; (3) differences in the expression levels of different genes; (4) if a single-cell transcriptomics dataset contains single-cell data from multiple batches, there is a batch effect between different batches of data; S4 includes the following steps: S41. Obtain the number of genes, total gene expression, and mitochondrial gene ratio for each cell in the single-cell transcriptomics dataset. In this solution, through Python In Scanpy of sc.pp.calculate_qc_metrics The function calculates the number of genes, total gene expression, and mitochondrial gene ratio for each cell. S42. Based on the characteristics of the single-cell transcriptomics dataset, set the threshold for the number of genes, the threshold for the total number of gene expression, and the threshold for the proportion of mitochondrial genes. Filter out cells with fewer than the number of genes, cells with a total number of gene expression less than the threshold for the total number of gene expression, and cells with a proportion of mitochondrial genes greater than the threshold for the proportion of mitochondrial genes. Because the experimental and sequencing conditions vary greatly across different single-cell transcriptomics datasets, the filtering thresholds will also differ. Therefore, in practice, the filtering thresholds should be determined based on the specific characteristics of the single-cell transcriptomics dataset. Scanpy or Seurat Guidelines for cell filtration; S43. Set a cell expression threshold and filter out genes that are below the cell expression threshold; in this embodiment, the cell expression threshold is set to 50, that is, genes expressed in less than 50 cells are filtered out. At the same time, the gene number threshold is set to 250, that is, cells with less than 250 expressed genes are filtered out. S44, Remove diploid cells; In this embodiment, using Scanpy of sc.pp.scrublet The function removes diploid samples; S45. If the single-cell transcriptomics dataset includes single-cell data from multiple batches, then all cells are projected into the latent space using the single-cell integration algorithm to obtain a low-dimensional embedding. In this scheme, the single-cell integration algorithm includes Harmony , RPCA , scVIThese methods effectively remove batch effects, and the projected common latent space is typically 50-dimensional, allowing the resulting low-dimensional embeddings to be further used for dimensionality reduction visualization or other machine learning or deep learning algorithms; however, despite Harmony , RPCA , scVI While single-cell integration algorithms can correct for batch effects in cell embedding in the low-dimensional latent space, they cannot correct for the original gene expression matrix. Scanpy / Seurat The built-in method normalizes based on sequencing depth, but this normalization method is not sensitive to batch effects and technical noise, and does not make full use of the batch-effect-free embedding provided by the single-cell integration algorithm. S46. Calculate the rank of each gene in the cell; The formula for calculating the rank of each gene in the cell is as follows: , in, Indicates gene In the Ranking of individual cells Indicates the first A ranking function for genes in each cell. Indicating the target gene expression matrix, the first... Genes in a cell ; S47. Based on low-dimensional embedding, cosine similarity is used to find neighboring cells of the cell, and a cell set of the cell is constructed based on the cell and its neighboring cells. The gene-specific score of the cell is then calculated based on the cell set. In this scheme, through... Scanpy of sc.pp.neighbors The function finds the neighboring cells of a cell, with a default value of 50 neighboring cells. The expression for calculating the gene-specific fraction in the cell is as follows: , in, Indicates the first Genes in a cell Specificity score, This represents the number of neighboring cells of a cell. Indicates the first Each cell, Indicates the first A collection of cells consisting of an individual cell and its neighboring cells. Indicates gene In the Ranking of individual cells This represents the total number of cells in a single-cell transcriptomics dataset. Indicates gene In the Ranking in individual cells; S48. If the gene specificity fraction in a cell is less than 1, or / gene The expression rate in the cell set of cells is less than that of genes. When the expression ratio in all cells is considered, the gene specificity score in the cell is set to 0. In this scheme, when the gene specificity score in the cell is less than 1, it is considered that the gene expression in the cell is not specific. When the gene expression ratio in the cell set is less than the gene expression ratio in all cells, even if the gene specificity score in the cell is greater than 1, it is considered as technical noise. Thus, while filtering technical noise, the sparsity of the gene specificity score matrix is ​​also guaranteed. S49. Exponentially scale the gene-specificity scores in cells to obtain scale-corrected gene-specificity scores in cells, and construct a gene-specificity score matrix based on the scale-corrected gene-specificity scores in cells. ,in, This indicates the number of genes in a single-cell transcriptomics dataset. The expression for calculating the gene-specific fraction in the scale-corrected cell is as follows: , in, Indicates the scaled-corrected number of... Genes in a cell The specificity score.

[0024] In this scheme, the gene specificity score can not only directly eliminate the effects of batch effect and technical noise at the gene expression matrix level, but also highlight the weight of genes specifically expressed in certain cell types and reduce the weight of widely expressed genes, so as to fit the single-cell analysis task of brain function.

[0025] S5. Based on the gene set and gene-specific score matrix, construct the spatial permutation test gene set, calculate the significance probability of association between a single cell and the spatial brain phenotype, and find cells that are significantly associated with the target spatial brain phenotype through multiple comparison correction. Having constructed a gene set, assigned gene weights to each gene within the set, and built a gene-specific score matrix, this scheme aims to further evaluate whether the gene set is enriched in a specific cell type and to perform analysis at the single-cell level. Python Bag scDRS Based on the existing process, improvements were made by using a spatial permutation test model to guide the selection of the control gene set in order to overcome false positives caused by spatial autocorrelation. S5 includes the following steps: S51, Regarding the first Wheel space permutation test, using partial least squares regression model fitting and And obtain the same as according to the methods in S33~S36. Spatial permutation test for equal number of genes in a gene set ; S52. According to the methods in S37~S38, assign Gene weight of each gene ,in, for Genes in; It does not contain any meaningful biological information, but it can reflect false positive genes that are more consistent with spatial autocorrelation patterns, therefore An empirical distribution that can be used to estimate the single-cell enrichment fraction of brain function; S53, according to and The original enrichment score at the single-cell level and the control score corresponding to each round of spatial permutation test were calculated, as well as the probability of significant association between single cells and spatial brain phenotypes. In this scheme, the original enrichment score at the single-cell level and the control score corresponding to each round of spatial permutation test reflect the degree of association between cells and target spatial brain phenotypes, as well as the degree of association with spatial autocorrelation patterns, respectively; the correlation between single cells and... The original enrichment scores and the significance probability of association between individual cells and spatial brain phenotypes are used. However, a single-cell transcriptomics dataset contains tens of thousands of cells. Therefore, multiple comparison corrections must be performed on the significance probabilities of association between multiple cells and spatial brain phenotypes. Individual cells are often divided into different clusters based on cell type or the tissue structure of cell type. If there are many cells with low significance probabilities of association between a certain cell type and spatial brain phenotype, then cells in this cell type are more likely to be true positives, while significant cells scattered in a certain cell type are more likely to be false positives. Therefore, this approach uses tissue-level false discovery rate correction. S53 includes the following steps: S531. Calculate the original enrichment score of each cell in the single-cell transcriptomics dataset in turn. The expression for calculating the original enrichment fraction of a single cell is as follows: , in, This represents the first cell in the single-cell transcriptomics dataset. The original enrichment fraction of each cell; S532, Regarding the first Wheel space permutation test, based on Calculated A control score; in this protocol, the control score can be used to control false positives; The formula for calculating the control score is as follows: in, Indicates the first In the test of wheel space permutation, the first The control score corresponding to each cell Indicates the first Genes in a cell Specificity score; S533. Standardize the original enrichment score and control score of a single cell to obtain standardized original enrichment score and control score; in this scheme, the standardization is achieved by first aligning according to the mean and then normalizing the weights. The standardized original enrichment scores and control scores are calculated using the following expressions: , , in, This represents the first cell in the single-cell transcriptomics dataset. The standardized raw enrichment fraction of each cell. Indicates the first In the test of wheel space permutation, the first The standardized control score corresponding to each cell; S534. Calculate the mean and variance of the empirical distribution obtained from the spatial permutation test, and align the standardized original enrichment score and control score to the standard normal distribution to obtain the original enrichment score and control score under the standard normal distribution. The calculation expressions for the original enrichment score and control score under the standard normal distribution are as follows: , , in, This represents the first cell in the single-cell transcriptomics dataset. The original enrichment fraction of each cell under the standard normal distribution Indicates the first In the test of wheel space permutation, the first The control score corresponding to each cell under the standard normal distribution This represents the mean of the empirical distribution obtained from the spatial permutation test. This represents the variance of the empirical distribution obtained from the spatial permutation test; S535. Based on the original enrichment fraction and control fraction under the standard normal distribution, calculate the significance probability of the association between a single cell and the spatial brain phenotype. The expression for calculating the significance probability of the association between a single cell and a spatial brain phenotype is as follows: , in, Indicates the first The probability of a significant association between individual cells and spatial brain phenotypes. This indicates the number of times the event occurred.

[0026] S54. Based on the cell clustering characteristics, perform multiple comparison correction on the significance probability of individual cells, and find cells that are significantly related to the target spatial brain phenotype.

[0027] S54 includes the following steps: S541, Setting the significance level The number of cell types is The significance probability of a single cell level without multiple comparison correction was tested on a cell type basis to obtain the number of cells that passed and failed the hypothesis test. The formulas for calculating the number of cells that passed and failed the hypothesis test are as follows: , , in, This represents the number of cells that passed the hypothesis test. This represents the number of cells that failed the hypothesis test, where, , Indicates the first Cell types; S542. For each cell type, based on the number of cells in that cell type that pass and fail the hypothesis test, assign weights to the significance probabilities of cells in that cell type with the spatial brain phenotype to obtain weighted significance probabilities. The expression for calculating the weighted significance probability is as follows: , in, Indicates the first The first of the cell types The weighted significance probability of each cell Indicates the first The first of the cell types The probability of significant association between individual cells and spatial brain phenotype; in this scheme, if =0, then The value of is denoted as infinity; S543. Perform multiple contrast correction on each weighted significance probability to obtain the corrected single-cell significance probability. Cells that can still pass the hypothesis test after multiple contrast correction are regarded as cells that are significantly related to the target space brain phenotype.

[0028] S543 includes the following steps: S5431. Based on the weighted significance probability of cells in each cell type, extract the weighted significance probability of each cell in the single-cell transcriptomics dataset. ; S5432, according to The cells in the single-cell transcriptomics dataset are reordered in ascending order to obtain the sorting index of each cell in the single-cell transcriptomics dataset. ; S5433, based on and The significance probability of each cell in the single-cell transcriptomics dataset after correction was calculated. The expression for calculating the significance probability of each cell after correction is as follows: , in, This represents the first cell in the single-cell transcriptomics dataset. The significance probability of individual cells after correction This represents the first cell in the single-cell transcriptomics dataset. The weighted significance probability of each cell; S5434. Based on the corrected significance probability of each cell, find the sorting index that meets the criteria for significant cells. And ranked the top single-cell transcriptomics datasets These cells were identified as cells significantly associated with the target spatial brain phenotype. The calculation expression for the significant cell condition is as follows: , in, This indicates taking the maximum value. It means that.

[0029] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for analyzing single-cell brain function based on multimodal whole-brain omics data association, characterized in that, Includes the following steps: S1. Preprocess probes and transcriptome samples from the spatial transcriptomics database, standardize and filter each gene, and map the spatial brain phenotype of interest onto the standard brain atlas to obtain the target gene expression matrix and the target spatial brain phenotype vector. S2. Perform partial least squares regression iteration based on the target gene expression matrix and the target space brain phenotype vector, and calculate the variable projection importance score of the gene. S3. Examine the correlation between the variable projection importance score of genes and the spatial autocorrelation pattern through zero distribution, sort and screen the genes, construct a gene set, and assign gene weights to each gene in the gene set. S4. Preprocess the single-cell transcriptomics dataset to obtain low-dimensional embeddings, and calculate the gene-specific score in each cell based on the neighboring cells of each cell and the ranking of each gene in the cell, and construct a gene-specific score matrix. S5. Based on the gene set and gene-specific score matrix, construct a spatial permutation test gene set based on the spatial permutation test, calculate the significance probability of association between a single cell and the spatial brain phenotype, and find cells that are significantly associated with the target spatial brain phenotype through multiple comparison correction.

2. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 1, characterized in that, S1 includes the following steps: S11. Use the Allen Brain Atlas as a spatial transcriptomics database at the whole-brain scale. S12. Preprocess probes and transcriptome samples from spatial transcriptomics databases; S13. Based on the scaling-enhanced sigmoid function, the expression of each gene is standardized according to the interval of each gene expression to obtain standardized gene expression. S14. Generate a gene expression matrix based on standardized gene expression. ,in, Indicates belonging to, Represents the set of real numbers. Indicates the number of transcriptome samples. Indicates the number of genes after standardization; S15, according to Transcriptome samples from spatial transcriptomics databases are assigned to standard brain atlases to generate gene expression matrices under the standard brain atlases. ,in, Indicates standard brain atlas labels; S16, calculate separately The stability of variation in each gene; The expression for calculating the variation stability of genes in S16 is as follows: , in, Indicates gene The stability of the variation, This indicates the number of sources of brain data. Indicates gene Sources of brain data The spatial representation vector in the text, Indicates gene Sources of brain data The spatial representation vector in the text, express and The Pearson correlation coefficient between them; S17, to The variation stability of each gene in the matrix is ​​sorted from largest to smallest, and the top 50% of genes are filtered out to obtain the target gene expression matrix. ; S18. Map the spatial brain phenotype of interest onto a standard brain atlas to obtain the target spatial brain phenotype vector. .

3. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 2, characterized in that, S2 includes the following steps: S21, will Let it be denoted as the independent variable matrix and will Denote as spatial brain phenotype vector ,in, This indicates the number of brain regions in a standard brain atlas. This indicates the number of genes remaining after filtering. S22. Based on the 10-fold cross-validation method, and Exploration yields the target optimal number of components ; S23, based on and The number of components is The partial least squares regression iterations are used to obtain the parameters of the partial least squares regression model, which include gene weights, loadings, and regression coefficients. S24. Calculate the variable projection importance score of the gene based on the parameters of the partial least squares regression model. .

4. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 3, characterized in that, S3 includes the following steps: S31. Treat the human cerebral cortex as a sphere and rotate the spatial brain phenotype on the surface of the sphere to obtain the zero distribution. ,in, Indicates the number of space permutation checks; S32. Use the zero distribution to check the probability that the partial least squares regression model predicts spatial brain phenotype vectors better than spatial brain phenotypes, and calculate the significance probability of the spatial permutation test. S33. Use the null distribution to examine the significance probability between the variable projection importance score and the spatial autocorrelation pattern of each gene in the partial least squares regression model; S34. Sort all genes in ascending order of their significance probability at the gene level. If genes have equal significance probabilities, sort them in descending order of their variable projection importance scores. This yields the gene ranking results. ; S35. Selecting through cross-validation Center front Construct a gene set from individual genes; S36. To study genes more closely associated with the positive side of the spatial brain phenotype, additional screening is required. Genes with a value greater than 0 are added to the gene set, and after screening, the resulting gene set is denoted as . ; S37. If it is necessary to study genes on the negative side of the spatial brain phenotype, additional screening is required. Genes with a value less than 0 are selected, and the resulting gene set is denoted as [genes whose value is less than 0]. ; S38, based on and , respectively Each gene in the gene is assigned a gene weight; The formula for calculating the gene weight is as follows: in, Indicates gene Gene weights.

5. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 4, characterized in that, The method for finding neighboring cells of a cell in S4 to calculate the gene-specific score in that cell, and constructing the gene-specific score matrix, includes the following steps: S46. Calculate the rank of each gene in the cell; The formula for calculating the rank of each gene in the cell is as follows: , in, Indicates gene In the Ranking of individual cells Indicates the first A ranking function for genes in each cell. Indicating the target gene expression matrix, the first... Genes in a cell ; S47. Based on low-dimensional embedding, use cosine similarity to find neighboring cells of the cell, and construct a cell set of the cell based on the cell and its neighboring cells, and calculate the gene specificity score of the cell based on the cell set. The expression for calculating the gene-specific fraction in the cell is as follows: , in, Indicates the first Genes in a cell Specificity score, This represents the number of neighboring cells of a cell. Indicates the first Each cell, Indicates the first A collection of cells consisting of an individual cell and its neighboring cells. Indicates gene In the Ranking of individual cells This represents the total number of cells in a single-cell transcriptomics dataset. Indicates gene In the Ranking in individual cells; S48. If the gene specificity fraction in a cell is less than 1, or / gene The expression rate in the cell set of cells is less than that of genes. When considering the expression ratio across all cells, the gene-specific fraction in the cell is set to 0. S49. Exponentially scale the gene-specificity scores in cells to obtain scale-corrected gene-specificity scores in cells, and construct a gene-specificity score matrix based on the scale-corrected gene-specificity scores in cells. ,in, This indicates the number of genes in a single-cell transcriptomics dataset. The expression for calculating the gene-specific fraction in the scale-corrected cell is as follows: , in, Indicates the scaled-corrected number of... Genes in a cell The specificity score.

6. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 5, characterized in that, S5 includes the following steps: S51, Regarding the first Wheel space permutation test, using partial least squares regression model fitting and And obtain the same as according to the methods in S33~S36. Spatial permutation test for equal number of genes in a gene set ; S52. According to the methods in S37~S38, assign Gene weight of each gene ,in, for Genes in; S53, according to and The original enrichment score at the single-cell level and the control score corresponding to each round of spatial permutation test were calculated, as well as the probability of significant association between single cells and spatial brain phenotypes. S54. Based on the cell clustering characteristics, perform multiple comparison correction on the significance probability of individual cells, and find cells that are significantly related to the target spatial brain phenotype.

7. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 6, characterized in that, The expression for calculating the significance probability of the association between a single cell and a spatial brain phenotype in S53 is as follows: , in, Indicates the first The probability of a significant association between individual cells and spatial brain phenotypes. Indicates the number of times it occurs. This represents the first cell in the single-cell transcriptomics dataset. The original enrichment fraction of each cell under the standard normal distribution Indicates the first In the test of wheel space permutation, the first The control score corresponding to each cell under the standard normal distribution.

8. The method for single-cell analysis of brain function based on multimodal whole-brain omics data association according to claim 7, characterized in that, S54 includes the following steps: S541, Setting the significance level The number of cell types is The significance probability of a single cell level without multiple comparison correction was tested on a cell type basis to obtain the number of cells that passed and failed the hypothesis test. The formulas for calculating the number of cells that passed and failed the hypothesis test are as follows: , , in, This represents the number of cells that passed the hypothesis test. This represents the number of cells that failed the hypothesis test, where, , Indicates the first Cell types; S542. For each cell type, based on the number of cells in that cell type that pass and fail the hypothesis test, assign weights to the significance probabilities of cells in that cell type with the spatial brain phenotype to obtain weighted significance probabilities. The expression for calculating the weighted significance probability is as follows: , in, Indicates the first The first of the cell types The weighted significance probability of each cell Indicates the first The first of the cell types The probability of significant association between individual cells and spatial brain phenotype; in this scheme, if =0, then The value of is denoted as infinity; S543. Perform multiple contrast correction on each weighted significance probability to obtain the corrected single-cell significance probability. Cells that can still pass the hypothesis test after multiple contrast correction are regarded as cells that are significantly related to the target space brain phenotype.