Cell state analysis method and system based on single cell transcriptome Pearson residual error

Through single-cell sequencing technology and Pearson residual analysis, combined with microbial genome metabolism model, metabolic flow is dynamically adjusted, which solves the unreliable problem of single-cell metabolic state analysis in traditional methods, and realizes the precise characterization and diversity quantification of microbial cell state.

CN120356528APending Publication Date: 2025-07-22SUN YAT SEN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325369.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional gene expression analysis methods cannot accurately reflect the open characteristics of single cells and external substance exchange, and the low depth of single cells sequencing data leads to high noise in functional gene expression, which cannot accurately distinguish between true expression differences and sequencing errors, resulting in unreliable cell status analysis results.

Method used

The number of transcripts of the microbial cell population was obtained through single-cell sequencing technology, Pearson's residuals were calculated, combined with the microbial genomic metabolism model and the coordinated regulation of the metabolic grid structure, dynamically adjust the metabolic flow, generate individualized metabolic flow vectors, and analyze cell state diversity through similarity values.

Benefits of technology

The precise characterization of the metabolic diversity of microbial single cells in an open system is achieved, which eliminates the noise caused by sequencing depth differences, improves the comparability of gene expression data, accurately reflects the differences in metabolic status of single cells, and provides quantitative tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356528A_ABST
    Figure CN120356528A_ABST
Patent Text Reader

Abstract

The invention discloses a single-cell transcriptome Pearson residual error-based cell state analysis method and system, and the method comprises the steps: collecting the transcript number of each functional gene of each microbial cell in a microbial cell population according to a single-cell sequencer, and obtaining microbial single-cell transcriptome data; calculating a Pearson residual error corresponding to each functional gene according to the microbe single cell transcriptome data and a Pearson residual error algorithm; calculating an initial metabolism flow solution in the microbial genome metabolism model; calculating a corresponding weighting coefficient for each biochemical reaction according to each Pearson residual error and the coordinated regulation metabolism grid structure of each functional gene; according to each weighting coefficient, correcting the initial metabolic flow of each biochemical reaction of each microbial cell, and generating a corresponding metabolic flow vector; and calculating a similarity value between each metabolic flow vector and the initial metabolic flow solution, and outputting a cell state analysis result of the microbial cell population. According to the invention, accurate characterization of microbial single-cell metabolism diversity in an open system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics and relates to a method and system for analyzing cell states based on Pearson residuals of single-cell transcriptomes. Background Art

[0002] The diversity of metabolic states of microorganisms in different environments is a key characteristic for their adaptation to complex ecosystems. Single-cell sequencing technology provides gene expression data at the single-cell level. Combining flux balance analysis (FBA) of genome-scale metabolic models (GEMs), it is theoretically possible to infer metabolic activity from gene expression and thereby characterize cell states.

[0003] However, traditional FBA is based on the assumption of a closed system and cannot reflect the open characteristics of material exchange between single cells and the outside world. At the same time, the low depth of single-cell sequencing data leads to high noise in the expression of functional genes. When directly used for calculating metabolic fluxes, it is impossible to accurately distinguish true expression differences from sequencing errors, resulting in unreliable cell state analysis results. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present application provides a method and system for analyzing cell states based on Pearson residuals of single-cell transcriptomes, realizing accurate characterization of the metabolic diversity of single microbial cells in an open system.

[0005] To achieve the above object, in a first aspect, the present invention provides a method for analyzing cell states based on Pearson residuals of single-cell transcriptomes, including:

[0006] Collecting the transcript counts of each functional gene of each microbial cell in a microbial cell population according to a single-cell sequencer to obtain single-cell transcriptome data of the microorganisms;

[0007] Calculating the Pearson residuals corresponding to each functional gene of each microbial cell according to the single-cell transcriptome data of the microorganisms and a preset Pearson residual algorithm;

[0008] Calculating an initial metabolic flux solution in a preset microbial genome-scale metabolic model; wherein, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome-scale metabolic model;

[0009] Calculating corresponding weighting coefficients for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction;

[0010] Correcting the initial metabolic fluxes corresponding to each biochemical reaction of each microbial cell according to each of the weighting coefficients to generate a metabolic flux vector for each microbial cell;

[0011] Calculate the similarity values between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis result of the microbial cell population.

[0012] Compared with the prior art, the embodiments of the present application have the following beneficial effects: By using single-cell sequencing technology to accurately obtain the number of functional gene transcripts of each cell in the microbial population, it provides the original data basis for subsequent calculation of Pearson residuals and metabolic model correction; By calculating the Pearson residuals of the microbial single-cell transcriptome data, the noise caused by differences in sequencing depth is effectively eliminated, and the comparability of gene expression data is improved; Solving the initial metabolic flux solution based on the microbial genome metabolic model provides a standardized benchmark for subsequent correction; Calculating the weighting coefficient according to the Pearson residuals and the co-regulation metabolic grid structure, dynamically adjusting the metabolic flux, breaking through the limitation of traditional flux balance analysis relying on a closed system; Generating an individualized metabolic flux vector by correcting the initial metabolic flux, accurately reflecting the metabolic state differences of single cells; Finally, analyzing the cell state diversity through similarity values provides a quantitative tool for the study of metabolic heterogeneity in complex microbial communities.

[0013] In some embodiments of the first aspect of the present application, calculating the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm includes:

[0014] Calculating the corresponding Pearson residuals according to the number of transcripts of each functional gene of each microbial cell in the microbial single-cell transcriptome data; wherein, the calculation method is as follows:

[0015] Where Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, x g,j represents the number of transcripts of the functional gene g of the microbial cell j, μ g is the average number of transcripts of the functional gene g of all microbial cells in the sample; σ g is the standard deviation of the number of transcripts of the functional gene g of all microbial cells in the sample.

[0016] Compared with the prior art, the above embodiments have the following beneficial effects: By specifically defining the calculation of Pearson residuals, the normalization processing of gene expression levels between single cells is realized, avoiding the deviation caused by uneven distribution of transcript numbers between samples, enhancing the comparability of the expression activities of functional genes between different cells, and providing high-signal-to-noise input data for the subsequent construction of weighting coefficients.

[0017] In some embodiments of the first aspect of the present application, calculating the initial metabolic flux solution in the preset microbial genome metabolic model includes:

[0018] Construct a homogeneous linear equation system according to the stoichiometric matrix in the microbial genome metabolic model; wherein, the homogeneous linear equation system is: S·V = 0; where S represents the stoichiometric matrix and V represents the initial metabolic flux solution;

[0019] Solve the homogeneous linear equation system through a preset objective function and constraint conditions to obtain the initial metabolic flux solution.

[0020] Compared with the prior art, the above embodiments have the following beneficial effects: constructing a homogeneous linear equation system based on the stoichiometric matrix and solving the initial metabolic flux solution ensures the mathematical rigor of the metabolic flux calculation. At the same time, optimizing the selection of the solution through the objective function and constraint conditions (such as maximizing biomass synthesis) makes the initial solution more in line with the actual metabolic characteristics of microorganisms, providing a reasonable starting point for subsequent correction.

[0021] In some embodiments of the first aspect of the present application, calculating the corresponding weighted coefficient for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction includes:

[0022] Calculate the corresponding first weight coefficient for each functional gene of each microbial cell according to a preset weight calculation algorithm and each of the Pearson residuals; wherein, the weight calculation algorithm is as follows:

[0023] wherein, Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, W g,j represents the first weight coefficient corresponding to the functional gene g of the microbial cell j, K represents the magnification relationship between gene expression and metabolic flux, and sgn() represents the sign function;

[0024] Calculate the corresponding weighted coefficient for each biochemical reaction of each microbial cell respectively according to each of the first weight coefficients and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction.

[0025] Compared with the prior art, the above embodiments have the following beneficial effects: using a piecewise function to calculate the weighted coefficient of the functional gene, dynamically adjusting the weight according to the absolute value of the Pearson residual, smoothing the response to gene expression changes through an exponential function when the residual is small, and enhancing the sensitivity to significant differences through linear scaling combined with the sign function when the residual is large, making the weighted coefficient more suitable for the non-linear relationship between gene expression and metabolic flux, and improving the flexibility of correction.

[0026] In some embodiments of the first aspect of the present application, calculating the corresponding weighted coefficients for each biochemical reaction of each microbial cell according to each of the first weight coefficients and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction includes:

[0027] Calculating the first weighted coefficient of each reaction path according to the co-regulation metabolic grid structure of each functional gene in each reaction path in each biochemical reaction and the first weight coefficient;

[0028] Among them, the specific calculation of the first weighted coefficient of each reaction path is as follows: If a certain reaction path requires multiple tandem functional genes to drive simultaneously, then the minimum value among the first weight coefficients corresponding to each tandem functional gene is used as the first weighted coefficient of the corresponding reaction path; If a certain reaction path only requires any one of multiple parallel functional genes to drive independently, then the average value of the first weight coefficients corresponding to each parallel functional gene is used as the first weighted coefficient of the corresponding reaction path;

[0029] Calculating the weighted coefficients corresponding to each biochemical reaction of each microbial cell according to each of the first weighted coefficients; among them, the calculation method of the weighted coefficient is as follows:

[0030] F r,j =∏ p∈r W p,j ; where F r,j represents the weighted coefficient corresponding to the biochemical reaction r of cell j, and W p,j represents the first weighted coefficient corresponding to the reaction path p in the biochemical reaction r of cell j.

[0031] Compared with the prior art, the above embodiments have the following beneficial effects: Differentially processing the weighted coefficients for the co-regulation structures of tandem and parallel functional genes: Taking the minimum weight for tandem genes to ensure the stability of the key path, and taking the average weight for parallel genes to balance the influence of redundant paths, avoiding the metabolic flux distortion that may be caused by a single processing method, and making the corrected metabolic flux closer to the real biological scenario.

[0032] In some embodiments of the first aspect of the present application, correcting the initial metabolic flux corresponding to each biochemical reaction of each microbial cell according to each of the weighted coefficients to generate the metabolic flux vector of each microbial cell includes:

[0033] Calculating the metabolic flux vector of each microbial cell according to the initial metabolic flux solution, each of the weighted coefficients, and the stoichiometric numbers corresponding to each biochemical reaction in the microbial genome metabolic model; among them, the calculation is as follows:

[0034] F j =∑ r∈M (Sr ·V r ·F r,j );where S r represents the stoichiometric coefficients of each compound in the biochemical reaction r in the microbial genome metabolic model M, V r represents the initial metabolic flux corresponding to the biochemical reaction r in the initial metabolic flux solution, F j represents the metabolic flux vector of cell j.

[0035] Compared with the prior art, the above embodiments have the following beneficial effects: By combining the initial metabolic flux solution, the weighting coefficient, and the stoichiometric coefficients to generate a metabolic flux vector, the global metabolic model is fused with the single-cell specific expression data. The corrected vector not only retains the topological information of the genomic metabolic network but also integrates the gene expression characteristics of individual cells, achieving accurate characterization of the metabolic state under an open system.

[0036] In some embodiments of the first aspect of the present application, calculating the similarity values between each of the metabolic flux vectors and the initial metabolic flux solution and outputting the cell state analysis result of the microbial cell population includes:

[0037] Calculating the similarity values between the corresponding metabolic flux vectors and the initial metabolic flux solution for each microbial cell according to a preset similarity value calculation algorithm;

[0038] Outputting the cell state analysis result of the microbial cell population according to the data distribution of each of the similarity values.

[0039] Compared with the prior art, the above embodiments have the following beneficial effects: By calculating the similarity values between the metabolic flux vectors and the initial solution, the complex metabolic state differences are converted into quantifiable statistical indicators, and at the same time, the analysis result is output based on the data distribution, making the evaluation of cell state diversity more intuitive and providing a clear statistical basis for experimental verification or subsequent analysis.

[0040] In some embodiments of the first aspect of the present application, calculating the similarity values between the corresponding metabolic flux vectors and the initial metabolic flux solution for each microbial cell according to a preset similarity value calculation algorithm includes:

[0041] Calculating the Pearson correlation coefficient between the corresponding metabolic flux vectors and the initial metabolic flux solution for each microbial cell according to a preset Pearson correlation coefficient algorithm as the similarity value; where the Pearson correlation coefficient algorithm is as follows:

[0042] Sim j represents the similarity value between the metabolic flux vector of cell j and the initial metabolic flux solution, C represents the total number of compounds c in the microbial genome metabolic model, F j,cRepresents the corrected flux of compound c in the metabolic flux vector of cell j, represents the average value of the corrected fluxes of all compounds in the metabolic flux vector of cell j, V c represents the flux of compound c in the initial metabolic flux solution V, represents the average value of the fluxes of all compounds in the initial metabolic flux solution V.

[0043] Compared with the prior art, the above embodiments have the following beneficial effects: The Pearson correlation coefficient is used to quantify the similarity value, and the normalization processing of covariance and standard deviation is used to eliminate the influence of dimensions, ensuring the comparability between the fluxes of different compounds and making the evaluation results of cell state diversity statistically significant and repeatable.

[0044] In some embodiments of the first aspect of the present application, the outputting the cell state analysis result of the microbial cell population according to the data distribution of each similarity value includes:

[0045] Calculating the corresponding average value according to each Pearson correlation coefficient;

[0046] Characterizing and outputting the cell state analysis result of the microbial cell population according to the average value.

[0047] Compared with the prior art, the above embodiments have the following beneficial effects: By calculating the average value of the similarity values and outputting the analysis result accordingly, the interpretation process of complex data is simplified. The smaller the average value, the greater the difference in the metabolic states between cells, which intuitively reflects the heterogeneity degree of the microbial community and provides a direct basis for environmental response or function prediction.

[0048] In the second aspect, the present invention also provides a cell state analysis system based on Pearson residuals of single-cell transcriptome, including: a data acquisition module, a residual calculation module, an initial metabolic flux solution calculation module, a weighted coefficient calculation module, a correction calculation module, and a result output module;

[0049] Among them, the data acquisition module is used to collect the transcript numbers of each functional gene of each microbial cell in the microbial cell population according to a single-cell sequencer to obtain microbial single-cell transcriptome data;

[0050] The residual calculation module is used to calculate the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm;

[0051] The initial metabolic flux solution calculation module is used to calculate the initial metabolic flux solution in a preset microbial genome metabolic model; wherein, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome metabolic model;

[0052] The weighted coefficient calculation module is configured to calculate corresponding weighted coefficients for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction;

[0053] The correction calculation module is configured to correct the initial metabolic fluxes corresponding to each biochemical reaction of each microbial cell according to each of the weighted coefficients, and generate a metabolic flux vector of each microbial cell;

[0054] The result output module is configured to calculate the similarity value between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis result of the microbial cell population.

[0055] Compared with the prior art, the embodiments of the present application have the following beneficial effects: accurately obtaining the functional gene transcript number of each cell in a microbial population through single-cell sequencing technology provides a raw data basis for subsequent calculation of Pearson residuals and metabolic model correction; effectively eliminating the noise caused by differences in sequencing depth by calculating the Pearson residuals of microbial single-cell transcriptome data improves the comparability of gene expression data; solving the initial metabolic flux solution based on the microbial genome metabolic model provides a standardized benchmark for subsequent correction; calculating weighted coefficients according to Pearson residuals and co-regulatory metabolic grid structure, dynamically adjusting metabolic fluxes, breaking through the limitation that traditional flux balance analysis relies on a closed system; generating individualized metabolic flux vectors by correcting the initial metabolic fluxes, accurately reflecting the differences in single-cell metabolic states; finally, analyzing the cell state diversity through similarity values provides a quantitative tool for the study of metabolic heterogeneity in complex microbial communities. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 : is a schematic flowchart of a method for analyzing cell state based on Pearson residuals of single-cell transcriptome provided in some embodiments of the present invention.

[0057] Figure 2 : is a schematic structural diagram of a system for analyzing cell state based on Pearson residuals of single-cell transcriptome provided in some embodiments of the present invention.

[0058] Figure 3 : is a schematic diagram of the calculation principle of an initial metabolic flux solution provided in some embodiments of the present invention.

[0059] Figure 4 : is a schematic diagram of the calculation principle of correcting an initial metabolic flux to obtain a metabolic flux vector provided in some embodiments of the present invention.

[0060] Figure 5: Prediction results of the diversity of Escherichia coli cell states under different carbon sources (glucose (G) and sodium acetate (A)) and culture times provided in some embodiments of the present invention. Detailed implementation manners

[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0062] Embodiment 1:

[0063] Please refer to Figure 1 , a cell state analysis method based on Pearson residuals of single-cell transcriptome provided in the embodiment of the present invention, including steps S1 to S6:

[0064] Step S1: According to a single-cell sequencer, collect the transcript numbers of each functional gene of each microbial cell in a microbial cell population to obtain microbial single-cell transcriptome data.

[0065] In specific implementation, the microbial single-cell transcriptome data can be stored in the form of a matrix for convenient subsequent calculation, where each element in the matrix is the transcript number of each functional gene of each microbial cell.

[0066] In this embodiment, step S1 accurately obtains the transcript numbers of functional genes of each cell in the microbial population through single-cell sequencing technology, providing a raw data basis for subsequent calculation of Pearson residuals and metabolic model correction.

[0067] Step S2: Calculate the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm.

[0068] Further, the step S2 can be implemented through the following preferred implementation manners, specifically as follows:

[0069] According to the transcript numbers of each functional gene of each microbial cell in the microbial single-cell transcriptome data, calculate the corresponding Pearson residuals; where the calculation method is as follows:

[0070] Where Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, x g,j represents the transcript number of the functional gene g of the microbial cell j, μ gis the average number of transcripts of the functional gene g of all microbial cells in the sample; σ g is the standard deviation of the number of transcripts of the functional gene g of all microbial cells in the sample.

[0071] In specific implementation, the single-cell transcriptome data of microorganisms is stored in the form of a matrix. When calculating the Pearson residuals, this matrix can be directly input into the SCTransform function of the Seurat package in R language to directly obtain the corresponding Pearson residual matrix. Each element in this Pearson residual matrix is the single-cell transcriptome Pearson residual of each functional gene of each microbial cell.

[0072] In this preferred embodiment, step S2 realizes the normalization processing of the gene expression levels between single cells by specifically defining the calculation of Pearson residuals, avoids the deviation caused by the uneven distribution of the number of transcripts between samples, enhances the comparability of the expression activities of functional genes between different cells, and provides input data with high signal-to-noise ratio for the construction of subsequent weighting coefficients.

[0073] Step S3: Calculate the initial metabolic flux solution in the preset microbial genome metabolic model; wherein, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome metabolic model.

[0074] In specific implementation, the microbial genome metabolic model can annotate the microbial genome through the RASTtk v1.073 module of the KBase platform, and then be constructed through the MS2-Build Prokaryotic Metabolic Models with OMEGGA module, which belongs to the prior art and will not be elaborated here.

[0075] Further, as Figure 3 shown in a schematic diagram of the calculation principle of an initial metabolic flux solution, step S3 can be implemented through the following preferred implementation manner, including steps S31 - S32, specifically as follows:

[0076] S31: Construct a homogeneous linear equation system according to the stoichiometric matrix in the microbial genome metabolic model; wherein, the homogeneous linear equation system is: S·V = 0; where S represents the stoichiometric matrix and V represents the initial metabolic flux solution;

[0077] S32: Solve the homogeneous linear equation system through a preset objective function and constraint conditions to obtain the initial metabolic flux solution.

[0078] Specifically, when solving, the Run Flux Balance Analysis module of the KBase platform can be used for solving to improve efficiency. In this initial flux solution, it contains the initial metabolic fluxes corresponding to each biochemical reaction r, providing a starting point for subsequent corrections.

[0079] In this preferred embodiment, a homogeneous linear equation system is constructed based on the stoichiometric matrix and the initial metabolic flux solution is solved, ensuring the mathematical rigor of the metabolic flux calculation. At the same time, the selection of the solution is optimized through the objective function and constraint conditions (such as maximizing biomass synthesis), making the initial solution more in line with the actual metabolic characteristics of microorganisms and providing a reasonable starting point for subsequent corrections.

[0080] Step S4: According to each of the Pearson residuals and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction, calculate the corresponding weighted coefficients for each biochemical reaction of each microbial cell.

[0081] Further, step S4 can be implemented through the following preferred implementation manner, specifically including steps S41 - S42, as follows:

[0082] S41: According to a preset weight calculation algorithm and each of the Pearson residuals, calculate the corresponding first weight coefficient for each functional gene of each microbial cell; wherein, the weight calculation algorithm is as follows:

[0083] wherein, Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, W g,j represents the first weight coefficient corresponding to the functional gene g of the microbial cell j, K represents the magnification relationship between gene expression and metabolic flux, and sgn() represents the sign function.

[0084] In specific implementation, 2 can be used as a preferred choice for the value of K.

[0085] S42: According to each of the first weight coefficients and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction, calculate the corresponding weighted coefficients for each biochemical reaction of each microbial cell respectively.

[0086] In this preferred embodiment, step S4 uses a piecewise function to calculate the weighted coefficients of functional genes, dynamically adjusts the weights according to the absolute value of the Pearson residual, smooths the response to gene expression changes through an exponential function when the residual is small, and enhances the sensitivity to significant differences through linear scaling combined with the sign function when the residual is large, making the weighted coefficients more conform to the non-linear relationship between gene expression and metabolic flux and improving the flexibility of correction.

[0087] Further, the step S42 can be implemented by the following preferred embodiments, including steps S421 - S422, specifically as follows:

[0088] S421: Calculate the first weighted coefficient of each reaction path according to the co - regulatory metabolic grid structure of each functional gene in each reaction path in each biochemical reaction and the first weight coefficient.

[0089] Among them, the calculation of the first weighted coefficient of each reaction path is specifically as follows: If a certain reaction path requires multiple tandem functional genes to drive simultaneously, then take the minimum value among the first weight coefficients corresponding to each tandem functional gene as the first weighted coefficient of the corresponding reaction path; if a certain reaction path only requires any one of multiple parallel functional genes to drive independently, then take the average value of the first weight coefficients corresponding to each parallel functional gene as the first weighted coefficient of the corresponding reaction path.

[0090] S422: Calculate the weighted coefficient corresponding to each biochemical reaction of each microbial cell according to each of the first weighted coefficients; the calculation method of the weighted coefficient is as follows:

[0091] F r,j =∏ p∈r W p,j ; where F r,j represents the weighted coefficient corresponding to the biochemical reaction r of cell j, and W p,j represents the first weighted coefficient corresponding to the reaction path p in the biochemical reaction r of cell j.

[0092] In this preferred embodiment, S42 differentially processes the weighted coefficients for the co - regulatory structures of tandem and parallel functional genes: taking the minimum weight for tandem genes to ensure the stability of the key path, and taking the average weight for parallel genes to balance the influence of redundant paths, avoiding the metabolic flux distortion that may be caused by a single processing method, and making the corrected metabolic flux closer to the real biological scenario.

[0093] Step S5: Correct the initial metabolic flux corresponding to each biochemical reaction of each microbial cell according to each of the weighted coefficients to generate the metabolic flux vector of each microbial cell.

[0094] Further, as Figure 4 shown in a schematic diagram of the calculation principle for correcting the initial metabolic flux to obtain the metabolic flux vector, the step S5 can be implemented by the following preferred embodiments, specifically as follows:

[0095] Calculate the metabolic flux vector of each microbial cell according to the initial metabolic flux solution, each of the weighted coefficients, and the stoichiometric number corresponding to each biochemical reaction in the microbial genome metabolic model; the calculation is as follows: F j= ∑ r∈M (S r ·V r ·F r,j ); where S r represents the stoichiometric coefficients of each compound in the biochemical reaction r in the microbial genome metabolic model M, V r represents the initial metabolic flux corresponding to the biochemical reaction r in the initial metabolic flux solution, and F j represents the metabolic flux vector of cell j.

[0096] In this preferred embodiment, step S5 generates a metabolic flux vector by combining the initial metabolic flux solution, the weighting coefficient, and the stoichiometric coefficients, fuses the global metabolic model with the single-cell specific expression data. The corrected vector not only retains the topological information of the genomic metabolic network but also integrates the gene expression characteristics of individual cells, achieving an accurate characterization of the metabolic state under an open system.

[0097] Step S6: Calculate the similarity values between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis result of the microbial cell population.

[0098] Further, step S6 can be implemented through the following preferred implementation manner, including steps S61 - S62, specifically as follows:

[0099] S61: According to the preset similarity value calculation algorithm, calculate the similarity values between the corresponding metabolic flux vectors and the initial metabolic flux solutions for each microbial cell;

[0100] S62: Output the cell state analysis result of the microbial cell population according to the data distribution of each of the similarity values.

[0101] In this preferred embodiment, steps S61 - S62 convert the complex metabolic state differences into quantifiable statistical indicators by calculating the similarity values between the metabolic flux vectors and the initial solutions, and output the analysis result based on the data distribution, making the evaluation of cell state diversity more intuitive and providing a clear statistical basis for experimental verification or subsequent analysis.

[0102] Further, step S61 can be implemented through the following preferred implementation manner, specifically as follows:

[0103] According to the preset Pearson correlation coefficient algorithm, calculate the Pearson correlation coefficient between the corresponding metabolic flux vectors and the initial metabolic flux solutions for each microbial cell as the similarity value; where the Pearson correlation coefficient algorithm is as follows:

[0104] Sim jrepresents the similarity value between the metabolic flux vector of cell j and the initial metabolic flux solution. C represents the total number of compounds c in the microbial genome metabolic model, and F j,c represents the corrected flux of compound c in the metabolic flux vector of cell j, represents the average value of the corrected fluxes of all compounds in the metabolic flux vector of cell j, V c represents the flux of compound c in the initial metabolic flux solution V, represents the average value of the fluxes of all compounds in the initial metabolic flux solution V.

[0105] In this preferred embodiment, step S61 quantifies the similarity value using the Pearson correlation coefficient, and eliminates the influence of dimensions through the normalization process of covariance and standard deviation, ensuring the comparability between the fluxes of different compounds, and making the evaluation results of cell state diversity statistically significant and repeatable.

[0106] Furthermore, the step S62 can be implemented through the following preferred implementation manner, including steps S621 - S622, specifically as follows:

[0107] S621: Calculate the corresponding average value according to each of the Pearson correlation coefficients;

[0108] S622: Characterize and output the cell state analysis result of the microbial cell population according to the average value.

[0109] In this preferred embodiment, steps S621 - S622 simplify the interpretation process of complex data by calculating the average value of the similarity values and outputting the analysis result accordingly. The smaller the average value, the greater the difference in the metabolic states between cells, intuitively reflecting the degree of heterogeneity of the microbial community, and providing a direct basis for environmental response or function prediction.

[0110] For example, as an embodiment, as Figure 5 shown in the prediction results of the cell state diversity of a certain Escherichia coli under different carbon sources (glucose (G) and sodium acetate (A)) and culture times. From the overall trend, as the culture time increases, the cell state diversity shows a gradually increasing trend; at the same time, the type of carbon source and the conversion of carbon sources (such as glucose first and then sodium acetate, etc.) also have a significant impact on the cell state diversity. Different combinations of carbon sources and culture times will cause Escherichia coli cells to exhibit different characteristics and changing trends in terms of state diversity.

[0111] In summary, compared with the prior art, the above embodiments of the present application have the following beneficial effects: By using single-cell sequencing technology, the transcript number of functional genes of each cell in the microbial population is accurately obtained, providing the original data basis for subsequent calculation of Pearson residuals and metabolic model correction; By calculating the Pearson residuals of microbial single-cell transcriptome data, the noise caused by differences in sequencing depth is effectively eliminated, improving the comparability of gene expression data; Solving the initial metabolic flux solution based on the microbial genome metabolic model provides a standardized benchmark for subsequent correction; Calculating the weighting coefficient according to the Pearson residuals and the co-regulation metabolic grid structure, dynamically adjusting the metabolic flux, breaking through the limitation that traditional flux balance analysis depends on a closed system; Generating an individualized metabolic flux vector by correcting the initial metabolic flux, accurately reflecting the differences in single-cell metabolic states; Finally, analyzing the diversity of cell states through similarity values provides a quantitative tool for studying the metabolic heterogeneity of complex microbial communities.

[0112] Embodiment 2:

[0113] Please refer to Figure 2 , based on the same inventive concept, a cell state analysis system based on Pearson residuals of single-cell transcriptome disclosed in an embodiment of the present invention includes: a data acquisition module M1, a residual calculation module M2, an initial metabolic flux solution calculation module M3, a weighting coefficient calculation module M4, a correction calculation module M5, and a result output module M6;

[0114] Among them, the data acquisition module M1 is used to collect the transcript number of each functional gene of each microbial cell in the microbial cell population according to a single-cell sequencer, so as to obtain microbial single-cell transcriptome data.

[0115] In this embodiment, the data acquisition module M1 accurately obtains the transcript number of functional genes of each cell in the microbial population through single-cell sequencing technology, providing the original data basis for subsequent calculation of Pearson residuals and metabolic model correction.

[0116] The residual calculation module M2 is used to calculate the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm.

[0117] Further, the residual calculation module M2 can be implemented through the following preferred implementation manner, specifically: According to the transcript number of each functional gene of each microbial cell in the microbial single-cell transcriptome data, calculate the corresponding Pearson residuals; The calculation method is as follows:

[0118] Where Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, and x g,jThe transcript count of functional gene g of microbial cell j, μ g is the average transcript count of functional gene g of all microbial cells in the sample; σ g is the standard deviation of the transcript counts of functional gene g of all microbial cells in the sample.

[0119] In this preferred embodiment, the residual calculation module M2 realizes the normalization processing of gene expression levels between single cells by specifically defining the calculation of Pearson residuals, avoids the deviation caused by uneven distribution of transcript counts between samples, enhances the comparability of the expression activities of functional genes between different cells, and provides high-signal-to-noise input data for the construction of subsequent weighting coefficients.

[0120] The initial metabolic flux solution calculation module M3 is used to calculate the initial metabolic flux solution in a preset microbial genome metabolic model; wherein, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome metabolic model.

[0121] Further, the initial metabolic flux solution calculation module M3 includes: an equation construction unit and an equation solving unit;

[0122] Among them, the equation construction unit is used to construct a homogeneous linear equation system according to the stoichiometric matrix in the microbial genome metabolic model; wherein, the homogeneous linear equation system is: S·V = 0; where S represents the stoichiometric matrix and V represents the initial metabolic flux solution;

[0123] The equation solving unit is used to solve the homogeneous linear equation system through a preset objective function and constraint conditions to obtain the initial metabolic flux solution.

[0124] In this preferred embodiment, the initial metabolic flux solution calculation module M3 constructs a homogeneous linear equation system based on the stoichiometric matrix and solves the initial metabolic flux solution, ensuring the mathematical rigor of the metabolic flux calculation. At the same time, the selection of the solution is optimized through the objective function and constraint conditions (such as maximizing biomass synthesis), making the initial solution more in line with the actual metabolic characteristics of microorganisms and providing a reasonable starting point for subsequent correction.

[0125] The weighting coefficient calculation module M4 is used to calculate the corresponding weighting coefficients for each biochemical reaction of each microbial cell according to each Pearson residual and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction.

[0126] Further, the weighting coefficient calculation module M4 includes: a gene weight coefficient calculation unit and a reaction weighting coefficient calculation unit;

[0127] Among them, the gene weight coefficient calculation unit is used to calculate the corresponding first weight coefficient for each functional gene of each microbial cell according to a preset weight calculation algorithm and each Pearson residual; among them, the weight calculation algorithm is as follows:

[0128] Among them, Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, and W g,j represents the first weight coefficient corresponding to the functional gene g of the microbial cell j, K represents the magnification relationship between gene expression and metabolic flux, and sgn() represents the sign function.

[0129] The reaction weighting coefficient calculation unit is used to calculate the corresponding weighting coefficient for each biochemical reaction of each microbial cell according to each first weight coefficient and the co-regulated metabolic grid structure of each functional gene corresponding to each biochemical reaction.

[0130] In this preferred embodiment, the gene weight coefficient calculation unit calculates the weighting coefficient of the functional gene by using a piecewise function, dynamically adjusts the weight according to the absolute value of the Pearson residual, smoothly responds to gene expression changes through an exponential function when the residual is small, and enhances the sensitivity to significant differences through linear scaling combined with the sign function when the residual is large, making the weighting coefficient more conform to the non-linear relationship between gene expression and metabolic flux, and improving the flexibility of correction.

[0131] Further, the reaction weighting coefficient calculation unit includes: a path weighting coefficient calculation sub-unit and a reaction weight calculation sub-unit;

[0132] Among them, the path weighting coefficient calculation sub-unit is used to calculate the first weighting coefficient of each reaction path according to the co-regulated metabolic grid structure of each functional gene in each reaction path in each biochemical reaction and the first weight coefficient;

[0133] Among them, the calculation of the first weighting coefficient of each reaction path is specifically: if a reaction path requires multiple serially-connected functional genes to drive simultaneously, the minimum value among the first weight coefficients corresponding to each serially-connected functional gene is used as the first weighting coefficient of the corresponding reaction path; if a reaction path only requires any one of multiple parallel functional genes to drive independently, the average value of the first weight coefficients corresponding to each parallel functional gene is used as the first weighting coefficient of the corresponding reaction path;

[0134] The reaction weight calculation sub-unit is used to calculate the weighting coefficient corresponding to each biochemical reaction of each microbial cell according to each first weighting coefficient; among them, the calculation method of the weighting coefficient is as follows:

[0135] F r,j =∏p∈r W p,j ; where F r,j represents the weighting coefficient corresponding to the biochemical reaction r of cell j, and W p,j represents the first weighting coefficient corresponding to the reaction path p in the biochemical reaction r of cell j.

[0136] In this embodiment, the path weighting coefficient calculation subunit differentially processes the weighting coefficients for the co-regulation structures of tandem and parallel functional genes: taking the minimum weight for tandem genes to ensure the stability of the key path, and taking the average weight for parallel genes to balance the influence of redundant paths, avoiding the distortion of metabolic flux that may be caused by a single processing method, and making the corrected metabolic flux closer to the real biological scenario.

[0137] The correction calculation module M5 is used to correct the initial metabolic flux corresponding to each biochemical reaction of each microbial cell according to each of the weighting coefficients, and generate a metabolic flux vector for each microbial cell.

[0138] Furthermore, the correction calculation module M5 can be implemented by the following preferred implementation method, specifically: calculating the metabolic flux vector for each microbial cell according to the initial metabolic flux solution, each of the weighting coefficients, and the stoichiometric numbers corresponding to each biochemical reaction in the microbial genome metabolic model; where the calculation is as follows: F j = ∑ r∈M (S r ·V r ·F r,j ); where S r represents the stoichiometric numbers of each compound in the biochemical reaction r in the microbial genome metabolic model M, V r represents the initial metabolic flux corresponding to the biochemical reaction r in the initial metabolic flux solution, and F j represents the metabolic flux vector of cell j.

[0139] In this preferred embodiment, the correction calculation module M5 generates a metabolic flux vector by combining the initial metabolic flux solution, the weighting coefficients, and the stoichiometric numbers, fusing the global metabolic model with the single-cell specific expression data. The corrected vector not only retains the topological information of the genome metabolic network but also integrates the gene expression characteristics of individual cells, achieving an accurate characterization of the metabolic state under an open system.

[0140] The result output module M6 is used to calculate the similarity value between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis result of the microbial cell population.

[0141] Furthermore, the result output module M6 includes: a similarity value calculation unit and an analysis and output unit;

[0142] Among them, the similarity value calculation unit is configured to calculate the similarity values of the corresponding metabolic flux vectors and the initial metabolic flux solutions for each microbial cell according to a preset similarity value calculation algorithm;

[0143] The analysis and output unit is configured to output the cell state analysis result of the microbial cell population according to the data distribution of each similarity value.

[0144] In this preferred embodiment, the result output module M6 converts the complex metabolic state differences into quantifiable statistical indicators by calculating the similarity values between the metabolic flux vectors and the initial solutions, and at the same time outputs the analysis results based on the data distribution, making the evaluation of cell state diversity more intuitive and providing a clear statistical basis for experimental verification or subsequent analysis.

[0145] Further, the similarity value calculation unit can be implemented by the following preferred implementation manner, specifically: calculating the Pearson correlation coefficient between the corresponding metabolic flux vector and the initial metabolic flux solution for each microbial cell according to a preset Pearson correlation coefficient algorithm as the similarity value; wherein, the Pearson correlation coefficient algorithm is as follows:

[0146] Sim j represents the similarity value between the metabolic flux vector of cell j and the initial metabolic flux solution, C represents the total number of compounds c in the microbial genome metabolic model, F j,c represents the corrected flux of compound c in the metabolic flux vector of cell j, represents the average value of the corrected fluxes of all compounds in the metabolic flux vector of cell j, V c represents the flux of compound c in the initial metabolic flux solution V, represents the average value of the fluxes of all compounds in the initial metabolic flux solution V.

[0147] In this preferred embodiment, the similarity value calculation unit quantifies the similarity value using the Pearson correlation coefficient, and eliminates the influence of dimensions through the normalization processing of covariance and standard deviation, ensuring the comparability between the fluxes of different compounds and making the evaluation results of cell state diversity statistically significant and repeatable.

[0148] Further, the analysis and output unit includes: an average value calculation subunit and a characterization output subunit;

[0149] Among them, the average value calculation subunit is configured to calculate the corresponding average value according to each Pearson correlation coefficient;

[0150] The characterization output subunit is configured to characterize and output the cell state analysis result of the microbial cell population according to the average value.

[0151] In this preferred embodiment, the analysis and output unit simplifies the interpretation process of complex data by calculating the average value of the similarity values and outputting the analysis result accordingly. The smaller the average value, the greater the difference in the metabolic states between cells, which intuitively reflects the degree of heterogeneity of the microbial community and provides a direct basis for environmental response or function prediction.

[0152] In summary, compared with the prior art, the embodiments of the present application have the following beneficial effects: accurately obtaining the number of functional gene transcripts of each cell in the microbial population through single-cell sequencing technology, providing an original data basis for subsequent calculation of Pearson residuals and metabolic model correction; effectively eliminating the noise caused by differences in sequencing depth by calculating the Pearson residuals of microbial single-cell transcriptome data, improving the comparability of gene expression data; solving the initial metabolic flux solution based on the microbial genome metabolic model, providing a standardized benchmark for subsequent correction; calculating the weighting coefficient according to the Pearson residuals and the co-regulation metabolic grid structure, dynamically adjusting the metabolic flux, breaking through the limitation of traditional flux balance analysis relying on a closed system; generating an individualized metabolic flux vector by correcting the initial metabolic flux, accurately reflecting the differences in single-cell metabolic states; and finally analyzing the diversity of cell states through similarity values, providing a quantitative tool for the study of metabolic heterogeneity of complex microbial communities.

[0153] For the specific working processes of the above-described modules, reference may be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein. The division of the modules is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system.

[0154] The above specific embodiments have further elaborated the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for analyzing cell states based on Pearson residuals of single-cell transcriptomes, characterized in that Including: According to a single-cell sequencer, collect the transcript numbers of each functional gene of each microbial cell in a microbial cell population to obtain microbial single-cell transcriptome data; Calculate the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm; Calculate the initial metabolic flux solution in a preset microbial genome metabolic model; wherein, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome metabolic model; Calculate the corresponding weighting coefficients for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction; Correct the initial metabolic fluxes corresponding to each biochemical reaction of each microbial cell according to each of the weighting coefficients to generate the metabolic flux vectors of each microbial cell; Calculate the similarity values between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis results of the microbial cell population.

2. The cell state analysis method based on Pearson residuals of single-cell transcriptome according to claim 1, characterized in that, The calculating the Pearson residuals corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm includes: Calculate the corresponding Pearson residuals according to the transcript numbers of each functional gene of each microbial cell in the microbial single-cell transcriptome data; wherein, the calculation method is as follows: where Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, and x g,j represents the transcript count of the functional gene g of the microbial cell j, and μ g is the average transcript count of the functional gene g of all microbial cells in the sample; and σ g is the standard deviation of the transcript counts of the functional gene g of all microbial cells in the sample.

3. The method for analyzing cell state based on Pearson residuals of single-cell transcriptome according to claim 2, wherein The calculating the initial metabolic flux solution in a preset microbial genome metabolic model includes: Construct a homogeneous linear equation system according to the stoichiometric matrix in the microbial genome metabolic model; wherein, the homogeneous linear equation system is: S·V = 0; where S represents the stoichiometric matrix and V represents the initial metabolic flux solution; Solve the homogeneous linear equation system through a preset objective function and constraint conditions to obtain the initial metabolic flux solution.

4. A method for analyzing cell states based on Pearson residuals of single-cell transcriptomes according to claim 1, characterized in that, The calculating the corresponding weighting coefficients for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction includes: Calculate the corresponding first weighting coefficients for each functional gene of each microbial cell according to a preset weight calculation algorithm and each of the Pearson residuals; wherein, the weight calculation algorithm is as follows: Among them, Z g,j represents the Pearson residual corresponding to the functional gene g of the microbial cell j, W g,j represents the first weight coefficient corresponding to the functional gene g of the microbial cell j, K represents the magnification relationship between gene expression and metabolic flux, and sgn() represents the sign function; Calculate the corresponding weighting coefficients for each biochemical reaction of each microbial cell respectively according to each of the first weighting coefficients and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction.

5. The cell state analysis method based on Pearson residuals of single-cell transcriptome according to claim 4, characterized in that The calculating the corresponding weighting coefficients for each biochemical reaction of each microbial cell respectively according to each of the first weighting coefficients and the co-regulation metabolic grid structure of each functional gene corresponding to each biochemical reaction includes: Calculate the first weighting coefficients of each reaction path according to the co-regulation metabolic grid structure of each functional gene in each reaction path in each biochemical reaction and the first weighting coefficients; Among them, the specific method for calculating the first weighting coefficient of each reaction path is as follows: If a certain reaction path requires multiple functionally genes in series to drive simultaneously, then the minimum value among the first weight coefficients corresponding to each of the functionally genes in series is used as the first weighting coefficient of the corresponding reaction path; If a certain reaction path only requires any one of the multiple functionally genes in parallel to drive independently, then the average value of the first weight coefficients corresponding to each of the functionally genes in parallel is used as the first weighting coefficient of the corresponding reaction path; According to each of the first weighting coefficients, the weighting coefficients corresponding to each biochemical reaction of each microbial cell are calculated; Among them, the calculation method of the weighting coefficient is as follows: F r,j = ∏ p∈r W p,j ; where F r,j represents the weighted coefficient corresponding to the biochemical reaction r of cell j, and W p,j represents the first weighted coefficient corresponding to the reaction path p in the biochemical reaction r of cell j.

6. The method for analyzing cell states based on Pearson residuals of single-cell transcriptomes according to claim 5, wherein The step of correcting the initial metabolic flux corresponding to each biochemical reaction of each microbial cell according to each of the weighting coefficients to generate the metabolic flux vector of each microbial cell includes: According to the initial metabolic flux solution, each of the weighting coefficients, and the stoichiometric numbers corresponding to each biochemical reaction in the microbial genome metabolic model, the metabolic flux vector of each microbial cell is calculated; Among them, the calculation is as follows: F j = ∑ r∈M (S r · V r · F r,j ); where S r represents the stoichiometric coefficient of each compound in the biochemical reaction r in the microbial genome metabolic model M, V r represents the initial metabolic flux corresponding to the biochemical reaction r in the initial metabolic flux solution, and F j represents the metabolic flux vector of cell j.

7. A method for analyzing cell states based on Pearson residuals of single-cell transcriptomes according to any one of claims 1-6, characterized in that, The step of calculating the similarity value between each of the metabolic flux vectors and the initial metabolic flux solution and outputting the cell state analysis result of the microbial cell population includes: According to a preset similarity value calculation algorithm, the similarity value between the corresponding metabolic flux vector and the initial metabolic flux solution is calculated for each microbial cell; According to the data distribution of each of the similarity values, the cell state analysis result of the microbial cell population is output.

8. The cell state analysis method based on Pearson residuals of single-cell transcriptome according to claim 7, characterized in that The step of calculating the similarity value between the corresponding metabolic flux vector and the initial metabolic flux solution for each microbial cell according to a preset similarity value calculation algorithm includes: According to a preset Pearson correlation coefficient algorithm, the Pearson correlation coefficient between the corresponding metabolic flux vector and the initial metabolic flux solution is calculated for each microbial cell as the similarity value; Among them, the Pearson correlation coefficient algorithm is as follows: Sim j represents the similarity value between the metabolic flux vector of cell j and the initial metabolic flux solution, C represents the total number of compounds c in the microbial genome metabolic model, F j,c represents the corrected flux of compound c in the metabolic flux vector of cell j, represents the average value of the corrected fluxes of all compounds in the metabolic flux vector of cell j, V c represents the flux of compound c in the initial metabolic flux solution V, represents the average value of the fluxes of all compounds in the initial metabolic flux solution V.

9. A method for analyzing cell states based on Pearson residuals of single-cell transcriptomes according to claim 8, characterized in that, The step of outputting the cell state analysis result of the microbial cell population according to the data distribution of each of the similarity values includes: According to each of the Pearson correlation coefficients, the corresponding average value is calculated; According to the average value, the cell state analysis result of the microbial cell population is characterized and output.

10. A cell state analysis system based on Pearson residuals of single-cell transcriptomes, characterized in that, It includes: A data acquisition module, a residual calculation module, an initial metabolic flux solution calculation module, a weighting coefficient calculation module, a correction calculation module, and a result output module; Among them, the data acquisition module is used to collect the transcript numbers of each functional gene of each microbial cell in the microbial cell population according to a single-cell sequencer to obtain microbial single-cell transcriptome data; The residual calculation module is used to calculate the Pearson residual corresponding to each functional gene of each microbial cell according to the microbial single-cell transcriptome data and a preset Pearson residual algorithm; The initial metabolic flux solution calculation module is used to calculate the initial metabolic flux solution in a preset microbial genome metabolic model; Among them, the initial metabolic flux solution includes the initial metabolic fluxes corresponding to all biochemical reactions in the microbial genome metabolic model; The weighted coefficient calculation module is used to calculate corresponding weighted coefficients for each biochemical reaction of each microbial cell according to each of the Pearson residuals and the co-regulatory metabolic grid structure of each functional gene corresponding to each biochemical reaction; The correction calculation module is used to correct the initial metabolic fluxes corresponding to each biochemical reaction of each microbial cell according to each of the weighted coefficients, and generate a metabolic flux vector for each microbial cell; The result output module is used to calculate the similarity value between each of the metabolic flux vectors and the initial metabolic flux solution, and output the cell state analysis result of the microbial cell population.