Single cell big data analysis system and method
By integrating multiple modules and automated processes in single-cell big data analysis system, the problem of algorithms and tool docking in single-cell data analysis is solved, and efficient and automated data processing and analysis is achieved.
Patent Information
- Application Number
- CN202510129665.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-05
AI Technical Summary
The prior art is difficult to achieve seamless connection between algorithms and tools in single-cell big data analysis, resulting in difficulty in ensuring analysis efficiency and accuracy.
A single-cell big data analysis system is proposed, including high-throughput sequencing data acquisition module, single-cell public data module, batch correction module, intelligent annotation module, huge statistics module, personalized analysis module and visual output module. Through the combination and automated process of these modules, seamless processing and analysis of data can be achieved.
It has realized the automation and efficiency of single-cell big data analysis, and can quickly process 500,000 single-cell data, significantly improving analysis efficiency and reducing R&D investment.
Smart Images

Figure CN120089197A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of high-throughput sequencing data analysis intelligent technology, specifically, a system and method for single-cell big data analysis. Background Art
[0002] Single-cell RNA sequencing (scRNA-seq) can independently detect each cell in the same sample, such as up to 10,000 cells, and generate massive, high-dimensional data sets. According to literature reports, there are hundreds of single-cell data algorithms and processes to process these data. However, how to select a suitable algorithm has become one of the most critical challenges in the current single-cell data analysis process. In addition, single-cell data analysis is highly dependent on programming environments such as R or Python. Although these two languages have their own advantages in statistical computing and machine learning, data exchange and integration between them is not easy.
[0003] Therefore, ensuring seamless integration between different algorithms and analysis tools while maintaining efficiency and accuracy, and achieving comprehensive analysis of single-cell data has become a technical problem that needs to be urgently solved in this field. Summary of the invention
[0004] The purpose of the present invention is to provide a single cell big data analysis system. System and method for single cell big data analysis
[0005] Another object of the present invention is to provide a single-cell big data analysis method.
[0006] To achieve the above purpose, the technical solution of the present invention is:
[0007] A single-cell big data analysis system, including a high-throughput sequencing data acquisition module, a single-cell public data module, a batch correction module, an intelligent annotation module, a massive statistics module, a personalized analysis module, and a visualization output module, characterized in that:
[0008] The high-throughput sequencing data acquisition module downloads the reference genome FASTA file of the required species from the Ensembl, UCSC or NCBI public database;
[0009] The single-cell public data module includes public single-cell data of dozens of species such as humans, mice, pigs, rhesus monkeys, elks, fruit flies, silkworms, zebrafish, salamanders, rice, and Arabidopsis;
[0010] The batch correction module includes algorithmic processes developed based on MNN, graph-based, and Deep-learn, as well as cross-species analysis processes;
[0011] The intelligent annotation module includes constructing hundreds of tissue-specific intelligent annotation models using the AysenseBio2.0 annotation system;
[0012] The massive statistics module includes annotation statistics plotting, inter-group difference plotting, and complex analysis plotting of annotation X grouping, GSEA, and GSVA analysis;
[0013] The personalized analysis module includes automatically completing dozens of analyses such as pseudotime, cell-cell communication, deconvolution, and transcription factor analysis;
[0014] The visualization output module includes the output of various interactive analysis graphs such as clustering graphs, bubble graphs, heat maps, and communication difference graphs.
[0015] The single-cell big data analysis system of the present invention can also be further implemented by the following technical measures.
[0016] For the aforementioned single-cell big data analysis system, the high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visualization output module are connected in sequence, and the output data of the previous module matches the input data of the next module.
[0017] For the aforementioned single-cell big data analysis system, the processes of the high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visualization output module can run automatically.
[0018] For the aforementioned high-throughput sequencing data acquisition module, it also includes converting BCL files to FASTQ format files using the mkfastq command, performing sequence alignment of the FASTQ format files with the corresponding reference genome, and converting them into count matrix files.
[0019] A single-cell big data analysis method, characterized by including the following steps:
[0020] a. The high-throughput sequencing data acquisition module includes the following steps:
[0021] (1) Download the reference genome FASTA file of the required species from the Ensembl, UCSC, or NCBI public database, cooperate with the corresponding GTF file, and use the mkref command to create a new reference genome database;
[0022] (2) Use the mkfastq command to convert BCL files into FASTQ format files, and perform data quality control and data filtering on low-quality reads;
[0023] (3) Align the FASTQ format files with the corresponding reference genome; use the CellRanger alignment software to read the FASTQ data in each sample directory, align each reads sequence to the reference genome database, and decode the cell barcode recognition code and the unique feature recognition sequence;
[0024] (4) Convert to a count matrix file; after obtaining the decoded cell barcode recognition code, count the decoded cell barcode recognition code and the unique feature recognition sequence, and construct a gene-cell expression matrix;
[0025] b. Single-cell public database module, including public single-cell data of dozens of species such as human, mouse, pig, rhesus monkey, elk, Drosophila, silkworm, zebrafish, salamander, rice, Arabidopsis thaliana, etc. The public single-cell data is uniformly processed into standard rds format data and undergoes standardized cell annotation. The processing method includes the following steps;
[0026] Download and organize various matrix data from GEO and Single Cell Portal databases, and through a unified normalization and correction algorithm, sum the UMI numbers assigned to specific cells as the total sequencing depth cell attribute, and use this cell attribute in a negative binomial error distribution and logarithmic link function regression model to construct a regression model to correct the sequencing depth differences between cells and standardize the data:
[0027] log(E(xi)) = β0 + β1log10m;
[0028] Where: xi is the UMI count vector of gene i, m is the molecular vector assigned to the cell, that is, mj = i xi j. The solution to the regression model problem is a set of parameters: intercept β0 and slope β1. The dispersion parameter θ of the underlying distribution is unknown, and this parameter is estimated by fitting a negative binomial regression model;
[0029] c. Batch correction module, including algorithmic processes developed based on MNN, graph-based, and Deep-learn, as well as cross-species analysis processes:
[0030] (1) Calculate the similarity matrix between all cells within each batch. Let X be an n×p matrix, where each row represents the expression profile data of a cell. The similarity matrix S is defined as:
[0031] S = exp(-l[z;-rllP / 202)
[0032] Where: σ is a parameter that controls the rate of distance decay;
[0033] (2) By calculating the similarity matrix between all cells within each batch, determine the most similar set of anchor points in each batch. Let Sb represent the similarity matrix between the cells in the b-th batch, and define a cost function C:
[0034]
[0035] Where: A is a set consisting of |A| anchor points, Ab represents the set of anchor points in the b-th batch, and λ is a regularization coefficient used to penalize the case of too many anchor points. The goal is to minimize the cost function C;
[0036] Map each non-anchor cell to the low-dimensional space where its nearest anchor point is located. Let x represent the expression profile data of the i-th cell, and ak represent the expression profile of the k-th anchor point;
[0037] d. Intelligent annotation module, including using the AysenseBio2.0 annotation system to build hundreds of tissue-specific intelligent annotation models, including the following steps:
[0038] (1) Input an scRNA-seq dataset containing known cell types as the training set. Each cell sample has an expression profile, that is, the expression level of each gene;
[0039] (2) Select discriminative genes from the training data as features. These genes have significant expression differences between different cell types;
[0040] (3) Suppose there are n cell types, or labels. For each label i, train a regression model, which is expressed as:
[0041]
[0042] Where: y i is the binary output of the i-th label, which is 0 or 1;
[0043] x is the input feature vector, which is the gene expression level;
[0044] w i is the weight vector related to the i-th class;
[0045] b i is the bias term, the intercept;
[0046] is the logistic function used to map the result of the linear combination to the probability interval of 0 to 1;
[0047] (4) After establishing the regression model, for a new input matrix x, use the regression model to calculate a probability value for each cell label, return the probability distribution of all labels, and input relevant parameters again to complete statistical annotation;
[0048] e. Mass statistics module, including annotation statistics plotting, between-group difference plotting, and complex analysis plotting of annotation X grouping, GSEA, and GSVA analysis, including the following steps:
[0049] (1) After annotating all cell type subgroups, continue to statistically analyze the differences in all cell groups among multiple groups, including using t-tests, multi-group variance analysis, and the RO / E algorithm for between-group difference comparison;
[0050] (2) For each between-group difference, calculate gene set variation analysis. Let f(x) represent the estimated kernel density of the gene expression value x in a certain gene set in the sample. The formula for calculating the positive enrichment score ES+ is:
[0051] where: F(x) is the cumulative distribution function, s represents a specific sample. For the negative enrichment score ES-, it is the low-expression region, and the enrichment score ES is the difference between ES+ and ES-;
[0052] (3) For each annotation result and each group, form an expression matrix again and analyze:
[0053] where: xi and yi are two variables respectively;
[0054] and represent the expression means of the disease group and the control group;
[0055] ∑ represents the summation symbol, and sums over all sample points;
[0056] f. Personalized analysis module, including automatically completing dozens of analyses such as pseudotime, cell-cell communication, deconvolution, and transcription factor analysis;
[0057] Personalized statistics module, which builds a workflow to automatically complete grouped pseudotime, between-group communication difference analysis, deconvolution of Bulk-seq, and subgroup-specific transcription factor analysis, and at the same time performs parallel processing to greatly optimize the analysis process;
[0058] g. Visual output module, including the output of various interactive analysis graphs such as clustering graphs, bubble graphs, heat maps, and communication difference graphs. Group the data through the hierarchical clustering algorithm, calculate the differential genes of specific groups, draw bubble graphs and heat maps, select specific subpopulations for cell - cell communication analysis, calculate the differences in communication data according to the groups, and finally obtain communication difference pictures to complete the analysis for single - cell big data.
[0059] The single - cell big data analysis method of the present invention can also be further implemented by adopting the following technical measures.
[0060] The foregoing method, wherein the parameter value for controlling the distance attenuation rate is 0.1 or 0.01.
[0061] The foregoing method, wherein the intelligent annotation module outputs dozens of annotation results at one time.
[0062] The foregoing method, wherein the massive statistics module calculates the between - group differences for each output result.
[0063] The foregoing method, wherein the two variables are the observed values of the disease group and the control group.
[0064] After adopting the above - mentioned technical solutions, the single - cell big data analysis system and method of the present invention have the following advantages:
[0065] 1. It includes numerous public databases, uses the latest algorithms to complete clustering annotation work, gives multiple accurate annotation results at the same time, and performs massive statistics on each intelligent annotation result, turning the complex single - cell data analysis results into multiple - choice questions, greatly optimizing the analysis content and improving the analysis efficiency;
[0066] 2. It has established a domestic single - cell data analysis platform with visualization and process flow, realizes the drawing and selection of intelligent charts, realizes the re - mining of research results at the CNS level, helps scientific research personnel reduce R & D investment, and realizes the efficient and reasonable utilization of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 It is the final clustering schematic diagram formed by batch correction, finding the best anchor points for the embodiment of the present invention;
[0068] Figure 2 It is the schematic diagram of annotation details obtained by the intelligent annotation system of the embodiment of the present invention for annotating cell subpopulations;
[0069] Figure 3 It is the ratio statistical data and functional clustering data between two groups for the embodiment of the present invention, including the schematic diagram of between - group difference comparison using algorithms such as T - test, multi - group analysis of variance, and RO / E, etc.;
[0070] Figure 4 Schematic diagram of the function analysis of GSVA according to an embodiment of the present invention;
[0071] Figure 5 Schematic diagrams of dozens of analyses such as pseudotime, cell-cell communication, deconvolution, and transcription factor analysis automatically completed in the personalized analysis module according to an embodiment of the present invention;
[0072] Figure 6 Schematic diagram of the output of various interactive analysis graphs such as clustering graphs, bubble graphs, heat maps, and communication difference graphs according to an embodiment of the present invention;
[0073] Figure 7 It is a flowchart of the process for single-cell big data analysis of the present invention. Detailed implementation manners
[0074] The present invention will be further described below in conjunction with embodiments and their accompanying drawings.
[0075] Embodiment 1
[0076] Now please refer to Figure 7 , Figure 7 It is a flowchart of the single-cell big data analysis system of the present invention. The single-cell big data analysis system of the present invention includes a high-throughput sequencing data acquisition module, a single-cell public data module, a batch correction module, an intelligent annotation module, a massive statistics module, a personalized analysis module, and a visualization output module.
[0077] The high-throughput sequencing data acquisition module downloads the reference genome FASTA file of the required species from public databases such as Ensembl, UCSC, or NCBI;
[0078] The single-cell public data module includes public single-cell data of dozens of species such as humans, mice, pigs, rhesus monkeys, elk, fruit flies, silkworms, zebrafish, salamanders, rice, and Arabidopsis thaliana;
[0079] The batch correction module includes algorithmic processes developed based on MNN, graph-based, and Deep-learn, as well as cross-species analysis processes;
[0080] The intelligent annotation module includes using the AysenseBio2.0 annotation system to construct hundreds of tissue-specific intelligent annotation models;
[0081] The massive statistics module includes annotation statistics mapping, inter-group difference mapping, and complex analysis mapping of annotation X grouping, GSEA, and GSVA analysis;
[0082] The personalized analysis module includes automatically completing dozens of analyses such as pseudotime, cell-cell communication, deconvolution, and transcription factor analysis;
[0083] The visualization output module includes the output of various interactive analysis graphs such as clustering graphs, bubble graphs, heat maps, and communication difference graphs.
[0084] The high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visualization output module are connected in sequence, and the output data of the previous module matches the input data of the next module.
[0085] The processes of the high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visualization output module can run automatically.
[0086] The high-throughput sequencing data acquisition module also includes converting BCL files into FASTQ format files using the mkfastq command, performing sequence alignment of the FASTQ format files with the corresponding reference genome, and converting them into count matrix files.
[0087] Example 2
[0088] Now please refer to Figures 1-6 , the single-cell big data analysis method of the present invention includes the following steps:
[0089] a. The high-throughput sequencing data acquisition module, including constructing a reference genome database, data quality control, data filtering, sequence alignment, quantitative analysis, and matrix generation, includes the following steps:
[0090] (1) Download the reference genome FASTA file of the required species and the corresponding GTF (Gene Transfer Format) file from the Ensembl, UCSC, or NCBI public database, and use the mkref command to create a new reference genome database;
[0091] (2) Use the mkfastq command to convert BCL files into FASTQ files, and perform data quality control and data filtering on low-quality reads;
[0092] (3) Perform sequence alignment of the FASTQ format files with the corresponding reference genome;
[0093] (4) Convert it into a count-based matrix file; Figure 1 For the embodiment of the present invention, after batch correction, the best anchor points are found to form the final clustering schematic diagram.
[0094] b. Single-cell public database module, including public single-cell data of dozens of species such as human, mouse, rat, pig, rhesus monkey, elk, Drosophila, silkworm, zebrafish, salamander, rice, Arabidopsis thaliana, etc. Since the data formats in the public literature are inconsistent, it is necessary to uniformly process the data, all processed into standard rds format data, and perform standardized cell annotation. The method includes the following steps;
[0095] Download and organize various matrix data from databases such as GEO and Single Cell Portal, and use a unified normalization and correction algorithm, specifically:
[0096] Sum the UMI (unique molecular identifiers) numbers assigned to specific cells as the total sequencing depth attribute, and use this cell attribute in the negative binomial error distribution and logarithmic link function regression model to construct a regression model to correct the sequencing depth differences between cells and standardize the data:
[0097] log(E(xi)) = β0 + β1log10m
[0098] where xi is the UMI count vector of gene i, and m is the molecular vector assigned to the cell, that is, mj = i xij. The solution to this regression problem is a set of parameters: intercept β0 and slope β1. The dispersion parameter θ of the underlying NB distribution is also unknown and needs to be estimated from the data, using the NB parameterization form with mean μ and variance μ + μ2θ;
[0099] c. Batch correction module includes algorithm processes developed based on MNN (mutual nearest neighbor method), graph-based, Deep-learn (deep learning method), etc., as well as cross-species analysis processes.
[0100] 1. Calculate the similarity matrix between all cells within each batch. Assume X is an n×p matrix, where each row represents the expression profile data of a cell. The similarity matrix S can be defined as:
[0101] S ij = exp(-||x i - x j || 2 / 2σ 2 )
[0102] where σ is a parameter that controls the rate of distance decay, usually taking a relatively small value, such as 0.1 or 0.01,
[0103] 2. Next, it is necessary to determine the most similar set of anchor points in each batch. This can be accomplished by calculating the similarity matrix between all cells within each batch. Let \(S_b\) represent the similarity matrix between the cells in the \(b\)-th batch, then a cost function \(C\) can be defined as follows:
[0104] where \(A\) is a set consisting of \(|A|\) anchor points, \(A_b\) represents the set of anchor points in the \(b\)-th batch, and \(\lambda\) is a regularization coefficient used to penalize the case of too many anchor points. The goal is to minimize the cost function \(C\).
[0105] Map each non-anchor cell to the low-dimensional space where its nearest anchor point is located. Specifically, let \(x_i\) represent the expression profile data of the \(i\)-th cell, and \(a_k\) represent the expression profile of the \(k\)-th anchor point;
[0106] d. The intelligent annotation module includes hundreds of tissue-specific intelligent annotation models constructed and the self-developed AysenseBio2.0 annotation system. Figure 2 This is a schematic diagram of the annotation details obtained by the annotation system of the embodiment of the present invention for annotating cell subsets.
[0107] 4. Input an scRNA-seq dataset containing known cell types as the training set. Each cell (sample) has its expression profile, that is, the expression level of each gene.
[0108] 5. Select discriminative genes from the training data as features. These genes have significant expression differences between different cell types
[0109] 6. Suppose there are \(n\) cell types (or labels), then for each label \(i\), train a regression model. This model can be expressed as:
[0110]
[0111] · \(y\) i is the binary output (0 or 1) of the \(i\)-th class label
[0112] · \(x\) is the input feature vector (for example, gene expression level).
[0113] · \(w\) i is the weight vector associated with the \(i\)-th class.
[0114] · \(b\) i is the bias term (intercept)
[0115] · is the logistic function (sigmoid function) used to map the result of the linear combination to the probability interval of (0, 1).
[0116] 4. After establishing the model, for a new input matrix x, use the model to calculate a probability value for each cell label, return the probability distribution of all labels, and input the relevant parameters again to complete the statistical annotation.
[0117] e. The massive statistics module includes annotation statistics plotting, between-group difference plotting, and complex analysis plotting of annotation X grouping, GSEA, and GSVA analysis.
[0118] 1. After obtaining the annotation results of a subgroup or all subgroups, continue to perform various groupings of the statistical data, including using algorithms such as T-test, one-way ANOVA, and RO / E for between-group difference comparison. Particularly importantly, the annotation module outputs dozens of annotation results at once, and the massive statistics module calculates the between-group differences for each output result.
[0119] 2. For each between-group difference, Gene Set Variation Analysis will be calculated. If f(x) represents the kernel density estimate of the gene expression values x within a certain gene set in a sample, then for the positive enrichment score ES+, the calculation method is as follows:
[0120]
[0121] where F(x) is the cumulative distribution function and s represents a specific sample. For the negative enrichment score ES-, the low-expression region will be considered. The final enrichment score ES may be the difference between ES+ and ES-.
[0122] 3. For each annotation result and each grouping, an expression matrix will be formed again and relevant analyses will be performed:
[0123] · Among them, xi and yi are two variables, such as the observed values of the disease group and the control group,
[0124] · and represent the expression means of the disease group and the control group
[0125] · ∑ represents the summation symbol, meaning summing over all sample points.
[0126] f. The personalized analysis module includes automatically completing dozens of analyses such as pseudotime ordering, cell-cell communication, deconvolution, and transcription factor analysis.
[0127] The personalized statistics module has currently built a workflow that can automatically complete grouped pseudotime ordering, between-group communication difference analysis, deconvolution of Bulk-seq, subgroup-specific transcription factor analysis, etc., and at the same time consider parallel processing to greatly optimize the analysis process.
[0128] g. Visual output module, including the output of various interactive analysis graphs such as clustering graphs, bubble graphs, heatmaps, and communication difference graphs. Group the data through the hierarchical clustering algorithm, calculate the differential genes of specific groups, and draw bubble graphs and heatmaps. Select specific subpopulations again for cell-cell communication analysis, calculate the differences in communication data according to the groups, and finally obtain communication difference pictures.
[0129] The present invention has substantial features and significant technological progress. The single-cell big data analysis system and method of the present invention encompass numerous public databases, use the latest algorithms to complete the clustering annotation work, and at the same time give multiple accurate and intelligent annotation results. For each annotation result, a huge amount of statistics are completed. These steps turn the complex single-cell data analysis results into multiple-choice questions, greatly optimizing the analysis content and improving the analysis efficiency. And a domestic single-cell data analysis platform with visualization and process flow is established, realizing the drawing and selection of intelligent charts, and realizing the re-mining of research results at the CNS level. Through these efforts, it assists scientific researchers in reducing R & D investment and realizing the efficient and reasonable utilization of data.
[0130] The single-cell big data analysis system and method of the present invention include a high-throughput sequencing data acquisition module, a matrix generation module, a batch correction module, a single-cell public data module, an intelligent annotation module, a huge amount of statistics module, a personalized analysis module, and a visual output module. The above modules are connected in the analysis order, and the output data of the previous module matches the input data of the next module. The processes of multiple sets of modules can run automatically. After testing, the data analysis of 500,000 single cells can be completed in only 2 days. The present invention uses a modular operation scheme to establish a domestic single-cell data analysis system with visualization and process flow, greatly reducing the difficulty of single-cell data analysis and improving the analysis efficiency of researchers.
[0131] The above embodiments are only for illustrating the present invention, rather than limiting the present invention. Those skilled in the relevant technical fields can also make various transformations or changes without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the present invention and should be defined by each claim.
Claims
1. A single-cell big data analysis system, comprising a high-throughput sequencing data acquisition module, a single-cell public data module, a batch correction module, an intelligent annotation module, a massive statistics module, a personalized analysis module, and a visualization output module, characterized in that: The high-throughput sequencing data acquisition module downloads the reference genome FASTA file of the required species from Ens embl, UCSC or NCBI public database; The single-cell public data module includes public single-cell data of dozens of species including humans, mice, pigs, rhesus monkeys, elks, fruit flies, silkworms, zebrafish, salamanders, rice, and Arabidopsis; The batch correction module includes an algorithm flow developed based on MNN, graph-based, and Deep-learn, as well as a cross-species analysis flow; The intelligent annotation module includes using the AysenseBio2.0 annotation system to construct hundreds of tissue-specific intelligent annotation models; The massive statistics module includes annotation statistical mapping, group difference mapping, and annotation X grouping complex analysis mapping, GSEA, and GSVA analysis; The personalized analysis module includes dozens of analyses such as pseudo-time series, cell-to-cell communication, deconvolution, and transcription factor analysis that are automatically completed; The visualization output module includes the output of multiple interactive analysis diagrams such as cluster diagrams, bubble diagrams, heat maps, and communication difference diagrams.
2. The single cell big data analysis system according to claim 1, characterized in that: The high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visualization output module are connected in order, and the output data of the previous module matches the input data of the next module.
3. The single cell big data analysis system according to claim 1, characterized in that: The high-throughput sequencing data acquisition module, matrix generation module, batch correction module, single-cell public data module, intelligent annotation module, massive statistics module, personalized analysis module, and visual output module, the processes of each module can run automatically.
4. The single cell big data analysis system according to claim 1, characterized in that: The high-throughput sequencing data acquisition module also includes converting the BCL file into a FASTQ format file using the mkfastq command, performing sequence alignment on the FASTQ format file and the corresponding reference genome, and converting the FASTQ format file into a count matrix file.
5. A single cell big data analysis method as claimed in claim 1, characterized in that The following steps are involved: a. High-throughput sequencing data acquisition module, including the following steps: (1) Download the reference genome FASTA file of the required species from the Ensembl, UCSC or NCBI public database, match it with the corresponding GTF file, and use the mkref command to create a new reference genome database; (2) Use the mkfas tq command to convert the BCL file into a FASTQ format file and perform data quality control and data filtering on low-quality reads; (3) Align the FASTQ format file with the corresponding reference genome; use CellRanger alignment software to read the FASTQ data in each sample directory, align each read sequence to the reference genome database, and decode the cell barcode and unique feature identification sequence; (4) converting to a count matrix file; after obtaining the decoded cell barcode, counting the decoded cell barcode and the unique feature identification sequence to construct a gene-cell expression matrix; b. Single-cell public database module, including public single-cell data of dozens of species such as human, mouse, pig, rhesus monkey, elk, fruit fly, silkworm, zebrafish, salamander, rice, and Arabidopsis. The public single-cell data are uniformly processed into standard RDS format data and standardized cell annotation is performed. The processing method includes the following steps; Download and organize various matrix data from GEO and Single Cell Portal databases, and use a unified normalization correction algorithm to sum up all UMI numbers assigned to specific cells as the total sequencing depth cell attribute. Use the cell attribute in the negative binomial error distribution and log link function regression model to build a regression model, correct for differences in sequencing depth between cells, and standardize the data: log(E(xi))=β0+β1log 10m; Where: xi is the UMI count vector of gene i, m is the molecule vector assigned to the cell, i.e. mj=i xi j, the solution to the regression model problem is a set of parameters: intercept β0 and slope β1, the dispersion parameter θ of the underlying distribution is unknown, and this parameter is estimated by fitting a negative binomial regression model; c. Batch correction module, including algorithmic processes developed based on MNN, graph-based, Deep-learn, and cross-species analysis processes: (1) Calculate the similarity matrix between all cells in each batch. Let X be an n×p matrix, where each row represents the expression profile data of a cell. The similarity matrix S is defined as: S ij =exp(-||x i -x j || 2 / 2σ 2 ) Among them: σ is a parameter that controls the distance decay rate; (2) By calculating the similarity matrix between all cells in each batch, determine the most similar set of anchor points in each batch. Let Sb represent the similarity matrix between cells in the bth batch and define a cost function C: Where: A is a set of |A| anchor points, Ab represents the anchor point set of the bth batch, λ is a regularization coefficient used to penalize the situation where there are too many anchor points, and the goal is to minimize the cost function C; Map each non-anchor cell to the low-dimensional space of its nearest anchor point. Let x represent the expression profile data of the i-th cell and ak represent the expression profile of the k-th anchor point. d. Intelligent annotation module, including the use of AysenseBio2.0 annotation system to build hundreds of tissue-specific intelligent annotation models, including the following steps: (1) Input a scRNA-seq dataset containing known cell types as a training set. Each cell sample has an expression profile, that is, the expression level of each gene; (2) selecting discriminative genes from the training data as features, wherein the genes have significant expression differences between different cell types; (3) Assume n cell types, or labels, and for each label i, train a regression model, which is expressed as: Where: y i is the binary output of the i-th label, which is 0 or 1; x is the input feature vector, which is the gene expression level; w i is the weight vector associated with the i-th class; b i is the bias term, intercept; It is a logistic function used to map the result of linear combination to the probability interval of 0 and 1; (4) After the regression model is established, for a new input matrix x, a probability value is calculated for each cell label using the regression model, and the probability distribution of all labels is returned, and the relevant parameters are input again to complete the statistical annotation; e. Massive statistics module, including annotation statistical plotting, group difference plotting, and complex analysis plotting of annotation X grouping, GSEA, GSVA analysis, including the following steps: (1) After annotating all cell type subpopulations, continue to count the difference information of all cell types in multiple groups, including using T test, multi-group analysis of variance, and RO / E algorithm to compare the differences between groups; (2) For each intergroup difference, gene set variation analysis was calculated. Let f(x) represent the estimated kernel density of gene expression value x in a gene set in the sample. The positive enrichment score ES+ was calculated using the formula: Where: F(x) is the cumulative distribution function, s represents a specific sample, for the negative enrichment score ES-, it is a low expression region, and the enrichment score ES is the difference between ES+ and ES-; (3) For each annotation result and each grouping, an expression matrix is formed again and analyzed: Among them: xi and yi are two variables; and represents the disease group and the mean expression value of the control group; ∑ represents the summation symbol, and all sample points are summed; f. Personalized analysis module, including automatic completion of dozens of analyses such as pseudo-time series, cell-to-cell communication, deconvolution, and transcription factor analysis; The personalized statistics module builds a workflow to automatically complete group pseudo-time series, inter-group communication difference analysis, bulk-seq deconvolution, and subgroup-specific transcription factor analysis, while performing parallel processing to greatly optimize the analysis process; g. Visual output module, including the output of various interactive analysis diagrams such as cluster diagrams, bubble diagrams, heat maps, and communication difference diagrams. The data are grouped through a hierarchical clustering algorithm, and the differential genes of specific groups are calculated. Bubble diagrams and heat maps are drawn, and specific subpopulations are selected for intercellular communication analysis. The communication data are differentially calculated according to the groups, and finally the communication difference pictures are obtained for single-cell big data analysis.
6. The single cell big data analysis method according to claim 5, characterized in that: The parameter value for controlling the distance attenuation speed is 0.1 or 0.
01.
7. The single cell big data analysis method according to claim 5, characterized in that: The intelligent annotation module outputs dozens of annotation results at one time.
8. The single cell big data analysis method according to claim 5, characterized in that: The massive statistics module calculates the inter-group differences for each output result.
9. The single cell big data analysis method according to claim 5, characterized in that: The two variables are the observed values of the disease group and the control group.
Citation Information
Patent Citations
Analysis method of spatial transcriptome sequencing data
CN112522371A
Construction method of modularized single cell rapid analysis system
CN113963747A
Single cell cross-species cell type identification method
CN115064220A
Single cell sequencing code-free analysis system
CN118116467A
Scrnaseq analysis systems
US20230420078A1