Method, device, instrument and storage medium for detecting cellular composition of biological samples

Single-cell transcriptome sequencing and bioinformatics analysis address the limitations of current cell type identification methods, providing low-cost, high-throughput, and accurate detection of cellular composition with reduced false positives/negatives.

JP2025528861APending Publication Date: 2025-09-02ZHEJIANG HUODE BIOENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025508994
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-16
Filing Date
2023-08-15
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Current cell type identification methods are limited in analyzing unknown samples, have high costs, low throughput, and are prone to false positives/negatives due to reliance on specific molecular markers, making high-throughput cell identification and quantification difficult.

Method used

A method involving single-cell transcriptome sequencing, generating a cellular gene expression matrix, and performing single-cell bioinformatics analysis to determine cell types, with optional steps for annotation, data normalization, and trajectory analysis to enhance accuracy.

Benefits of technology

Enables one-step, low-cost, high-throughput detection of cellular composition with improved accuracy by exploring cellular characteristics and reducing false judgments, allowing for data accumulation and iterative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528861000001_ABST
    Figure 2025528861000001_ABST
Patent Text Reader

Abstract

This application discloses a method, device, electronic device, and readable storage medium for detecting the cellular composition of a biological sample, which are applicable to the biomedical technology field. The method includes: performing single-cell transcriptome sequencing on a test biological sample to obtain single-cell sequencing results; analyzing the single-cell sequencing results to generate a cellular gene expression matrix; and performing single-cell bioinformatics analysis on the cellular gene expression matrix to determine the cell types contained in the test biological sample. This application enables the detection of the cellular composition of a biological sample in one step with low cost, high throughput, and accurate quantification.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of biomedical technology, and in particular to methods, devices, electronic devices and readable storage media for detecting cellular constituents of biological samples. [Background technology]

[0002] With the development of regenerative medicine technologies, stem cells hold promising prospects for applications in areas such as the construction of organoid disease models, research into the mechanisms of new drugs, and the treatment of diseases for which no specific cure is available and cell transplantation is required. Stem cells are a group of cells with the ability to self-renew and differentiate. Among them, embryonic stem cells and induced pluripotent stem cells are collectively referred to as pluripotent stem cells because they have the potential to differentiate into any cell type in the body. Cell products derived from stem cells used in these fields require highly sensitive and accurate cell type identification to ensure that they are primarily composed of the expected differentiated cells, that the cell composition is stable and controllable, and that they have batch-to-batch uniformity and overall functional reproducibility. This ensures that cells do not deviate from the expected direction of differentiation and that unintended differentiated cells that are harmful to the human body are not generated, thereby ensuring the safety and efficacy of drugs, particularly in disease treatment.

[0003] Currently, commonly used techniques for cell type identification include quantitative real-time PCR (qPCR), digital PCR (dPCR), fluorescence in situ hybridization (FISH), immunofluorescence (IF), and flow cytometry (FC). Among these, qPCR and dPCR identify cell types based on the expression levels of functional cell-specific gene target amplification assays. FISH identifies cell types by hybridizing fluorescently labeled specific nucleic acid probes with corresponding target DNA or RNA molecules in cells and observing the fluorescent signal with a fluorescence microscope or confocal laser scanner. Both IF and FC qualitatively or quantitatively detect cell types by identifying cell-specific antigenic substances (molecular markers) through antigen-antibody reactions using fluorescently labeled specific antibodies. All of the above methods require the design, manufacture, or procurement of specific primers or antibodies based on specific genes or molecular markers of known cell types, and can only detect known cell types. However, if the cell type cannot be predicted in advance or if there are no highly specific molecular markers or antibodies for a certain cell type, these methods are difficult to use to detect unknown cell types. Furthermore, all of these detection methods have difficulty simultaneously analyzing multiple cell types in the same sample, and the information obtained is very limited, making high-throughput cell identification and quantification impossible. Furthermore, because all of these methods rely on the detection of a few specific gene or protein targets in functional cells to determine cell type, the limited number of detection targets and heterogeneous or nonspecific expression of targets in different cellular states may result in some false positives or false negatives, and the post-detection data cannot be deeply investigated or analyzed.

[0004] In summary, all related technologies, whether protein-based or nucleic acid-based target detection methods, have drawbacks such as difficulty in analyzing unknown samples, high cost and low throughput of detection by matrix screening, and being heavily influenced by the specificity of limited molecular markers, antibodies, and probes.

[0005] In view of this, achieving one-step detection of sample cellular composition with low cost, high throughput and accurate quantification is a technical challenge that must be overcome by those skilled in the art. Summary of the Invention

[0006] The present application provides a method, device, electronic device and readable storage medium for detecting cellular constituents of a biological sample, which enables one-step detection of cellular constituents of a biological sample with low cost, high throughput and accurate quantification.

[0007] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions.

[0008] Embodiments of the present invention, on the one hand, provide a method for detecting a cellular composition of a biological sample, comprising: Obtaining single-cell sequencing results by performing single-cell transcriptome sequencing on the test biological sample; generating a cellular gene expression matrix by analyzing the single-cell sequencing results; and determining the cell types contained in the test biological sample by performing single-cell bioinformatics analysis on the cellular gene expression matrix.

[0009] Optionally, after performing single-cell bioinformatics analysis on the cell gene expression matrix described above, The method further includes mapping functional cellular gene expression matrix information onto the cellular gene expression matrix to perform initial cellular annotation.

[0010] Optionally, after mapping the functional cellular gene expression matrix information to the cellular gene expression matrix, generating and displaying a functional cell target gene expression level and cell expression ratio diagram in response to a drawing command for a functional cell target gene expression cluster diagram; The method further includes generating annotation confirmation information including the cell type to which each cell cluster belongs in response to the annotation confirmation result input command.

[0011] Optionally, performing single-cell bioinformatics analysis on the cell gene expression matrix; generating gene splicing data by analyzing the single-cell sequencing results; and determining whether the annotated cell clusters conform to a biological developmental trajectory by performing a trajectory analysis of the annotated cell clusters based on the gene splicing data.

[0012] Optionally, after determining the cell type contained in the test biological sample, The method further includes calculating the cellular composition ratio and generating a detection result of the cellular composition ratio of the test biological sample.

[0013] Optionally, determining the cell types contained in the test biological sample by performing single-cell bioinformatics analysis on the cellular gene expression matrix includes: and performing single-cell transcriptional data analysis on the cellular gene expression matrix in an interactive computing environment to obtain cell types contained in the test biological sample.

[0014] Optionally, determining the cell types contained in the test biological sample by performing single-cell bioinformatics analysis on the cellular gene expression matrix includes: Invoke a cellular gene calculation equation to calculate the proportion of cellular mitochondrial genes in the cellular gene expression matrix, the total number of genes detected from the cells, the total number of gene fragments detected from the cells, the total number of fragments detected from the genes, and the total number of cells; In response to the filtering parameter setting command, the cell filtering relational formula and the gene filtering relational formula are respectively called, and the genes and cells whose detection quality does not satisfy the preset quality condition are filtered out, thereby obtaining the target cell gene data; In response to a data normalization command, perform a data normalization process on the target cellular genetic data, and perform a dimension reduction process on the normalized data; In response to a cell clustering processing command, performing a cell clustering process on the dimension-reduced data to obtain cell sub-cluster information.

[0015] Optionally, after obtaining the target cellular genetic data, The method further includes calling a cell cycle evaluation relational expression to determine the cell cycle of each cell in the target cell genetic data.

[0016] Optionally, performing a data normalization process on the cellular gene expression matrix in response to the data normalization process command comprises: calling a normalization relation to normalize the cellular gene expression matrix; Logarithmically transforming the standardized data by invoking the logarithmic transformation formula; and invoking an aberrant gene removal relation to remove aberrantly over-expressed genes from the log-transformed data.

[0017] Optionally, the test biological sample includes multiple batches of biological samples, and the single-cell sequencing result includes multiple single-cell sequencing results carrying batch information. After determining the cell types contained in the test biological sample, The method further includes generating batch-to-batch stability results of cellular constituent ratios by analyzing the cellular constituent ratio data of each batch of biological sample.

[0018] Optionally, the test biological sample is obtained by sampling the same biological sample at multiple time points, and the single-cell sequencing results include single-cell sequencing results of the same biological sample at multiple time points. After determining the cell types contained in the test biological sample, For each biological sample at each time point, obtaining cell composition ratio data determined by the current biological sample; The method further includes generating information on changes in the cellular composition ratios by performing time series analysis on the cellular composition ratio data of the biological sample at different time points.

[0019] Embodiments of the present invention, on the other hand, provide an apparatus for detecting a cellular composition of a biological sample, the apparatus comprising: a sequencing module used to perform single-cell transcriptome sequencing on the test biological sample to obtain single-cell sequencing results; a data analysis module used to generate a cellular gene expression matrix by analyzing the single-cell sequencing results; and a cell type determination module used to determine the cell type contained in the test biological sample by performing single-cell bioinformatics analysis on the cell gene expression matrix.

[0020] An embodiment of the present invention provides an electronic device including a processor, which, when executing a computer program stored in a memory, is used to implement the steps of the method for detecting the cellular composition of a biological sample described in any one of the preceding claims.

[0021] Finally, an embodiment of the present invention provides a readable storage medium having a computer program stored therein, the computer program implementing the steps of the method for detecting a cellular composition of a biological sample according to any one of the preceding claims when executed by a processor.

[0022] The advantages of the technical solution provided by this application are as follows: By performing bioinformatics analysis based on the single-cell sequencing results of the test sample, cellular characteristic information can be deeply explored, and the gene expression transcription level of each cell's state at the time of detection can be presented at the single-cell level. This effectively improves the detection stability of the entire biological sample, helps avoid false positive or false negative judgment errors, and effectively improves detection accuracy. Furthermore, the entire analysis is performed by a modular program, which does not require high operator experience, effectively reduces the detection cost of cellular composition, and meets the requirements for detection stability and reproducibility. At the same time, the detected data can be gradually accumulated and, combined with the iterative development of the technology, can be repeatedly mined, analyzed, and utilized, thereby further improving the detection accuracy of cellular composition of biological samples.

[0023] Furthermore, embodiments of the present invention provide corresponding realization devices, electronic devices and readable storage media for the method of detecting cellular composition of a biological sample, which make the method more practical, and the devices, electronic devices and readable storage media have corresponding advantages.

[0024] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the present disclosure. [Brief explanation of the drawings]

[0025] In order to more clearly explain the technical solutions in the embodiments of the present invention or the related art, the following will briefly introduce the drawings used in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative efforts.

[0026] [Figure 1] 1 is a flowchart of a method for detecting cellular constituents of a biological sample. [Figure 2] 1 is a structural diagram of one particular embodiment of an apparatus for detecting the cellular composition of a biological sample. [Figure 3] FIG. 1 is a structural diagram of a particular embodiment of an electronic device. [Figure 4] 1 is a flowchart of another method for detecting cellular constituents of a biological sample. [Figure 5] FIG. 10 is a violin diagram of the total number of genes detected before the sample cells are filtered. [Figure 6] FIG. 10 is a violin diagram of the total number of gene fragments detected before the sample cells are filtered. [Figure 7] FIG. 10 is a violin diagram of the mitochondrial gene detection ratio before the sample cells are filtered. [Figure 8] FIG. 10 is a violin diagram of the total gene detection counts after sample cells are filtered. [Figure 9] FIG. 10 is a violin diagram of the total number of gene fragments detected after sample cells are filtered. [Figure 10] FIG. 10 is a violin diagram of mitochondrial gene detection ratios after sample cells are filtered. [Figure 11] This is a cell cycle distribution map obtained by reducing the dimension of the data. [Figure 12] This is an annotation diagram of functional cell matrix mapping cells after dimensionality reduction of the data. [Figure 13] 10 is a pie chart of the composition ratio of functional cell matrix mapping annotation cells. [Figure 14] This is a cell cluster diagram obtained by reducing the dimension of the data. [Figure 15] 1 is a bubble chart of cell clustering gene expression. [Figure 16] Annotation diagram of target gene expression corrected cells after dimensionality reduction of data. [Figure 17] This is a diagram of cell trajectories obtained by reducing the dimension of the data. [Figure 18] 10 is a pie chart of cell annotation composition ratios ascertained by biological developmental trajectories. [Figure 19] 1 is a pie chart of cell cycle ratios. [Figure 20] 1 is a bar graph of cell cycle composition ratio. [Figure 21] 1 is a bar graph of cell constituent cycle ratios. [Figure 22] This is a cell annotation diagram in which the dimensions of the data were reduced using the SingleR software based on the multi-cell annotation database. [Figure 23] 1 is a pie chart of the cell annotation composition ratio based on the multi-cell annotation database of the SingleR software. DETAILED DESCRIPTION OF THE INVENTION

[0027] In order to allow those skilled in the art to better understand the aspects of the present invention, the following will further describe the present invention in detail in combination with drawings and specific embodiments. It is clear that the described embodiments are only a part of the embodiments of the present invention, and are not all of them. Based on the embodiments in the present invention, all other embodiments that can be obtained by those skilled in the art without requiring creative efforts shall fall within the protection scope of the present invention.

[0028] The terms "comprises" and "having" and variations thereof in the specification, claims, and drawings of this application are intended to be non-exclusive inclusive. For example, a process, method, system, product, or apparatus of a series of steps or units is not limited to the listed steps or units, but may also include unlisted steps or units.

[0029] The expression "method for detecting the cellular composition of a biological sample" in the specification, claims, and drawings of this application refers to a method for determining the cell types contained in a biological sample and / or a method for generating a detection result of the cellular composition ratio of a biological sample.

[0030] Having described the technical solutions of the embodiments of the present invention, various non-limiting embodiments of the present application are described in detail below.

[0031] First, referring to FIG. 1, FIG. 1 is a flowchart of a method for detecting the cellular composition of a biological sample provided by an embodiment of the present invention, which may include the following steps S101 to S103.

[0032] S101: Single-cell transcriptome sequencing is performed on the biological sample to be tested, thereby obtaining single-cell sequencing results.

[0033] In this example, the test sample may be a dissociated cell sample or a dissociated tissue sample, but neither of these will affect the implementation of this application. At the single-cell level, genomes or transcriptomes are extracted and amplified, followed by high-throughput sequencing analysis to obtain single-cell sequencing results. The single-cell sequencing results are used to represent the genetic information of a single cell, i.e., the gene structure and gene expression state specific to the single cell. In this example, any software capable of analyzing and comparing single-cell transcriptome sequencing data to obtain a cellular gene expression matrix may be used, including software that does not affect the implementation of this application, such as Cell ranger on the 10X GENOMICS platform, CeleScope software, and STARsolo software. Increasing the amount of single-cell sequencing data refers to sequencing technology that can perform high-depth, i.e., more than 100G, single-cell transcriptome sequencing on the biological sample being tested, in order to further improve the performance of single-cell sequencing, where G is the unit of sequencing data volume, and 1G represents 1 billion bases. For example, 1 million cell samples or tissues can be used to dissociate them into single cells, and then 100G deep single-cell transcriptome sequencing can be performed using the 10X GENOMICS sequencing platform.

[0034] S102: Generate a cellular gene expression matrix by analyzing single-cell sequencing results.

[0035] In this step, a comparison index can be constructed based on the single-cell sequencing results, referencing a human genome database such as the GRCh38 version and a gene annotation database such as the GENECODE database annotation file. This comparison index is then used in the sequence comparison process, i.e., converting the sequence file into a cellular gene expression matrix. Furthermore, for single-cell transcriptome data, a genome comparison and cellular quantification can be used to obtain an expression matrix of the single-cell sequencing data. After filtering the expression matrix of the single-cell sequencing data, a cellular gene expression matrix can be obtained.

[0036] S103: Determine the cell types contained in the biological sample under test by performing single-cell bioinformatics analysis on the cellular gene expression matrix.

[0037] In this step, to improve the accuracy and efficiency of detecting the cellular composition of biological samples, the single-cell bioinformatics analysis process can be performed in an interactive computing environment. Performing the entire data analysis process in an interactive computing environment allows parameters to be adjusted in real time according to the results of each step, thereby achieving the goal of flexible processing. In this example, single-cell bioinformatics analysis refers to a series of data analysis operations on the cellular gene expression matrix using biological analysis tools, ultimately aiming to identify the cell types contained in the biological sample being tested. In this step, existing single-cell transcriptional data analysis functions, such as Scanpy, can be called to perform single-cell transcriptional data analysis on the cellular gene expression matrix in an interactive computing environment, thereby obtaining functional cellular gene expression matrix information. Alternatively, the cellular gene expression matrix can be analyzed using Scanpy on the visual operation platform of Python Jupyter Notebook, or using Seurat software on the R language platform.

[0038] The technical solution provided in the examples of this application can perform bioinformatics analysis based on the single-cell sequencing results of the test sample to deeply explore cellular characteristic information, and can present the gene expression transcription level of each cell state at the time of detection at the single-cell level, which effectively improves the detection stability of the entire biological sample, helps avoid false positive or false negative judgment errors, and effectively improves detection accuracy. Furthermore, the entire analysis is performed by a modular program, which does not require high operator experience, effectively reduces the detection cost of cellular composition, and meets the requirements for detection stability and reproducibility. At the same time, the detected data can be gradually accumulated and, combined with the iterative development of the technology, can be repeatedly mined, analyzed, and utilized, thereby further improving the detection accuracy of cellular composition of biological samples.

[0039] To further improve detection accuracy and avoid inaccurate final cell detection results due to false positive or false negative results caused by an operator using a small number of targets, the present application, based on the above embodiment, can rapidly identify the cell types contained in the test biological sample using functional cell gene expression matrix mapping before determining the cell types contained in the test biological sample in step S103, thereby further improving detection accuracy. Furthermore, based on the above embodiment, after performing single-cell bioinformatics analysis on the cell gene expression matrix, initial cell type annotation can be performed by mapping the functional cell gene expression matrix information to the cell gene expression matrix. After obtaining information related to cell type annotation, the cell composition ratio is calculated based on the cell annotation information, and a cell composition ratio detection result for the test biological sample is generated.

[0040] Single-cell transcriptome sequencing provides unlabeled data. To determine the cell type of the biological sample being tested that corresponds to the single-cell transcriptome sequencing data, a similarity comparison with existing labeled data can be performed. In this step, the annotated functional cell gene expression matrix information is mapped to a cell gene expression matrix corresponding to the original single-cell sequencing results, allowing for rapid identification of the cell type. The functional cells used for the functional cell gene expression matrix information can be any cells selected by a skilled artisan for a more practical application scenario. The functional cell gene expression matrix information is obtained by performing single-cell transcriptome sequencing and single-cell sequencing result analysis on the functional cells. Here, the functional cell gene expression matrix information can be mapped to a cell gene expression matrix using software such as scGCN or singleR. Naturally, other annotation techniques can also be used to map the functional cell gene expression matrix information to a cell gene expression matrix.

[0041] After initial annotation of cell types, the proportion of each type of cell in the biological sample to be tested is calculated based on the results of the initial annotation, thereby obtaining the proportion of each type of cell in the biological sample to be tested. The detection result of cell composition proportion may indicate the types of cells contained in the biological sample to be tested, or may directly represent the data of the cell composition proportion of the biological sample to be tested. The detection result of cell composition proportion may be in any format, such as a Word document, PDF text, or Excel spreadsheet, and may also be exported as an image, audio video, etc., without affecting the realization of the present application.

[0042] In the above embodiment, in the process of performing bioinformatics analysis of single-cell sequencing data, in order to quickly identify and further improve the detection accuracy of cell types by mapping functional cell gene expression matrices, this embodiment, based on the above embodiment, after performing the step of "mapping functional cell gene expression matrix information onto the cell gene expression matrix and performing initial cell annotation" in the above embodiment, may further include generating and displaying functional cell target gene expression levels and cell expression ratio diagrams, such as bubble charts, scatter plots, trajectory diagrams, heat maps, etc., in response to a drawing command for a functional cell target gene expression cluster diagram, and generating annotation confirmation information including the cell type to which each cell cluster belongs in response to an annotation confirmation result input command.

[0043] In this embodiment, any drawing software can be adopted to draw the target gene expression cluster diagram of the functional cells used in the above embodiment, and the accuracy of the cell type annotation in the above embodiment can be confirmed based on the target gene expression cluster diagram obtained through artificial experience. Through the human-computer interaction module, the cell type to which each cell cluster contained in the biological sample to be tested belongs can be further input into the system, and the system will generate annotation confirmation information for the initial cell annotation based on the received information.

[0044] In this embodiment, functional cellular gene expression matrix mapping is used to rapidly identify the cell, and in combination with multi-target gene cross-validation, the cell type, developmental state, and cell cycle characteristics can be further confirmed, thereby avoiding the false positive or false negative results caused by using a small number of targets in the above technology and further improving the accuracy of detecting the cellular composition of the biological sample. Correspondingly, the cellular composition proportion detection results of the above embodiment can be verified based on the annotation confirmation information, and optionally the cellular composition proportions can be recalculated based on the annotation confirmation information, and then the cellular composition proportion detection results of the tested biological sample can be updated.

[0045] In order to achieve accurate qualitative and quantitative detection of cell types contained in a biological sample and quantitative detection of the cellular composition ratio of the biological sample being tested, the method may further include generating gene splicing data by analyzing the single-cell sequencing results based on the above embodiment, and performing trajectory analysis of the annotated cell clusters based on the gene splicing data to determine whether the annotated cell clusters conform to biological developmental trajectories.

[0046] In this example, the sequence comparison file can be processed using any RNA kinetic analysis technology to obtain gene splicing data. The gene splicing data can then be used to analyze the trajectory of cell development. In this step, any biological development trajectory analysis software that performs trajectory analysis on annotated cell clusters, such as scvelo, monocle, or CellRank, can be used to perform trajectory analysis on the test biological sample. The trajectory analysis verification results for the test biological sample can be obtained to determine whether they match the biological development trajectory of the corresponding type of cell cluster.

[0047] Correspondingly, in order to verify the accuracy of the cell composition ratio detection results in the above-mentioned embodiments, the cell type annotation results may be further verified based on the target gene expression, i.e., the cell type annotation results may be verified using annotation confirmation information and trajectory analysis, the cell composition ratio may be recalculated, and finally, the cell composition ratio detection results of the biological sample to be tested that have already been generated may be updated based on the currently calculated cell composition ratio.

[0048] In this example, trajectory analysis is used to determine whether cell subpopulations conform to biological developmental trajectories, and the proportion of each type of cell is calculated, enabling accurate qualitative and quantitative detection of the cell types in the sample and quantitative detection of the cellular composition of the biological sample being tested.

[0049] Based on the above embodiment, in order to further enhance the richness of the cell composition detection results and improve the user experience, the cell composition detection results of the biological sample to be tested in this embodiment may include more data. Optionally, this embodiment may further include calling a plot function according to the proportion information of each type of cell to generate a pie chart of cell type distribution and a bar chart of cell cycle distribution of each type, and generating a cell composition proportion detection result based on the pie chart of cell type distribution and the bar chart of cell cycle distribution of each type.

[0050] In this embodiment, the plot function may include a plot function for generating a pie chart of cell type distribution, and may further include a plot function for generating a bar chart of each type of cell cycle distribution. Of course, in addition to the proportion of each type of cell in the biological sample being tested, other information such as trajectory analysis verification results, functional cell target gene expression levels and cell expression ratio diagrams, annotation confirmation information, cell type distribution pie charts, and cell cycle distribution bar charts can be appropriately selected according to actual needs, and none of these will affect the realization of the present application.

[0051] In the above-described example, the method for performing the step of "confirming the cell type contained in the test biological sample by performing single-cell bioinformatics analysis on the cell gene expression matrix" is not limited, and the present example further provides an optional embodiment, Invoke a cellular gene calculation equation to calculate the proportion of cellular mitochondrial genes in the cellular gene expression matrix, the total number of genes detected from the cells, the total number of gene fragments detected from the cells, the total number of fragments detected from the genes, and the total number of cells; In response to the filtering parameter setting command, the cell filtering relational formula and the gene filtering relational formula are respectively called, and the genes and cells whose detection quality does not satisfy the preset quality condition are filtered out, thereby obtaining the target cell gene data; Calling a cell cycle evaluation equation and determining the cell cycle of each cell included in the target cell genetic data; In response to a data normalization command, perform data normalization on the target cell gene data and perform dimension reduction on the normalized data; In response to a cell clustering processing command, performing a cell clustering process on the dimension-reduced data to obtain cell subpopulation information.

[0052] In this embodiment, any clustering algorithm may be used to perform the cell clustering process, for example, clustering is performed directly using a clustering function, and the clustering function may be, for example, Leiden, Louvain, etc., and the present application is not limited thereto.

[0053] The above embodiment does not limit how the data normalization operation is performed, and the embodiment further provides an optional embodiment of data normalization, Invoking the normalization equation to normalize the cellular gene expression matrix; Logarithmically transforming the standardized data by invoking the logarithmic transformation formula; and invoking an aberrant gene removal relation to remove aberrantly over-expressed genes from the log-transformed data.

[0054] Based on the above embodiment, in order to further enhance the richness of the cell type detection report and improve the user experience, based on the above embodiment, the detection result report of the biological sample to be tested in this embodiment may further include more data, and optionally, the cell composition proportion detection result may include target cell genetic data, dimension reduction processing data, cell classification information, etc., that is, the cell composition proportion detection result can be generated according to the biological development trajectory verification result, target cell genetic data, dimension reduction processing data, cell classification information, and the proportion of each type of cell.

[0055] In addition, to further explore the stability of the cellular composition of biological samples from different batches or the changes in the cellular composition of biological samples at different time points, based on the above embodiment, this embodiment can also perform integrated analysis, verify batch-to-batch cellular stability, time series analysis of cells from the same batch at different culture times, and monitor changes in cellular composition. In this embodiment, the tested biological sample may include multiple batches of biological samples, and correspondingly, the single-cell sequencing result will include multiple single-cell sequencing results carrying batch information. After generating the detection result report for the tested biological sample, the cellular composition data of each batch of biological sample is analyzed to generate a batch-to-batch cellular composition stability result, thereby integrating the analyses of multiple batches of detected samples to verify the stability of the cellular composition of batches.

[0056] The same test biological sample is obtained at multiple time points, and the single-cell sequencing results include single-cell sequencing results for the same biological sample at multiple time points. According to the method of the above embodiment, the biological sample collected at each time point is processed to obtain data on the cellular composition ratio determined for the biological sample at the current time point, and the data on the cellular composition ratio at different time points is analyzed in a time series to generate information on changes in the cellular composition ratio. Changes in the cellular composition ratio are monitored by performing time series analysis on multiple detections for the same batch.

[0057] In this example, the verification results of cell stability for multiple batches are combined with the time series analysis results of cells from the same biological sample at different culture times, thereby further improving the detection accuracy of the cellular composition ratio of the biological sample being tested and further exploring the stability of the cellular composition of the biological sample or changes in the cellular development process.

[0058] It should be noted that there is no strict order to the execution of steps in this application, and as long as they follow a logical order, these steps may be executed simultaneously or in a certain predetermined order. Figure 1 is merely an example and does not mean that the steps can only be executed in that order.

[0059] The present invention provides a corresponding apparatus for a method for detecting the cellular composition of a biological sample, thereby making the method more practical. Hereinafter, the apparatus can be described from the perspective of functional modules and from the perspective of hardware. The apparatus for detecting the cellular composition of a biological sample provided by the present invention will be described below. The apparatus for detecting the cellular composition of a biological sample and the method for detecting the cellular composition of a biological sample described below can be referred to correspondingly.

[0060] From the viewpoint of functional modules, referring to FIG. 2, FIG. 2 is a structural diagram of an apparatus 200 for detecting cellular components of a biological sample provided by one embodiment of the present invention, which includes: A sequencing module 201 is used to perform single-cell transcriptome sequencing on the test biological sample to obtain single-cell sequencing results; a data analysis module 202 used to generate a cellular gene expression matrix by analyzing the single-cell sequencing results; and a cell type determination module 203 used to determine the cell type contained in the biological sample being tested by performing single-cell bioinformatics analysis on the cellular gene expression matrix.

[0061] Optionally, in some embodiments of this example, the device comprises: It may further include an annotation module that is used to map the functional cell gene expression matrix information to a cell gene expression matrix and perform initial cell annotation.

[0062] As an optional embodiment of the above example, the device may further include an annotation confirmation module used to generate and display a functional cell target gene expression level and cell expression ratio diagram in response to a drawing command of a functional cell target gene expression cluster diagram, and to generate annotation confirmation information including the cell type to which each cell cluster belongs in response to an annotation confirmation result input command.

[0063] As another optional embodiment of the above example, the device may further include a trajectory analysis module used for generating gene splicing data by analyzing single-cell sequencing results, and determining whether the annotated cell clusters conform to a biological developmental trajectory by performing trajectory analysis of the annotated cell clusters based on the gene splicing data.

[0064] As another optional embodiment of the above example, the device may further include a result generation module, for example, used to calculate the cellular composition ratio and generate a cellular composition ratio detection result for the biological sample being tested.

[0065] Optionally, in some other embodiments of this example, the device may further include a verification module used to generate and display a functional cell target gene expression level and cell expression ratio diagram in response to a drawing command of a target gene expression cluster diagram of a functional cell, and to generate cell type annotation accuracy information of the annotated cell cluster in response to an accuracy information input command.

[0066] Optionally, in some other embodiments of this example, the cell type determination module 203 may further be used to call a single-cell transcriptional data analysis function to perform single-cell transcriptional data analysis on a cell gene expression matrix in an interactive computing environment to obtain functional cell gene expression matrix information.

[0067] In any embodiment of the above example, the cell type determination module 203 may further be used to: call a cell gene calculation relational expression to calculate the mitochondrial genes, the number of cell genes, and the total number of gene fragments detected per cell in the cell gene expression matrix; call a cell filtering relational expression and a gene filtering relational expression, respectively, in response to a filtering parameter setting command, to filter out genes and cells whose detection quality does not satisfy a preset quality condition, and obtain target cell gene data; call a cell cycle evaluation relational expression to determine the cell cycle of each cell included in the target cell gene data; perform data normalization processing on the target cell gene data and perform dimensionality reduction processing on the normalized data in response to a data normalization processing command; and perform cell clustering processing on the dimensionality reduction processing data to obtain cell subpopulation information in response to a cell clustering processing command.

[0068] As another optional embodiment of the above example, the cell type determination module 203 may further be used to invoke a standardization relational expression to perform a standardization process on the cell gene expression matrix, invoke a logarithmic transformation relational expression to logarithmically transform the standardized data, and invoke an abnormal gene removal relational expression to remove abnormally highly expressed genes from the logarithmically transformed data.

[0069] Optionally, in some other embodiments of this embodiment, the device may further include a result verification module, and the result verification module may include the following stability verification unit and change verification unit.

[0070] The stability verification unit is used to generate batch-to-batch stability results of cell constituent ratios by analyzing the cell constituent ratio data of each batch of biological sample, the tested biological sample includes multiple batches of biological samples, and the single-cell sequencing results include multiple single-cell sequencing results carrying batch information.

[0071] The variation verification unit is used to obtain cell composition ratio data determined by the current biological sample for each time point, and to generate variation information of cell composition ratio by time series analysis of the cell composition ratio data of the biological sample at different time points. The test biological sample is the same biological sample sampled at multiple time points, and the single-cell sequencing results include single-cell sequencing results of the same biological sample at multiple time points.

[0072] In the device for detecting the cellular composition ratio of a biological sample described in the embodiments of the present invention, the functions of each functional module can be specifically realized according to the methods in the method embodiments, and the specific realization process can be referred to the relevant descriptions in the method embodiments, and will not be repeated here.

[0073] As can be seen from the above, embodiments of the present invention enable one-step detection of cellular composition of biological samples with low cost, high throughput, and accurate quantification.

[0074] The device for detecting the cellular composition of a biological sample is described in terms of functional modules, and the present application also provides an electronic device described in terms of hardware. Figure 3 is a schematic diagram showing the structure of an electronic device provided by one embodiment of the present application. As shown in Figure 3, the electronic device includes a memory 30 for storing a computer program and a processor 31 for implementing the steps of the method for detecting the cellular composition of a biological sample described in any one of the above examples when executing the computer program.

[0075] Here, the processor 31 may include one or more processor cores, such as a 4-core processor or an 8-core processor. The processor 31 may also be a controller, microcontroller, microprocessor, or other data processing chip. The processor 31 may be implemented in at least one hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). The processor 31 may include a main processor and a coprocessor. The main processor, also referred to as a CPU (Central Processing Unit), processes data in a wake-up state, and the coprocessor is a low-power processor that processes data in a standby state. In some embodiments, the processor 31 may be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content displayed by the screen. In some embodiments, the processor 31 may also include an AI (Artificial Intelligence) processor used to process computational operations related to machine learning.

[0076] The memory 30 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 30 may further include non-volatile memory, such as one or more disk storage devices or flash memory storage devices, in addition to high-speed random access memory. In some embodiments, the memory 30 may be an internal storage device of the electronic device, such as a server hard disk. In other embodiments, the memory 30 may be an external storage device of the electronic device, such as a plug-in hard disk installed in the server, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash card. Furthermore, the memory 30 may include both an internal storage device and an external storage device of the electronic device. The memory 30 may be used not only to store application software and various data installed in the electronic device, such as program code for a process of executing a method for detecting cellular constituents of a biological sample, but also to temporarily store output data or data to be output. In this embodiment, the memory 30 is used to store at least the following computer program 301, which, after being loaded and executed by the processor 31, can realize the relevant steps of the method for detecting the cellular composition of a biological sample disclosed in any one of the above embodiments. Furthermore, resources stored in the memory 30 may include an operating system 302, data 303, etc., and the storage may be temporary or permanent. Here, the operating system 302 may include Windows, Unix, Linux, etc. The data 303 may include, but is not limited to, data corresponding to the detection result of the cellular composition of the biological sample.

[0077] In some embodiments, the electronic device may further include a screen 32, an input / output interface 33, a communication interface 34 (also called a network interface), a power supply 35, and a communication bus 36. Here, the screen 32 and the input / output interface 33, such as a keyboard, belong to a user interface, and optional user interfaces may include standard wired interfaces, wireless interfaces, etc. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch panel liquid crystal display, an OLED (organic light-emitting diode) touch screen, etc. The display, also called a screen or display unit, is used to display information processed within the electronic device or to display a visualized user interface. Optionally, the communication interface 34 may include a wired interface and / or a wireless interface, such as a Wi-Fi interface or a Bluetooth interface, and is typically used to establish a communication connection between the electronic device and other electronic devices. The communication bus 36 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one bold line is shown in Figure 3, but this does not mean that there is only one bus or only one type of bus.

[0078] Those skilled in the art will appreciate that the structure shown in FIG. 3 is not intended to limit the electronics and may include more or fewer components than those shown, such as sensors 37 that implement various functions.

[0079] In the electronic devices described in the embodiments of the present invention, the functions of each functional module can be specifically realized according to the methods in the method embodiments, and the specific realization process can be referred to the relevant descriptions of the method embodiments, and will not be repeated here.

[0080] As can be seen from the above, embodiments of the present invention enable one-step detection of cellular composition of biological samples with low cost, high throughput, and accurate quantification.

[0081] It should be understood that the method for detecting cellular components of a biological sample in the above embodiments may be implemented in the form of a software functional unit and stored in a computer-readable storage medium when sold or used as a standalone product. Based on this understanding, the technical solution of the present application that essentially contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product, which is stored in a storage medium and executes all or part of the steps of the method in each embodiment of the present application. The storage medium includes various media capable of storing program code, such as USB memory, mobile hard disk, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, multimedia card, card-type memory (such as SD or DX memory), magnetic memory, removable disk, CD-ROM, magnetic disk, or optical disk.

[0082] Based on this, an embodiment of the present invention further provides a readable storage medium having a computer program stored therein, the computer program, when executed by a processor, implementing the steps of the method for detecting the cellular composition of a biological sample according to any one of the preceding embodiments.

[0083] Finally, in order to make those skilled in the art understand the technical solution of the present application more clearly, the present application also presents a schematic embodiment in conjunction with FIG. The method may include dissociating one million cell samples or tissues from the biological sample to be tested into single cells, performing 150G deep sequencing of the NPC cells using a 10X GENOMICS sequencing platform, and obtaining single-cell test results. Bioinformatics analysis is performed on the single-cell sequencing results, and the bioinformatics analysis step may include:

[0084] A1: The downstream format data from the 10X GENOMICS sequencing platform, i.e., fastq single-cell data, is obtained, and the mkref function, i.e., the reference single-cell comparison database construction function, of the single-cell sequencing analysis software Cell ranger is used to reference the human GRCh38 (i.e., human reference genome version number) version genome and construct a comparison index with the latest version of the annotation file in the GENECODE (i.e., gene annotation information database) database. The count function of the Cell ranger software and its default parameter analysis are used to obtain an expression matrix of the single-cell sequencing data. The output data from the Cell ranger software, i.e., the expression matrix of the single-cell sequencing data, is quality-checked. If the quality check is successful, a cellular gene expression matrix is ​​generated, along with a bam file for sequence comparison. If the quality check is unsuccessful, single-cell transcriptome sequencing is performed again on the biological sample being tested using the 10X GENOMICS sequencing platform. Alternatively, this can be implemented, for example, by the following computer program: Cell ranger count --id=hNPC --fastqs=DATA_DIR --transcriptome=GRCh38-Genecode --localcores=16 --localmem=80.

[0085] A2: The run10x function of the Velocyto software is used to process the BAM file generated from the Cell Ranger analysis, i.e., to analyze the splicing rate of cellular genes, and obtain a cellular gene splicing rate LOOM file. The LOOM file is generated by the Velocyto software and stores the number of spliced / unspliced ​​transcripts, which is then used to analyze cell developmental trajectories. Alternatively, this can be implemented using, for example, the following computer programs: velocyto run10x -m human38_rmsk.gtf DATA_DIR genes.gtf

[0086] A3: To optimize parameters in a timely manner, the functions for each of the following steps can be compiled directly on the Python platform, or these functions can be integrated into the backend. After compiling the functions for each of the following steps in advance, single-cell transcriptome data can be analyzed using Scanpy based on the Python platform.

[0087] A31: Use the function scvelo.read to read the loom file (e.g., velo.loom) generated by the velocyto software. Alternatively, this can be implemented, for example, by the following computer program: ldata=scvelo.read(velo.loom,cache=True)

[0088] A32: Use the function read_cellranger to read the cell gene expression matrix file (e.g., filtered_feature_bc_matrix.h5) analyzed, processed, and filtered by the Cellranger software. Alternatively, this can be implemented, for example, by the following computer program: count_data=read_Cell ranger(filtered_feature_bc_matrix.h5,gex_ only=True,add_sample_id=False,cache=False)

[0089] A33: Use the function annotate_qc_metrics to calculate the mitochondrial genes and ribosomal genes of cells and their detection ratios in the cells. Alternatively, this can be realized, for example, by the following computer program: count_data=annotate_qc_metrics(count_data)

[0090] A34: Using the functions filtering_cell and filtering_gene, specific filtering parameters can be set to perform data quality control and filtering, where the filtering parameters can be specifically analyzed and changed according to the data. That is, using filtering_cell and filtering_gene, the filtering parameters can be adjusted according to the data to exclude genes and cells with low detection quality. Optionally, this can be realized, for example, by the following computer program: count_data=filtering_cell(count_data,min_t_count=1500,max_t_count=150000,min_t_gene=1500,max_t_gene=10000,min_mito_percent=0,max_mito_percent=0.5) count_data=filtering_gene(count_data,min_cells=10)

[0091] A35: The function annotate_cellcycle is used to evaluate and confirm the cell cycle, which can optionally be realized, for example, by the following computer program: annotate_cellcycle(count_data)

[0092] A36: The cell gene expression matrix is ​​normalized using the scanpy.pp.normalize_total function, the normalized data is logarithmically transformed using the scanpy.pp.log1p function, and then the data is calculated and analyzed using the scanpy.pp.recipe_zheng17 function to remove abnormally highly expressed genes. Alternatively, this can be implemented, for example, by the following computer program: scanpy.pp.normalize_total(count_data,exclude_highly_expressed=True,max_fraction=0.05,inplace=True) scanpy.pp.log1p(count_data) scanpy.pp.recipe_zheng17(count_data,n_top_genes=3000,log=False,plot=False,copy=False)

[0093] A37: Use the scanpy.tl.pca function to perform principal component analysis on the data obtained in the previous step, and use the UMAP function or tSNE function to perform dimensionality reduction on the data obtained in the previous step. Alternatively, this can be realized by, for example, the following computer program. scanpy.tl.pca(count_data) scanpy.pl.pca_overview(count_data) cpm_nml_umap=umap_.UMAP(n_neighbors=30,min_dist=0.9,n_components=2,random_state=42).fit_transform(count_data.obsm[“X_pca ”])

[0094] A38: Unsupervised clustering of the data obtained in the previous step can be performed using the scanpy.pp.neighbors function or the Leiden or Louvain algorithm, where the resolution can be adjusted depending on the subpopulation. Alternatively, this can be realized, for example, by the following computer program: scanpy.pp.neighbors(count_data,n_neighbors=50,n_pcs=30,use_rep=“X_pca”,knn=True,random_state=0,method='umap',metric='euclidean') scanpy.tl.leiden(count_data,resolution=2,key_added=“leiden2”)

[0095] A4: Using scGCN software, functional cell gene expression matrix information is mapped to detection data such as filtered_feature_bc_matrix.hd5 to rapidly identify cell types and calculate the proportion of each type of cell.

[0096] A5: Draw a diagram of functional cell target gene expression levels and cell expression ratios to confirm the accuracy of cell type annotation.

[0097] A6: Use scvelo software to perform trajectory analysis of the annotated cell clusters to determine whether the annotation information matches the biological development trajectory. Optionally, this can be implemented, for example, by the following computer program: adata_combine=scv.utils.merge(count_data,ldata) scvelo.pp.filter_and_normalize(adata_combine,min_shared_counts=10,n_top_genes=5000) scvelo.pp.moments(adata_combine,n_pcs=30,n_neighbors=30) scvelo.tl.recover_dynamics(adata_combine,n_jobs=36) scvelo.tl.velocity(adata_combine,mode='dynamical',n_jobs=36) scvelo.tl.velocity_graph(adata_combine, n_jobs=36) scvelo.pl.velocity_embedding_stream(adata_combine,basis='umap',color='cell_type_annot2',X=adata_combine.obsm[“X_umap”])

[0098] A7: Calculate the proportion of each type of cell, and use the self-created functions cell_proportion_pieplo and cell_proportion_barplot to plot a pie chart of cell type distribution and a bar chart of cell cycle distribution for each type. Alternatively, this can be realized, for example, by the following computer program: cell_proportion_pieplot(count_data,x=“cell_type_annot2”,save=“Cell_tp_prpie.pdf”) cell_proportion_barplot(count_data,x=“cell_type_annot2”,y=“phase”,color=colors_leiden_2,save=“Cell_tp_prbar.pdf”)

[0099] A8: Finally, create a code-free PDF version of your analysis report with one click.

[0100] In this step, pandoc, TeXLive, and nbextensions are pre-installed on the analysis server, and an analysis report of the current file can be automatically created by selecting "PDF via LaTeX (.pdf)" via "Download as" in the "File" menu bar of the Jupyter analysis file.

[0101] In order to verify the technical solution provided by this embodiment, the present application performs corresponding processing according to the above method based on a specific sample, as shown in Figures 5 to 23, Figure 5 is a violin chart of the total number of gene detections before the sample cells are filtered, where the vertical axis indicates the total number of genes detected; Figure 6 is a violin chart of the total number of gene fragment detections before the sample cells are filtered, where the vertical axis indicates the total number of gene fragment detections; Figure 7 is a violin chart of the mitochondrial gene detection ratio before the sample cells are filtered, where the vertical axis indicates the cellular mitochondrial gene detection ratio; Figure 8 is a violin chart of the total number of gene detections after the sample cells are filtered, where the vertical axis indicates the total number of genes detected; Figure 9 is a violin chart of the total number of gene fragment detections after the sample cells are filtered, where the vertical axis indicates the total number of gene fragment detections; and Figure 10 is a violin chart of the mitochondrial gene detection ratio after the sample cells are filtered, where the vertical axis indicates the cellular mitochondrial gene detection ratio.

[0102] FIG. 11 is a cell cycle distribution map obtained by reducing the dimension of the data, where the horizontal axis indicates dimension 1 and the vertical axis indicates dimension 2. FIG. 12 is an annotation map of functional cell matrix mapping cells obtained by reducing the dimension of the data, where the horizontal axis indicates dimension 1 and the vertical axis indicates dimension 2. FIG. 13 is a pie chart of the composition ratio of functional cell matrix mapping annotation cells. FIG. 14 is a cell cluster map obtained by reducing the dimension of the data, where the horizontal axis indicates dimension 1 and the vertical axis indicates dimension 2. FIG. 15 is a bubble chart of cell clustering target gene expression, where the horizontal axis indicates cell Figure 16 is an annotation diagram of target gene expression corrected cells after dimensional reduction of the data, where the horizontal axis represents dimension 1 and the vertical axis represents dimension 2; Figure 17 is a cell trajectory diagram after dimensional reduction of the data; Figure 18 is a pie chart of the cell annotation composition ratio confirmed by the biological development trajectory; Figure 19 is a pie chart of the cell cycle ratio; Figure 20 is a bar chart of the cell cycle composition ratio, where the horizontal axis represents the cell cycle and the vertical axis represents the cell proportion; Figure 21 is a bar chart of the cell composition cycle ratio, where the horizontal axis represents the cell type and the vertical axis represents the cell proportion.

[0103] In addition, the single-cell transcriptome sequencing data of this example was rapidly annotated using the SingleR software based on the multicellular annotation database (Human Primary Cell Atlas Data and Blueprint Encode Data). As shown in Figures 22 and 23, Figure 22 is a cell annotation diagram in which the data was dimension-reduced based on the multicellular annotation database using SingleR software. Here, the horizontal axis represents dimension 1 and the vertical axis represents dimension 2. Figure 23 is a pie chart of the cell annotation composition ratio based on the multicellular annotation database using SingleR software. As can be seen from the results, the information annotated using this method is generally inaccurate. While the sample in this example is composed of neural progenitor cells and their differentiated neurons, the SingleR annotation results do not contain any neural progenitor cells at all, and instead show clusters annotated with other tissues and organ cells. In contrast, in the present invention, as shown in Figures 12 and 13, actual cell types can be quickly and accurately annotated. Furthermore, as shown in Figures 15 and 16, annotation accuracy is improved by multi-target gene verification. Multi-target gene confirmation makes the pericyte annotation ratio more accurate, correcting the bias of high-speed mapping annotation. Finally, cell trajectories are confirmed, and as shown in Figure 17, the progression from neural progenitor cells to immature neurons and then to terminally differentiated cells (pericytes, ependymal cells, glutamatergic neurons, and γ-aminobutyric acidergic neurons) conforms to the biological developmental trajectory, confirming accurate annotation.

[0104] As can be seen from the above, embodiments of the present invention have high throughput, high accuracy, high resolution, and high reproducibility, and are capable of detecting the cellular composition of a biological sample in one step at low cost, with high throughput and accurate quantification.

[0105] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts of each embodiment can be mutually referenced. As for the hardware disclosed in the embodiments, including the apparatus and electronic equipment, it corresponds to the method disclosed in the embodiments, so the description thereof will be simplified, and it is sufficient to refer to the method part for the relevant part.

[0106] It is obvious to those skilled in the art that the units and algorithm steps of each embodiment described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly explain the interchangeability of hardware and software, the configurations and steps of each embodiment have been generally described by function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to achieve the described functions for each specific application, but such embodiments should not be considered outside the scope of the present invention.

[0107] The method, device, electronic device, and readable storage medium for detecting cellular components of a biological sample provided by the present application have been described in detail above. Specific examples are used in this specification to explain the principles and embodiments of the present invention. However, the description of the examples is merely for understanding the method of the present invention and its core concept. It should be noted that a person skilled in the art can make some improvements and modifications to the present application without departing from the principles of the present invention, and these improvements and modifications are also included in the scope of the claims of the present application.

Claims

1. Obtaining single-cell sequencing results by performing single-cell transcriptome sequencing on the test biological sample; generating a cellular gene expression matrix by analyzing the single-cell sequencing results; and determining the cell types contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cellular gene expression matrix.

2. After performing single cell bioinformatics analysis on the cell gene expression matrix described above, The method for detecting the cellular composition of a biological sample according to claim 1, further comprising mapping functional cellular gene expression matrix information onto the cellular gene expression matrix to perform initial cell annotation.

3. After mapping the functional cell gene expression matrix information onto the cell gene expression matrix, generating and displaying a functional cell target gene expression level and cell expression ratio diagram in response to a drawing command for a functional cell target gene expression cluster diagram; The method for detecting the cellular composition of a biological sample described in claim 2, further comprising generating annotation confirmation information including the cell type to which each cell cluster belongs in response to an annotation confirmation result input command.

4. After performing single cell bioinformatics analysis on the cell gene expression matrix described above, generating gene splicing data by analyzing the single-cell sequencing results; 4. The method for detecting the cellular composition of a biological sample according to claim 1, further comprising: determining whether the annotated cell clusters fit a biological developmental trajectory by performing a trajectory analysis of the annotated cell clusters based on the gene splicing data.

5. After determining the cell type contained in the biological sample to be tested, The method for detecting the cellular composition of a biological sample according to any one of claims 1 to 4, further comprising calculating the cellular composition ratio and generating a detection result of the cellular composition ratio of the biological sample to be tested.

6. performing single-cell bioinformatics analysis on the cellular gene expression matrix to determine the cell types contained in the test biological sample; A method for detecting the cellular composition of a biological sample described in any one of claims 1 to 5, characterized in that it includes obtaining the cell types contained in the test biological sample by performing single-cell transcriptional data analysis on the cellular gene expression matrix in an interactive computing environment.

7. The method of determining the cell types contained in the test biological sample by performing single-cell bioinformatics analysis on the cell gene expression matrix described above includes: Invoke a cellular gene calculation equation to calculate the proportion of cellular mitochondrial genes in the cellular gene expression matrix, the total number of genes detected from the cells, the total number of gene fragments detected from the cells, the total number of fragments detected from the genes, and the total number of cells; In response to the filtering parameter setting command, the cell filtering relational formula and the gene filtering relational formula are respectively called, and the genes and cells whose detection quality does not satisfy the preset quality condition are filtered out, thereby obtaining the target cell gene data; In response to a data normalization command, perform a data normalization process on the target cellular genetic data, and perform a dimension reduction process on the normalized data; The method for detecting the cellular composition of a biological sample described in claim 6, characterized in that it includes performing a cell clustering process on the dimension-reduced data in response to a cell clustering process command to obtain cell sub-cluster information.

8. After obtaining the target cellular genetic data described above, The method for detecting the cellular composition of a biological sample according to claim 7, further comprising: invoking a cell cycle evaluation relational expression to determine the cell cycle of each cell in the target cell genetic data.

9. performing a data normalization process on the cellular gene expression matrix in response to the data normalization process command described above, calling a normalization relation to normalize the cellular gene expression matrix; Logarithmically transforming the standardized data by invoking the logarithmic transformation formula; 8. The method for detecting the cellular composition of a biological sample according to claim 7, further comprising: invoking an abnormal gene removal relational expression to remove abnormally highly expressed genes from the log-transformed data.

10. The test biological sample includes multiple batches of biological samples, and the single-cell sequencing results include multiple single-cell sequencing results carrying batch information. After determining the cell types contained in the test biological sample as described above, The method for detecting the cellular composition of a biological sample according to any one of claims 1 to 5, further comprising generating a stability result of the cellular composition ratio between batches by analyzing the cellular composition ratio data of each batch of biological sample.

11. The test subject biological sample is the same biological sample sampled at multiple time points, and the single cell sequencing results include single cell sequencing results of the same biological sample at multiple time points. After determining the cell types contained in the test subject biological sample as described above, For each biological sample at each time point, obtaining cell composition ratio data determined by the current biological sample; The method for detecting the cellular composition of a biological sample according to any one of claims 1 to 5, further comprising generating information on changes in the cellular composition ratio by performing time series analysis on data on the cellular composition ratio of the biological sample at different time points.

12. a sequencing module used to perform single-cell transcriptome sequencing on the test biological sample to obtain single-cell sequencing results; a data analysis module used to generate a cellular gene expression matrix by analyzing the single-cell sequencing results; and a cell type determination module used to determine the cell type contained in the test biological sample by performing single-cell bioinformatics analysis on the cellular gene expression matrix.

13. An electronic device comprising a processor and a memory, wherein the processor, when executing a computer program stored in the memory, implements the steps of the method for detecting the cellular composition of a biological sample according to any one of claims 1 to 11.

14. A readable storage medium having stored thereon a computer program, the computer program being characterized in that, when executed by a processor, the computer program implements the steps of the method for detecting the cellular composition of a biological sample according to any one of claims 1 to 11.