Biological sample cell composition detection method, device, equipment and storage medium

By using single-cell transcriptome sequencing and bioinformatics analysis, the low throughput and false positive problems of cell composition detection in biological samples in existing technologies have been solved, achieving low-cost, high-throughput, and accurate quantitative cell composition detection, and improving the stability and accuracy of the detection.

CN116825184BActive Publication Date: 2026-03-17SHANGHAI HUOJIANDE BIOPHARMACEUTICAL CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210982600.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2026-03-17
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Existing technologies are insufficient for low-cost, high-throughput, accurate, and quantitative detection of cellular composition in biological samples, especially when cell types are unknown or specific molecular markers are lacking. Conventional methods carry the risk of false positives or false negatives and are difficult to perform simultaneous analysis of multiple cell types.

Method used

By performing single-cell transcriptome sequencing on the biological samples to be tested, a cell gene expression matrix is ​​generated, and single-cell bioinformatics analysis is performed. Combined with functional cell gene expression matrix mapping and biological developmental trajectory analysis, cell types and composition ratios are determined.

Benefits of technology

It enables low-cost, high-throughput, accurate, and quantitative detection of cellular composition in biological samples, improving the stability and accuracy of detection, reducing the risk of false positives and false negatives, and supporting the reproducibility and stability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116825184B_ABST
    Figure CN116825184B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, electronic device, and readable storage medium for detecting the cellular composition of biological samples, applicable to the field of biomedical technology. The method includes performing single-cell transcriptome sequencing on the biological sample to be tested to obtain single-cell sequencing results; generating a cell gene expression matrix by analyzing the single-cell sequencing results; and determining the cell types contained in the biological sample by performing single-cell bioinformatics analysis on the cell gene expression matrix. This application enables low-cost, high-throughput, accurate, and quantitative one-step detection of the cellular composition of biological samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical technology, and in particular to a method, apparatus, electronic device, and readable storage medium for detecting the cell composition of biological samples. Background Technology

[0002] With the development of regenerative medicine technology, stem cells have promising applications in organoid disease model construction, novel drug mechanism research, and diseases requiring cell transplantation for which there are currently no specific drugs. Stem cells are a type of cell with self-renewal and differentiation capabilities. Embryonic stem cells and induced pluripotent stem cells are collectively referred to as pluripotent stem cells because they have the potential to differentiate into all types of cells in the human body. Cell products derived from stem cells used in the above fields require sensitive and accurate cell type identification to confirm that the cell products are mainly composed of cells of the intended differentiation, with stable and controllable cell composition, batch-to-batch homogeneity, and overall functional reproducibility. Especially in disease treatment, it is crucial to ensure that the cell differentiation direction does not deviate and that no unintended differentiated cells harmful to the human body are present, thus guaranteeing the safety and efficacy of the drugs.

[0003] Currently, commonly used methods for cell type identification include qPCR (Quantitative Real-time PCR), dPCR (Digital PCR), FISH (Fluorescence in situ hybridization), IF (Immunofluorescence technique), and FC (Flow Cytometry). qPCR and dPCR determine cell type by targeting and amplifying specific genes in functional cells to analyze their expression levels. FISH uses fluorescently labeled specific nucleic acid probes to hybridize with corresponding target DNA or RNA molecules within cells, and the cell type is determined by observing the fluorescence signal under a fluorescence microscope or confocal laser scanner. IF and FC both use fluorescently labeled specific antibodies to identify cell-specific antigens (molecular markers) through antigen-antibody reactions, thereby achieving qualitative or quantitative detection of cell type. All of these methods require specific primers or markers based on known cell type-specific genes or molecular markers. Specific antibodies must be designed, manufactured, or procured before detection of known cell types can be performed. When cell types cannot be predicted in advance, or when certain cell types lack highly specific molecular markers or antibodies, it is difficult to detect unknown cell types using the methods described above. Furthermore, these methods struggle to analyze multiple cell types simultaneously on the same sample, yielding very limited information and failing to achieve high-throughput cell identification and quantification. Moreover, since cell type determination relies on detecting a few specifically expressed gene or protein targets in functional cells, the limited number of targets, heterogeneity in target expression under different cell states, and non-specific expression can lead to false positives or false negatives, hindering in-depth exploratory analysis of the data.

[0004] In summary, all related technologies, whether protein-based or nucleic acid-based target detection methods, have drawbacks such as difficulty in analyzing unknown samples, high cost and low throughput of matrix screening detection, and significant dependence on limited molecular markers and antibody / probe specificity.

[0005] Therefore, how to achieve low-cost, high-throughput, and precise quantitative analysis of the composition of sample cells in one step is a technical problem that needs to be solved by professionals in this field. Summary of the Invention

[0006] This application provides a method, apparatus, electronic device, and readable storage medium for detecting the cell composition of biological samples, which can achieve low-cost, high-throughput, accurate, and quantitative one-step detection of the cell composition of biological samples.

[0007] To address the aforementioned technical problems, the embodiments of the present invention provide the following technical solutions:

[0008] One embodiment of the present invention provides a method for detecting the cellular composition of biological samples, comprising:

[0009] Single-cell transcriptome sequencing was performed on the biological sample to be tested to obtain single-cell sequencing results;

[0010] By analyzing the single-cell sequencing results, a cell gene expression matrix is ​​generated.

[0011] By performing single-cell bioinformatics analysis on the cell gene expression matrix, the cell types contained in the biological sample to be tested can be determined.

[0012] Optionally, after performing single-cell bioinformatics analysis on the cell gene expression matrix, the method further includes:

[0013] The functional cell gene expression matrix information is mapped to the cell gene expression matrix for preliminary cell annotation.

[0014] Optionally, after mapping the functional cell gene expression matrix information to the cell gene expression matrix, the method further includes:

[0015] In response to the command to draw a cluster map of functional cell target gene expression, generate and display a map of functional cell target gene expression levels and cell expression ratios;

[0016] In response to the annotation confirmation input command, generate annotation confirmation information containing the cell type to which each cell group belongs.

[0017] Optionally, after performing single-cell bioinformatics analysis on the cell gene expression matrix, the method further includes:

[0018] Gene splicing data is generated by analyzing the single-cell sequencing results;

[0019] Based on the gene splicing data, trajectory analysis was performed on the annotated cell groups to confirm whether the cell annotation groups fit biological developmental trajectories.

[0020] Optionally, after determining the cell types contained in the biological sample to be tested, the method further includes:

[0021] The cell composition ratio is statistically analyzed to generate the cell composition ratio detection result of the biological sample to be tested.

[0022] Optionally, the step of determining the cell types contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cell gene expression matrix includes:

[0023] In an interactive computing environment, single-cell transcription data analysis is performed on the cell gene expression matrix to obtain the cell types contained in the biological sample to be tested.

[0024] Optionally, the step of determining the cell types contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cell gene expression matrix includes:

[0025] Call the cell gene calculation formula to statistically analyze the proportion of mitochondrial genes in the cell gene expression matrix, the total number of genes detected in the cells, the total number of gene fragments detected in the cells, the total number of gene fragments detected, and the total number of cells detected.

[0026] In response to the filter parameter setting command, the cell filter relation and gene filter relation are called respectively to filter out genes and cells whose detection quality does not meet the preset quality conditions, and obtain the target cell gene data.

[0027] In response to the data normalization processing instruction, the target cell gene data is normalized, and the normalized data is then subjected to dimensionality reduction processing.

[0028] In response to the cell clustering instruction, the data that has been reduced in dimensionality is subjected to cell clustering to obtain cell group information.

[0029] Optionally, after obtaining the target cell gene data, the method further includes:

[0030] The cell cycle assessment relation is invoked to determine the cell cycle of each cell in the target cell gene data.

[0031] Optionally, the response data normalization processing instruction performs data normalization processing on the cell gene expression matrix, including:

[0032] The standardized relation is used to standardize the cell gene expression matrix;

[0033] Call the logarithmic transformation formula to perform a logarithmic transformation on the standardized data;

[0034] Use the abnormal gene removal relation to remove abnormally overexpressed genes from log-transformed data.

[0035] Optionally, the biological sample to be tested includes multiple batches of biological samples, and the single-cell sequencing results include multiple sets of single-cell sequencing results carrying batch information; after determining the cell types contained in the biological sample to be tested, the method further includes:

[0036] By analyzing the cell composition ratio data of each batch of biological samples, the stability results of cell composition ratio between batches are generated.

[0037] Optionally, the biological sample to be tested is sampled from multiple time points of the same biological sample, and the single-cell sequencing results include single-cell sequencing results from multiple time points of the same biological sample; after determining the cell types contained in the biological sample to be tested, the method further includes:

[0038] For each biological sample at a given time point, obtain the cell composition ratio data determined for the biological sample at that current time.

[0039] By performing time-series analysis on the cell composition ratio data of biological samples at different times, information on changes in cell composition ratio is generated.

[0040] Another embodiment of the present invention provides a biological sample cell composition detection device, comprising:

[0041] The sequencing module is used to perform single-cell transcriptome sequencing on the biological sample to be tested, and obtain single-cell sequencing results.

[0042] The data analysis module is used to generate a cell gene expression matrix by analyzing the single-cell sequencing results;

[0043] The cell type determination module is used to determine the cell type contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cell gene expression matrix.

[0044] This invention also provides an electronic device, including a processor, which executes a computer program stored in a memory to implement the steps of the biological sample cell composition detection method as described in any of the preceding claims.

[0045] Finally, this embodiment of the invention also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the biological sample cell composition detection method as described in any of the preceding claims.

[0046] The advantages of the technical solution provided in this application are that bioinformatics analysis based on single-cell sequencing results of the sample to be tested can deeply mine cellular characteristic information and present the gene expression and transcription levels of each cell at the time of detection at the single-cell level. This can effectively improve the stability of the entire biological sample detection, help avoid false positives or false negatives, and effectively improve detection accuracy. In addition, the entire analysis is carried out through a modular procedure, which does not require high operator experience, effectively reducing the detection cost of cell composition, and also meeting the requirements of detection stability and reproducibility. At the same time, the gradually accumulated detection data can be repeatedly mined and analyzed in conjunction with the iterative development of technology to further improve the detection accuracy of cell composition of biological samples.

[0047] Furthermore, embodiments of the present invention also provide corresponding implementation devices, electronic devices, and readable storage media for the detection method of cell composition of biological samples, further making the method more practical, and the devices, electronic devices, and readable storage media have corresponding advantages.

[0048] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this disclosure. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of the present invention or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 A flowchart illustrating a method for detecting the cellular composition of biological samples is shown.

[0051] Figure 2 A structural diagram showing a specific embodiment of a biological sample cell composition detection device is displayed.

[0052] Figure 3 A structural diagram of one specific embodiment of the electronic device is shown;

[0053] Figure 4 A flowchart illustrating another method for detecting the cellular composition of biological samples is shown.

[0054] Figure 5 A violin plot showing the total number of genes detected in the sample cells before filtering;

[0055] Figure 6 A violin plot showing the total number of gene fragments detected in the sample cells before filtering;

[0056] Figure 7 A violin plot showing the proportion of mitochondrial gene detection before filtration of sample cells;

[0057] Figure 8 A violin plot showing the total number of genes detected after filtering the sample cells;

[0058] Figure 9 A violin plot showing the total number of gene fragments detected after filtering the sample cells;

[0059] Figure 10 A violin plot showing the proportion of mitochondrial gene detection after cell filtration is displayed.

[0060] Figure 11 The data shows a dimensionality-reduced cell cycle distribution map;

[0061] Figure 12 The data dimensionality reduction function cell matrix mapping cell annotation diagram is displayed;

[0062] Figure 13 A pie chart showing the proportions of cell composition mapped to the functional cell matrix is ​​displayed.

[0063] Figure 14 The data shows a dimensionality-reduced cell clustering diagram;

[0064] Figure 15 A bubble diagram of gene expression in cell clusters is displayed;

[0065] Figure 16 This displays a cell annotation diagram showing the data dimensionality reduction and target gene expression correction.

[0066] Figure 17 The data is displayed as a dimensionality-reduced cell trajectory plot;

[0067] Figure 18 A pie chart showing the proportions of cell annotations confirming biological developmental trajectories;

[0068] Figure 19 A pie chart showing the cell cycle proportions is displayed;

[0069] Figure 20 A bar chart showing the proportions of cell cycle components is displayed;

[0070] Figure 21 A bar chart showing the proportion of cell cycle components is displayed;

[0071] Figure 22 The data dimensionality reduction is shown using a cell annotation graph from a multi-cell annotation database based on the SingleR software.

[0072] Figure 23 A pie chart showing the proportion of cell annotation components in a multi-cell annotation database based on the SingleR software is displayed. Detailed Implementation

[0073] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0074] The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.

[0075] The term "a method for detecting the cell composition of a biological sample" in the specification, claims, and accompanying drawings of this application refers to a method for determining the cell types contained in a biological sample and / or the cell composition ratio of the generated biological sample.

[0076] After introducing the technical solutions of the embodiments of the present invention, the various non-limiting embodiments of this application will be described in detail below.

[0077] First see Figure 1 , Figure 1 This is a flowchart illustrating a method for detecting the cellular composition of a biological sample according to an embodiment of the present invention. The embodiment of the present invention may include the following:

[0078] S101: Perform single-cell transcriptome sequencing on the biological sample to be tested to obtain single-cell sequencing results.

[0079] In this embodiment, the sample to be tested can be a cell dissociation sample or a tissue dissociation sample, neither of which affects the implementation of this application. By extracting, amplifying, and performing high-throughput sequencing analysis on the genome or transcriptome at the single-cell level, single-cell sequencing results are obtained. These results represent the genetic information of a single cell, or in other words, the unique gene structure and gene expression state of a single cell. This embodiment can employ any software capable of analyzing and comparing single-cell transcriptome sequencing data to obtain a cell gene expression matrix, such as Cellranger on the 10X GENOMICS platform, or software such as CeleScope or STARsolo, none of which affect the implementation of this application. Increasing the amount of single-cell sequencing data refers to sequencing technology that enables high-depth single-cell transcriptome sequencing (greater than 100G) of biological samples to be tested, in order to further improve single-cell sequencing performance. Here, G represents the size of the sequencing data, with 1G representing 1 billion base pairs. For example, 1 million cell samples or tissues can be dissociated into single cells and then the 10X GENOMICS sequencing platform can be used to perform 100G-depth single-cell transcriptome sequencing.

[0080] S102: Generate a cell gene expression matrix by analyzing single-cell sequencing results.

[0081] In this step, an alignment index can be constructed based on single-cell sequencing results by referencing human genome data such as the GRCh38 version and gene annotation databases such as GENECODE. This alignment index is used during sequence alignment, specifically in the process of converting sequence files into a cellular gene expression matrix. Further, a single-cell sequencing data expression matrix can be obtained by performing genome alignment and cell quantification on the single-cell transcriptome data. After processing and filtering the single-cell sequencing data expression matrix, the cellular gene expression matrix is ​​obtained.

[0082] S103: By performing single-cell bioinformatics analysis on the cell gene expression matrix, the cell types contained in the biological sample to be tested are determined.

[0083] In this step, to improve the accuracy and efficiency of cellular composition detection in biological samples, the single-cell bioinformatics analysis process can be executed in an interactive computing environment. By executing the entire data analysis process in an interactive computing environment, parameters can be adjusted in a timely manner based on the results of each step, achieving flexible processing. The single-cell bioinformatics analysis in this embodiment refers to performing a series of data analysis operations on the cell gene expression matrix using biological analysis methods, with the ultimate goal of obtaining the cell types contained in the biological sample to be tested. This step can be achieved by calling existing single-cell transcription data analysis functions such as Scanpy in an interactive computing environment to perform single-cell transcription data analysis on the cell gene expression matrix, thereby obtaining functional cell gene expression matrix information. Optionally, the cell gene expression matrix can be analyzed using Scanpy on the Python Jupyter Notebook visualization platform, or it can be analyzed using the Seurat software on the R language platform.

[0084] In the technical solution provided by this invention, bioinformatics analysis based on single-cell sequencing results of the sample to be tested can deeply mine cellular characteristic information. It can present the gene expression and transcription levels of each cell at the time of detection at the single-cell level, effectively improving the overall stability of biological sample testing and helping to avoid false positives or false negatives, thus significantly improving detection accuracy. Furthermore, the entire analysis is performed through a modular procedure, requiring less experience from operators, effectively reducing the cost of cellular composition detection, and meeting the requirements for detection stability and reproducibility. Simultaneously, the gradually accumulated detection data can be repeatedly mined and analyzed in conjunction with iterative technological developments, further improving the accuracy of cellular composition detection in biological samples.

[0085] To further improve detection accuracy and avoid false positives or false negatives caused by operators using a limited number of targets, leading to inaccurate final cell detection results, based on the above embodiments, this application can also use functional cell gene expression matrix mapping for rapid identification before determining the cell types contained in the biological sample to be tested in step S103, i.e., further improving detection accuracy. Based on the above embodiments, after performing single-cell bioinformatics analysis on the cell gene expression matrix, the functional cell gene expression matrix information can be mapped to the cell gene expression matrix for preliminary cell type annotation. After obtaining cell type annotation information, the cell composition ratio can be statistically analyzed based on the cell annotation information to generate the cell composition ratio detection result of the biological sample to be tested.

[0086] Single-cell transcriptome sequencing yields unlabeled data. To determine the cell type of the biological sample corresponding to the single-cell transcriptome sequencing data, a similarity comparison with existing labeled data can be performed to identify the cell type. In this step, the annotated functional cell gene expression matrix information can be mapped onto the cell gene expression matrix corresponding to the original single-cell sequencing results for rapid cell type identification. The functional cells used in the functional cell gene expression matrix information can be any type of cell selected by those skilled in the art for a more practical application scenario. The functional cell gene expression matrix information is obtained by analyzing the single-cell transcriptome sequencing and single-cell sequencing results. Software such as scGCN and singleR can be used to map the functional cell gene expression matrix information onto the cell gene expression matrix. Of course, other types of annotation techniques can also be used to map the functional cell gene expression matrix information onto the cell gene expression matrix.

[0087] After initial annotation of cell types, based on the initial annotation results, the proportion of each cell type in the biological sample to be tested is obtained by statistically analyzing the percentage of each cell type. The cell composition ratio detection results can be used to indicate the types of cells contained in the biological sample to be tested or directly represent the cell composition ratio data of the biological sample to be tested. The cell composition ratio detection results can be in any document format, such as a Word document, PDF text, or Excel spreadsheet. Of course, they can also be exported as images, audio, video, etc., without affecting the implementation of this application.

[0088] In the above embodiments, functional cell gene expression matrix mapping is used for rapid identification during bioinformatics analysis of single-cell sequencing data. To further improve the accuracy of cell type detection, based on the above embodiments, this embodiment, after performing the step "mapping the functional cell gene expression matrix information to the cell gene expression matrix for preliminary cell annotation" in the above embodiments, may further include:

[0089] In response to commands to draw cluster maps of functional cell target gene expression, generate and display maps showing the expression levels of functional cell target genes and the proportion of expression in different cells, such as bubble charts, scatter plots, trajectory plots, and heatmaps. In response to commands to input annotation confirmation results, generate annotation confirmation information including the cell type to which each cell group belongs.

[0090] In this embodiment, any drawing software can be used to draw the target gene expression cluster map of the functional cells used in the above embodiment. The accuracy of the cell type annotation in the above embodiment is confirmed based on the human experience target gene expression cluster map. Through the human-computer interaction module, the cell type of each cell group in the biological sample to be tested can be further input into the system. The system generates annotation confirmation information for the preliminary cell annotation based on the received information.

[0091] This embodiment, while using functional cell gene expression matrix mapping for rapid identification, also incorporates multi-target gene cross-validation to further confirm cell type, developmental state, and cell cycle characteristics. This avoids false positives or false negatives caused by the use of a limited number of targets in the aforementioned techniques, further improving the accuracy of cell composition detection in biological samples. Accordingly, the cell composition ratio detection results of the above embodiment are verified based on the annotation confirmation information. Optionally, the cell composition ratio can be recalculated based on the annotation confirmation information, and then the cell composition ratio detection results of the biological sample to be tested can be updated.

[0092] To achieve accurate qualitative and quantitative detection of cell types contained in a biological sample and quantitative detection of the cell composition ratio of the biological sample to be tested, based on the above embodiments, the method may further include:

[0093] Gene splicing data is generated by analyzing the single-cell sequencing results; based on the gene splicing data, trajectory analysis is performed on the annotated cell groups to confirm whether the annotated cell groups fit the biological developmental trajectory.

[0094] This embodiment can employ any RNA rate analysis technique to process the sequence alignment file, obtaining gene splicing data. This gene splicing data is used to analyze cell developmental trajectories. This step can utilize any biological developmental trajectory analysis software to perform trajectory analysis on the annotated cell groups, such as Scvelo, Monocle, or CellRank, to analyze the trajectory of the biological sample to be tested, obtaining the trajectory analysis verification results and determining whether they fit the biological developmental trajectory of the corresponding cell group.

[0095] Accordingly, in order to verify the accuracy of the cell composition ratio detection results of the above embodiments, the cell type annotation results can be further confirmed based on the target gene expression, i.e., the annotation confirmation information and trajectory analysis to verify the cell type annotation results, the cell composition ratio can be counted again, and finally the cell composition ratio detection results of the generated biological sample to be tested can be updated based on the currently counted cell composition ratio.

[0096] This embodiment uses trajectory analysis to confirm whether cell clusters fit biological developmental trajectories and counts the proportion of each cell type to achieve accurate qualitative and quantitative detection of sample cell types and quantitative detection of the cell composition ratio of the biological sample to be tested.

[0097] Based on the above embodiments, in order to further improve the richness of cell composition detection results and enhance the user experience, the cell composition detection results of the biological sample to be tested in this embodiment may include more data. Optionally, this embodiment may also include the following:

[0098] Based on the proportion information of each cell type, a plotting function is called to generate a pie chart of cell type distribution and a bar chart of cell cycle distribution for each cell type; based on the pie chart of cell type distribution and the bar chart of cell cycle distribution for each cell type, the cell composition ratio detection results are generated.

[0099] In this embodiment, the plotting function may include a plotting function for generating a pie chart of cell type distribution, and a plotting function for generating bar charts of cell cycle distribution for each cell type. Of course, in addition to the proportions of each cell type in the biological sample to be tested, other information in the cell composition ratio detection results, such as trajectory analysis verification results, expression levels of functional cell target genes and cell expression ratios, annotation confirmation information, cell type distribution pie charts, and bar charts of cell cycle distribution for each cell type, can be flexibly selected according to actual needs, and this does not affect the implementation of this application.

[0100] The above embodiments do not limit how to perform the step "determine the cell types contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cell gene expression matrix". This embodiment also provides an optional implementation method, which may include the following:

[0101] Call the cell gene calculation formula to statistically analyze the proportion of mitochondrial genes in the cell gene expression matrix, the total number of genes detected in the cells, the total number of gene fragments detected in the cells, the total number of gene fragments detected, and the total number of cells detected.

[0102] In response to the filter parameter setting command, the cell filter relation and gene filter relation are called respectively to filter out genes and cells whose detection quality does not meet the preset quality conditions, and obtain the target cell gene data.

[0103] Use the cell cycle assessment relation to determine the cell cycle of each cell in the target cell gene data;

[0104] In response to the data normalization processing instruction, the target cell gene data is normalized, and the normalized data is then subjected to dimensionality reduction processing.

[0105] In response to the cell clustering instruction, the data that has been reduced in dimensionality is subjected to cell clustering to obtain cell group information.

[0106] In this embodiment, any clustering algorithm can be used to perform cell clustering, such as using a clustering function to perform clustering directly. For example, clustering functions such as Leiden and Louvain can be used. This application does not impose any limitations on this.

[0107] The above embodiments do not limit how to perform data normalization. This embodiment also provides an optional implementation of data normalization, which may include the following:

[0108] The standardized relation is used to standardize the cell gene expression matrix;

[0109] Call the logarithmic transformation formula to perform a logarithmic transformation on the standardized data;

[0110] Use the abnormal gene removal relation to remove abnormally overexpressed genes from log-transformed data.

[0111] Based on the above embodiments, in order to further improve the richness of the cell type detection report and enhance the user experience, the detection result report of the biological sample to be tested in this embodiment may further include more data. Optionally, the cell composition ratio detection result may also include target cell gene data, dimensionality reduction data, cell classification information, etc., that is, the cell composition ratio detection result can be generated based on the biological developmental trajectory verification result, target cell gene data, dimensionality reduction data, cell classification information and the proportion of each type of cell.

[0112] Furthermore, to further explore the stability of cell composition ratios in different batches of biological samples or the changes in cell composition ratios in biological samples at different time points, based on the above embodiments, this embodiment can also perform integrated analysis to verify cell stability between batches, conduct time-series analysis on cells from the same batch at different culture times, and monitor changes in cell composition ratios. This embodiment may include the following:

[0113] The biological sample to be tested may include multiple batches of biological samples. Correspondingly, the single-cell sequencing results include multiple sets of single-cell sequencing results carrying batch information. After generating the test result report of the biological sample to be tested, the cell composition ratio data of each batch of biological samples is analyzed to generate the stability results of the cell composition ratio between batches, thereby realizing the verification of the stability of the cell composition ratio between batches by integrating the analysis of multiple batches of test samples.

[0114] For the same biological sample to be tested, the sample can be acquired at multiple time points. Single-cell sequencing results include single-cell sequencing results of the same biological sample at multiple time points. The biological sample collected at each time point is processed according to the method in the above embodiment to obtain the cell composition ratio data of the biological sample at the current time. By performing time-series analysis on the cell composition ratio data at different times, information on changes in cell composition ratio is generated. By conducting time-series analysis through multiple tests in the same batch, the changes in cell composition ratio can be monitored.

[0115] This embodiment, by verifying the stability of cells from multiple batches and combining the temporal analysis results of cells cultured at different times in the same biological sample, can further improve the accuracy of detecting the cell composition ratio of the biological sample to be tested, and can also further explore the stability of cell composition or changes in cell development process of the biological sample.

[0116] It should be noted that there is no strict order of execution for the steps in this application. As long as they conform to a logical order, these steps can be executed simultaneously or in a certain preset order. Figure 1 This is just an illustrative example and does not mean that this is the only possible execution order.

[0117] This invention also provides a corresponding device for detecting the cell composition of biological samples, further enhancing the practicality of the method. The device can be described from both a functional module perspective and a hardware perspective. The following describes the biological sample cell composition detection device provided by this invention; the biological sample cell composition detection device described below corresponds to the biological sample cell composition detection method described above.

[0118] From the perspective of functional modules, see Figure 2 , Figure 2 This is a structural diagram of a biological sample cell composition detection device 200 provided in an embodiment of the present invention, which may include:

[0119] Sequencing module 201 is used to perform single-cell transcriptome sequencing on the biological sample to be tested, and obtain single-cell sequencing results;

[0120] Data analysis module 202 is used to generate a cell gene expression matrix by analyzing single-cell sequencing results;

[0121] Cell type determination module 203 is used to determine the cell types contained in the biological sample to be tested by performing single-cell bioinformatics analysis on the cell gene expression matrix.

[0122] Optionally, in some embodiments of this example, the above-described apparatus may further include:

[0123] The annotation module is used to map the functional cell gene expression matrix information to the cell gene expression matrix for preliminary cell annotation.

[0124] As an optional implementation of the above embodiments, the device may further include an annotation confirmation module, which is used to generate and display a functional cell target gene expression level and cell expression ratio map in response to a functional cell target gene expression clustering map drawing instruction; and to generate annotation confirmation information containing the cell type to which each cell group belongs in response to an annotation confirmation result input instruction.

[0125] As another optional implementation of the above embodiments, the device may further include a trajectory analysis module, used to generate gene splicing data by analyzing single-cell sequencing results; and to perform trajectory analysis on the annotated cell groups based on the gene splicing data to confirm whether the cell annotation groups fit biological developmental trajectories.

[0126] As another optional implementation of the above embodiments, the above device may further include a result generation module for statistically analyzing cell composition ratios and generating cell composition ratio detection results for the biological sample to be tested.

[0127] Optionally, in some other embodiments of this example, the above-mentioned device may further include a verification module, which is used to generate and display a functional cell target gene expression level and cell expression ratio map in response to a functional cell target gene expression clustering map drawing instruction; and to generate cell type annotation accuracy information of annotated cell groups in response to an accuracy information input instruction.

[0128] Optionally, in some other embodiments of this example, the cell type determination module 203 may be further used to: call a single-cell transcription data analysis function to perform single-cell transcription data analysis on the cell gene expression matrix in an interactive computing environment to obtain functional cell gene expression matrix information.

[0129] As an optional implementation of the above embodiments, the cell type determination module 203 can also be used to: call the cell gene calculation formula to calculate the number of mitochondrial genes, cell genes, and the total number of gene fragments detected in each cell in the cell gene expression matrix; respond to the filter parameter setting instruction to call the cell filtering formula and the gene filtering formula respectively to filter out genes and cells whose detection quality does not meet the preset quality conditions, and obtain target cell gene data; call the cell cycle assessment formula to determine the cell cycle of each cell in the target cell gene data; respond to the data normalization processing instruction to perform data normalization processing on the target cell gene data, and perform dimensionality reduction processing on the normalized data; respond to the cell clustering processing instruction to perform cell clustering processing on the dimensionality reduction data, and obtain cell clustering information.

[0130] As another optional implementation of the above embodiments, the cell type determination module 203 can be further used to: call the standardization relation to standardize the cell gene expression matrix; call the logarithmic transformation relation to perform logarithmic transformation on the standardized data; and call the abnormal gene removal relation to remove abnormally high-expressed genes in the logarithmically transformed data.

[0131] Optionally, in some other embodiments of this example, the above-mentioned device may further include a result verification module, which may include:

[0132] The stability verification unit is used to generate inter-batch cell composition ratio stability results by analyzing the cell composition ratio data of each batch of biological samples; the biological samples to be tested include biological samples from multiple batches, and the single-cell sequencing results include multiple sets of single-cell sequencing results carrying batch information.

[0133] The change verification unit is used to acquire the cell composition ratio data of the biological sample at each time point. By performing time-series analysis on the cell composition ratio data of biological samples at different times, it generates information on changes in cell composition ratio. The biological sample to be tested is sampled from multiple time points of the same biological sample, and the single-cell sequencing results include single-cell sequencing results from multiple time points of the same biological sample.

[0134] The functions of each functional module of the biological sample cell composition ratio detection device described in the embodiments of the present invention can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above method embodiments, which will not be repeated here.

[0135] As can be seen from the above, the embodiments of the present invention can achieve low-cost, high-throughput, accurate, and quantitative one-step detection of the cellular composition of biological samples.

[0136] The biological sample cell composition detection device mentioned above is described from the perspective of functional modules. Furthermore, this application also provides an electronic device, which is described from the perspective of hardware. Figure 3 This is a schematic diagram of the structure of the electronic device provided in one embodiment of this application. For example... Figure 3 As shown, the electronic device includes a memory 30 for storing a computer program; and a processor 31 for executing the computer program to implement the steps of the biological sample cell composition detection method as described in any of the above embodiments.

[0137] The processor 31 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 31 may also be a controller, microcontroller, microprocessor, or other data processing chip. The processor 31 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 31 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 31 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 31 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0138] The memory 30 may include one or more computer-readable storage media, which may be non-transitory. The memory 30 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the memory 30 may be an internal storage unit of an electronic device, such as a server hard drive. In other embodiments, the memory 30 may be an external storage device of an electronic device, such as a plug-in hard drive on a server, a smart media card (SMC), a secure digital card (SD), a flash card, etc. Furthermore, the memory 30 may include both internal and external storage units of the electronic device. The memory 30 can be used not only to store application software and various types of data installed on the electronic device, such as code for programs executing the biological sample cell composition detection method, but also to temporarily store data that has been output or will be output. In this embodiment, the memory 30 is used to store at least the following computer program 301, which, after being loaded and executed by the processor 31, can implement the relevant steps of the biological sample cell composition detection method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 30 may also include an operating system 302 and data 303, and the storage method may be temporary storage or permanent storage. The operating system 302 may include Windows, Unix, Linux, etc. The data 303 may include, but is not limited to, data corresponding to the detection results of biological sample cell composition.

[0139] In some embodiments, the aforementioned electronic device may further include a display screen 32, an input / output interface 33, a communication interface 34 (or network interface), a power supply 35, and a communication bus 36. The display screen 32 and input / output interface 33, such as a keyboard, are user interfaces; optional user interfaces may also include standard wired interfaces, wireless interfaces, etc. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a display screen or display unit, used to display information processed in the electronic device and to display a visual user interface. The communication interface 34 may optionally include a wired interface and / or a wireless interface, such as a Wi-Fi interface, a Bluetooth interface, etc., typically used to establish communication connections between the electronic device and other electronic devices. The communication bus 36 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0140] Those skilled in the art will understand that Figure 3 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, such as sensors 37 that perform various functions.

[0141] The functions of each functional module of the electronic device described in the embodiments of the present invention can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above method embodiments, which will not be repeated here.

[0142] As can be seen from the above, the embodiments of the present invention can achieve low-cost, high-throughput, accurate, and quantitative one-step detection of the cellular composition of biological samples.

[0143] It is understood that if the biological sample cell composition detection method in the above embodiments is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, removable disk, CD-ROM, magnetic disk or optical disk, and other media capable of storing program code.

[0144] Based on this, embodiments of the present invention also provide a readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the biological sample cell composition detection method described in any of the above embodiments.

[0145] Finally, in order to make the technical solution of this application clearer to those skilled in the art, this application also incorporates... Figure 4 A schematic example is given, which may include the following:

[0146] One million cell samples or tissue samples were obtained from the biological sample to be tested and dissociated into single cells. The NPC cells were then subjected to 150G deep sequencing using a 10X GENOMICS sequencing platform to obtain single-cell test results. Bioinformatics analysis was performed on the single-cell sequencing results. The bioinformatics analysis steps may include the following:

[0147] A1: Obtain the FASTQ sequence data from the 10X GENOMICS sequencing platform. Use the mkref function (reference sequence alignment database construction function) of the Cell Ranger software for single-cell sequencing analysis, referencing the human GRCh38 (a version of the human reference genome) and the latest version of the GENECODE (gene annotation information database) database to construct an alignment index. Use the Cell Ranger software's count function and its default parameters to analyze and obtain the single-cell sequencing data expression matrix. Perform quality checks on the Cell Ranger software output data, i.e., the single-cell sequencing data expression matrix. If the quality check passes, generate the cell gene expression matrix and the sequence alignment BAM file. If the quality check fails, re-sequencing the biological sample on the 10X GENOMICS sequencing platform using single-cell transcriptome sequencing. Alternatively, it can be implemented, for example, by the following computer program: Cell rangercount --id=hNPC --fastqs=DATA_DIR --transcriptome=GRCh38-Genecode --localcores=16 --localmem=80.

[0148] A2: Use the run10x function of the Velocyto software to process the BAM file generated by Cell Ranger analysis, that is, analyze the cell gene splicing rate and obtain the cell gene splicing rate LOOM file. The LOOM file is a software generated by Velocyto to store the number of spliced / unspliced ​​transcripts, which is subsequently used to analyze cell developmental trajectory. Optionally, this can be achieved, for example, through the following computer program:

[0149] velocyto run10x-m human38_rmsk.gtf DATA_DIR genes.gtf

[0150] A3: The functions in the following steps can be pre-compiled directly on the Python platform for easy parameter adjustment and optimization. Alternatively, these functions can be embedded in the background for import. After pre-compiling the functions in the following steps, Scanpy is used to analyze single-cell transcriptome data based on the Python platform:

[0151] A31: Use the function `scvelo.read` to read the loom file generated by the Velocyto software, for example, it can be represented as `velo.loom`. Alternatively, this can be achieved, for example, through the following computer program:

[0152] ldata=scvelo.read(velo.loom,cache=True)

[0153] A32: Use the function `read_cellranger` to read the filtered cell gene expression matrix file processed by the Cell Ranger software, for example, represented as `filtered_feature_bc_matrix.h5`. Optionally, this can be implemented, for example, by the following computer program:

[0154] count_data=read_Cell ranger(filtered_feature_bc_matrix.h5, gex_only=True, add_sample_id=False, cache=False)

[0155] A33: Use the function `annotate_qc_metrics` to identify mitochondrial and ribosomal genes in cells and calculate their detection rates. Optionally, this can be achieved, for example, through the following computer program:

[0156] count_data=annotate_qc_metrics(count_data)

[0157] A34: Use the functions `filtering_cell` and `filtering_gene` to set specific filtering parameters for data quality control. These parameters can be modified based on the specific data analysis. In other words, `filtering_cell` and `filtering_gene` can be used to adjust the filtering parameters according to the data to filter out genes and cells with poor detection quality. Alternatively, this can be achieved, for example, through the following computer program:

[0158] count_data=filtering_cell(count_data, min_t_count=1500, max_t_count=150000, min_t_gene=1500, max_t_gene=10000, min_mito_percent=0, max_mito_percent=0.5)

[0159] count_data=filtering_gene(count_data,min_cells=10)

[0160] A35: Use the function `annotate_cellcycle` to assess and confirm the cell cycle. Optionally, this can be achieved, for example, through the following computer program:

[0161] annotate_cellcycle(count_data)

[0162] A36: The cell gene expression matrix is ​​standardized using the `scanpy.pp.normalize_total` function, the `scanpy.pp.log1p` function performs a logarithmic transformation on the standardized data, and then the `scanpy.pp.recipe_zheng17` function is used to calculate and analyze the data, removing abnormally overexpressed genes. Alternatively, this can be achieved, for example, using the following computer program:

[0163] scanpy.pp.normalize_total(count_data,exclude_highly_expressed=True,max_fraction=0.05,inplace=True)

[0164] scanpy.pp.log1p(count_data)

[0165] scanpy.pp.recipe_zheng17(count_data,n_top_genes=3000, log=False, plot=False, copy=False)

[0166] A37: Use the `scanpy.tl.pca` function to perform principal component analysis on the data obtained in the above steps, and use the `UMAP` or `tSNE` function to perform dimensionality reduction on the data obtained in the above steps. Optionally, this can be achieved, for example, through the following computer program:

[0167] scanpy.tl.pca(count_data)

[0168] scanpy.pl.pca_overview(count_data)

[0169] cpm_nml_umap=umap_.UMAP(n_neighbors=30, min_dist=0.9, n_components=2, random_state=42).fit_transform(count_data.obsm["X_pca"])

[0170] A38: Perform unsupervised clustering on the data obtained in the above steps using the `scanpy.pp.neighbors` function or the Leiden or Louvain algorithms, where the resolution can be adjusted according to the clustering. Alternatively, this can be achieved, for example, through the following computer program:

[0171] scanpy.pp.neighbors(count_data, n_neighbors=50, n_pcs=30, use_rep="X_pca", knn=True, random_state=0, method='umap', metric='euclidean')

[0172] scanpy.tl.leiden(count_data,resolution=2,key_added="leiden2")

[0173] A4: Use scGCN software to map the functional cell gene expression matrix information to detection data such as filtered_feature_bc_matrix.hd5 data to quickly identify cell types and statistically analyze the composition ratio of each type of cell.

[0174] A5: Draw a graph showing the expression levels of functional cell target genes and the proportion of cell expression to confirm the accuracy of cell type annotation.

[0175] A6: Use Scvelo software to perform trajectory analysis on the annotated cell groups to confirm whether the annotation information fits the biological developmental trajectory. Alternatively, this can be achieved, for example, using the following computer program:

[0176] adata_combine=scv.utils.merge(count_data,ldata)

[0177] scvelo.pp.filter_and_normalize(adata_combine,min_shared_counts=10,n_top_genes=5000)

[0178] scvelo.pp.moments(adata_combine,n_pcs=30,n_neighbors=30)

[0179] scvelo.tl.recover_dynamics(adata_combine,n_jobs=36)

[0180] scvelo.tl.velocity(adata_combine,mode='dynamical',n_jobs=36)

[0181] scvelo.tl.velocity_graph(adata_combine,n_jobs=36)

[0182] scvelo.pl.velocity_embedding_stream(adata_combine,basis='umap',color='cell_type_annot2',X=adata_combine.obsm["X_umap"])

[0183] A7: Calculate the proportion of each cell type, and use the custom functions `cell_proportion_pieplo` and `cell_proportion_barplot` to plot the cell type distribution as a pie chart and the cell cycle distribution as a bar chart for each cell type. Alternatively, this can be achieved, for example, through the following computer program:

[0184] cell_proportion_pieplot(count_data,x="cell_type_annot2",save="Cell_tp_prpie.pdf")

[0185] cell_proportion_barplot(count_data, x="cell_type_annot2", y="phase", color="colors_leiden_2, save="Cell_tp_prbar.pdf")

[0186] A8: The final click generates a PDF version of the analysis report without code.

[0187] In this step, you can pre-install pandoc, TeXLive, and nbextensions on the analysis server. In Jupyter's analysis file menu bar, select Download as PDF via LaTeX (.pdf) under File, and the analysis report for the current file will be automatically generated.

[0188] To verify the technical solution provided in this embodiment, this application processed a specific sample according to the above method, such as... Figures 5-23 As shown, Figure 5 A violin plot of the total number of gene tests before cell filtering, where the vertical axis represents the total number of gene tests. Figure 6A violin plot showing the total number of gene fragments detected before filtering the sample cells, where the vertical axis represents the total number of gene fragments detected. Figure 7 A violin plot showing the proportion of mitochondrial gene detection before cell filtration, where the vertical axis represents the proportion of mitochondrial gene detection in cells. Figure 8 A violin plot of the total number of gene tests after filtering sample cells, where the vertical axis represents the total number of gene tests; Figure 9 A violin plot showing the total number of gene fragments detected after filtering the sample cells, where the vertical axis represents the total number of gene fragments detected. Figure 10 A violin plot showing the proportion of mitochondrial gene detection after filtering sample cells, where the vertical axis represents the proportion of mitochondrial gene detection in cells.

[0189] Figure 11 A cell cycle distribution diagram for dimensionality reduction of the data, where the horizontal axis represents dimension 1 and the vertical axis represents dimension 2; Figure 12 The cell matrix is ​​used to map cell annotation diagrams for dimensionality reduction of data, where the horizontal axis represents dimension 1 and the vertical axis represents dimension 2. Figure 13 Annotate the cell composition ratio pie chart for the functional cell matrix mapping; Figure 14 Dimensionality reduction cell clustering diagram for the data, where the horizontal axis represents dimension 1 and the vertical axis represents dimension 2; Figure 15 This is a bubble chart of target gene expression for cell clustering, where the horizontal axis represents cell groups and the vertical axis represents genes. Figure 16 Cell annotation diagram for dimensionality reduction of target gene expression. The horizontal axis represents dimension 1 and the vertical axis represents dimension 2. Figure 17 Dimensionality reduction of cell trajectory plot; Figure 18 Pie chart to confirm the proportion of cell annotation components for biological developmental trajectories; Figure 19 A pie chart showing the proportions of the cell cycle; Figure 20 This is a bar chart showing the proportions of cell cycle components, where the horizontal axis represents the cell cycle and the vertical axis represents the proportion of cells. Figure 21 This is a bar chart showing the cell cycle proportions, where the horizontal axis represents cell type and the vertical axis represents cell proportion.

[0190] Simultaneously, the single-cell transcriptome sequencing data of this embodiment were rapidly annotated using the SingleR software based on a multi-cell annotation database (Human Primary Cell Atlas Data and Blueprint Encode Data), such as... Figures 22-23 As shown, Figure 22 Dimensionality reduction of data is achieved using a cell annotation graph from a multi-cell annotation database based on the SingleR software, where the horizontal axis represents dimension 1 and the vertical axis represents dimension 2. Figure 23This is a pie chart showing the proportion of cell annotations based on the SingleR software multi-cell annotation database. The results show that the information annotated by this method is largely inaccurate. The example sample consists of neural progenitor cells and their differentiated neurons, but the SingleR annotation results completely exclude neural progenitor cells and even include annotation groups for cells from other tissues and organs. In contrast, this invention can quickly and accurately annotate real cell types, such as... Figures 12-13 Presenting results; and improving annotation accuracy through multi-target gene validation, such as... Figures 15-16 The results were presented, and after confirmation with multiple target genes, the annotation of pericyte proportions was particularly accurate, correcting the annotation bias of the rapid mapping. Finally, cell trajectory confirmation was performed, such as... Figure 17 The results show that the development from neural progenitor cells to immature neurons and then to terminally differentiated cells (pericytes, ependymal cells, glutamatergic neurons, and GABAergic neurons) is consistent with the biological developmental trajectory, confirming the accuracy of the annotation.

[0191] As can be seen from the above, the embodiments of the present invention have high throughput, high accuracy, high resolution and high reproducibility, and can realize low cost, high throughput and accurate quantitative one-step detection of the cellular composition of biological samples.

[0192] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the hardware disclosed in the embodiments, including devices and electronic equipment, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0193] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0194] The foregoing provides a detailed description of a method, apparatus, electronic device, and readable storage medium for detecting the cell composition of biological samples. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are merely illustrative of the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from the principles of the invention, and these improvements and modifications also fall within the scope of protection of the claims.

Claims

1. A method for detecting the cellular composition of a biological sample, characterized in that, The method comprises the following steps: single-cell transcriptome sequencing is performed on a biological sample to be detected to obtain single-cell sequencing results; a cell gene expression matrix is generated by analyzing the single-cell sequencing results; the cell types contained in the biological sample to be detected are determined by performing single-cell bioinformatics analysis on the cell gene expression matrix; functional cell gene expression matrix information is mapped to the cell gene expression matrix for preliminary cell annotation; the functional cell gene expression matrix information is obtained by performing single-cell transcriptome sequencing and single-cell sequencing result analysis on functional cells of known cell types; trajectory analysis is performed on the annotated cell clusters based on gene splicing data to determine whether the cell annotation clusters fit the biological development trajectory; wherein the single-cell bioinformatics analysis on the cell gene expression matrix comprises: single-cell transcript data analysis is performed on the cell gene expression matrix in an interactive computing environment to obtain the cell types contained in the biological sample to be detected.

2. The method of claim 1, wherein the biological sample is selected from the group consisting of blood, plasma, serum, urine, saliva, and cerebrospinal fluid. After the functional cell gene expression matrix information is mapped to the cell gene expression matrix, the method further comprises: in response to a functional cell target gene expression clustering instruction, generating and displaying a functional cell target gene expression amount and cell expression proportion diagram; in response to an annotation confirmation result input instruction, generating annotation confirmation information containing the cell types of each cell cluster.

3. The method according to claim 1 or 2, wherein After the single-cell bioinformatics analysis on the cell gene expression matrix, the method further comprises: gene splicing data is generated by analyzing the single-cell sequencing results.

4. The method according to any one of claims 1 to 3, wherein After the cell types contained in the biological sample to be detected are determined, the method further comprises: the cell composition proportion of the biological sample to be detected is generated by counting the cell composition proportion.

5. The method of claim 1, wherein the biological sample is selected from the group consisting of blood, plasma, serum, urine, saliva, and cerebrospinal fluid. The determination of the cell types contained in the biological sample to be detected comprises: a cell gene calculation relationship is called to count the cell mitochondrial gene proportion, the total number of genes detected by the cell, the total number of gene fragments detected by the cell, the total number of fragments detected by the gene, and the total number of cells detected in the cell gene expression matrix; in response to a filtering parameter setting instruction, a cell filtering relationship and a gene filtering relationship are called respectively to filter out genes and cells whose detection quality does not meet the preset quality condition to obtain target cell gene data; in response to a data normalization processing instruction, the target cell gene data is subjected to data normalization processing, and the normalized data is subjected to dimension reduction processing; in response to a cell clustering processing instruction, the dimension reduction processing data is subjected to cell clustering processing to obtain cell clustering information.

6. The method of claim 4, wherein the biological sample is blood. After the target cell gene data is obtained, the method further comprises: a cell cycle evaluation relationship is called to determine the cell cycle of each cell in the target cell gene data.

7. The method according to claim 4, wherein The data normalization processing of the cell gene expression matrix in response to the data normalization processing instruction comprises: a standardization relationship is called to perform standardization processing on the cell gene expression matrix; a logarithmic conversion relationship is called to perform logarithmic conversion on the standardized processing data; Call the abnormal gene removal relationship formula to remove the abnormal high expression genes in the logarithmic conversion data.

8. The method according to any one of claims 1 to 5, wherein The to-be-detected biological sample includes multiple batches of biological samples, and the single-cell sequencing result includes multiple groups of single-cell sequencing results carrying batch information; after the cell types contained in the to-be-detected biological sample are determined, the method further includes: By analyzing the cell composition proportion data of each batch of biological samples, a batch-to-batch cell composition proportion stability result is generated.

9. The method according to any one of claims 1 to 5, wherein The to-be-detected biological sample is a same biological sample sampled at multiple time points, and the single-cell sequencing result includes single-cell sequencing results of the same biological sample at multiple time points; after the cell types contained in the to-be-detected biological sample are determined, the method further includes: For each time point of the biological sample, the cell composition proportion data of the current time biological sample is obtained; By performing time series analysis on the cell composition proportion data of the biological samples at different time points, cell composition proportion change information is generated.

10. A biological sample cell composition detection device, characterized by, The method comprises: a sequencing module configured to perform single-cell transcriptome sequencing on a to-be-detected biological sample to obtain a single-cell sequencing result; a data analysis module configured to generate a cell gene expression matrix by analyzing the single-cell sequencing result; a cell type determination module configured to determine cell types contained in the to-be-detected biological sample by performing single-cell bioinformatics analysis on the cell gene expression matrix; an annotation module configured to map functional cell gene expression matrix information to the cell gene expression matrix to perform preliminary cell annotation; the functional cell gene expression matrix information is obtained by performing single-cell transcriptome sequencing and single-cell sequencing result analysis on functional cells of known cell types; a trajectory analysis module configured to perform trajectory analysis on the annotated cell cluster based on gene splicing data to confirm whether the cell annotation cluster fits the biological development trajectory; The cell type determination module is further configured to perform single-cell transcriptome data analysis on the cell gene expression matrix in an interactive computing environment to obtain the cell types contained in the to-be-detected biological sample.

11. An electronic device, comprising: The method comprises a processor and a memory, and the processor is configured to implement the steps of the biological sample cell composition detection method according to any one of claims 1 to 9 when executing a computer program stored in the memory.

12. A readable storage medium, characterized by, The computer program stored on the readable storage medium is executed by the processor to implement the steps of the biological sample cell composition detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for analyzing cell track based on single-cell sequencing data and electronic equipment

    CN111951892A

  • Cell function annotation method, device, equipment and medium

    CN114496099A