Single-cell RNA sequencing data generation method and device, equipment and storage medium
Through the multi-task joint learning method of artificial intelligence, single-cell RNA sequencing data is generated using data perturbation and mask reconstruction, which solves the problem of low generation quality in the existing technology and achieves higher quality data generation and analysis effects.
Patent Information
- Application Number
- CN202510464271.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing single-cell RNA sequencing data generation methods are difficult to generate high-quality data. Traditional statistical methods assume that the distribution is not in line with reality, the stability of the generation adversarial network training is poor, and it is difficult to mine data feature information, resulting in unsatisfactory generation results and low generalization ability.
Using an artificial intelligence-based sequencing data generation model, target single-cell RNA sequencing data is generated through data perturbation and mask reconstruction, including comparison learning and mask learning.
The quality of the generated single-cell RNA sequencing data is improved, the overall statistical characteristics and global characteristics of the data are maintained, and the accuracy of cell clustering, rare cell recognition and batch effect correction is improved.
Smart Images

Figure CN120340618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and storage medium for generating single-cell RNA sequencing data. Background Art
[0002] Single-cell RNA sequencing (scRNA-seq) technology is a breakthrough in the field of bioinformatics. This technology can concurrently capture comprehensive genetic information within a single cell, providing higher-resolution data support for revealing cell heterogeneity and diversity.
[0003] However, due to the inherent characteristics of single-cell RNA sequencing, such as high dimensionality, high sparsity, high noise, non-linearity, and frequent missing events, the relationships between variables are complex and non-linear, making the precise clustering of cells and downstream analysis tasks complex and difficult. Summary of the Invention
[0004] The purpose of this application is to provide a method, device, equipment and storage medium for generating single-cell RNA sequencing data, aiming to improve the data quality of the generated single-cell RNA sequencing data.
[0005] An embodiment of this application provides a method for generating single-cell RNA sequencing data, including: Obtain the single-cell sequencing data to be processed; Input the single-cell sequencing data into a pre-trained sequencing data generation model to generate target single-cell RNA sequencing data; the sequencing data generation model is obtained through multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrastive learning and masked learning; The step of generating the target single-cell RNA sequencing data includes: Perform data perturbation on the single-cell sequencing data to obtain perturbed data; Perform masked reconstruction on the single-cell sequencing data to obtain reconstructed data; Concatenate the perturbed data and the reconstructed data to obtain the target single-cell sequencing data.
[0006] In one embodiment, performing data perturbation on the single-cell sequencing data includes: Randomly sort the gene expression values in the single-cell sequencing data to obtain re-sorted data; Randomly replace the gene expression values in the single-cell sequencing data with the gene expression values at the corresponding positions in the re-sorted data to obtain the perturbed data.
[0007] In one embodiment, the step of randomly replacing the gene expression value of each gene in the single-cell sequencing data with the gene expression value at the corresponding position in the randomly sorted sequencing data includes: Randomly mask the gene expression values in the single-cell sequencing data, and replace the masked part in the single-cell sequencing data with the gene expression value at the corresponding position in the re-sorted data to obtain the perturbed data.
[0008] In one embodiment, the step of performing mask reconstruction on the single-cell sequencing data includes: Randomly mask the gene expression values in the single-cell sequencing data, and predict the masked part in the single-cell sequencing data to obtain the reconstructed data.
[0009] In one embodiment, the step of splicing the perturbed data and the reconstructed data includes: Reduce the dimensions of the perturbed data and the reconstructed data and then splice them to obtain the spliced data; Convert the data form of the spliced data to be the same as the data form of the single-cell sequencing data to obtain the target single-cell sequencing data.
[0010] In one embodiment, the sequencing data generation model includes a data perturbation network obtained by contrastive learning based on the single-cell sample sequencing data. The contrastive learning based on the single-cell sample sequencing data includes: Randomly sort the gene expression values in the single-cell sample sequencing data of the target cell to obtain multiple sample re-sorted data; Randomly replace the gene expression values in the single-cell sample sequencing data of the target cell with the gene expression values at the corresponding positions in the sample re-sorted data to obtain the sample perturbed data; Perform contrastive learning on the similarity distance between the single-cell sample sequencing data of the target cell and the sample perturbed data and the similarity distance between the single-cell sample sequencing data of the target cell and the single-cell sample sequencing data of non-target cells to obtain the data perturbation network.
[0011] In one embodiment, the sequencing data generation model includes a mask reconstruction network obtained by mask learning based on the single-cell sample sequencing data. The mask learning based on the single-cell sample sequencing data includes: Randomly mask the gene expression values in the single-cell sample sequencing data, and predict the masked part in the single-cell sample sequencing data to obtain the sample reconstructed data; Perform mask learning on the similarity distance between the single-cell sample sequencing data and the sample reconstructed data to obtain the mask reconstruction network.
[0012] An embodiment of the present application also provides a single-cell RNA sequencing data generation device, including: A first module, configured to obtain single-cell sequencing data to be processed; A second module, configured to input the single-cell sequencing data into a pre-trained sequencing data generation model to generate target single-cell RNA sequencing data; the sequencing data generation model is obtained by performing multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrast learning and masked learning; The steps of generating the target single-cell RNA sequencing data include: Performing data perturbation on the single-cell sequencing data to obtain perturbed data; Performing masked reconstruction on the single-cell sequencing data to obtain reconstructed data; Concatenating the perturbed data and the reconstructed data to obtain the target single-cell sequencing data.
[0013] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned single-cell RNA sequencing data generation method is implemented.
[0014] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned single-cell RNA sequencing data generation method is implemented.
[0015] The beneficial effects of the present application: After inputting the single-cell sequencing data to be processed into the pre-trained sequencing data generation model, the pre-trained sequencing data generation model performs data perturbation and masked reconstruction on the single-cell sequencing data to be processed, and then concatenates the perturbed data obtained by data perturbation and the masked data obtained by masked reconstruction to generate target single-cell sequencing data. Since the sequencing data generation model is obtained by performing contrast learning and masked learning based on single-cell sample sequencing data, the generated target single-cell sequencing data not only maintains the overall statistical characteristics of the single-cell sequencing data but also maintains the global features and local dependence features of the single-cell sequencing data, effectively improving the data quality of the generated target single-cell RNA sequencing data. Description of the Drawings
[0016] Figure 1 is a flowchart of the single-cell RNA sequencing data generation method provided by an embodiment of the present application.
[0017] Figure 2 is a schematic diagram of the performance evaluation of the single-cell RNA sequencing data generation method provided by an embodiment of the present application in terms of cell clustering.
[0018] Figure 3 It is a schematic diagram for evaluating the performance of the single-cell RNA sequencing data generation method provided by the embodiments of the present application in identifying rare cell types.
[0019] Figure 4 It is a schematic diagram for evaluating the performance of the single-cell RNA sequencing data generation method provided by the embodiments of the present application in batch effect correction.
[0020] Figure 5 It is a schematic diagram for evaluating the performance of the single-cell RNA sequencing data generation method provided by the embodiments of the present application in single-cell RNA sequencing data analysis.
[0021] Figure 6 It is a schematic structural diagram of the single-cell RNA sequencing data generation device provided by the embodiments of the present application.
[0022] Figure 7 It is a schematic hardware structure diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0024] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown can be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the description, claims and drawings are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.
[0026] First, several terms involved in the present application are analyzed: Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information processes of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0027] Ribonucleic Acid (RNA) is a carrier of genetic information that exists in biological cells, as well as in some viruses and viroids. RNA is a long-chain molecule formed by the condensation of ribonucleotides through phosphodiester bonds. A ribonucleotide molecule consists of a phosphate, a ribose, and a base. The bases of RNA are mainly of 4 types, namely A (adenine), G (guanine), C (cytosine), and U (uracil). Among them, U (uracil) replaces T (thymine) in DNA. The main role of ribonucleic acid in the body is to guide protein synthesis. A single human cell contains approximately 10 pg of RNA (and approximately 7 pg of DNA). Compared with DNA, RNA is diverse in types, has a smaller molecular weight, and shows large variations in content. RNA can be classified into messenger RNA and non-coding RNA according to different structures and functions. Non-coding RNA is divided into non-coding large RNA and non-coding small RNA. Non-coding large RNA includes ribosomal RNA and long non-coding RNA. Non-coding small RNA includes transfer RNA, ribozymes, small molecule RNA, etc. Small molecule RNA (20 - 300 nt) includes miRNA, SiRNA, piRNA, scRNA, snRNA, snoRNA, etc. Bacteria also have small molecule RNA (50 - 500 nt).
[0028] Polymerase Chain Reaction (PCR) is a molecular biology technique used to amplify specific DNA fragments. It can be regarded as a special DNA replication outside the body. The biggest feature of PCR is that it can greatly increase trace amounts of DNA. PCR utilizes the fact that DNA denatures into single strands at a high temperature of 95°C in vitro, and at a low temperature (usually around 60°C), primers bind to the single strands according to the principle of base complementary pairing. Then, the temperature is adjusted to the optimal reaction temperature of DNA polymerase (around 72°C), and DNA polymerase synthesizes complementary strands along the direction from phosphate to pentose (5'-3').
[0029] Currently, a common problem in biomedical research is that due to the lack of available biological samples, high costs, or ethical issues, the research samples are too few, resulting in unstable analysis results and even situations where other researchers cannot reproduce the results. Therefore, using generative methods to augment single-cell RNA sequencing data can well solve this series of problems and at the same time offer great hope for the real generation and enhancement of other biomedical data types.
[0030] Compared with other data, single-cell RNA sequencing data has unique properties. The target object obtains single cells from biological tissues, extracts RNA from them, reverse transcribes it into cDNA, amplifies the data using PCR, and finally obtains sequencing data using a high-throughput sequencer. After aligning with the known genome, a count matrix can be obtained. However, the single-cell RNA sequencing data obtained in this way has disadvantages such as a large data dimension, a large number of missing values, and overdispersion, which increase the difficulty in subsequent application processes.
[0031] Currently, many generative methods have emerged for the characteristics of single-cell RNA sequencing data. Among them, scDesign2 classifies different cells according to judgment conditions and assumes corresponding distributions, such as Poisson distribution, negative binomial distribution, zero-inflated negative binomial distribution, etc. It estimates the parameters in the distribution by maximizing the likelihood function, generates random numbers corresponding to the distribution, and completes data generation; scGAN makes a unified restriction on the total expression of cell data and uses a fully connected generative adversarial network to generate corresponding single-cell RNA sequencing data.
[0032] However, current single-cell RNA sequencing data generation methods all have significant defects. Traditional statistical methods require specific distribution assumptions for cell expression, but often the expression distributions of different cell types are different, and the assumed distributions are difficult to match the actual situation, so they cannot generate high-quality data; for generative methods based on generative adversarial networks, due to the challenging training stability of the generative adversarial network model, it is difficult to find a way to achieve Nash equilibrium during the adversarial training process, and it is easy to have a situation where the loss function oscillates continuously and is difficult to converge. Even when the loss function converges, there may be a mode collapse situation, making it difficult to complete the mining of the characteristic information of single-cell RNA sequencing data, thus unable to correctly grasp the true situation of the data distribution. The generation effect is often not ideal, and the generalization ability is not high, and it cannot efficiently and accurately complete the generation task.
[0033] Based on this, the embodiments of the present application provide a single-cell RNA sequencing data generation method, device, equipment, and storage medium, aiming to improve the data quality of the generated single-cell RNA sequencing data.
[0034] The single-cell RNA sequencing data generation method, apparatus, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the single-cell RNA sequencing data generation method in the embodiments of the present application is described.
[0035] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0036] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] The single-cell RNA sequencing data generation method provided by the embodiments of the present application relates to biometric technology in the field of artificial intelligence technology. The single-cell RNA sequencing data generation method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the single-cell RNA sequencing data generation method, etc., but is not limited to the above forms.
[0038] This application can be used in many computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0039] Figure 1 It is a flowchart of the single-cell RNA sequencing data generation method provided by an embodiment of this application. Figure 1 The method in may include but is not limited to steps S110 to S120.
[0040] Step S110, obtain the single-cell sequencing data to be processed.
[0041] Specifically, the single-cell sequencing data to be processed can be a single-cell RNA sequencing data extracted from an scRNA-seq dataset. Good clustering effects on the scRNA-seq dataset are of great guiding significance to bioinformatics researchers, not only in discovering new cell subtypes, but also in the prevention and treatment of cancer and other diseases.
[0042] The single-cell sequencing data to be processed is presented in the form of a matrix. By extracting features from the original single-cell sequencing data, the single-cell sequencing data to be processed in the form of a matrix is constructed, where the rows of the matrix represent gene features and the columns of the matrix represent cell samples.
[0043] Step S120, input the single-cell sequencing data into a pre-trained sequencing data generation model to generate the target single-cell RNA sequencing data.
[0044] Among them, the sequencing data generation model is obtained through multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrast learning and masked learning.
[0045] Specifically, the RNA sequencing data generation model is mainly composed of a data perturbation network and a masked reconstruction network. Among them, the data perturbation network is obtained through contrastive learning based on single-cell sample sequencing data, and the masked reconstruction network is obtained through masked learning based on single-cell sample sequencing data. Based on the target single-cell RNA sequencing data generated by the sequencing data generation model, the global structure and local dependence relationship of the single-cell sequencing data are used to mine the characteristic information and distribution of cell expression, and the generation task is completed efficiently and robustly. In this embodiment, after the single-cell sequencing data is input into the RNA sequencing data generation model, the single-cell sequencing data is respectively subjected to data perturbation and masked reconstruction in the RNA sequencing data generation model, and the obtained perturbed data and reconstructed data are spliced and output as the target single-cell sequencing data.
[0046] In the embodiment of the present application, single-cell RNA sequencing is performed in batches using single-cell sequencing technology to obtain multiple single-cell sample sequencing data. The single-cell sequencing technology can adopt existing 10X genomics Chromium technology and Smart-Seq2 technology. The 10X genomics Chromium technology is based on the 10X Genomics platform, which can isolate and label 5000 - 10000 single cells at one time and can perform detection at the single-cell level. This method is based on microfluidics-based approaches, has similar molecular biology principles to the Smart-Seq technology, and uses template conversion technology, but has different cell capture and throughput from the Smart-Seq technology. The droplet-based method encapsulates a single cell in a small oil droplet (containing barcode and RT primer) for reverse transcription into cDNA, and then the oil droplet breaks to release the cDNA for unified library construction, increasing the experimental throughput, but requires specialized experimental equipment. The Smart-Seq2 technology is an improvement based on the Smart-Seq technology. The Smart-Seq technology is based on high-fidelity reverse transcriptase, template conversion, and pre-amplification to increase the cDNA yield and thus obtain full-length transcripts.
[0047] Specifically, step S120 includes the following steps S121 to S123.
[0048] Step S121: Perturb the single-cell sequencing data to obtain perturbed data.
[0049] Specifically, the single-cell sequencing data is input into the data perturbation network of the sequencing data generation model, and the data perturbation network is used to slightly adjust the single-cell sequencing data and generate perturbed data. Since the data perturbation network is obtained by contrastive learning based on single-cell sample sequencing data, the data perturbation network can distinguish the single-cell sequencing data of different cells and the perturbed data obtained by perturbing the single-cell sequencing data of the same cell, so that the generated perturbed data maintains the overall statistical characteristics of the single-cell sequencing data.
[0050] In one embodiment, data perturbation of the single-cell sequencing data includes: randomly sorting the gene expression values in the single-cell sequencing data to obtain re-sorted data; randomly replacing the gene expression values in the single-cell sequencing data with the gene expression values at the corresponding positions in the re-sorted data to obtain perturbed data.
[0051] Randomly sorting the gene expression values in the single-cell sequencing data can be to randomly rearrange the gene expression values in each column of the single-cell sequencing data (i.e., the expression values of each gene) to generate a data in the form of a randomly sorted matrix, which is the re-sorted data. During the random sorting process, the distribution of the gene expression values in the single-cell sequencing data is retained, but the original relationship between cells is disrupted. Randomly replacing the gene expression values in the single-cell sequencing data with the gene expression values at the corresponding positions in the re-sorted data can be to randomly mark the gene expression values in the single-cell sequencing data, and then replace the marked gene expression values in the single-cell sequencing data with the gene expression values at the same positions in the re-sorted data, thereby generating perturbed data.
[0052] In one embodiment, randomly replacing the gene expression value of each gene in the single-cell sequencing data with the gene expression value at the corresponding position in the randomly sorted sequencing data includes: randomly masking the gene expression values in the single-cell sequencing data, and replacing the masked part in the single-cell sequencing data with the gene expression value at the corresponding position in the re-sorted data to obtain perturbed data.
[0053] Randomly masking the gene expression values in the single-cell sequencing data can be to generate a mask matrix according to the Bernoulli distribution. The elements of the mask matrix indicate whether to mask the gene expression values in the single-cell sequencing data, and the gene expression values in the single-cell sequencing data are randomly masked according to the mask matrix. Replacing the masked part in the single-cell sequencing data with the gene expression value at the corresponding position in the re-sorted data can be to replace the masked gene expression value in the single-cell sequencing data with the gene expression value at the same position as the masked gene expression value in the re-sorted data. For example, the third gene expression value in the first column of the single-cell sequencing data is masked and replaced with the third gene expression value in the first column of the re-sorted data. In this way, after traversing all the gene expression values in the single-cell sequencing data, perturbed data is generated.
[0054] The expression of the mask matrix is: , where indicates whether to mask the gene expression value of the th column in the single-cell sequencing data for the th gene, indicates that the gene expression value is masked, indicates that the gene expression value is not masked, indicates the masking probability, indicates the Bernoulli distribution.
[0055] The expression for replacing the masked gene expression values in the single-cell sequencing data is: , where is the gene expression value of the th column in the perturbed data for the th gene, is the gene expression value of the th column in the single-cell sequencing data for the th gene, is the gene expression value of the th column in the re-ordered data for the th gene.
[0056] Step S122, perform mask reconstruction on the single-cell sequencing data to obtain reconstructed data.
[0057] Specifically, input the single-cell sequencing data into the mask reconstruction network of the sequencing data generation model, and use the mask reconstruction network to perform mask prediction and data reconstruction on the single-cell sequencing data to obtain reconstructed data. Since the mask reconstruction network is obtained by mask learning based on single-cell sample sequencing data, the mask reconstruction network can learn the global structure and local dependence relationship of the single-cell sequencing data, so that the generated reconstructed data maintains the global characteristics and local dependence characteristics of the single-cell sequencing data.
[0058] In one embodiment, performing mask reconstruction on the single-cell sequencing data includes: randomly masking the gene expression values in the single-cell sequencing data, and predicting the masked part in the single-cell sequencing data to obtain reconstructed data.
[0059] Randomly mask the gene expression values in single-cell sequencing data, which can be to generate a mask matrix. The elements of the mask matrix indicate whether to mask the gene expression values in single-cell sequencing data, and randomly mask the gene expression values in single-cell sequencing data according to the mask matrix. Predict the masked part in single-cell sequencing data, which can be to first encode the masked data to obtain a low-dimensional masked data representation, then use a pre-trained mask predictor to perform mask prediction on the low-dimensional masked data representation to obtain mask prediction data, and finally decode the mask prediction data to obtain reconstructed data.
[0060] Step S123, splice the perturbed data and the reconstructed data to obtain the target single-cell sequencing data.
[0061] Specifically, process the perturbed data and the reconstructed data into the same data form and then splice them, and process the spliced data into the required data form to obtain the target single-cell sequencing data.
[0062] In one embodiment, splicing the perturbed data and the reconstructed data includes: performing dimensionality reduction mapping on the perturbed data and the reconstructed data and then splicing them to obtain spliced data; converting the data form of the spliced data to be the same as the data form of the single-cell sequencing data to obtain the target single-cell sequencing data.
[0063] Performing dimensionality reduction mapping on the perturbed data and the reconstructed data and then splicing them can be to respectively encode the perturbed data and the reconstructed data and input the encoded data into a mapping head to map them to a low-dimensional latent space, and then splice the mapped data of the perturbed data and the reconstructed data to obtain spliced data. Converting the data form of the spliced data to be the same as the data form of the single-cell sequencing data can be to input the spliced data into a decoder for decoding, so that the data form of the spliced data is converted to be the same as the data form of the single-cell sequencing data to obtain the target single-cell sequencing data.
[0064] In one embodiment, contrastive learning based on single-cell sample sequencing data includes: randomly sorting the gene expression values in the single-cell sample sequencing data of the target cell to obtain multiple sample reordering data; randomly replacing the gene expression values in the single-cell sample sequencing data of the target cell with the gene expression values at the corresponding positions in the sample reordering data to obtain sample perturbed data; performing contrastive learning on the similarity distance between the single-cell sample sequencing data of the target cell and the sample perturbed data and the similarity distance between the single-cell sample sequencing data of the target cell and the single-cell sample sequencing data of non-target cells to obtain a data perturbation network.
[0065] Based on the above embodiments, optionally, the calculation formula of the contrastive learning loss is: , where, is the contrast learning loss, is the feature representation of the i-th sample perturbation dataset of the target cell, is the feature representation of the single-cell sample sequencing data of the target cell, is the feature representation of the sample perturbation data, is the feature representation of the single-cell sample sequencing data of non-target cells, is the temperature parameter, where i, j, and k are all positive integers, and N is the total number of sample perturbation data.
[0066] In one embodiment, mask learning is performed based on single-cell sample sequencing data, including: randomly masking the gene expression values in the single-cell sample sequencing data, predicting the masked part in the single-cell sample sequencing data to obtain sample reconstruction data; performing mask learning on the similarity distance between the single-cell sample sequencing data and the sample reconstruction data to obtain a mask reconstruction network.
[0067] The single-cell RNA sequencing data generation method provided by the embodiments of the present application inputs the to-be-processed single-cell sequencing data into a pre-trained sequencing data generation model. Then, the pre-trained sequencing data generation model performs data perturbation and mask reconstruction on the to-be-processed single-cell sequencing data, and then splices the perturbation data obtained by data perturbation and the mask data obtained by mask reconstruction to generate target single-cell sequencing data. Since the sequencing data generation model is obtained by performing contrast learning and mask learning based on single-cell sample sequencing data, the generated target single-cell sequencing data not only maintains the overall statistical characteristics of the single-cell sequencing data but also maintains the global features and local dependence features of the single-cell sequencing data, effectively improving the data quality of the generated target single-cell RNA sequencing data.
[0068] Refer to Figure 2, to verify the performance of the single-cell RNA sequencing data generation method (scCMA) provided in the embodiments of the present application in cell clustering, a comparative experiment was conducted on it and seven existing single-cell clustering methods (including scGNN, scDeepCluster, scGCNClustering, scDCC, DCA+K-means, scNAME, and scMAE) on 13 real scRNA-seq datasets from public platforms. The adjusted Rand index (ARI) and normalized mutual information (NMI) were used as the main evaluation indicators to effectively evaluate the actual performance of various clustering algorithms. Among them, ARI and NMI evaluate the alignment between the derived clustering labels and the actual cell labels by providing quantitative measurements. The higher the scores of the two indicators, the more accurate the clustering results. Finally, the sequencing data generation model (scCMA) provided in the embodiments of the present application obtained the highest average ranking of ARI and NMI in 13 datasets, and obtained the highest NMI and ARI scores in 11 and 10 datasets respectively. The experimental results prove the excellent performance of the sequencing data generation model (scCMA) provided in the embodiments of the present application.
[0069] See Figure 3 , in terms of rare cell type identification, the recall rate, precision rate, and F1 score, which are widely used metrics, were used to evaluate the rare cell type identification capabilities of the above-mentioned various models. It can be seen that the single-cell RNA sequencing data generation method (scCMA) provided in the embodiments of the present application performs relatively prominently in terms of precision rate and F1 score. The experimental results show that the single-cell RNA sequencing data generation method (scCMA) provided in the embodiments of the present application has advantages in accuracy and comprehensive performance, and can achieve the suppression of technical noise without losing rare cell signals.
[0070] See Figure 4, in terms of batch effect correction, a dataset containing batch effects, i.e., the Mouse1 dataset, was used. In addition to using common metrics such as ARI, NMI, and Cell-type ASW to evaluate the clustering results, the ASW score of the batch (Batch ASW) was also calculated to evaluate the degree of batch effect removal. Ultimately, the higher the Batch ASW score, the better the mixing effect, thereby reflecting the improvement of batch removal effect. The experimental results show that the single-cell RNA sequencing data generation method (scCMA) provided by the embodiments of the present application achieved the highest scores in both the Cell-type ASW and Batch ASW metrics. It is worth noting that the single-cell RNA sequencing data generation method provided by the embodiments of the present application obtained a high score of 0.998 in Batch ASW, indicating that when dealing with data with batch effects, the single-cell RNA sequencing data generation method provided by the embodiments of the present application can achieve satisfactory results.
[0071] Refer to Figure 5 , in the analysis of single-cell RNA sequencing (scRNA-seq) data, dimensionality reduction and visualization are the core steps to reveal cell heterogeneity, identify cell types and states. The clustering results of the single-cell RNA sequencing data generation method (scCMA) provided by the embodiments of the present application were compared with three widely used dimensionality reduction methods, PCA, t-SNE, and UMAP, in the Human1 and Human_kidney datasets, and the results were visually represented. It can be seen that the single-cell RNA sequencing data generation method (scCMA) provided by the embodiments of the present application can accurately distinguish different cell subsets while obtaining the highest ASW score, which is superior to PCA, t-SNE, and UMAP.
[0072] Please refer to Figure 6 , the embodiments of the present application also provide a single-cell RNA sequencing data generation device, which can implement the above single-cell RNA sequencing data generation method. The device includes: The first module 601 is used to obtain the single-cell sequencing data to be processed; The second module 602 is used to input the single-cell sequencing data into a pre-trained sequencing data generation model to generate target single-cell RNA sequencing data; the sequencing data generation model is obtained by multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrast learning and masked learning; The steps of generating target single-cell RNA sequencing data include: Perform data perturbation on the single-cell sequencing data to obtain perturbed data; Mask the single-cell sequencing data for reconstruction to obtain reconstructed data; Concatenate the perturbed data and the reconstructed data to obtain the target single-cell sequencing data.
[0073] The specific implementation of the single-cell RNA sequencing data generation device is basically the same as the specific embodiments of the above single-cell RNA sequencing data generation method, and will not be elaborated here.
[0074] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment.
[0075] Next, refer to Figure 7 to describe the electronic device 700 according to this embodiment of the present disclosure. Figure 7 The electronic device 700 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0076] As Figure 7 shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include but are not limited to: at least one processing unit 710, at least one storage unit 720, a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710), a display unit 740, etc.
[0077] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the above single-cell RNA sequencing data generation method section of this specification.
[0078] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 7201 and / or a cache storage unit 7202, and may further include a read-only storage unit (ROM) 7203.
[0079] The storage unit 720 may also include a program / utilities 7204 having a set (at least one) of program modules 7205. Such program modules 7205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.
[0080] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0081] The electronic device 700 can also communicate with one or more external devices 700' (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. The network adapter 760 can communicate with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0082] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned single-cell RNA sequencing data generation method is implemented.
[0083] For the single-cell RNA sequencing data generation method, device, equipment and storage medium provided by the embodiments of the present application, after the to-be-processed single-cell sequencing data is input into the pre-trained sequencing data generation model, the pre-trained sequencing data generation model performs data perturbation and mask reconstruction on the to-be-processed single-cell sequencing data, and then splices the perturbed data obtained by data perturbation and the mask data obtained by mask reconstruction to generate the target single-cell sequencing data. Since the sequencing data generation model is obtained by contrast learning and mask learning based on single-cell sample sequencing data, the generated target single-cell sequencing data not only maintains the overall statistical characteristics of the single-cell sequencing data but also maintains the global characteristics and local dependence characteristics of the single-cell sequencing data, effectively improving the data quality of the generated target single-cell RNA sequencing data.
[0084] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above methods according to the embodiments of the present disclosure.
[0085] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0086] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0087] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and distributed in one or more devices that are uniquely different from this embodiment. The modules of the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0088] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A method for generating single-cell RNA sequencing data, characterized in that, Including: Obtain single-cell sequencing data to be processed; Input the single-cell sequencing data into a pre-trained sequencing data generation model to generate target single-cell RNA sequencing data; The sequencing data generation model is obtained through multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrastive learning and masked learning; The step of generating the target single-cell RNA sequencing data includes: Perform data perturbation on the single-cell sequencing data to obtain perturbed data; Perform masked reconstruction on the single-cell sequencing data to obtain reconstructed data; Concatenate the perturbed data and the reconstructed data to obtain the target single-cell sequencing data.
2. The method for generating single-cell RNA sequencing data according to claim 1, wherein The performing data perturbation on the single-cell sequencing data includes: Randomly sort the gene expression values in the single-cell sequencing data to obtain re-sorted data; Randomly replace the gene expression values in the single-cell sequencing data with the gene expression values at the corresponding positions in the re-sorted data to obtain the perturbed data.
3. The method for generating single-cell RNA sequencing data according to claim 2, wherein The randomly replacing the gene expression value of each gene in the single-cell sequencing data with the gene expression value at the corresponding position in the randomly sorted sequencing data includes: Perform random masking on the gene expression values in the single-cell sequencing data, and replace the masked part in the single-cell sequencing data with the gene expression value at the corresponding position in the re-sorted data to obtain the perturbed data.
4. The method for generating single-cell RNA sequencing data according to claim 1, wherein The performing masked reconstruction on the single-cell sequencing data includes: Perform random masking on the gene expression values in the single-cell sequencing data, and predict the masked part in the single-cell sequencing data to obtain the reconstructed data.
5. The method for generating single-cell RNA sequencing data according to claim 1, wherein The concatenating the perturbed data and the reconstructed data includes: Perform dimensionality reduction mapping on the perturbed data and the reconstructed data and then concatenate them to obtain concatenated data; Convert the data form of the concatenated data to be the same as the data form of the single-cell sequencing data to obtain the target single-cell sequencing data.
6. The method for generating single-cell RNA sequencing data according to claim 1, wherein The sequencing data generation model includes a data perturbation network obtained through contrastive learning based on the single-cell sample sequencing data. The contrastive learning based on the single-cell sample sequencing data includes: Randomly sort the gene expression values in the single-cell sample sequencing data of the target cell to obtain multiple sample re-sorted data; Randomly replace the gene expression values in the single-cell sample sequencing data of the target cell with the gene expression values at the corresponding positions in the sample re-sorted data to obtain the sample perturbed data; Perform contrastive learning on the similarity distance between the single-cell sample sequencing data of the target cell and the sample perturbed data and the similarity distance between the single-cell sample sequencing data of the target cell and the single-cell sample sequencing data of non-target cells to obtain the data perturbation network.
7. The method for generating single-cell RNA sequencing data according to claim 1, wherein The sequencing data generation model includes a masked reconstruction network obtained through masked learning based on the single-cell sample sequencing data. The masked learning based on the single-cell sample sequencing data includes: Perform random masking on the gene expression values in the single-cell sample sequencing data, and predict the masked part in the single-cell sample sequencing data to obtain the sample reconstructed data; Mask learning is performed on the similarity distance between the single-cell sample sequencing data and the sample reconstruction data to obtain the mask reconstruction network.
8. A single-cell RNA sequencing data generation device, characterized in that, It includes: A first module for obtaining single-cell sequencing data to be processed; A second module for inputting the single-cell sequencing data into a pre-trained sequencing data generation model to generate target single-cell RNA sequencing data; the sequencing data generation model is obtained by performing multi-task joint learning based on single-cell sample sequencing data, and the multi-task joint learning includes contrast learning and mask learning; The step of generating the target single-cell RNA sequencing data includes: Performing data perturbation on the single-cell sequencing data to obtain perturbed data; Performing mask reconstruction on the single-cell sequencing data to obtain reconstruction data; Concatenating the perturbed data and the reconstruction data to obtain the target single-cell sequencing data.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the single-cell RNA sequencing data generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the single-cell RNA sequencing data generation method according to any one of claims 1 to 7.