Sequencing data imputation method, apparatus, device, and storage medium
Patent Information
- Application Number
- CN202610852263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,由于测序过程中的技术限制,单细胞转录组测序数据中通常存在较高比例的缺失值,这些缺失值会严重影响下游的数据分析,增加数据分析的难度,甚至可能导致研究者获得错误的生物学结论
[0010]本申请提供一种测序数据插补方法、装置、设备及存储介质,本申请通过获取待处理的单细胞转录组测序数据,并确定所述单细胞转录组测序数据的数据缺失位点;利用预设的编码器,对所述单细胞转录组测序数据进行特征提取处理,得到所述单细胞转录组测序数据的第一特征数据;获取所述单细胞转录组测序数据的第二特征数据,所述第二特征数据用于表征所述单细胞转录组测序数据的生物学先验知识;利用预设的条件扩散模型,对所述第一特征数据和所述第二特征数据进行处理,以生成所述单细胞转录组测序数据对应的目标特征数据;基于所述数据缺失位点和所述目标特征数据,从而能够准确地对单细胞转录组测序数据中的缺失值进行插补,进而有效提高单细胞转录组测序数据的完整性,实现单细胞转录组测序数据的基因表达的精确恢复。
Smart Images

Figure CN122842702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a sequencing data interpolation method, apparatus, device, and storage medium. Background Technology
[0002] Single-cell RNA sequencing (scRNA-seq) is a technology that enables high-throughput analysis of gene expression at the single-cell level. This technology achieves high-resolution transcriptome sequencing at the single-cell level, allowing researchers to deeply analyze the heterogeneity of cell populations, identify functionally diverse subpopulations, and reveal key characteristics of gene expression based on the obtained single-cell transcriptome sequencing data.
[0003] However, due to technical limitations in the sequencing process, single-cell transcriptome sequencing data often contain a high proportion of missing values. These missing values can seriously affect downstream data analysis, increase the difficulty of data analysis, and may even lead researchers to obtain incorrect biological conclusions.
[0004] Therefore, the problem of improving the integrity of single-cell transcriptome sequencing data urgently needs to be solved. Summary of the Invention
[0005] The main objective of this application is to provide a sequencing data interpolation method, apparatus, device, and storage medium that can accurately interpolate missing values in single-cell transcriptome sequencing data, thereby effectively improving the integrity of single-cell transcriptome sequencing data and achieving precise recovery of gene expression from single-cell transcriptome sequencing data.
[0006] In a first aspect, this application provides a sequencing data interpolation method, comprising: Acquire the single-cell transcriptome sequencing data to be processed, and determine the missing data sites in the single-cell transcriptome sequencing data; Using a preset encoder, feature extraction processing is performed on the single-cell transcriptome sequencing data to obtain the first feature data of the single-cell transcriptome sequencing data; A second feature data is obtained from the single-cell transcriptome sequencing data, and the second feature data is used to characterize the biological prior knowledge of the single-cell transcriptome sequencing data. Using a preset conditional diffusion model, the first feature data and the second feature data are processed to generate target feature data corresponding to the single-cell transcriptome sequencing data; Based on the missing data sites and the target feature data, the single-cell transcriptome sequencing data are subjected to data imputation processing.
[0007] Secondly, this application also provides a sequencing data interpolation device, the sequencing data interpolation device comprising: The first acquisition module is used to acquire the single-cell transcriptome sequencing data to be processed and to determine the missing data sites in the single-cell transcriptome sequencing data. The feature extraction module is used to perform feature extraction processing on the single-cell transcriptome sequencing data using a preset encoder to obtain the first feature data of the single-cell transcriptome sequencing data. The second acquisition module is used to acquire second feature data of the single-cell transcriptome sequencing data, and the second feature data is used to characterize the biological prior knowledge of the single-cell transcriptome sequencing data. The data generation module is used to process the first feature data and the second feature data using a preset conditional diffusion model to generate target feature data corresponding to the single-cell transcriptome sequencing data. The data imputation module is used to perform data imputation processing on the single-cell transcriptome sequencing data based on the missing data sites and the target feature data.
[0008] Thirdly, this application also provides a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the sequencing data interpolation method as described above.
[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the sequencing data interpolation method as described above.
[0010] This application provides a sequencing data imputation method, apparatus, device, and storage medium. The method involves acquiring single-cell transcriptome sequencing data to be processed and identifying missing data sites in the single-cell transcriptome sequencing data; using a preset encoder, performing feature extraction processing on the single-cell transcriptome sequencing data to obtain first feature data; acquiring second feature data from the single-cell transcriptome sequencing data, which characterizes the biological prior knowledge of the single-cell transcriptome sequencing data; processing the first feature data and the second feature data using a preset conditional diffusion model to generate target feature data corresponding to the single-cell transcriptome sequencing data; and accurately imputing missing values in the single-cell transcriptome sequencing data based on the missing data sites and the target feature data, thereby effectively improving the integrity of the single-cell transcriptome sequencing data and achieving precise recovery of gene expression from the single-cell transcriptome sequencing data. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This application provides a schematic flowchart of a sequencing data interpolation method according to an embodiment of the present application. Figure 2 for Figure 1 A flowchart illustrating the sub-steps of the sequencing data interpolation method; Figure 3 A schematic block diagram of a sequencing data interpolation device provided in this application embodiment; Figure 4 for Figure 3 A schematic block diagram of a submodule of the sequencing data interpolation device; Figure 5 This is a schematic block diagram of a computer device provided in an embodiment of this application.
[0013] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0016] Single-cell RNA sequencing (scRNA-seq) is a technology that enables high-throughput analysis of gene expression at the single-cell level. This technology achieves high-resolution transcriptome sequencing at the single-cell level, allowing researchers to deeply analyze cell population heterogeneity, identify functionally diverse subpopulations, and reveal key gene expression characteristics based on the obtained single-cell transcriptome sequencing data. However, due to inherent limitations of this technology, such as relatively low RNA capture and sequencing efficiency, some genes may not be effectively detected, leading to spurious zero values in the expression matrix. This phenomenon is defined as a missing value event in scRNA-seq data. In other words, single-cell transcriptome sequencing data obtained through scRNA-seq typically contains a high proportion of missing values, resulting in high sparsity. These missing values can severely impact downstream data analysis, increase the difficulty of data analysis, and may even lead researchers to obtain erroneous biological conclusions.
[0017] To improve the integrity of single-cell transcriptome sequencing data, existing sequencing data interpolation methods can be broadly classified into the following four categories: The first category is imputation methods based on statistical models. These methods estimate the true expression level by assuming that the expressed data follows a specific distribution and modeling zero inflation and technical noise, thereby achieving denoising and imputation of sparse data. However, these methods are limited by the assumed expressed data. If the assumed distribution deviates significantly from the actual data distribution, it may lead to inaccurate imputation results or even introduce new biases.
[0018] The second category is similarity-based imputation methods. These methods utilize the idea of data smoothing, performing smooth imputation based on the similarity between cells and genes using a similarity matrix. They typically adjust all expression values (including technical zero, biological zero, and observed non-zero values). However, these methods may lead to over-smoothing that eliminates natural intercellular randomness, and may also exacerbate the differences in expression levels of dissimilar genes in cells, introducing new biases.
[0019] The third category is interpolation methods based on matrix theory. These methods decompose the observed gene expression matrix into low-dimensional subspaces to eliminate noise. However, matrix decomposition in these methods is difficult to fully capture complex nonlinear biological signals, easily ignores local features, and is not accurate enough for the expression of rare cell types or local variations.
[0020] The fourth category is deep learning-based imputation methods. These methods use deep learning algorithms to learn the potential patterns and regularities of gene expression from large amounts of sequencing data, thereby predicting and filling in missing values. Compared with the three categories mentioned above, deep learning-based imputation methods have stronger learning and generalization capabilities. However, these methods often require a large amount of training data, the model construction and parameter tuning process is relatively complex, the training process is unstable, and it is prone to failure to converge or inaccurate generation of abnormal samples. In addition, the generated results of the model are prone to pattern collapse and cannot fully cover the diversity of biological data.
[0021] In summary, all existing interpolation methods have limitations. Therefore, the problem of improving the integrity of single-cell transcriptome sequencing data remains to be solved.
[0022] This application provides a sequencing data interpolation method, apparatus, device, and storage medium that can accurately interpolate missing values in single-cell transcriptome sequencing data, thereby effectively improving the integrity of single-cell transcriptome sequencing data and achieving precise restoration of gene expression in single-cell transcriptome sequencing data. The sequencing data interpolation method can be applied to terminal devices or servers. The terminal device can be an electronic device such as a mobile phone, tablet computer, laptop computer, desktop computer, personal digital assistant, or wearable device; the server can be a single server or a server cluster composed of multiple servers. The following explanation uses the application of the sequencing data interpolation method to a server as an example.
[0023] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0024] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a sequencing data interpolation method provided in an embodiment of this application.
[0025] like Figure 1 As shown, the sequencing data interpolation method includes steps S101 to S105.
[0026] Step S101: Obtain the single-cell transcriptome sequencing data to be processed and determine the missing data sites in the single-cell transcriptome sequencing data.
[0027] The single-cell transcriptome sequencing data to be processed can be a gene expression matrix obtained through single-cell RNA sequencing technology (scRNA-seq). In this gene expression matrix, rows represent different cells, columns represent different genes, and elements represent the expression level of the corresponding gene in the corresponding cell. Data missing sites refer to the locations in the gene expression matrix where gene expression levels are missing. At data missing sites, gene expression levels can be zero or specific markers.
[0028] It is important to emphasize that not all zero values in single-cell transcriptome sequencing data are considered missing data sites. That is, some zero values in single-cell transcriptome sequencing data represent genuine gene non-expression states, reflecting actual biological gene silencing, while others are due to limitations in current technology, i.e., technical omissions. Therefore, it is necessary to accurately distinguish between genuine gene non-expression states and technically caused omissions to ensure the accuracy of subsequent imputation results. The identification of missing data sites in single-cell transcriptome sequencing data can be achieved by comparing the gene expression distribution with that of other cells in the same sequencing batch, or by using statistical models to estimate a background noise threshold, identifying zero values below this threshold as potential missing data sites, and so on.
[0029] In one embodiment, determining missing data sites in single-cell transcriptome sequencing data includes: obtaining the cell type corresponding to the single-cell transcriptome sequencing data; determining a target sample dataset corresponding to the single-cell transcriptome sequencing data from multiple preset sample datasets based on the cell type corresponding to the single-cell transcriptome sequencing data; wherein each sample dataset corresponds to a different cell type, and the sample dataset includes multiple sample single-cell transcriptome sequencing data; determining multiple candidate sites from the single-cell transcriptome sequencing data; wherein the gene expression value of the candidate sites is zero; comparing the gene expression value of each candidate site with the gene expression value of the corresponding site in the multiple sample sequencing data in the target sample dataset; if the gene expression value of the candidate site is inconsistent with the gene expression value of the corresponding site in the multiple sample sequencing data in the target sample dataset, then the candidate site is determined to be a missing data site in the single-cell transcriptome sequencing data.
[0030] Cell type refers to different cell population categories classified according to differences in developmental stage, physiological function, molecular characteristics, etc. The pre-defined sample datasets can be obtained by annotating and classifying cell types from a large amount of single-cell transcriptome sequencing data. Each sample dataset corresponds to one cell type, and each dataset contains single-cell transcriptome sequencing data from multiple biological samples within the corresponding cell type. These single-cell transcriptome sequencing data relatively fully reflect the gene expression distribution characteristics and other features of the corresponding cell type.
[0031] Specifically, the gene expression values of each candidate site are compared with the gene expression values of the corresponding sites in the sequencing data of multiple samples in the target sample dataset. If the corresponding site in the target sample dataset has non-zero expression, it indicates that the gene has expression potential in this cell type. Therefore, it is determined that the corresponding candidate site has missing data, thus identifying the candidate site as a missing data site. Conversely, if the corresponding site in the target sample dataset generally has zero values, consistent with the gene expression values of the candidate sites, it indicates that the zero value of the candidate site reflects the biological gene silencing phenomenon.
[0032] Understandably, gene expression patterns in single cells of the same cell type typically exhibit high similarity, while different single cells show relatively significant differences. By using sequencing data from multiple samples in the target dataset as reference samples, and comparing the gene expression values of each candidate locus with the corresponding loci in the sequencing data of multiple samples in the target dataset, we can more accurately distinguish between true biological gene silencing and technical missing data in single-cell transcriptome sequencing data, effectively improving the accuracy of identifying missing data sites in single-cell transcriptome sequencing data.
[0033] Step S102: Using a preset encoder, feature extraction processing is performed on the single-cell transcriptome sequencing data to obtain the first feature data of the single-cell transcriptome sequencing data.
[0034] The preset encoder can be used to capture high-dimensional nonlinear feature relationships in single-cell transcriptome data to obtain a core, compact and information-rich representation of single-cell transcriptome sequencing data in a low-dimensional space, thereby obtaining the first feature data of single-cell transcriptome sequencing data.
[0035] For example, the preset encoder is the encoder in the autoencoder. The encoder Composed of a K-layer feedforward neural network, it is used to map high-dimensional gene expression profiles to low-dimensional latent representations. Let the encoder... The input is the raw gene expression matrix of single-cell transcriptome sequencing data. ∈ × ,in For cell number, For each element in the matrix, representing the number of genes... Indicates the first The first cell The expression values of each gene.
[0036] Using encoder The high-dimensional input is transformed through a layer-by-layer nonlinear transformation. Mapping to a low-dimensional latent space: Specifically, each layer of the encoder performs feature transformation through a non-linear activation function (ReLU or Sigmoid, etc.) to obtain the features corresponding to each layer. ,in, and The first The weight matrix and bias vector of the layer, The activation function is used. Furthermore, each layer sequentially transforms the output of the previous layer, ultimately yielding a low-dimensional latent representation that fully preserves the key biological information from the single-cell transcriptome sequencing data through the output layer. This low-dimensional potential representation This is the first feature data.
[0037] Understandably, acquiring first-feature data that can characterize the expression patterns of each gene in single-cell transcriptome sequencing data and the intrinsic relationships between cells can help improve the accuracy of subsequent sequencing data imputation.
[0038] In one embodiment, the encoder training process includes: acquiring multiple first sample sequencing data; randomly determining at least one first point from each first sample sequencing data and deleting the gene expression value of each first point to obtain multiple second sample sequencing data; wherein the gene expression value of the first point is not zero; inputting the multiple second sample sequencing data into the encoder for feature extraction processing to obtain multiple first sample feature data; using a preset encoder to reconstruct each first sample feature data to obtain multiple reconstructed sequencing data; and updating the encoder model parameters based on the multiple second sample sequencing data and the multiple reconstructed sequencing data until the encoder converges.
[0039] By randomly determining at least one first point and deleting the gene expression values of each first point, the common data missing phenomenon in real single-cell transcriptome sequencing data can be simulated.
[0040] For example, the encoder includes a multi-layer feedforward neural network. When updating the encoder's model parameters, the likelihood function is used to calculate the difference between the second sample sequencing data and the reconstructed sequencing data, and the weight matrix and bias vector of each layer of the encoder are iteratively optimized through the backpropagation algorithm.
[0041] It is understandable that by training the encoder, it can learn to extract robust feature representations from partial observation data, thereby helping to improve the accuracy of subsequent sequencing data imputation.
[0042] Step S103: Obtain the second feature data of the single-cell transcriptome sequencing data. The second feature data is used to characterize the biological prior knowledge of single-cell transcriptome sequencing data. The biological prior knowledge may include information such as developmental stage, cell type, tissue origin, experimental conditions, or known marker gene expression patterns. It may be annotation information related to the sequencing data to be interpolated retrieved from public databases or published literature. For example, typical expression characteristics of the corresponding cell type can be obtained from a cell atlas database, or gene expression markers of a specific developmental stage can be extracted from developmental trajectory studies, etc.
[0043] It is understandable that second feature data can reflect the inherent characteristics of cells in a specific biological context. Therefore, by obtaining the second feature data of single-cell transcriptome sequencing data, effective and reliable guidance information and constraints can be provided for the subsequent generation of target feature data for sequencing data interpolation, thereby helping to improve the accuracy and rationality of sequencing data interpolation.
[0044] Step S104: Using a preset conditional diffusion model, process the first feature data and the second feature data to generate target feature data corresponding to the single-cell transcriptome sequencing data.
[0045] Among them, the conditional diffusion model is used to associate and fuse the original data distribution information contained in the first feature data with the biological prior knowledge carried by the second feature data, thereby generating feature data that retains the statistical characteristics of the original data and conforms to the known biological prior knowledge.
[0046] In one embodiment, the conditional diffusion model includes a noise-adding network and a noise-removing network. The noise-adding network is used to add noise to the input data, that is, the noise-adding network performs a forward diffusion process. The noise-removing network is used to predict and remove noise from the noise-adding data after processing by the noise-adding network, that is, the noise-removing network performs a reverse noise-removing process to restore the original distribution characteristics of the data.
[0047] In one embodiment, such as Figure 2 As shown, step S104 includes sub-steps S1041 to S1042.
[0048] Sub-step S1041: Using a noisy network, the first feature data is gradually subjected to noise embedding processing at multiple preset time steps to obtain noisy feature data.
[0049] The forward diffusion process follows a Markov chain structure. During forward diffusion using a noisy network, the noisy data at each time step is based on the data from the previous time step. That is, the state of the first feature data after noise embedding at each time step depends on the previous state. Finally, after multiple time steps of progressive noise superposition, the first feature data gradually loses its original distribution characteristics, eventually transforming into noisy feature data that is approximately pure noise at the largest time step.
[0050] In one embodiment, a noisy network is used to progressively embed noise into the first feature data at multiple preset time steps to obtain noisy feature data. This includes: calculating the noise embedding data corresponding to each time step based on the noise data and preset noise intensity; and progressively superimposing and embedding the noise embedding data into the first feature data in ascending order of each time step to obtain noisy feature data.
[0051] The noise data corresponding to each time step is obtained by random sampling from a noise sample set that conforms to a standard normal distribution. The noise sample set contains multiple noise samples.
[0052] For example, following the order of time steps from smallest to largest, i.e. starting from t=0, the noise embedding process of the noisy network on the first feature data follows the following formula:
[0053] in, This is the first feature data. This is the noise embedding data corresponding to the first feature data at time step. , , , Indicates at time step The innermost small positive value representing the noise level, when T When large enough, It follows a Gaussian distribution.
[0054] Sub-step S1042: Using a denoising network, the second feature data is used as a guiding condition to progressively predict the noise data of the noisy feature data at each time step, and based on each noise data, the noisy feature data is denoised to obtain the target feature data.
[0055] The denoising network is used to perform the reverse denoising process of the conditional diffusion model. The denoising network can be a deep neural network model, such as a multilayer perceptron (MLP), convolutional neural network (CNN), U-Net network, Transformer network, etc.
[0056] For example, a DiT (Diffusion Transformer) denoising network based on the Transformer concept is used. Specifically, the input of the denoising network includes noisy feature data at the current time step, time step information, and second feature data, while the output includes the noise prediction result corresponding to the current time step. The time step information can be input into the denoising network through time step encoding so that the denoising network can perceive the current diffusion stage.
[0057] It is understandable that by using a denoising network to progressively predict noise at each time step during the backdiffusion process and performing denoising processing step by step under the guidance of the second feature data, the noisy feature data can be gradually restored from a high-noise state to target feature data that conforms to the target distribution. This improves the generation accuracy, stability, and consistency of the target feature data with the second feature data, reduces noise interference, and enhances the quality of feature representation.
[0058] In one embodiment, the second feature data is used as a guiding condition to progressively predict the noise data of the noisy feature data at each time step, and the noisy feature data is denoised based on each noise data to obtain the target feature data. This includes: using the second feature data as a guiding condition to predict the noise data of the noisy feature data at the largest time step; and denoising the noisy feature data based on the noise data at the largest time step to obtain the candidate feature data corresponding to the largest time step; predicting the noise data at the next time step using the candidate feature data corresponding to the previous time step in descending order of time steps, and denoising the candidate feature data corresponding to the previous time step, until the noise data and candidate feature data at the smallest time step are determined; and using the candidate feature data corresponding to the smallest time step as the target feature data.
[0059] In the reverse denoising process of the diffusion model, the decreasing order of time steps corresponds to the gradual reduction of noise level.
[0060] For example, the denoising network starts from the time step. It starts with random Gaussian noise, and then undergoes an inverse denoising process at the transition time step. The noise input is corrected through iterations, ultimately generating biologically meaningful potential cell embeddings. As shown in the following formula:
[0061] in, hour, ,otherwise Each embedding vector For a given cell, the set of potential representations generated by all cells This refers to the target feature data.
[0062] Understandably, by using the second feature data as a guiding condition to constrain noise prediction at each time step of diffusion denoising, the final target feature data can not only retain the effective information in the noisy feature data, but also simultaneously integrate the key information in the two types of guiding features, effectively improving the accuracy and reliability of the target feature data.
[0063] Step S105: Based on missing data sites and target feature data, perform data imputation processing on single-cell transcriptome sequencing data.
[0064] The target feature data contains gene expression values corresponding to missing data sites. Therefore, based on the location index of the missing data sites, the gene expression values at the corresponding positions can be directly extracted from the target feature data as imputation values, and these imputation values can be filled into the missing data sites of the single-cell transcriptome sequencing data.
[0065] It should be noted that the target feature data includes not only gene expression values corresponding to missing sites but also gene expression values corresponding to other non-missing sites. These gene expression values for non-missing sites may be consistent with the actual gene expression values, or they may have slight differences due to the generation characteristics of the conditional diffusion model. Therefore, when performing data imputation on single-cell transcriptome sequencing data, only the imputation values for missing sites can be replaced, while the original gene expression values for non-missing sites are retained. This ensures that the imputed single-cell transcriptome sequencing data not only fully preserves the accuracy of the original experimental measurements but also effectively restores the technically missing gene expression information.
[0066] In one embodiment, data imputation processing is performed on single-cell transcriptome sequencing data based on missing data sites and target feature data, including: decoding the target feature data to obtain simulated single-cell transcriptome data; aligning the simulated single-cell transcriptome data with the single-cell transcriptome sequencing data to identify target sites in the simulated single-cell transcriptome data corresponding to the missing data sites in the single-cell transcriptome sequencing data; extracting the gene expression values of the target sites in the simulated single-cell transcriptome data to obtain the target gene expression values; and performing data imputation processing on the single-cell transcriptome sequencing data based on the missing data sites and the target gene expression values.
[0067] This can be achieved by using a decoder within a pre-defined autoencoder to decode the target feature data, mapping it from a low-dimensional latent space back to a high-dimensional gene expression space. This decoder can be jointly trained with the encoder, learning the mapping relationship from latent features to the original data distribution by minimizing reconstruction error. The decoder's network architecture can adopt a symmetrical design to the encoder. For example, if the encoder uses a multi-layer fully connected network with batch normalization and activation functions, the decoder can correspondingly use transposed fully connected layers or upsampling structures of the same number, gradually expanding the feature dimension to the same gene quantity dimension as the original sequencing data.
[0068] For example, target feature data The decoder in the trained autoencoder is used to obtain a complete gene expression profile, which is then compared with single-cell transcriptome sequencing data. Alignment, based on an indicator matrix used to characterize missing data sites. Impute missing values to obtain the imputed data. The specific formula is as follows:
[0069] Understandably, in the above data imputation process, for data missing sites in single-cell transcriptome sequencing data, the gene expression value of the corresponding site in the target feature data is directly used as the imputation result; for other sites in single-cell transcriptome sequencing data that are not marked as data missing, their gene expression values are retained unchanged. In this way, the integrity of the original observation data before data imputation of single-cell transcriptome sequencing data is not destroyed, and the imputed gene expression values are highly consistent with the real biological signals in terms of global distribution.
[0070] The sequencing data imputation method provided in the above embodiments acquires single-cell transcriptome sequencing data to be processed and determines the missing data sites in the single-cell transcriptome sequencing data; uses a preset encoder to perform feature extraction processing on the single-cell transcriptome sequencing data to obtain first feature data of the single-cell transcriptome sequencing data; acquires second feature data of the single-cell transcriptome sequencing data, which is used to characterize the biological prior knowledge of the single-cell transcriptome sequencing data; uses a preset conditional diffusion model to process the first feature data and the second feature data to generate target feature data corresponding to the single-cell transcriptome sequencing data; based on the missing data sites and the target feature data, it can accurately imput missing values in the single-cell transcriptome sequencing data, thereby effectively improving the integrity of the single-cell transcriptome sequencing data and realizing the accurate recovery of gene expression in the single-cell transcriptome sequencing data.
[0071] Please refer to Figure 3 , Figure 3 This is a schematic block diagram of a sequencing data interpolation device provided in an embodiment of this application.
[0072] like Figure 3 As shown, the sequencing data interpolation device 200 includes: The first acquisition module 201 is used to acquire the single-cell transcriptome sequencing data to be processed and to determine the missing data sites in the single-cell transcriptome sequencing data for interpolation.
[0073] The feature extraction module 202 is used to perform feature extraction processing on the sequencing data imputation single-cell transcriptome sequencing data using a preset encoder, so as to obtain the first feature data of the sequencing data imputation single-cell transcriptome sequencing data.
[0074] The second acquisition module 203 is used to acquire the second feature data of sequencing data interpolated from single-cell transcriptome sequencing data. The second feature data of sequencing data interpolation is used to characterize the biological prior knowledge of sequencing data interpolated from single-cell transcriptome sequencing data.
[0075] The data generation module 204 is used to process the first feature data and the second feature data of the sequencing data interpolation using a preset conditional diffusion model, so as to generate the target feature data corresponding to the sequencing data interpolated from the single-cell transcriptome sequencing data.
[0076] The data imputation module 205 is used to perform data imputation processing on single-cell transcriptome sequencing data based on missing data sites and target feature data of sequencing data imputation.
[0077] In one embodiment, the first acquisition module 201 is further configured to: The process involves: obtaining the cell types corresponding to the single-cell transcriptome sequencing data imputed from the sequencing data; determining the target sample dataset corresponding to the single-cell transcriptome sequencing data imputed from multiple pre-defined sample datasets, where each sample dataset corresponds to a different cell type and includes single-cell transcriptome sequencing data from multiple samples; identifying multiple candidate sites for sequencing data imputation from the single-cell transcriptome sequencing data imputed from the sequencing data imputed from the candidate sites, where the gene expression value of each candidate site is zero; comparing the gene expression value of each candidate site with the gene expression value of the corresponding site in the sequencing data of multiple samples in the target sample dataset; and identifying the candidate site as a missing data site in the single-cell transcriptome sequencing data imputed if the gene expression value of the candidate site is inconsistent with the gene expression value of the corresponding site in the sequencing data of multiple samples in the target sample dataset.
[0078] In one embodiment, the feature extraction module 202 is further configured to: train the encoder, wherein the encoder training process includes: Acquire multiple first sample sequencing data; randomly select at least one first locus from each first sample sequencing data and delete the gene expression values of each first locus to obtain multiple second sample sequencing data; wherein the gene expression value of the first locus is not zero; input the multiple second sample sequencing data into the encoder for feature extraction processing to obtain multiple first sample feature data; use a preset decoder to reconstruct each first sample feature data to obtain multiple reconstructed sequencing data; based on the multiple second sample sequencing data and the multiple reconstructed sequencing data, update the encoder model parameters until the encoder converges.
[0079] In one embodiment, the conditional diffusion model includes a noisy network and a denoising network, such as Figure 4 As shown, the data generation module 204 includes: The noise-adding submodule 2041 is used to perform noise embedding processing on the first feature data step by step at multiple preset time steps using a noise-adding network to obtain noisy feature data.
[0080] The denoising submodule 2042 is used to use a denoising network to predict the noise data of the noisy feature data at each time step by using the second feature data as a guiding condition, and to perform denoising processing on the noisy feature data based on each noise data to obtain the target feature data.
[0081] In one embodiment, the noise-adding submodule 2041 is further configured to: Based on the noise data corresponding to each time step and the preset noise intensity, the noise embedding data corresponding to each time step is calculated. The noise data corresponding to each time step is randomly sampled from a noise sample set that conforms to a standard normal distribution. The noise sample set contains multiple noise samples. The noise embedding data is gradually superimposed and embedded into the first feature data in the order of each time step from small to large to obtain the noisy feature data.
[0082] In one embodiment, the noise reduction submodule 2042 is further configured to: Using the second feature data as a guiding condition, the noisy data of the noisy feature data at the largest time step is predicted; and based on the noisy data at the largest time step, the noisy feature data is denoised to obtain the candidate feature data corresponding to the largest time step; according to the order of each time step from largest to smallest, the noisy data at the next time step is predicted using the candidate feature data corresponding to the previous time step, and the candidate feature data corresponding to the previous time step is denoised, until the noisy data and candidate feature data at the smallest time step are determined; the candidate feature data corresponding to the smallest time step is taken as the target feature data.
[0083] In one embodiment, the data interpolation module 205 is further configured to: The target feature data is decoded to obtain simulated single-cell transcriptome data. The simulated single-cell transcriptome data is then aligned with the single-cell transcriptome sequencing data to identify the target sites in the simulated single-cell transcriptome data that correspond to the missing data sites in the single-cell transcriptome sequencing data. The gene expression values of the target sites in the simulated single-cell transcriptome data are extracted to obtain the target gene expression values. Based on the missing data sites and the target gene expression values, the single-cell transcriptome sequencing data is imputed.
[0084] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus and its modules and units can be referred to the corresponding processes in the aforementioned sequencing data interpolation method embodiments, and will not be repeated here.
[0085] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 5 It runs on the computer device shown.
[0086] Please see Figure 5 , Figure 5 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0087] like Figure 3 As shown, the computer device 300 includes a processor 301, a memory 302, and a network interface connected via a system bus 303. The memory 302 may include a storage medium and internal memory. The storage medium may be non-volatile or volatile.
[0088] The processor 301 provides computing and control capabilities to support the operation of the entire computer device 300.
[0089] The storage medium may store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor 301 to perform any sequencing data interpolation method.
[0090] The internal memory provides an environment for the execution of computer programs stored in the storage medium. When the computer program is executed by the processor 301, the processor 301 can execute any sequencing data interpolation method.
[0091] Network interfaces are used for network communication, such as sending assigned tasks.
[0092] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0093] It should be understood that processor 301 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0094] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: Acquire the single-cell transcriptome sequencing data to be processed and identify the missing data sites in the single-cell transcriptome sequencing data; Using a pre-defined encoder, feature extraction processing is performed on single-cell transcriptome sequencing data to obtain the first feature data of the single-cell transcriptome sequencing data; Second feature data is obtained from single-cell transcriptome sequencing data. The second feature data is used to characterize the biological prior knowledge of single-cell transcriptome sequencing data. Using a pre-defined conditional diffusion model, the first feature data and the second feature data are processed to generate target feature data corresponding to single-cell transcriptome sequencing data; Data imputation was performed on single-cell transcriptome sequencing data based on missing data sites and target feature data.
[0095] In one embodiment, the processor, when determining data missing sites in single-cell transcriptome sequencing data, is configured to: Obtain the cell type corresponding to single-cell transcriptome sequencing data; Based on the cell type corresponding to the single-cell transcriptome sequencing data, the target sample dataset corresponding to the single-cell transcriptome sequencing data is determined from multiple pre-set sample datasets; wherein, each sample dataset corresponds to a different cell type, and the sample dataset includes multiple sample single-cell transcriptome sequencing data. Multiple candidate sites were identified from single-cell transcriptome sequencing data; among them, the gene expression value of the candidate sites was zero. The gene expression values of each candidate site are compared with the gene expression values of the corresponding sites in the sequencing data of multiple samples in the target sample dataset; If the gene expression value of a candidate site is inconsistent with the gene expression value of the corresponding site in the sequencing data of multiple samples in the target sample dataset, then the candidate site is determined to be a missing data site in the single-cell transcriptome sequencing data.
[0096] In one embodiment, the processor is further configured to implement: Acquire multiple first sample sequencing data; randomly select at least one first locus from each first sample sequencing data and delete the gene expression values of each first locus to obtain multiple second sample sequencing data; wherein the gene expression value of the first locus is not zero; input the multiple second sample sequencing data into the encoder for feature extraction processing to obtain multiple first sample feature data; use a preset decoder to reconstruct each first sample feature data to obtain multiple reconstructed sequencing data; based on the multiple second sample sequencing data and the multiple reconstructed sequencing data, update the encoder model parameters until the encoder converges.
[0097] In one embodiment, the conditional diffusion model includes a noise-adding network and a noise-reducing network. When the processor processes the first feature data and the second feature data using the preset conditional diffusion model to generate target feature data corresponding to single-cell transcriptome sequencing data, it is used to: Using a noise-adding network, the first feature data is progressively subjected to noise embedding at multiple preset time steps to obtain noisy feature data. Using a denoising network, the second feature data is used as a guiding condition to progressively predict the noise data of the noisy feature data at each time step. Based on each noise data, the noisy feature data is denoised to obtain the target feature data.
[0098] In one embodiment, when the processor progressively performs noise embedding processing on the first feature data at multiple preset time steps to obtain noisy feature data, it is used to: Based on the noise data corresponding to each time step and the preset noise intensity, the noise embedding data corresponding to each time step is calculated. The noise data corresponding to each time step is randomly sampled from a noise sample set that conforms to a standard normal distribution. The noise sample set contains multiple noise samples. The noise embedding data is gradually superimposed and embedded into the first feature data in the order of each time step from small to large to obtain the noisy feature data.
[0099] In one embodiment, when the processor uses the second feature data as a guiding condition to progressively predict the noise data of the noisy feature data at each time step, and performs denoising processing on the noisy feature data based on each noise data to obtain the target feature data, it is used to implement: Using the second feature data as a guiding condition, the noisy data of the noisy feature data at the largest time step is predicted; and based on the noisy data at the largest time step, the noisy feature data is denoised to obtain the candidate feature data corresponding to the largest time step; according to the order of each time step from largest to smallest, the noisy data at the next time step is predicted using the candidate feature data corresponding to the previous time step, and the candidate feature data corresponding to the previous time step is denoised, until the noisy data and candidate feature data at the smallest time step are determined; the candidate feature data corresponding to the smallest time step is taken as the target feature data.
[0100] In one embodiment, when the processor performs data imputation processing on single-cell transcriptome sequencing data based on missing data sites and target feature data, it is used to: The target feature data is decoded to obtain simulated single-cell transcriptome data. The simulated single-cell transcriptome data is then aligned with the single-cell transcriptome sequencing data to identify the target sites in the simulated single-cell transcriptome data that correspond to the missing data sites in the single-cell transcriptome sequencing data. The gene expression values of the target sites in the simulated single-cell transcriptome data are extracted to obtain the target gene expression values. Based on the missing data sites and the target gene expression values, the single-cell transcriptome sequencing data is imputed.
[0101] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the computer equipment described above can be referred to the corresponding process in the aforementioned sequencing data interpolation method embodiments, and will not be repeated here.
[0102] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0103] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, and the method implemented when the program instructions are executed can refer to various embodiments of the sequencing data interpolation method of this application.
[0104] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0105] Those skilled in the art will understand that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CDs, etc. ROM, digital multifunction disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0106] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0107] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. It should be noted that, herein, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0108] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. The above descriptions are merely specific implementations of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A sequencing data interpolation method, characterized in that, include: Acquire the single-cell transcriptome sequencing data to be processed, and determine the missing data sites in the single-cell transcriptome sequencing data; Using a preset encoder, feature extraction processing is performed on the single-cell transcriptome sequencing data to obtain the first feature data of the single-cell transcriptome sequencing data; A second feature data is obtained from the single-cell transcriptome sequencing data, and the second feature data is used to characterize the biological prior knowledge of the single-cell transcriptome sequencing data. Using a preset conditional diffusion model, the first feature data and the second feature data are processed to generate target feature data corresponding to the single-cell transcriptome sequencing data; Based on the missing data sites and the target feature data, the single-cell transcriptome sequencing data are subjected to data imputation processing.
2. The sequencing data interpolation method as described in claim 1, characterized in that, The conditional diffusion model includes a noisy network and a denoising network; The process of using a preset conditional diffusion model to process the first feature data and the second feature data to generate target feature data corresponding to the single-cell transcriptome sequencing data includes: Using the aforementioned noise-adding network, the first feature data is progressively subjected to noise embedding processing at multiple preset time steps to obtain noisy feature data; Using the denoising network, the second feature data is used as a guiding condition to progressively predict the noise data of the noisy feature data at each time step, and based on each noise data, the noisy feature data is denoised to obtain the target feature data.
3. The sequencing data interpolation method as described in claim 2, characterized in that, The step of progressively performing noise embedding processing on the first feature data at multiple preset time steps to obtain noisy feature data includes: Based on the noise data corresponding to each time step and the preset noise intensity, the noise embedding data corresponding to each time step is calculated; wherein, the noise data corresponding to each time step is obtained by random sampling from a noise sample set that conforms to a standard normal distribution, and the noise sample set contains multiple noise samples; Following the order of each time step from smallest to largest, the noise embedding data are gradually superimposed and embedded into the first feature data to obtain noisy feature data.
4. The sequencing data interpolation method as described in claim 2, characterized in that, The step of using the second feature data as a guiding condition to progressively predict the noise data of the noisy feature data at each time step, and performing denoising processing on the noisy feature data based on each noise data to obtain the target feature data, includes: Using the second feature data as a guiding condition, predict the noise data of the noisy feature data at the maximum time step; and Based on the noisy data at the largest time step, the noisy feature data is denoised to obtain the candidate feature data corresponding to the largest time step. According to the order of each time step from large to small, the candidate feature data corresponding to the previous time step is used to predict the noise data in the next time step, and the candidate feature data corresponding to the previous time step is denoised until the noise data and candidate feature data in the smallest time step are determined. The candidate feature data corresponding to the smallest time step is taken as the target feature data.
5. The sequencing data interpolation method as described in claim 1, characterized in that, The process of determining the missing data sites in the single-cell transcriptome sequencing data includes: Obtain the cell type corresponding to the single-cell transcriptome sequencing data; Based on the cell type corresponding to the single-cell transcriptome sequencing data, a target sample dataset corresponding to the single-cell transcriptome sequencing data is determined from multiple preset sample datasets; wherein, each sample dataset corresponds to a different cell type, and the sample dataset includes multiple sample single-cell transcriptome sequencing data. The plurality of candidate sites were determined from the single-cell transcriptome sequencing data; wherein the gene expression value of the candidate sites is zero; The gene expression values of each candidate site are compared with the gene expression values of the corresponding sites in the sequencing data of multiple samples in the target sample dataset; If the gene expression value of the candidate site is inconsistent with the gene expression value of the corresponding site in the sequencing data of multiple samples in the target sample dataset, then the candidate site is determined to be a data missing site in the single-cell transcriptome sequencing data.
6. The sequencing data interpolation method according to any one of claims 1-5, characterized in that, The data imputation process for the single-cell transcriptome sequencing data based on the missing data sites and the target feature data includes: The target feature data is decoded to obtain simulated single-cell transcriptome data; The simulated single-cell transcriptome data is aligned with the single-cell transcriptome sequencing data to identify target sites in the simulated single-cell transcriptome data that correspond to the data missing sites in the single-cell transcriptome sequencing data. Gene expression values of the target sites in the simulated single-cell transcriptome data are extracted to obtain the target gene expression values; Based on the missing data sites and the expression values of the target genes, the single-cell transcriptome sequencing data are imputed.
7. The sequencing data interpolation method according to any one of claims 1-5, characterized in that, The training process of the encoder includes: Obtain sequencing data from multiple first-sample samples; At least one first locus is randomly selected from the sequencing data of each first sample, and the gene expression value of each first locus is deleted to obtain multiple sequencing data of second samples; wherein the gene expression value of the first locus is not zero; Multiple second sample sequencing data are input into the encoder for feature extraction processing to obtain multiple first sample feature data. The feature data of each first sample are reconstructed using a preset decoder to obtain multiple reconstructed sequencing data. Based on multiple sequencing data of the second sample and multiple reconstructed sequencing data, the model parameters of the encoder are updated until the encoder converges.
8. A sequencing data interpolation device, characterized in that, The device includes: The first acquisition module is used to acquire the single-cell transcriptome sequencing data to be processed and to determine the missing data sites in the single-cell transcriptome sequencing data. The feature extraction module is used to perform feature extraction processing on the single-cell transcriptome sequencing data using a preset encoder to obtain the first feature data of the single-cell transcriptome sequencing data. The second acquisition module is used to acquire second feature data of the single-cell transcriptome sequencing data, and the second feature data is used to characterize the biological prior knowledge of the single-cell transcriptome sequencing data. The data generation module is used to process the first feature data and the second feature data using a preset conditional diffusion model to generate target feature data corresponding to the single-cell transcriptome sequencing data. The data imputation module is used to perform data imputation processing on the single-cell transcriptome sequencing data based on the missing data sites and the target feature data.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the sequencing data interpolation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the sequencing data interpolation method as described in any one of claims 1 to 7.