A single-cell sequencing data quality evaluation method
By creating a synthetic gene expression matrix and optimizing the function, the problem of quality assessment of single-cell sequencing data imputation algorithms in zero-sample scenarios was solved, achieving accurate quantification of data quality differences and making it suitable for zero-sample assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2026-04-02
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies cannot effectively assess the data quality differences between single-cell sequencing data imputation algorithms in zero-sample scenarios, making it impossible to use existing methods for evaluation under certain privacy or other issues.
By creating a synthetic gene expression matrix based on the statistical characteristics of real data, and using eigenvectors and optimization functions for multiple rounds of optimization, the data quality difference value D is calculated, thus realizing the zero-sample evaluation of the data quality difference in the imputation algorithm.
It achieves accurate quantification of data quality differences between imputation algorithms without the participation of real single-cell sequencing data, and is suitable for zero-sample scenarios.
Smart Images

Figure CN122290700A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics metrology, and proposes a method for measuring the data quality differences between different single-cell sequencing data imputation algorithms through quantitative indicators, particularly a method for evaluating algorithm quality based on high-precision imputation without the participation of real samples. Background Technology
[0002] Currently, in single-cell sequencing data processing, it is necessary to measure the data quality differences between different imputation algorithms to provide performance evaluation metrics for methods such as data domain adaptation, active learning, and semi-supervised learning. However, there is currently no zero-shot method for evaluating the data quality differences between imputation algorithms. Current methods generally apply performance evaluation metrics designed for a single algorithm to two algorithms, and then calculate the differences in these metrics to obtain the data quality differences between the algorithms. For example, in single-cell transcriptome imputation tasks, the difference can be evaluated by comparing the accuracy differences between two algorithms on the same dataset. However, this method requires the test dataset to be involved in the measurement of algorithm data quality differences. If the dataset is inaccessible due to privacy or other issues, this method cannot be used, thus failing to meet the requirements of zero-shot scenarios. Summary of the Invention
[0003] The purpose of this invention is to fill the gap in data quality assessment methods for zero-sample imputation algorithms by proposing a single-cell sequencing data quality assessment method. The technical solution is as follows: A method for assessing the quality of single-cell sequencing data: (1) Prepare two single-cell sequencing data imputation algorithms for which data quality differences need to be evaluated. , ; (2) Create a synthetic gene expression matrix based on the statistical characteristics of real data and perform normalization preprocessing on it to generate the imputation gene expression profile; (3) Input the normalized gene expression matrix into the algorithm respectively. , Obtain feature vectors , ; (4) Using eigenvectors , Set optimization function Then, the gene expression matrix is optimized in multiple rounds using an optimization function. Finally, the optimized matrix is denormalized to obtain the imputed gene expression profile. (5) Re-enter the obtained gene expression profile into the algorithm , Two sets of predicted feature vectors were obtained. , ; (6) Using eigenvectors , Calculate the data quality difference value D.
[0004] The beneficial effects of this invention are: (1) This invention is a specially designed interpolation algorithm for evaluating data quality differences, which can more accurately quantify data quality differences compared to current accuracy difference methods.
[0005] (2) This invention is a zero-sample method that can obtain data quality difference values without the need for real single-cell sequencing data to participate in the calculation. Attached Figure Description
[0006] Figure 1 This is an overall structural diagram of the present invention; Detailed Implementation
[0007] To make the technical solution of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. The present invention is implemented in specific steps: (1) Prepare two single-cell sequencing data imputation algorithms for which data quality differences need to be evaluated. , The single-cell sequencing data interpolation tasks served by the two algorithms need to be the same, that is, the interpolation target needs to be consistent. (2) Create A synthetic gene expression matrix X based on the statistical characteristics of real data, wherein... The number of cell types is represented by N, which represents the number of imputed data points in a single cell type. Then, the average value from single-cell sequencing technology is used. and standard deviation Normalize X to , used to generate interpolated gene expression profiles; (3) Gene expression matrix Input the algorithm respectively , The feature vectors are obtained, and then normalized using the softmax function to obtain the feature vectors. , ; in Represents the total number of cell types. Representative Algorithm Prediction Matrix The feature value belonging to cell type j; Representative Algorithm Prediction Matrix The feature value belonging to cell type j; (4) Using eigenvectors , Set optimization function : Where y represents the expression profile of the y-th type of imputed gene. Then, the matrix is optimized using an optimization function. Perform multiple rounds of optimization; in Represents the learning rate. This represents the matrix optimized up to round i. This process is repeated until the preset number of rounds is reached. Finally, the optimized matrix is denormalized to obtain the imputed gene expression profile. : (5) Re-input the interpolated gene spectrum belonging to the i-th class into the algorithm Re-entry algorithm , Two sets of predicted feature vectors were obtained. , ; in The expression profile of the first type of imputed genes was represented by the algorithm. The feature value when identified as cell type j. The expression profile of the i-th imputed gene was represented by the algorithm. The feature value when identified as cell type j; (6) Repeat step (5) until all are obtained. The interpolated gene expression profiles are predicted by the algorithm to be characteristic values for each cell type. Then, the average of these characteristic values is calculated to obtain the mean of the characteristic values of all interpolated gene expression profiles of cell type i when predicted to be cell type j. Utilize all categories The sum of the absolute values of the differences can be used to calculate the data quality variance value D:
Claims
1. A method of single-cell sequencing data quality assessment, the method comprising: The steps for evaluating the differences in data quality between interpolation algorithms include: (1) Prepare two single-cell sequencing data imputation algorithms that need to evaluate the difference in data quality , ; (2) Create a synthetic gene expression matrix based on the statistical characteristics of real data, and normalize it according to the characteristic parameters of single-cell sequencing technology to generate the imputation gene expression profile; (3) The normalized gene expression matrix is input into algorithms , to obtain feature vectors , ; (4) using eigenvectors , An optimization function O is set, and the gene expression matrix is optimized for multiple rounds using the optimization function, and finally the optimized matrix is denormalized to obtain the interpolated gene expression profile. (5) Re-enter the obtained gene expression profile into the algorithm , Two sets of predicted feature vectors were obtained. , ; (6) Using eigenvectors , Calculate the data quality difference value D.
2. The method for assessing the quality of single-cell sequencing data as described in claim 1, characterized in that... The method for obtaining the interpolated gene expression profile in step (4) is as follows: Using algorithms , eigenvectors of gene expression matrix X , Set the optimization function O: ; Where y represents the y-th type of interpolated gene spectrum. This represents the total number of features. Then, the matrix X is optimized in multiple rounds using an optimization function. ; in Represents the learning rate. This represents the matrix optimized up to round i. This process is repeated until the preset number of rounds is reached. Finally, the optimized matrix is denormalized to obtain the imputed gene expression profile.
3. The method for assessing the quality of single-cell sequencing data as described in claim 1, characterized in that: The method for calculating the data quality difference value in step (6) is as follows: Re-enter the interpolated gene expression profile into the algorithm. , The algorithm obtains the feature values of each imputed gene expression profile when it is predicted as a feature, and then calculates the average of these features to obtain the mean of the feature values of all imputed gene expression profiles of class I when they are predicted as class J. , ; ; ; in The expression profile of the i-th imputed gene was represented by the algorithm. The feature value when identified as class j. The expression profile of the i-th imputed gene was represented by the algorithm. The feature value when identified as class j. Then, using the features of all classes. The sum of the absolute values of the differences can be used to calculate the data quality variance value D: 。