Single cell RNA sequencing data recovery and feature extraction method and system based on collaborative modeling

By synergistically integrating variational autoencoders, generative adversarial networks, and compressed sensing models, and combining them with the Ivy optimization algorithm, the problems of expression recovery and feature extraction in single-cell RNA sequencing data analysis were solved, achieving high-quality data recovery and stable feature extraction, which is suitable for bioinformatics analysis of multiple datasets.

CN121306245APending Publication Date: 2026-01-09JILIN INST OF CHEM TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511455423.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing single-cell RNA sequencing data analysis suffers from problems such as insufficient expression recovery quality, fragmented modeling strategies, unstable sparse optimization, and lack of universality in feature extraction. It is difficult to simultaneously take into account the global expression structure and the local detail recovery of low-expression regions, and it lacks stability for cross-dataset validation.

Method used

We employ a combination of variational autoencoders and generative adversarial networks to construct a compressed sensing model and introduce the Ivy League optimization algorithm. Through the synergistic integration of multiple modeling mechanisms, we perform expression recovery and feature extraction. We combine high-variability screening and differential expression analysis, and cross-validate across multiple datasets to screen out stable key gene features.

Benefits of technology

It significantly improves the recovery accuracy and feature extraction stability of single-cell RNA sequencing data, is suitable for cell clustering and disease mechanism mining, provides a general and robust modeling tool, and improves data integrity and the reliability of key feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306245A_ABST
    Figure CN121306245A_ABST
Patent Text Reader

Abstract

The invention discloses a single-cell RNA sequencing data recovery and feature extraction method and system based on collaborative modeling, and relates to the technical field of bioinformatics and artificial intelligence. The method comprises a data acquisition step, a data preprocessing step, a recovery step, an optimization step, a first extraction step, a screening step and a second extraction step. According to the method, the data integrity and the downstream feature extraction accuracy are effectively improved, and the method is suitable for bioinformatics application scenes such as cancer heterogeneity analysis and candidate gene screening.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics and artificial intelligence, and particularly relates to a single-cell RNA sequencing data recovery and feature extraction method and system based on collaborative modeling. BACKGROUND

[0002] Single-cell RNA sequencing (scRNA-seq) is an important means for depicting cell transcriptome heterogeneity, and has been widely applied to tumor typing, developmental regulation, immune microenvironment and other biomedical research directions. Due to its ability to capture all gene expressions in a single cell at one time, scRNA-seq provides unprecedented analytical ability to reveal tissue complexity.

[0003] However, due to technical bottlenecks such as library preparation, amplification efficiency and sequencing depth, scRNA-seq data often exhibits severe sparsity and high Dropout ratio, that is, genes that originally exist are mismeasured as zero. This phenomenon not only weakens the true expression structure of the data, but also seriously interferes with downstream analysis tasks such as cell clustering, differential gene identification and pathway enrichment. Therefore, how to recover the Dropout region with high quality and reconstruct a reliable expression matrix has become one of the key challenges in the current single-cell analysis field.

[0004] To address the above problems, a variety of recovery strategies have been proposed. Traditional methods such as MAGIC, SAVER, etc. fill in the missing values through graph diffusion or Bayesian modeling, but have limitations in capturing nonlinear relationships and maintaining expression gradients. In recent years, deep generative models have been introduced into the expression recovery task. Among them, the Variational Autoencoder (VAE) realizes the modeling of complex expression distribution by constructing a latent space, and the Generative Adversarial Network (GAN) improves the authenticity of the generated results through a discriminant mechanism, and the combination of the two can achieve a good balance between global structure and local details.

[0005] On the other hand, compressed sensing (CS) theory is widely used in signal reconstruction. By constructing sparse dictionaries and sparse coefficient expression models, it can achieve high-fidelity reconstruction of the original signal with limited observations. This theory has been gradually introduced into scRNA-seq recovery tasks to capture the sparse structural features of low-expression regions. However, single sparse modeling often gets stuck in local optima. To address this, heuristic optimization algorithms such as genetic algorithms, particle swarm optimization, and Ivy algorithms have been introduced into the sparse coefficient solution process to improve convergence stability and global optimum capability.

[0006] After data restoration, the extraction of characteristic genes places further demands on model performance. Existing methods often directly perform hypervariable gene screening and differential expression analysis based on the original or preliminarily restored expression matrix, lacking unified standards and stability controls, particularly in extracting common features from multiple datasets. Furthermore, most current studies treat feature screening and data restoration as separate processes, making it difficult to construct a closed-loop process from expression reconstruction to key gene extraction.

[0007] Therefore, it can be seen that current scRNA-seq data analysis faces the following technical bottlenecks: 1. Insufficient quality of expression recovery: Existing methods have difficulty simultaneously recovering both the global expression structure and the local details of low-expression regions, and the recovered data still has biases; 2. Fragmented modeling strategies: Deep generative models and sparse optimization methods are often used in isolation, lacking a collaborative mechanism under a unified framework, which limits the accuracy and stability of reconstruction. 3. The sparse optimization process is unstable: Traditional compressed sensing solutions rely on initial conditions, are prone to getting trapped in local optima, and lack efficient global search strategies; 4. Feature extraction lacks universality: Most methods do not introduce cross-dataset validation or stability screening, resulting in a lack of robustness in the screening results, making it difficult to generalize to different samples or clinical application scenarios.

[0008] To address the aforementioned issues, this paper proposes a method and system for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling. This method framework integrates multiple modeling mechanisms while simultaneously considering the quality of expression recovery and the stability of feature extraction, aiming to improve the accuracy, robustness, and interpretability of scRNA-seq data analysis. This is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0009] In view of this, this invention provides a method and system for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling. It combines a variational autoencoder (VAE) with a generative adversarial network (GAN) to capture the nonlinear latent structure of gene expression and generate preliminary recovery results. Subsequently, a compressed sensing model is constructed, and the Ivy League optimization algorithm is introduced to iteratively update the sparse coefficient matrix, thereby further improving the recovery accuracy and structural consistency of the expression matrix. Based on the completed data recovery, the method continues to screen for highly variable genes and perform differential expression analysis, and cross-validates across multiple single-cell datasets, retaining only gene features that repeat in two or more datasets. Furthermore, a stable and sparse subset of key genes is screened using a Lasso regression model, enhancing the interpretability and generality of downstream analysis.

[0010] To achieve the above objectives, the present invention adopts the following technical solution: A method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling includes the following steps: S1 Data Acquisition Steps: Obtain the raw single-cell RNA sequencing expression matrix; S2 Data Preprocessing Steps: Perform cell and gene quality control and normalization on the original single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix; S3 Recovery Step: Input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix; S4 optimization steps: Establish a compressed sensing model for the first-stage recovered expression matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the second-stage optimized expression matrix; S5 First extraction step: In the second stage, high-variability genes are screened on the optimized expression matrix to extract a subset of genes with significant expression variations; S6 Screening Steps: Perform differential expression analysis based on cell grouping to screen for genes that show significant differences in expression among different populations from hypervariable genes; S7 Second extraction step: Differentially expressed genes from different datasets are cross-preserved to form a final subset of feature genes for downstream bioinformatics analysis tasks.

[0011] Optionally, the original single-cell RNA sequencing expression matrix in S1 is as described above. ,in g Indicates the number of genes. n Indicates the number of cells.

[0012] Optionally, the following details the quality control and normalization of the original single-cell RNA sequencing expression matrix in S2, as described above: Quality control: Removal from the original single-cell RNA sequencing expression matrix Low-quality cells with fewer than 200 expressed genes and low-complexity genes expressed in fewer than 3 cells were detected. Normalization: expression matrix of raw single-cell RNA sequencing Each expression value Perform logarithmic normalization to obtain the normalized representation matrix. The conversion formula is as follows:

[0013] in, This serves as the foundational representation data for subsequent modeling and feature extraction.

[0014] Optionally, the generative model in S3 includes a shared encoder, as described above. Generator Discriminator ; Shared encoder Used for normalized expression matrices Mapping to latent space representation Model its conditional distribution ; Generator , used to classify latent variables Map back to the expression matrix space to generate the first-stage recovered expression matrix. ; Discriminator Used to determine the first-stage recovery expression matrix With normalized expression matrix Improve the discriminative power and enhance the recovery of the expression matrix in the first stage. Consistency in expression.

[0015] The above method, optionally, can be used to enhance the generative model's response to the normalized representation matrix. To enhance the sensitivity of sparse region recovery, a weighting mechanism based on element value state is introduced into the reconstruction error term of the joint loss function. That is, higher reconstruction weight is given to the position with expression value of zero, so as to improve the ability to identify and recover missing regions. The weighted loss is in the following form: ,

[0016] Finally, the generative models jointly optimize the following loss function:

[0017] in, and is a hyperparameter used to control the weights of the fused variational autoencoder (VAE) and the generative adversarial network (GAN) branch during joint training.

[0018] Optionally, the compressed sensing sparse modeling process in S4 includes the following steps: The first-stage recovery expression matrix generated using the generative model in S3 As input, construct the compressed sensing model and build the dictionary matrix. With sparse coefficient matrix ,satisfy:

[0019] in, , where is the sparse representation dimension in compressed sensing; By minimizing the following bands Solving the sparse coefficient matrix of the objective function of the regularization term W :

[0020] in, It is the Frobenius norm. Regularization coefficients used to control sparsity; S4 uses the Ivy Algorithm to process the sparse coefficient matrix. W For global optimization, the iterative update formula is:

[0021] in, Indicates the first i Each individual, i.e., the current solution; The individual with the best current fitness; As a growth factor, it controls the growth rate towards the best neighbor; It is a random expansion factor; It is Gaussian noise; Finally passed D·W The second-stage optimized representation matrix after compressed sensing optimization is obtained. .

[0022] Optionally, the hypervariable gene screening process in S5 includes the following steps: Optimize the expression matrix for the second stage of S4 The expression variance of each gene in the dataset is calculated. All genes in each dataset are sorted in descending order of expression variance value, and the top [genes] are selected. N A set of highly variable genes constitutes a collection of highly variable genes. ; The current dataset has insufficient expressed genes. N At that time, all expressed genes were retained; finally, multiple hypervariable gene expression matrices were generated. This is used for subsequent differential expression analysis.

[0023] Optionally, the differentially expressed gene screening process in S6 includes the following steps: For each hypervariable gene expression matrix Dimensionality reduction was performed, and a KMeans clustering model was constructed based on its principal components to divide the cell samples into several populations; For each cluster and the remaining cell population, the Wilcoxon rank-sum test was used to analyze the expression differences of each gene, including the logarithmic fold change and the significance level after multiple hypothesis correction. Only genes that meet the expression range criteria are retained to form the differentially expressed gene set. : .

[0024] Optionally, the process of constructing a feature subset in S7 includes the following steps: S1 through S6 were executed on multiple single-cell scRNA-seq datasets to obtain the corresponding feature gene sets for each dataset. , , ; Multiple differentially expressed gene sets Cross-selection is performed to retain only genes that appear in at least two datasets, forming a stable feature set. This is used for subsequent biological modeling or classification analysis, and to stabilize the feature set. The expression is as follows: .

[0025] A single-cell RNA sequencing data recovery and feature extraction system based on collaborative modeling, used to implement the single-cell RNA sequencing data recovery and feature extraction method based on collaborative modeling as described above, comprising a data acquisition module, a data preprocessing module, a recovery module, an optimization module, a first extraction module, a screening module, and a second extraction module connected in sequence; The data acquisition module is used to acquire the raw single-cell RNA sequencing expression matrix; The data preprocessing module is used to perform cell and gene quality control and normalization on the raw single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix. The recovery module is used to input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix. The optimization module is used to establish a compressed sensing model for the first-stage recovery matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the optimized representation matrix for the second stage. The first extraction module is used to screen for highly variable genes on the optimized expression matrix and extract a subset of genes with significant expression variations. The screening module is used to perform differential expression analysis based on cell grouping, and to screen out genes that express significantly different values ​​among different populations from hypervariable genes; The second extraction module is used to cross-preserve differentially expressed genes from different datasets to form a final subset of feature genes for downstream bioinformatics analysis tasks.

[0026] As can be seen from the above technical solution, compared with the prior art, the present invention provides a method and system for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling, which has the following beneficial effects: 1) This invention is the first to fuse variational autoencoders and generative adversarial networks for scRNA-seq expression data recovery, effectively modeling the nonlinear structure of gene expression and improving the recovery ability of Dropout regions; 2) By introducing compressed sensing theory and the Ivy optimization algorithm to construct sparse representation, the sparsity representation accuracy and structural consistency of the recovery results are significantly improved. 3) After completing the expression matrix recovery, a feature extraction method integrating high-variability screening and differential expression analysis was proposed, and a cross-dataset cross-retention strategy was designed to effectively improve the stability of feature selection results; 4) It can effectively improve the integrity of single-cell sequencing data and the reliability of key feature extraction. It is suitable for scRNA-seq analysis workflows such as cell segmentation and disease mechanism mining, and provides a universal and robust modeling tool for tumor heterogeneity research and personalized medicine. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0028] Figure 1 This is a flowchart of a single-cell RNA sequencing data recovery and feature extraction method based on collaborative modeling disclosed in this invention; Figure 2 This is a schematic diagram of the principle framework of a single-cell RNA sequencing data recovery and feature extraction method based on collaborative modeling disclosed in this invention; Figure 3 This is a block diagram of a single-cell RNA sequencing data recovery and feature extraction system based on collaborative modeling disclosed in this invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0031] See Figure 1 and Figure 2 As shown, this invention discloses a method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling, comprising the following steps: S1 Data Acquisition Steps: Obtain the raw single-cell RNA sequencing expression matrix; S2 Data Preprocessing Steps: Perform cell and gene quality control and normalization on the original single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix; S3 Recovery Step: Input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix; S4 optimization steps: Establish a compressed sensing model for the first-stage recovered expression matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the second-stage optimized expression matrix; S5 First extraction step: In the second stage, high-variability genes are screened on the optimized expression matrix to extract a subset of genes with significant expression variations; S6 Screening Steps: Perform differential expression analysis based on cell grouping to screen for genes that show significant differences in expression among different populations from hypervariable genes; S7 Second extraction step: Differentially expressed genes from different datasets are cross-preserved to form a final subset of feature genes for downstream bioinformatics analysis tasks.

[0032] Furthermore, the original single-cell RNA sequencing expression matrix in S1 is as follows: ,in g Indicates the number of genes. n Indicates the number of cells.

[0033] Furthermore, the specific details of cell and gene quality control and normalization of the original single-cell RNA sequencing expression matrix in S2 are as follows: Quality control: Removal from the original single-cell RNA sequencing expression matrix Low-quality cells with fewer than 200 expressed genes and low-complexity genes expressed in fewer than 3 cells were detected. Normalization: expression matrix of raw single-cell RNA sequencing Each expression value Perform logarithmic normalization to obtain the normalized representation matrix. The conversion formula is as follows:

[0034] in, This serves as the foundational representation data for subsequent modeling and feature extraction.

[0035] Furthermore, the generative model in S3 includes: a shared encoder. Generator Discriminator ; Shared encoder Used for normalized expression matrices Mapping to latent space representation Model its conditional distribution ; Generator , used to classify latent variables Map back to the expression matrix space to generate the first-stage recovered expression matrix. ; Discriminator Used to determine the first-stage recovery expression matrix With the normalized expression matrix Improve the discriminative power and enhance the generation of the first-stage recovered expression matrix. Consistency in expression.

[0036] Furthermore, to enhance the generative model's response to the normalized representation matrix... To enhance the sensitivity of sparse region recovery, a weighting mechanism based on element value state is introduced into the reconstruction error term of the joint loss function. That is, higher reconstruction weight is given to the position with expression value of zero, so as to improve the ability to identify and recover missing regions. The weighted loss is in the following form: ,

[0037] Finally, the generative models jointly optimize the following loss function:

[0038] in, and is a hyperparameter used to control the weights of the fused variational autoencoder (VAE) and the generative adversarial network (GAN) branch during joint training.

[0039] Furthermore, the compressed sensing sparse modeling process in S4 includes: The first-stage recovery expression matrix generated using the generative model in S3 As input, construct the compressed sensing model and build the dictionary matrix. With sparse coefficient matrix ,satisfy:

[0040] in, , where is the sparse representation dimension in compressed sensing; By minimizing the following bands Solving the sparse coefficient matrix of the objective function of the regularization term W :

[0041] in, It is the Frobenius norm. Regularization coefficients used to control sparsity; S4 uses the Ivy Algorithm to process the sparse coefficient matrix. W For global optimization, the iterative update formula is:

[0042] in, Indicates the first i Each individual, i.e., the current solution; The individual with the best current fitness; As a growth factor, it controls the growth rate towards the best neighbor; It is a random expansion factor; Gaussian noise enhances the diversity of exploration.

[0043] Finally passed D·W The second-stage optimized representation matrix after compressed sensing optimization is obtained. .

[0044] Furthermore, the hypervariable gene screening process in S5 includes: Optimize the expression matrix for the second stage of S4 The expression variance of each gene in the dataset is calculated. All genes in each dataset are sorted in descending order of expression variance value, and the top [genes] are selected. N A set of highly variable genes constitutes a collection of highly variable genes. ; The current dataset has insufficient expressed genes. N At that time, all expressed genes were retained; finally, multiple hypervariable gene expression matrices were generated. This is used for subsequent differential expression analysis.

[0045] Furthermore, the differentially expressed gene screening process in S6 includes: For each hypervariable gene expression matrix Dimensionality reduction was performed, and a KMeans clustering model was constructed based on its principal components to divide the cell samples into several populations; For each cluster and the remaining cell population, the Wilcoxon rank-sum test was used to analyze the expression differences of each gene, including the logarithmic fold change and the significance level after multiple hypothesis correction. Only genes that meet the expression range criteria are retained to form the differentially expressed gene set. : .

[0046] Furthermore, the process of constructing feature subsets in S7 includes: S1 through S6 were executed on multiple single-cell scRNA-seq datasets to obtain the corresponding feature gene sets for each dataset. , , ; Multiple differentially expressed gene sets Cross-selection is performed to retain only genes that appear in at least two datasets, forming a stable feature set. This is used for subsequent biological modeling or classification analysis, and to stabilize the feature set. The expression is as follows: .

[0047] and Figure 1 and Figure 2 Corresponding to the method shown, this invention also discloses a single-cell RNA sequencing data recovery and feature extraction system based on collaborative modeling, used to achieve... Figure 1 and Figure 2 The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling is shown in the diagram below. Figure 3 As shown, it includes a data acquisition module, a data preprocessing module, a recovery module, an optimization module, a first extraction module, a filtering module, and a second extraction module connected in sequence. The data acquisition module is used to acquire the raw single-cell RNA sequencing expression matrix; The data preprocessing module is used to perform cell and gene quality control and normalization on the raw single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix. The recovery module is used to input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix. The optimization module is used to establish a compressed sensing model for the first-stage recovered expression matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the second-stage optimized expression matrix. The first extraction module is used to screen for highly variable genes on the expression matrix optimized in the second stage and extract a subset of genes with significant expression variations. The screening module is used to perform differential expression analysis based on cell grouping, and to screen out genes that express significantly different values ​​among different populations from hypervariable genes; The second extraction module is used to cross-preserve differentially expressed genes from different datasets to form a final subset of feature genes for downstream bioinformatics analysis tasks.

[0048] In a specific embodiment: A method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling is illustrated in the attached flowchart. Figure 2 As shown, it includes the following steps: Step 1: Preprocess the raw single-cell RNA sequencing expression matrix, including quality control and normalization of cells and genes.

[0049] In this embodiment, the original single-cell RNA sequencing expression matrix is ​​denoted as... ,in g Indicates the number of genes. n This indicates the number of cells. Each row of elements... Indicates the first j The first cell i The original expression values ​​of each gene. To improve data quality and ensure the stability and consistency of subsequent modeling, standardized quality control and normalization processing needs to be performed on the original data.

[0050] During cell quality control, cells with fewer than 200 expressed genes detected in the expression matrix were removed to exclude sequencing failures or biologically unrepresentative samples. During gene quality control, genes expressed in fewer than 3 cells were removed to avoid interference from low-complexity background noise on modeling accuracy. To maintain symbol consistency, the dimension of the quality-controlled expression matrix is ​​still denoted as [dimension number missing]. , among them g and n This indicates the number of genes and cells retained.

[0051] Log-normalization is performed on the quality-controlled expression matrix X to generate a normalized expression matrix. This transformation is used to mitigate the impact of inconsistent sequencing depths on expression intensity and to compress the dynamic range of the data. Its calculation formula is as follows:

[0052] Where ln represents the natural logarithm function, and the constant term 1 ensures that the logarithm is still valid even when the original value is 0. Normalized expression matrix As input to subsequent deep generative models, it establishes the foundation for modeling the latent structure of the expressive data.

[0053] This example selects three publicly available breast cancer single-cell RNA sequencing datasets from the public GEO database: GSE123358, GSE124989, and GSE147326. The aforementioned quality control and normalization procedures were independently performed on each dataset, resulting in a normalized expression matrix. Sufficient cell expression information was retained and used as initial data for each modeling process. Table 1 shows the dimensionality information of each dataset and the dimensionality transformation information after quality control.

[0054] Table 1. Dataset Dimensions Before and After Quality Control

[0055] Step 2: Input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and generative adversarial network (GAN) to obtain the first-stage recovered expression matrix.

[0056] To learn the nonlinear latent structure of the representation matrix and achieve preliminary Dropout recovery, this invention employs a deep generative model structure that integrates a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). The input is the normalized representation matrix processed in step 1. The output is the first-stage recovery expression matrix. .

[0057] In this example, X is used. Represents the complete cell-gene expression matrix. A column vector representing gene expression in a single cell sample.

[0058] The model structure consists of a shared encoder. Generator Discriminator Composition. The encoder encodes the input representation matrix into parameters of the latent variable distribution, i.e.:

[0059] in, Indicates mean, It is expressed as variance.

[0060] Latent variables are obtained by sampling from a normal distribution using parameterization techniques. :

[0061] Generator Receive latent variable z and output preliminary reconstructed expression matrix As the first-stage output of Dropout recovery, This represents the predicted expression value of the i-th gene in the j-th cell.

[0062] At the level of a single sample, the generator's output can be represented as:

[0063] Discriminator Used to distinguish generator matrices Is the real sample x derived from the real normalized representation matrix? This is to enhance the realism of the expression pattern in both local and overall structure.

[0064]

[0065] To improve the quality of representation recovery for Dropout regions, this implementation introduces multiple joint loss functions during model training. These functions, along with latent space regularization and adversarial modeling, jointly constrain the consistency of the generated representation and the realism of the reconstruction. The complete loss function structure is shown below:

[0066] The first term is the reconstruction error between the generator output and the true normalized matrix, which measures the model's ability to reconstruct the original expression values; the second term is the KL divergence term, which is used to restrict the distribution of latent variables. The specific expansion form is shown below, which is used to encourage the latent space to maintain consistency with the standard normal distribution.

[0067]

[0068] Meanwhile, to ensure that the generator can produce an expression matrix consistent with the distribution of real samples, a discriminator is introduced for adversarial constraints. The discriminator is trained to distinguish the differences between generated samples and real samples, while the generator attempts to maximize the discriminator's "false positive rate." This adversarial process forms the basis of GAN training, and its mathematical form is shown below.

[0069]

[0070] To achieve unified optimization, the three losses mentioned above are weighted and integrated into a final hybrid loss function, as shown below:

[0071] in and is a hyperparameter used to control the weights of each sub-item in joint training.

[0072] To enhance the recovery performance of the Dropout region, this implementation introduces a Dropout-aware weighting mechanism into the reconstruction loss, assigning higher weights to elements with values ​​of zero to highlight their recovery priority. The specific weighted loss is defined as follows: ,

[0073] in, The weight coefficient represents the positional weight of each element when calculating the loss function; this weight depends on whether that position in the original data is Dropout (i.e., its value is 0).

[0074] This mechanism enhances the model's ability to perceive Dropout regions and is an effective optimization for addressing technical deficiencies in single-cell sequencing.

[0075] After the final training is completed, the generator outputs the expression matrix. The results of the first phase of recovery will be used in subsequent compressed sensing modeling and sparse feature optimization processes.

[0076] Step 3: Establish a compressed sensing model for the first-stage recovered expression matrix, construct a sparse representation and observation matrix, and use the Ivy Algorithm to optimize the sparse coefficients, outputting the second-stage optimized expression matrix.

[0077] To further enhance the first-stage recovery expression matrix To maintain structural integrity and expressive compactness, this implementation method constructs a sparse modeling process based on compressed sensing theory and uses the Ivy Population Search algorithm to achieve global optimization of the sparse coefficient matrix, thereby obtaining a more stable expression matrix with biological interpretability.

[0078] In this process, the expression matrix is ​​first restored in the first stage. The model is in the form of a sparse dictionary product, that is:

[0079] in, Represents a sparse dictionary matrix. Let be the sparse coefficient matrix to be solved. To ensure expressive power under high compression ratios, this paper employs singular value decomposition (SVD) to concisely construct a high-quality dictionary matrix. Feature extraction is performed, and the first d principal components are used as a dictionary:

[0080] After obtaining the dictionary, in order to achieve expression compression, we further construct a compressed sensing measurement matrix. This is used to perform a linear projection of the expression matrix and form the observation matrix:

[0081] in, For the observation matrix, Let m be the compressed dictionary representation, where m represents the compressed dimension, satisfying... .

[0082] To initialize the sparse coefficient matrix W, we first perform preliminary encoding on the observation matrix Y using a Lasso regularization model, with the goal of minimizing the sum of the reconstruction error and the sparse terms:

[0083] Building upon this, to further improve the local accuracy of sparse representation, we introduce the Orthogonal Matching Pursuit (OMP) algorithm, which performs a process on each column of observed samples. Perform sparse coefficient initialization: ) Where k is a sparsity control parameter used to limit the number of non-zero elements, the resulting... This will serve as the initial population for global optimization.

[0084] To overcome the local optima limitation, we introduce the Ivy Algorithm to perform a population search optimization on the sparse coefficient matrix W. This algorithm optimizes each column in each iteration. Treat each i-th individual as a single entity and update it with perturbations based on the current population state. In the t-th iteration, the update rule for the i-th entity is as follows:

[0085] in, The individual with the best current fitness. To control the rate of approach for growth factors, Control the magnitude of search perturbations. Growth factors are used to accommodate structural differences between individuals. A dynamic adjustment mechanism is adopted:

[0086] To measure the quality of individual solutions, we define the fitness function as the Pearson correlation error between the reconstructed result and the original sample:

[0087] in, This represents the Pearson correlation coefficient. Iteration terminates when the error decrease is less than a threshold for two consecutive rounds or when the maximum number of iterations is reached. The final output is the globally optimal sparse coefficient matrix. .

[0088] Finally, the optimized sparse expression coefficients were used. The expression is restored, and the final expression matrix is ​​reconstructed to obtain the second-stage optimized expression matrix. The expression is as follows:

[0089] Second stage: Optimize the expression matrix The expression patterns of the initial recovery results were preserved, and the structural consistency of low-expression regions was further improved through compression mapping and sparse optimization, providing a higher quality input basis for subsequent feature selection and biological analysis.

[0090] Step 4: Screen for highly variable genes on the optimized expression matrix and extract a subset of genes with significant expression variations.

[0091] To further extract key genes with high expression variability and potential bioinformatics characteristics, this implementation method is based on the second-stage optimized expression matrix constructed in step 3. Perform high-variance gene screening. This screening process uses the expression variance of each gene in the entire sample as the evaluation index to sort and select genes, and prioritizes retaining a subset of genes with large fluctuations and significant differences in expression levels.

[0092] In practice, to improve the stability and universality of gene screening, variability assessment was performed independently on three single-cell datasets. For each dataset, the top N genes with the highest expression variance were selected from the structure-optimized expression matrix, where N is a preset screening upper limit, recommended to be 10,000. If the number of expressed genes in a dataset is less than this threshold, all expressed genes are retained for subsequent analysis. The above screening process was performed on the three single-cell datasets, and the screening results are shown in Table 2.

[0093] During feature selection, the extracted set of highly variable genes was uniformly mapped to a standard gene name space and cross-referenced. Specifically, highly variable genes that appeared repeatedly in at least two datasets were retained, forming a feature subset with cross-sample consistency and robustness. In this way, this embodiment ultimately obtained 735 candidate genes that were cross-selected. These genes showed high expression variability across multiple datasets and were suitable as the initial feature space for subsequent differential analysis and modeling.

[0094] Finally, hypervariable genes were screened in three single-cell datasets: GSE123358, GSE124989, and GSE147326. Based on the top 10,000 genes with the highest expression variance retained from each dataset, a corresponding hypervariable gene expression matrix set was constructed. Each of the above matrices serves as input data for subsequent differential expression analysis of its respective dataset, forming independent but structurally consistent feature subspaces, providing a foundation for group mining and cross-gene integration.

[0095] Table 2. Screening for hypervariable genes

[0096] Step 5: Perform differential expression analysis based on cell grouping to screen out genes with significant differences in expression among different populations from the hypervariable genes.

[0097] To further explore the systematic differences in expression levels among cell populations, this implementation method independently performs differential expression analysis in each dataset to identify differentially expressed genes (DEGs) associated with intercellular expression heterogeneity. The input used in this step is the hypervariable gene expression matrix extracted in step 4, denoted as... Each matrix corresponds to a dataset, with columns representing cell samples and genes exhibiting the highest behavioral variability.

[0098] Considering that some datasets lack clear clinical group labels, this implementation method employs unsupervised clustering to group cell samples for automated population segmentation. Specifically, Principal Component Analysis (PCA) is first used to reduce the dimensionality of each expression matrix to a lower-dimensional expression space to preserve the dominant variation structure. Within this expression space, the KMeans clustering algorithm is applied to divide the samples into several groups. The number of clusters, k, can be determined empirically or automatically. Each cluster represents a subpopulation of cells with similar expression patterns, forming structurally distinct sample labels.

[0099] After clustering, differential expression analysis was performed between each cluster. For each cluster, its cells were designated as the target group, and the remaining cells as the control group. The Wilcoxon rank-sum test was used to statistically compare all hypervariable genes, and the log2 fold change and adjusted p-value were calculated. In the analysis results, only genes that met both of the following conditions were retained as differentially expressed genes (DEGs): The magnitude of the change in expression satisfies:

[0100] The significance level is satisfied:

[0101] After completing the above analysis, an independent list of differentially expressed genes was generated for each dataset, denoted as follows: These differentially expressed genes represent key candidate molecules that may exhibit upregulated or repressed expression trends in specific cell subpopulations, and are an important basis for subsequent cross-integration and feature selection.

[0102] Step 6: Cross-preserve differentially expressed genes from different datasets, and then use the LASSO regression model to perform further feature selection to form the final subset of feature genes for downstream bioinformatics analysis tasks.

[0103] To further compress the feature space and improve the stability and generalization ability of key genes in different datasets, this implementation method performs cross-integration and feature selection operations on the results after completing the differential expression analysis of each dataset, so as to finally form a subset of key genes with representativeness and predictive power.

[0104] Specifically, let the differentially expressed gene sets of each dataset obtained in step 5 be respectively... This step first performs a cross-retention operation on these three gene sets, that is, retains differentially expressed genes that appear in at least two datasets, thereby constructing a set of gene features with consistent differential expression across multiple independent sample conditions:

[0105] After this cross-selection process, a total of 735 genes were obtained, denoted as... As a set of candidate characteristic genes, it provides stable and reliable basic variables to support subsequent risk modeling, biological mechanism analysis or prognostic assessment.

[0106] In this embodiment, the step of obtaining the original single-cell RNA sequencing expression matrix is ​​included before step 1.

[0107] For the system or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and relevant details can be found in the description of the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0108] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling, characterized in that, Includes the following steps: S1 Data Acquisition Steps: Obtain the raw single-cell RNA sequencing expression matrix; S2 Data Preprocessing Steps: Perform cell and gene quality control and normalization on the original single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix; S3 Recovery Step: Input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix; S4 optimization steps: Establish a compressed sensing model for the first-stage recovered expression matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the second-stage optimized expression matrix; S5 First extraction step: In the second stage, high-variability genes are screened on the optimized expression matrix to extract a subset of genes with significant expression variations; S6 Screening Steps: Perform differential expression analysis based on cell grouping to screen for genes that show significant differences in expression among different populations from hypervariable genes; S7 Second extraction step: Differentially expressed genes from different datasets are cross-preserved to form a final subset of feature genes for downstream bioinformatics analysis tasks.

2. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 1, characterized in that, The original single-cell RNA sequencing expression matrix in S1 is as follows: ,in g Indicates the number of genes. n Indicates the number of cells.

3. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 2, characterized in that, The specific details of cell and gene quality control and normalization of the original single-cell RNA sequencing expression matrix in S2 are as follows: Quality control: Removal from the original single-cell RNA sequencing expression matrix Low-quality cells with fewer than 200 expressed genes and low-complexity genes expressed in fewer than 3 cells were detected. Normalization: expression matrix of raw single-cell RNA sequencing Each expression value Perform logarithmic normalization to obtain the normalized representation matrix. The conversion formula is as follows: in, This serves as the foundational representation data for subsequent modeling and feature extraction.

4. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 3, characterized in that, The generative models in S3 include: shared encoders Generator Discriminator ; Shared encoder Used for normalized expression matrices Mapping to latent space representation Model its conditional distribution ; Generator , used to classify latent variables Map back to the expression matrix space to generate the first-stage recovered expression matrix. ; Discriminator Used to determine the first-stage recovery expression matrix With normalized expression matrix Improve the discriminative power and enhance the recovery of the expression matrix in the first stage. Consistency in expression.

5. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 4, characterized in that, To enhance the generative model's response to the normalized representation matrix To enhance the sensitivity of sparse region recovery, a weighting mechanism based on element value state is introduced into the reconstruction error term of the joint loss function. That is, higher reconstruction weight is given to the position with expression value of zero, so as to improve the ability to identify and recover missing regions. The weighted loss is in the following form: , Finally, the generative models jointly optimize the following loss function: in, and is a hyperparameter used to control the weights of the fused variational autoencoder (VAE) and the generative adversarial network (GAN) branch during joint training.

6. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 5, characterized in that, The compressed sensing sparse modeling process in S4 includes: The first-stage recovery expression matrix generated using the generative model in S3 As input, construct the compressed sensing model and build the dictionary matrix. With sparse coefficient matrix ,satisfy: in, , where is the sparse representation dimension in compressed sensing; By minimizing the following bands Solving the sparse coefficient matrix of the objective function of the regularization term W : in, It is the Frobenius norm. Regularization coefficients used to control sparsity; S4 uses the Ivy Algorithm to process the sparse coefficient matrix. W For global optimization, the iterative update formula is: in, Indicates the first i Each individual, i.e., the current solution; The individual with the best current fitness; As a growth factor, it controls the growth rate towards the best neighbor; It is a random expansion factor; It is Gaussian noise; Finally passed D·W The second-stage optimized representation matrix after compressed sensing optimization is obtained. .

7. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 6, characterized in that, The hypervariable gene screening process in S5 includes: Optimize the expression matrix for the second stage of S4 The expression variance of each gene in the dataset is calculated. All genes in each dataset are sorted in descending order of expression variance value, and the top [genes] are selected. N A set of highly variable genes constitutes a collection of highly variable genes. ; The current dataset has insufficient expressed genes. N At that time, all expressed genes were retained; finally, multiple hypervariable gene expression matrices were generated. This is used for subsequent differential expression analysis.

8. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 7, characterized in that, The differentially expressed gene screening process in S6 includes: For each hypervariable gene expression matrix Dimensionality reduction was performed, and a KMeans clustering model was constructed based on its principal components to divide the cell samples into several populations; For each cluster and the remaining cell population, the Wilcoxon rank-sum test was used to analyze the expression differences of each gene, including the logarithmic fold change and the significance level after multiple hypothesis correction. Only genes that meet the expression range criteria are retained to form the differentially expressed gene set. : 。 9. The method for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling according to claim 8, characterized in that, The process of constructing a feature subset in S7 includes: S1 through S6 were executed on multiple single-cell scRNA-seq datasets to obtain the corresponding feature gene sets for each dataset. , , ; Multiple differentially expressed gene sets Cross-selection is performed to retain only genes that appear in at least two datasets, forming a stable feature set. This is used for subsequent biological modeling or classification analysis, and to stabilize the feature set. The expression is as follows: 。 10. A system for single-cell RNA sequencing data recovery and feature extraction based on collaborative modeling, characterized in that, A method for recovering and extracting features from single-cell RNA sequencing data based on collaborative modeling as described in any one of claims 1-9, comprising a data acquisition module, a data preprocessing module, a recovery module, an optimization module, a first extraction module, a screening module, and a second extraction module connected in sequence; The data acquisition module is used to acquire the raw single-cell RNA sequencing expression matrix; The data preprocessing module is used to perform cell and gene quality control and normalization on the raw single-cell RNA sequencing expression matrix to obtain the preprocessed expression matrix. The recovery module is used to input the preprocessed expression matrix into the generative model that fuses the variational autoencoder (VAE) and the generative adversarial network (GAN) to obtain the first-stage recovered expression matrix. The optimization module is used to establish a compressed sensing model for the first-stage recovery matrix, construct a sparse representation and observation matrix, optimize the sparse coefficients using the Ivy Algorithm, and output the optimized representation matrix for the second stage. The first extraction module is used to screen for highly variable genes on the optimized expression matrix and extract a subset of genes with significant expression variations. The screening module is used to perform differential expression analysis based on cell grouping, and to screen out genes that express significantly different values ​​among different populations from hypervariable genes; The second extraction module is used to cross-preserve differentially expressed genes from different datasets to form a final subset of feature genes for downstream bioinformatics analysis tasks.