Single cell gene completion method based on sparse perception diffusion model

By employing a sparse-perceived diffusion model and utilizing the U-Net network and Chebyshev index for sparse-corrected resampling, the problem of missing genes in single-cell gene expression data was solved, achieving high-fidelity and biologically reliable gene expression matrix reconstruction.

CN121331231APending Publication Date: 2026-01-13CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511504490.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively fill in up to 20k missing genes in single-cell gene expression data, especially when crossing platforms or disease cohorts, where consistency cannot be guaranteed. Furthermore, image diffusion models cannot be directly transferred to gene expression data, leading to pseudo-zero expression and loss of biological associations.

Method used

A method based on a sparse sensing diffusion model is adopted. Through a U-Net network, a sparse spatial state encoder, and a dynamic radiation domain decoder, Chebyshev index is used for sparse correction resampling to construct local smoothing priors and sparse correction priors, thereby achieving high-fidelity completion of gene expression.

Benefits of technology

High-fidelity completion was achieved in extremely missing single-cell gene expression data. The generated expression profile conforms to biological laws and has global biological consistency and local sparsity fidelity, which is significantly better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331231A_ABST
    Figure CN121331231A_ABST
Patent Text Reader

Abstract

The invention discloses a single-cell gene completion method based on a sparse perception diffusion model, and belongs to the technical field of gene expression completion, and the method comprises the following steps: carrying out forward noise addition diffusion on an original single-cell gene expression profile to obtain a noise sample; constructing a sparse correction resampling diffusion model; and inputting a noise sample into the sparse correction resampling diffusion model, and outputting a complete single-cell gene expression matrix after gene completion through a multi-step reverse denoising process. The complementation normal form provided by the invention can generate originally missing gene entries in an expression profile, perceive and correct sparse distribution deviation, provide consistent and reliable over 30k gene input for a basic model, and improve the robustness of the expression profile. Double constraints of geometric rearrangement and sparse correction are adopted, gene expression integrity and sparse structure authenticity are considered, and high-fidelity reconstruction of extreme deletion expression is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of gene expression complementation, and particularly relates to a single-cell gene complementation method based on a sparse perception diffusion model. BACKGROUND

[0002] With the continuous expansion of the training scale of single-cell basic models, researchers are trying to explore the transcriptional regulation rules of more than 30k gene dimensions in hundreds of millions of cells and hundreds of cell types. However, the existing high-throughput sequencing data only covers 10-12k genes, and there are a large number of technical false zeros, especially in cancer and rare disease samples, which makes the expression of key regulatory genes be covered up, and historical data is difficult to be directly used for new generation models. If the missing genes cannot be complemented, the model will be difficult to learn more extensive cell type-specific expression, and thus limit the precision medicine and drug target discovery. In order to break through this bottleneck, an algorithm is urgently needed to complement the missing severe scRNA-seq data to 30k gene scale and correct the bias caused by sparsity.

[0003] The core problem that the existing method cannot deal with is not the ordinary technical zero value, but the complete missing of a large number of genes in the sequencing stage, which leads to a serious lack of data coverage, which is far beyond the conventional technical missing. Imputation algorithms, such as (a) in Figure 1 , assume that cells of the same type share the complete gene set within the same batch, rely on neighborhood smoothing or low-rank assumption, and infer the true value of the noise gene with the observed gene. However, when the method is used in cross-platform or cross-disease cohort, the "same distribution" premise is invalid. Joint reconstruction methods, such as (b) in Figure 1 , introduce external high-quality cell alignment distribution, but are still limited to "known complete gene set", that is, can only be modeled within the original 10-12k gene subspace. The key problem of early data is not drop-out, but that as high as 20k genes have never been sequenced. The existing algorithm cannot "make up" these genes, and cannot guarantee the consistency of cross-batch and cross-platform input. Therefore, it is urgent to realize the super-space gene complementation of large-scale single-cell expression data, to provide a unified and accurate data preprocessing framework for the next generation of 30k-scale single-cell basic models, so as to overcome the inherent bottleneck of the existing method in cross-platform and cross-batch gene expansion.

[0004] Although diffusion models perform well in image inpainting and seem to be applicable to "filling in missing genes", scRNA-seq is quite different from images in data structure and statistical properties: images are continuous pixel grids, while scRNA-seq is extremely sparse vectors with dimension up to 30,000 and more than 80% zero values. Image diffusion models rely on the statistical correlation of neighboring pixels, while gene expression data lacks natural spatial structure. Forcing reshaping into a two-dimensional grid will scatter gene entries, destroy their local structure, and even degenerate into an all-zero solution. Therefore, image diffusion models are difficult to migrate directly, and the fundamental obstacle lies in the completely different assumptions of local smoothness. In addition, the activation patterns of genes in different cells are highly heterogeneous, making the diffusion process prone to "all diffusion to zero" lazy solution, which cannot capture biological correlations and will generate pseudo-zero expression. SUMMARY

[0005] The present application aims at the above-mentioned deficiencies in the prior art, and provides a single-cell gene completion method based on a sparse-aware diffusion model, so as to solve the problem that the existing image diffusion model is difficult to migrate directly, and the fundamental obstacle lies in the completely different assumptions of local smoothness; in addition, the activation patterns of genes in different cells are highly heterogeneous, making the diffusion process prone to "all diffusion to zero" lazy solution, which cannot capture biological correlations and will generate pseudo-zero expression.

[0006] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows: A single-cell gene completion method based on a sparse-aware diffusion model, comprising the following steps: S1, forward noise diffusion is performed on the original single-cell gene expression profile to obtain a noise sample; S2, a sparse correction resampling diffusion model is constructed; S3, the noise sample is input into the sparse correction resampling diffusion model, and through a multi-step reverse denoising process, a complete single-cell gene expression matrix after gene completion is output.

[0007] Further, in the S2, the sparse correction resampling diffusion model comprises a U-Net network, a sparse space state encoder and a dynamic radiation domain decoder; The U-Net network is used for predicting the input noise; The sparse space state encoder and the dynamic radiation domain decoder are cooperatively driven, and Chebyshev index is introduced to perform embedding coding and denoising processing on the predicted noise.

[0008] Further, the definition of the Chebyshev index specifically comprises the following steps: A1, rearranging the original expression of the gene into a grid, and defining a discrete Chebyshev distance; In the formula, Represents the discrete Chebyshev distance; Indicates the coordinate position in the grid; Indicates the maximum number of radiation domain layers; A2. Based on the discrete Chebyshev distance, divide the concentric radiation layers. : , The total capacity of the concentric radiation layers satisfies: In the formula, Indicates the innermost layer; The radius of the concentric circle; This represents the maximum number of concentric circles. The number of genes; A3. Genes are arranged and ordered using a distribution function, so that the sparse correction resampling diffusion model can obtain a local smooth prior. A4. The Chebyshev index representation for any concentric radiation layer is as follows: In the formula, Indicates the first The Chebyshev indexing operation is performed on the radiative layer; Indicates the first in the radiation domain Layer data; This indicates that the layers are flattened in sequence. This indicates that data is taken from a specific coordinate position within the radiation domain.

[0009] Furthermore, in A3, the arrangement function is: Its output is , This represents the coordinate position of each gene expression in the arrangement. The genes are arranged and ordered using an arrangement function, including: Genes are sorted from smallest to largest according to the radius of their respective spheres; Genes located in the same layer are ordered by traversing clockwise according to the Manhattan method. After the genes are arranged by the arrangement function, highly variable genes will be compressed into Chebyshev neighborhoods: Thus restore Even if the sparse correction resampling diffusion model obtains a locally smooth prior, then... To satisfy the set of all coordinate positions that meet the local smoothness prior; The coordinates of any value within the neighborhood; It is the conditional probability density function; It is the set of neighborhood coordinates in the gene expression space of a single-cell transcriptome; This represents the coordinates of any gene expression location within the domain.

[0010] Furthermore, S3 specifically includes the following sub-steps: S31. Input the noise samples into the U-Net network for noise prediction to obtain the prediction data to be corrected; S32. Input the prediction data to be corrected into the sparse spatial state encoder, and embed the sparse distribution signal of the real gene expression into the encoder to obtain sparse embedding data. S33. The dynamic radiation domain decoder reverse-derives the gene expression vector of optimal transmission in the radiation domain based on the sparse embedded data aligned by the sparse spatial state encoder, and assigns the gene expression vector of optimal transmission to the prediction data to be corrected. S34. Resample the assigned prediction data to be corrected to obtain the noisy prediction data to be corrected. S35. Input the noise-added prediction data to be corrected into the U-Net network, and repeat S31~S34 until a complete single-cell gene expression matrix is ​​obtained.

[0011] Furthermore, in step S31, noise samples are input into the U-Net network to obtain inverse mask pseudo-prediction. Compared with real mask data Then, combined with the obtained prediction data to be corrected, it is expressed as: in, In the formula, This represents the predicted data to be corrected; This indicates a false prediction using the inverse mask; Represents element-wise product; Represents the actual mask data; Represents a binary mask; This represents the inverse mask.

[0012] Furthermore, S32 specifically includes: The corrected prediction data is input into the sparse spatial state encoder. Chebyshev index is used to calculate the expression mean of the uncorrected prediction data, inverse mask, and original single-cell gene expression profile in each radiation layer, and the predicted sparsity intensity, mask sparsity intensity, and prototype sparsity intensity matrix are calculated respectively. Then, the predicted sparsity intensity is aligned to the prototype sparsity intensity matrix under the projection constraint of the mask sparsity intensity, and thus sparse embedding data is obtained.

[0013] Furthermore, Chebyshev indexing is used to calculate the expression mean of the uncorrected prediction data and the inverse mask at each radiation layer, yielding the predicted sparsity intensity and the mask sparsity intensity, respectively, expressed as: In the formula, Indicates the predicted sparsity intensity; Indicates average operation; This represents the Chebyshev index.

[0014] Furthermore, the predicted sparse intensity is aligned to the prototype sparse intensity matrix under the projection constraint of the mask sparse intensity, where the projection constraint is specifically expressed as: In the formula, Represents the projection constraint loss function; Represents the sparse intensity vector In the absence of intensity vector The process of finding the expected value under constraints; Represents the sparse intensity vector In the absence of intensity vector Linear mapping process under constraints; Represents the prototype sparse intensity matrix; express The mean vector; express The covariance matrix; Represents the prototype sparse intensity matrix The first in One prototype center; Indicates the prototype index; This is the balance coefficient; Indicates Mahalanobis distance; Indicates the nearest neighbor distance.

[0015] Furthermore, S33 specifically includes: Treatment of corrective prediction data Perform layer decomposition and use Chebyshev indexing Get into the circle gene vectors and its mask vector Subsequently, the layers were calculated. gene vectors mean And extract the target mean of the same layer from sparse embedding data. ; The mean of the gene vector In the mask vector Under the constraint of the target mean By getting closer, the corrected gene vector can be obtained. For the corrected gene vector Perform the inverse mapping of the Chebyshev index. The optimal gene expression vector is obtained and then assigned to the prediction data to be corrected.

[0016] The single-cell gene completion method based on a sparse-sensing diffusion model provided by this invention has the following beneficial effects: 1. This invention, based on the Chebyshev distance decoupling model, ChebSCRD, achieves high-fidelity gene expression completion for the first time in extremely missing scRNA-seq scenarios. It introduces two complementary priors: first, a geometric rearrangement prior, which uses Chebyshev distance to decouple highly variable genes (HVGs) without parameters, mapping them to concentric radiative domains and sorting them by Manhattan adjacency within each loop to ensure consistent biological associations in the initial local neighborhood; second, a sparse correction prior, which integrates high-confidence sparsity outside the domain and sparsity within the domain during inversion, correcting the zero-channel gradient bias through sparse spatial mapping, effectively suppressing "lazy diffusion." Furthermore, ChebSCRD constructs a linear optimal transport space corrected by sparse spatial mapping and progressively corrects the global expression using the true distribution. Benefiting from the dual constraints of "geometric rearrangement" and "sparse correction," ChebSCRD balances gene expression integrity with the realism of sparse structure, achieving high-fidelity reconstruction of extremely missing expressions.

[0017] 2. This invention has undergone quantitative and qualitative evaluation on multiple benchmarks. The results show that in scenarios with more than 80% gene deletion, ChebSCRD significantly outperforms existing image restoration methods in terms of gene completion accuracy, cell expression quality, and downstream analysis indicators, and the generated expression profile is more consistent with biological laws.

[0018] 3. This invention proposes the first framework for large-scale gene deletion completion, which utilizes the resampling step of the diffusion model to perform sparse biological associations between implicit fusion genes, effectively guiding the generation of a wide range of deleted genes.

[0019] 3. This invention designs a sparse spatial state encoder that aligns the gene distribution in the prototype through differentiable optimization of dual distances, eliminates zero-channel gradient bias, and injects cell-type-specific sparse knowledge into resampling.

[0020] 4. This invention constructs a dynamic radiation domain decoder and designs a gene expression correction mechanism based on linear optimal transport of radiation domain layers to ensure that the generated expression profile has both global biological consistency and local sparsity fidelity. Attached Figure Description

[0021] Figure 1 For the comparison of single-cell expression correction paradigms, among which, Figure 1 (a) in the text is imputation: noise genes are corrected by neighborhood smoothing, but genes that were not originally captured cannot be generated; Figure 1 (b) in the figure represents reconstruction: alignment based on the reference map can remove noise, but is still limited to the existing gene range; Figure 1 (c) in the text is a completion: ChebSCRD, the method proposed in this invention.

[0022] Figure 2 This is a network architecture diagram of the sparse sensing diffusion model in Example 1.

[0023] Figure 3 This is a schematic diagram of the four dynamic radiation domain shrinkage patterns in Example 2.

[0024] Figure 4 This is a comparison of gene completion SSIM and time cost (unit: seconds / sample) under different shrinking strategies in Example 2.

[0025] Figure 5 This is a comparison of inference efficiency between methods under different random seeds in Example 2.

[0026] Figure 6 This is a violin diagram showing the complete expression of three highly expressed genes (top) and three low expressed genes (bottom) in Naive CD4 T cells in Example 2.

[0027] Figure 7 This is a flowchart of the single-cell gene completion method based on the sparse sensing diffusion model in Example 1. Detailed Implementation

[0028] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0029] Example 1 This embodiment's single-cell gene completion method based on a sparse-perceived diffusion model employs a stepwise sparse-corrected diffusion resampler for the Chebyshev radiation domain. This resampler is driven collaboratively by a Sparse Spatial State Encoder (SSSE) and a Dynamic Radiation Domain Decoder (DRDD). SSSE, leveraging the inherent characteristic of zero expansion, learns sparse representations of "real biological silence," correcting noise trajectories in real-time during the reverse diffusion process and suppressing the excessive generation of inertial zero values. DRDD, on the other hand, injects gene-gene biological arrangement priors and, relying on the sparse embeddings passed from SSSE, progressively rearranges and decodes a high-resolution expression spectrum that simultaneously achieves global biological consistency and local sparse fidelity. At certain intermediate nodes of inference, ChebSCRD calls the resampler once to continuously purify the latent states, ultimately outputting a sharp and biologically reliable complete gene expression matrix. The problem of single-cell gene expression completion in this embodiment can be formally defined as follows: Let the complete gene set of scRNA-seq data be... Its gene count is For a given cell, only a subset was observed. The expression, The unobserved gene set is , Let the observation representation vector be... The goal is to infer the expression of the missing gene. And concatenate them to obtain the complete vector. This allows its distribution to be as close as possible to the actual gene expression.

[0030] To more realistically simulate different degrees of sparsity in experiments, the original expression was... Apply masking operator ,in This represents the mask ratio. Specifically, a binary mask is constructed for each gene. ,like Then the expression of that gene will be preserved, if This is considered missing. The representation vector after masking is denoted as... .in This represents the Hadamard (element-wise) product. Therefore, the single-cell gene expression completion task can be viewed as learning a mapping. ,Right now To restore gene expression as completely and closely approximates its true distribution as much as possible. Based on this, refer to Figure 1 (c) and Figure 7 This embodiment specifically includes the following: S1. Perform forward noise diffusion on the original single-cell gene expression profile to obtain noise samples. ; S2. Construct the sparse correction resampling diffusion model ChebSCRD; refer to Figure 2 The sparse correction resampling diffusion model includes a U-Net network, a sparse spatial state encoder, and a dynamic radiation domain decoder. The U-Net network is used to predict the input noise. The sparse spatial state encoder and dynamic radiation domain decoder work together, and Chebyshev indexes are introduced to embed and denoise the predicted noise.

[0031] S3. Input the noise samples into the sparse correction resampling diffusion model, and output the complete single-cell gene expression matrix after gene completion through a multi-step reverse denoising process.

[0032] In some embodiments, the Chebyshev index is defined as follows: The local smoothness assumption upon which conventional image diffusion models rely for effectiveness Adjacent pixels are required to have statistical similarity. However, adjacent entries in a single-cell gene sequence do not share biological commonalities; if... Forcibly rearranged into a two-dimensional matrix in any order This assumption thus becomes invalid. To reconstruct an "interpretable" local domain within the diffusion framework, a parameter-free decoupling arrangement strategy based on Chebyshev distance is proposed, referencing... Figure 2 Specifically, it includes the following: A1. First, the original expression of genes. Perform grid rearrangement and define discrete Chebyshev distances; Specifically, let's set , Ensure that there are no empty spaces in the grid after rearrangement. Center the grid. Define the discrete Chebyshev distance with the origin as the origin. : In the formula, The height of the radiation field; The width of the radiation domain; This represents the maximum number of concentric circles. A2. Subsequently, based on the discrete Chebyshev distance... Divide into concentric radiation layers : , Its innermost ring capacity From the properties of arithmetic sequences, we get The total capacity satisfies the following expression, thus ensuring that all genes exactly fill the entire grid: In the formula, The radius is the concentric circle radius.

[0033] A3. Next, the gene arrangement is ordered using the arrangement function, so that the sparse correction resampling diffusion model can obtain a local smooth prior. Define the arrangement function Its output Determined according to the following rules: 1) Genes are classified according to the radius of their respective spheres. Sort by size from smallest to largest; 2) Within the same layer, the genes are sequenced by traversing clockwise according to the Manhattan method.

[0034] Due to the arrangement function Based solely on gene indexes and independent of expression values, this process is parameter-free decoupling. After arrangement, the highly variable gene HVG will be compressed into Chebyshev neighborhoods. ; Thus restore This allows the diffusion model to regain its local smooth prior; A4. Finally, for the reasoning period To access any sphere, pre-compute the Chebyshev index and its inverse mapping. The Chebyshev index is represented as: In the formula, Indicates the first The Chebyshev indexing operation is performed on the radiative layer; Indicates the first in the radiation domain Layer data; This indicates that the layers are flattened in sequence. This indicates that data at a specific coordinate location within the radiation domain is retrieved. It can be a gene matrix or a mask matrix; Indicates the coordinate position of the numerical value; It represents concentric radiating layers.

[0035] Accessing or restoring the entire concentric circle only requires Memory mapping is far less computationally intensive than a single reverse inference step. By utilizing the parameter-free decoupling of Chebyshev distance, a locally smoothed domain consistent with biological co-expression was constructed for the first time in a single-cell scenario for the diffusion model. This laid a geometrically consistent expression foundation for subsequent SSSE-DRDD collaborative resampling and significantly reduced computational and memory overhead.

[0036] Based on the defined Chebyshev index, see reference Figure 2 This step specifically includes the following sub-steps: S31. Input the noise samples into the U-Net network for noise prediction to obtain the prediction data to be corrected; Specifically, noisy samples are input into the U-Net network to obtain inverse mask pseudo-predictions. Compared with real mask data Then, combined with the obtained prediction data to be corrected, it is expressed as: in, In the formula, This represents the predicted data to be corrected; This indicates a false prediction using the inverse mask; Represents element-wise product; Represents the actual mask data; Represents a binary mask; This represents the inverse mask.

[0037] S32. Input the prediction data to be corrected into the sparse spatial state encoder, and embed the sparse distribution signal of the real gene expression into the encoder to obtain sparse embedding data. Although the decoupling of the radiation domain layers has rationally arranged co-expressed genes according to biological relationships, the single-cell expression matrix is ​​still filled with a large number of zero values, resulting in "holes" within the layers. In this highly sparse feature space, traditional diffusion models can easily slip into "lazy solutions," that is, tend to output pseudo-zero expression, thus making the sparsity of the predicted distribution far exceed the actual biological distribution, especially when the deletion ratio is high. To suppress this distortion, SSSE is introduced into the ChebSCRD framework. By explicitly characterizing the sparsity intensity of the cells to be completed and aligning it with the prototype cells, it provides fine-grained sparsity correction for subsequent completion. The processing of the Sparse Spatial State Encoder (SSSE) is as follows: Correction Prediction Data Chebyshev indexing is used in the input sparse spatial state encoder. Treatment of corrective prediction data Inverse mask Expression mean was calculated for each radiation layer, and the predicted sparsity intensity was calculated separately. Mask sparsity strength Specifically, it is expressed as: In the formula, This represents the maximum number of concentric circles.

[0038] Meanwhile, the training phase retains Each single-cell prototype sample was also indexed by the same Chebyshev index. The operation yields the prototype sparse intensity matrix. .

[0039] Subsequently, the predicted sparsity intensity is aligned to the prototype sparsity intensity matrix under the projection constraint of the mask sparsity intensity, thereby obtaining sparse embedded data.

[0040] In order to align the predicted sparse pattern with the prototype distribution, the projection constraint in this step is specifically expressed as follows: In the formula, Represents the projection constraint loss function; Represents the sparse intensity vector In the absence of intensity vector The process of finding the expected value under constraints; Represents the sparse intensity vector In the absence of intensity vector Linear mapping process under constraints; Represents the prototype sparse intensity matrix; express The mean vector; express The covariance matrix; Represents the prototype sparse intensity matrix The first in One prototype center; Indicates the prototype index; This is the balance coefficient; Represents Mahalanobis distance, responsible for global distribution alignment; This indicates the nearest neighbor distance and is used to enhance the detail of the local radiation domain layers. and Both factors drive the convergence of predicted sparse structures to true biological sparsity. The resulting sparse embeddings... Feedback to DRDD significantly reduced the generation of pseudo-zero expression and achieved sparse distribution consistency at both the global and local levels.

[0041] By establishing a differentiable prediction-to-prototype mapping in the feature space, SSSE effectively overcomes the inertia problem of parameter updates in diffusion models under high sparsity scenarios.

[0042] S33. The dynamic radiation domain decoder reverse-derives the gene expression vector of optimal transmission in the radiation domain based on the sparse embedded data aligned by the sparse spatial state encoder, and assigns the gene expression vector of optimal transmission to the prediction data to be corrected. Specifically, SSSE has established a sparsity mapping from the predicted distribution to the true distribution within each radiation domain layer using Chebyshev distance, while DRDD is based on the sparse embedding aligned by SSSE. The process involves inversely deriving the genes optimally transported by Kantorovich in the radiation domain. If these layer means are not explicitly "pulled back" during inference, local drift will rapidly spread to the complete gene expression vector, leading to global completion offset. Therefore, DRDD is proposed to use radiation domain layers as indices to ensure that the expression mean of each radiation domain layer in the inference chain remains consistent with the target mean from the outside in, without disrupting the sparse structure. The specific processing steps of the Dynamic Radiation Domain Decoder (SSSE) are as follows: Treatment of corrective prediction data Perform layer decomposition and use Chebyshev indexing Get into the circle gene vectors and its mask vector Subsequently, the layers were calculated. gene vectors mean And extract the target mean of the same layer from sparse embedding data. ; The mean of the gene vector In the mask vector Under the constraint of the target mean By moving closer to the target gene and modifying only the unmasked gene entries, the corrected gene vector can be obtained. ; For the corrected gene vector Perform the inverse mapping of the Chebyshev index. The optimal gene expression vector is obtained and then assigned to the prediction data to be corrected.

[0043] Among them, for "the mean of gene vectors" In the mask vector Under the constraint of the target mean "Closer proximity, while only modifying unmasked gene entries"—this embodiment transforms the alignment correction problem into a mean alignment optimization problem with boundary constraints, assuming an adjustable index set. Displacement = Introducing auxiliary variables and linearization ,get: In the formula, Let be the objective loss function for the mean alignment optimization problem with boundary constraints; Gene indexes in an adjustable index set; The total displacement corrected for the mean; This serves as a lower bound constraint for gene components; These are the corrected gene vector component values; These are the gene vector component values ​​before correction; This serves as an upper bound constraint for gene components; where, It can be set based on priors such as nonnegativity or upper limit of count.

[0044] Traditional linear programming time complexity To address the impact on inference time, this embodiment employs a gene tolerance-driven hierarchical adjustment strategy and a closed-loop global offset formula. This enables cross-sample batch processing. This involves minimizing the total displacement cost within the Kantorovich framework in a discrete one-dimensional space. Specifically, it defines the gene tolerance formula: In the formula, Gene tolerance.

[0045] according to This is an unmodifiable location. Set as .set up To be according to The sorting mapping for ascending order from smallest to largest, for the first The sum of the upper and lower boundaries of each gene is defined as: In the formula, For the front The sum of the lower bounds of each gene; For the sorted number The lower bound of a gene; For the front The sum of the upper bounds of each gene; For the sorted number The upper bound of each gene.

[0046] because Follow Increase and grow. Decreasing, thus the boundary can be accumulated to calculate its binary search boundary point: In the formula, To ensure that the cumulative upper and lower bounds mean is greater than or equal to The minimum number of genes; This represents a candidate boundary point in the binary search.

[0047] Utilizing the monotonicity mentioned above, only one binary search is needed to determine the minimum value. .forward One gene will be pushed to a position close to its boundary, and the remaining ones... Each gene shares the remaining mean migration amount. Given the given conditions, a closed-form solution for the global offset of a batch can be calculated: In the formula, The optimal global offset step size; For the front The sum of the lower bounds of each gene; For the front The sum of the upper bounds of each gene; To ensure that the cumulative upper and lower bounds mean is greater than or equal to The minimum number of genes; Total number of genes; Calculate the uniform step size for the remaining genes to be shifted in one go. Thus in Mean alignment is achieved within the total complexity.

[0048] In this embodiment, DRDD corrects the sparse distribution offset by tracing back the resampling nodes to an "anchor point" that satisfies statistical consistency, thereby continuously correcting the sparse distribution offset in a shrinking jump manner, and by re-injecting noise. Send the corrected sample back to the previous time step. This process effectively suppresses the spread of cumulative errors and maintains synchronous convergence between local sparsity and global expression distribution, providing a stable and accurate dynamic decoding mechanism for ChebSCRD in single-cell gene expression completion tasks.

[0049] S34. Resample the assigned prediction data to be corrected to obtain the noisy prediction data to be corrected. S35. Input the noise-added prediction data to be corrected into the U-Net network, and repeat S31~S34 until a complete single-cell gene expression matrix is ​​obtained.

[0050] Example 2 This embodiment is based on the ChebSCRD model framework—sparsely corrected resampling diffusion model—from Embodiment 1, and uses experiments to verify and evaluate it. Specifically, it includes the following: 1. The experimental setup is as follows: (1) Dataset; The training and evaluation data used in this embodiment are from the latest version of the DISCO database. This database currently integrates transcriptomics data from 21,346 human samples, covering over 135 million single cells across 39 tissues and 461 cell types. All raw data in DISCO comes from publicly available scRNA-seq studies, primarily stored in public repositories such as GeneExpression Omnibus (GEO), Sequence Read Archive (SRA), and ArrayExpress, as well as large international collaborative projects like Human Cell Atlas (HCA).

[0051] To verify the model's generalization ability, a scRNA-seq dataset covering multiple human tissues and organs was constructed, containing approximately 1.5 million cells across 38 tissues / organs. Of these, 1 million samples were used for training the diffusion model (sparse-corrected resampling diffusion model), and the remaining samples were divided into six test sets (named Comp1–Comp6). For each test set, a masking operator was used. Deletion simulations were performed on gene sequences, with the deletion rate set at six levels. This approach aims to closely approximate the distribution characteristics of high deletion scenarios in real-world scRNA-seq data. The data processing workflow strictly follows DISCO's recommended specifications and undergoes secondary quality control before being used for model training, validation, and comparative experiments. The data resources used in this embodiment can be obtained from the DISCO website: https: / / disco.bii.a- star.edu.sg / .

[0052] (2) Evaluation system; To evaluate the performance of the gene column completion algorithm proposed in this invention for the first time in the field of single-cell correction, it is necessary to first clarify that existing imputation methods are not suitable as a baseline. Whether based on statistical modeling or deep learning, imputation models rely on existing gene conditions in the population sample for inference, essentially only able to recover missing values ​​at the population level, and unable to directly complete arbitrary missing genes at the single-cell granularity. In contrast, the pre-trained diffusion model proposed in this invention can achieve single-cell-level completion without relying on population conditions; therefore, the two are not comparable in terms of design philosophy.

[0053] Based on this difference, several advanced diffusion models in the field of image inpainting were selected as comparisons, including DDNM, Repaint, DeqIR, and SPGD. These represent the latest advancements in diffusion models across different inpainting strategies: DDNM improves stability by decoupling noise prediction and the inverse diffusion process; Repaint achieves high-quality inpainting through multiple iterative resampling; DeqIR emphasizes an efficient denoising and inversion mechanism; and SPGD strikes a balance between global consistency and local detail fidelity. These methods are widely used as performance benchmarks in image inpainting tasks, thus providing a reasonable and challenging reference for the research in this embodiment.

[0054] To ensure the comprehensiveness and impartiality of the comparison, seven quantitative indicators covering both data reconstruction and downstream analysis were designed. Among them, structural similarity (SSIM), Pearson correlation coefficient (PCC), and mean squared error (MSE) are used to measure the quality of the gene expression matrix completion process at both the numerical and structural levels. Accuracy (ACC), F1 score (F1), adjusted RAND index (ARI), and normalized mutual information (NMI) are used to evaluate the potential application value of the completion results in downstream single-cell analysis (such as cell annotation and clustering). Through this comprehensive evaluation system, this invention can comprehensively test the model's performance in data recovery and biological applications.

[0055] (3) Implementation details; The diffusion model proposed in this invention employs a conditional U-Net architecture for gene expression synthesis, utilizing text embeddings generated by a frozen PubMedBERT encoder as conditional input. PubMedBERT was chosen because its pre-training on large-scale biomedical literature allows for more accurate capture of semantic features of cell type names and related biological concepts. The U-Net backbone consists of hierarchical downsampling and upsampling modules, and a cross-attention mechanism is introduced in the intermediate resolution layer to effectively fuse textual information. During training, the model runs on four Tesla V100 GPUs (32GB VRAM per GPU), with a learning rate set to [value missing]. The batch size is 128. A linear DDPM scheduler is used to progressively inject Gaussian noise into the gene matrix over 1000 time steps. The denoising process is learned by minimizing the mean squared error (MSE) loss, while gradient checkpointing is enabled to optimize memory usage. During the inference phase, [the following is used]. A step-by-step DDIM scheduler is used, and SSSE-DRDD resampling noise is injected every 5 steps in the inference chain to improve the robustness and diversity of the generated results. To ensure the consistency of gene spatial structure, this invention places 32,400 target genes in a gene arrangement function. The effect of mapping is The matrix representation is further divided into 90 radiation domain layers based on Chebyshev distance, thus preserving both global and local hierarchical information during training and inference.

[0056] 2. The quantitative results of the experiment are as follows; Table 1 shows the performance comparison of the gene completion method of this invention on six datasets with different missing rates (Comp1–Comp6). The results cover three core metrics: SSIM, MSE, and PCC, which are used to evaluate the structural consistency, numerical error, and correlation of the completion results, respectively. Overall, the method of this invention achieves the best performance on all metrics and datasets, demonstrating high stability and robustness.

[0057] In low missing rate scenarios (Comp1–Comp3, Under these conditions, the method of this invention outperforms almost all control methods in almost all metrics. For example, in Comp3, the model of this invention achieved an SSIM of 84.45, significantly surpassing the second-ranked DDNM (SSIM of 83.71). This indicates that, under relatively complete data conditions, the method of this invention can well preserve the global structure of the gene expression matrix and the correlations between genes.

[0058] As the missing rate increases to a moderate level (Comp4–Comp5), The performance of the comparison methods showed a significant decline, while the method of this invention remained superior. For example, in Comp5, the model of this invention outperformed all baseline methods on SSIM (78.96) and PCC (58.78), and also achieved the lowest MSE (10.42). This result highlights the role of the structural prior and spatial constraints imposed on SAD by this invention under moderately sparse conditions, enabling the model to excel in both detail reproduction and global consistency.

[0059] In extremely sparse scenarios (Comp6, Under these conditions, the performance gap widens further. The method of this invention still achieves an SSIM of 75.79 and a PCC of 58.91, significantly outperforming the best comparison method, DeqIR (SSIM 70.22, PCC 43.85). Simultaneously, the method of this invention maintains the lowest MSE value of 10.44, demonstrating that even under a missing value as high as 80%, the model can still complete highly reliable expression patterns.

[0060] In summary, the method of this invention not only outperforms current state-of-the-art (SOTA) methods in the field of image inpainting in terms of overall performance, but its advantages become more pronounced when the missing data rate is high. This trend fully demonstrates the effectiveness of combining the SSSE and DRDD modules, and verifies its strong robustness and potential application value in processing highly sparse single-cell data.

[0061] Table 1 compares the gene completion quantification results with those of SPGD, DeqIR, DDNM, and Repaint. In Table 1, the best results are indicated in bold, and the second-best results are indicated in red. underlined Indicated. The MSE indicator has been multiplied by When scaling, SSIM and PCC are expressed as percentages (%).

[0062] 3. Ablation Research (1) Module ablation; An ablation experiment was designed to evaluate the role of each component in ChebSCRD (SSIM index), as shown in Table 2, to experimentally assess the impact on gene rearrangement priors and sparse correction. The focus was on analyzing the influence of gene arrangement decoupled by Chebyshev distance and the resampling device constructed from SSSE-DRDD on performance. The baseline model did not use these modules and served as a comparison reference. The results showed that the lack of gene rearrangement priors (w / o) significantly reduced the effectiveness of sparse correction. High missing values ​​significantly reduce completion performance. Secondly, performance degradation also occurs when SSSE-DRDD is lacking for sparse distribution correction (w / o SC). Furthermore, experiments were conducted using... Minimizing the linear optimal transport solution in DRDD yields results that are roughly equivalent but less than ideal. Overall, the method in this invention enhances the model's ability to model gene expression characteristics through a sparse attention mechanism from local to global perspectives.

[0063] Table 2 - Ablation studies under different missing rates (SSIM index) (2) The impact of dynamic radiation domain update strategy; like Figure 3As shown, four dynamic shrinking strategies are compared. The horizontal axis represents the layer number of the Chebyshev radiation domain (1–D, where D is the outermost layer), and the vertical axis represents the time step of diffusion inference (0–T). Gray squares represent layers that have not been updated in the current step, and red squares represent layers that have been updated. Upper triangular / Lower triangular corrects multiple layers simultaneously at a specific time step, exhibiting a pattern of "increasing from outer to inner" and "decreasing from outer to inner," respectively. Main diagonal / Anti-diagonal corrects only one layer at a specific time step, with the former following an "outer → inner" order and the latter following an "inner → outer" order.

[0064] The first two strategies involve multiple concentric layers in each step, resulting in high computational costs and a tendency to overcorrect, which in turn reduces the completion effect. While Anti-diagonal has lower costs, it only corrects the inner layer initially, leading to insufficient resampling of genes and difficulty in alleviating sparsity. Maindiagonal, in Comp1–6 of Figure 4, combines completion accuracy and efficiency. Since the results are similar, normalized SSIM and PCC indices are used to enhance the comparison. After comprehensive consideration, the Maindiagonal strategy, which performs stepwise correction from the outer to the inner layer, is ultimately selected. Prioritizing the correction of the outer layer in the early stages allows for more efficient use of resampling to alleviate sparsity, resulting in stable performance and higher efficiency.

[0065] 4. Evaluation of completion efficiency To evaluate the stability and efficiency of ChebSCRD and various baseline methods under different random initialization conditions, experiments were conducted on the Comp6 dataset, which has the most severe gene deletions. Ten different random seeds were set, and the inference performance of each method was statistically analyzed. Figure 5 The diagram shows a comparison of the performance of various methods in terms of PCC and inference time (unit: seconds / sample), with the size of the cross indicating the SSIM metric.

[0066] It can be observed that Ours w / With Ours w / The proposed method exhibits high PCC and short inference time across multiple seeds, with stable SSIM distribution, demonstrating a good balance between generalization performance and efficiency. The Repaint method has significantly longer inference time than other methods, indicating higher computational cost. Traditional optimization methods such as DeqIR and DDNM show relatively low PCC and SSIM, and while their inference time is short, their generation quality is limited. The SPGD method has moderate PCC performance, but its inference time and SSIM scattering are more volatile, resulting in slightly weaker stability. This experiment verifies that the proposed method has superior stability and overall performance under different initial conditions, especially achieving a good balance between efficiency and gene completion consistency.

[0067] 5. Downstream biological analysis (1) Cell population discovery; To verify the application potential of completed data in downstream single-cell analysis, the performance of various gene completion algorithms in clustering and annotation tasks was evaluated on the Comp6 dataset, which has the highest missing data rate. As shown in Table 3, all completion methods brought significant performance improvements compared to the original data, especially the method proposed in this invention, which showed more prominent advantages in clustering and annotation metrics. The results show that the gene completion algorithm proposed in this invention can effectively recover more gene expression even when gene loss is severe, thereby significantly enhancing single-cell characterization capabilities and improving the recognition effect of specific expressions. Table 3 - Comparison of gene completion quantification results with methods such as SPGD, DeqIR, DDNM, and Repaint In Table 3, the best results are indicated in bold, and the second-best results are indicated in red. underlined Indicated. The MSE indicator has been multiplied by When scaling, SSIM and PCC are expressed as percentages (%).

[0068] (2) Gene-specific expression; On the Comp6 dataset, the completion effects of three highly expressed and three low-expressed genes in Naive CD4 T cells were compared. Figure 6 Based on the original sequencing values, the method of this invention is closest to the baseline among the six genes. In the upper part, high-expression genes generally exhibit "lazy zeros" in the four comparison methods, i.e., the zero region is significantly expanded; this invention significantly narrows this region, indicating that the SSSE proposed in this invention overcomes the inertia of the diffusion model under high zero expansion. In the lower part, the control method for low-expression genes is overly smoothed and increases noise, resulting in a raised violin plot; this invention has the shortest distribution and the lowest error, further demonstrating that the sparse correction of DRDD proposed in this invention can suppress false positives and retain true weak signals.

[0069] This invention integrates geometric rearrangement and sparse correction as dual priors into a diffusion network, reconstructing gene spatial associations while suppressing the "lazy zero" phenomenon, thus significantly reducing the uncertainty caused by sparse gene distribution during the completion process. Comprehensive evaluation shows that ChebSCRD comprehensively outperforms existing methods in terms of reconstruction quality, cell population differentiation, and marker gene recovery, while maintaining near real-time inference efficiency. The framework of this invention lays a high-fidelity data foundation for constructing large-scale single-cell basic models and advancing precision medicine research, and also opens up new research directions for sparse gene expression completion.

[0070] Although specific embodiments of the invention have been described in detail with reference to the accompanying drawings, this should not be construed as limiting the scope of protection of this patent. Various modifications and variations that can be made by a person skilled in the art without inventive effort within the scope described in the claims still fall within the scope of protection of this patent.

Claims

1. A single-cell gene completion method based on a sparse sensing diffusion model, characterized in that, Includes the following steps: S1. The original single-cell gene expression profile is subjected to forward noise diffusion to obtain a noise sample; S2. Construct a sparse correction resampling diffusion model; S3. Input the noise samples into the sparse correction resampling diffusion model, and output the complete single-cell gene expression matrix after gene completion through a multi-step reverse denoising process.

2. The single-cell gene completion method based on a sparse sensing diffusion model according to claim 1, characterized in that, In S2, the sparse correction resampling diffusion model includes a U-Net network, a sparse spatial state encoder, and a dynamic radiation domain decoder. The U-Net network is used to predict input noise; The sparse spatial state encoder and dynamic radiation domain decoder work together, and Chebyshev indexes are introduced to embed and denoise the predicted noise.

3. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 1, characterized in that, The definition of the Chebyshev index specifically includes the following steps: A1. Rearrange the original gene expression in a grid and define the discrete Chebyshev distance; In the formula, Represents the discrete Chebyshev distance; Indicates the coordinate position in the grid; Indicates the maximum number of radiation domain layers; A2. Based on the discrete Chebyshev distance, divide the concentric radiation layers. : , The total capacity of the concentric radiation layers satisfies: In the formula, Indicates the innermost layer; The radius of the concentric circle; This represents the maximum number of concentric circles. The number of genes; A3. Genes are arranged and ordered using a distribution function, so that the sparse correction resampling diffusion model can obtain a local smooth prior. A4. The Chebyshev index representation for any concentric radiation layer is as follows: In the formula, Indicates the first The Chebyshev indexing operation is performed on the radiative layer; Indicates the first in the radiation domain Layer data; This indicates that the layers are flattened in sequence. This indicates that data is taken from a specific coordinate position within the radiation domain.

4. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 3, characterized in that, In A3, the arrangement function is: Its output is , This represents the coordinate position of each gene expression in the arrangement. The genes are arranged and ordered using an arrangement function, including: Genes are sorted from smallest to largest according to the radius of their respective spheres; Genes located in the same layer are ordered by traversing clockwise according to the Manhattan method. After the genes are arranged by the arrangement function, highly variable genes will be compressed into Chebyshev neighborhoods: Thus restore Even if the sparse correction resampling diffusion model obtains a locally smooth prior, then... To satisfy the set of all coordinate positions that meet the local smoothness prior; The coordinates of any value within the neighborhood; It is the conditional probability density function; It is the set of neighborhood coordinates in the gene expression space of a single-cell transcriptome; This represents the coordinates of any gene expression location within the domain.

5. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 3, characterized in that, S3 specifically includes the following steps: S31. Input the noise samples into the U-Net network for noise prediction to obtain the prediction data to be corrected; S32. Input the prediction data to be corrected into the sparse spatial state encoder, and embed the sparse distribution signal of the real gene expression into the encoder to obtain sparse embedding data. S33. The dynamic radiation domain decoder reverse-derives the gene expression vector of optimal transmission in the radiation domain based on the sparse embedded data aligned by the sparse spatial state encoder, and assigns the gene expression vector of optimal transmission to the prediction data to be corrected. S34. Resample the assigned prediction data to be corrected to obtain the noisy prediction data to be corrected. S35. Input the noise-added prediction data to be corrected into the U-Net network, and repeat S31~S34 until a complete single-cell gene expression matrix is ​​obtained.

6. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 5, characterized in that, In step S31, noise samples are input into the U-Net network to obtain inverse mask pseudo-prediction. Compared with real mask data Then, combined with the obtained prediction data to be corrected, it is expressed as: in, In the formula, This represents the predicted data to be corrected; This indicates a false prediction using the inverse mask; Represents element-wise product; Represents the actual mask data; Represents a binary mask; This represents the inverse mask.

7. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 5, characterized in that, Specifically, S32 includes: The corrected prediction data is input into the sparse spatial state encoder. Chebyshev index is used to calculate the expression mean of the uncorrected prediction data, inverse mask, and original single-cell gene expression profile in each radiation layer, and the predicted sparsity intensity, mask sparsity intensity, and prototype sparsity intensity matrix are calculated respectively. Then, the predicted sparsity intensity is aligned to the prototype sparsity intensity matrix under the projection constraint of the mask sparsity intensity, and thus sparse embedding data is obtained.

8. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 7, characterized in that, Chebyshev indexing is used to calculate the expression mean of the uncorrected prediction data and the inverse mask at each radiation layer, yielding the predicted sparsity intensity and the mask sparsity intensity, respectively, expressed as follows: In the formula, Indicates the predicted sparsity intensity; Indicates average operation; This represents the Chebyshev index.

9. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 7, characterized in that, The predicted sparse intensity is aligned to the prototype sparse intensity matrix under the projection constraint of the mask sparse intensity, where the projection constraint is specifically expressed as: In the formula, Represents the projection constraint loss function; Represents the sparse intensity vector In the absence of intensity vector The process of finding the expected value under constraints; Represents the sparse intensity vector In the absence of intensity vector Linear mapping process under constraints; Represents the prototype sparse intensity matrix; express The mean vector; express The covariance matrix; Represents the prototype sparse intensity matrix The first in One prototype center; Indicates the prototype index; This is the balance coefficient; Indicates Mahalanobis distance; Indicates the nearest neighbor distance.

10. The single-cell gene completion method based on the sparse sensing diffusion model according to claim 5, characterized in that, Specifically, S33 includes: Treatment of corrective prediction data Perform layer decomposition and use Chebyshev indexing Get into the circle gene vectors and its mask vector Subsequently, the layers were calculated. gene vectors mean And extract the target mean of the same layer from sparse embedding data. ; The mean of the gene vector In the mask vector Under the constraint of the target mean By getting closer, the corrected gene vector can be obtained. For the corrected gene vector Perform the inverse mapping of the Chebyshev index. The optimal gene expression vector is obtained and then assigned to the prediction data to be corrected.