Single cell increment annotation method based on distribution and expression perception return visit

Multi-view gene expression alignment is performed through the diffusion model based on distribution-perceptual conditions and the knowledge distillation model of expression-perceptualization, combined with fuzzy guidance constraints, the problems of catastrophic forgetting and long-tail distribution in single-cell type annotation are solved, efficient incremental annotation is achieved, and the robustness and applicability of single-cell type annotation is improved.

CN120581072AActive Publication Date: 2025-09-02CHENGDU UNIV OF INFORMATION TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510688542.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-02
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

When faced with exponentially growing single-cell sequencing data, existing single-cell type annotation methods have problems such as catastrophic forgetting, difficulty in meeting the needs of diversified cell annotation in open environments, and long-tail distribution characteristics, resulting in limited promotion and application in precision medicine clinical practice.

Method used

Gene expression profile data of old cell samples were generated using a diffusion model based on distribution perception conditions, and multi-view gene expression alignment was performed through the expression perception knowledge distillation model, combined with fuzzy guidance constraints, and incremental annotation was achieved.

Benefits of technology

It effectively solves the core problems in single-cell long-tail incremental annotation, improves the performance and robustness of incremental annotation, can seamlessly integrate in continuous learning, and significantly improves the learning balance and efficiency of new and old cell types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120581072A_ABST
    Figure CN120581072A_ABST
Patent Text Reader

Abstract

The invention discloses a single cell increment annotation method based on distribution and expression perception return visit, and belongs to the technical field of single cell type annotation. Comprising the following steps: inputting a single cell expression profile sample to be processed into a diffusion model based on a distribution perception condition, and generating gene expression profile data of an old cell sample participating in diffusion model training; integrating the gene expression profile data of the old cell sample and the collected gene expression profile data of the new cell sample; and respectively inputting the integrated data into two different feature extractors in the knowledge distillation model for expressing perception so as to carry out multi-view gene expression alignment, and outputting a corresponding cell type annotation result. According to the method, the core problem of long-tail distribution and high-dimensional sparse single cell data in incremental annotation can be effectively solved, data playback and continuous learning can be efficiently realized, and thus the performance and robustness of incremental annotation are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of single cell type annotation, and in particular relates to a single cell incremental annotation method based on distribution and expression-aware revisiting. Background Art

[0002] Single-cell type annotation plays an important role in the study of targeted therapy of tumor microenvironment. However, with the exponential growth of single-cell sequencing data, existing methods face three core challenges: First, directly iteratively training existing models will lead to serious catastrophic forgetting problems, such as Figure 1 As shown in (a) in Figure 2; secondly, independent models based on batch training are difficult to meet the diverse cell annotation requirements in open environments, such as Figure 1 As shown in (b) of the figure; furthermore, the data categories exhibit a significant long-tail distribution. These limitations essentially reflect the inadequacy of existing methods in supporting incremental annotation of single-cell classes, severely restricting the widespread application of this technology in precision medicine clinical practice. Therefore, overcoming the scientific challenge of single-cell long-tail incremental annotation is not only of great theoretical value but also of urgent practical significance.

[0003] While other fields currently focus on incremental scenarios with balanced data distribution, single-cell gene expression, due to its high-dimensional sparsity, high specificity, and extreme imbalance between cell types, faces the severe challenge of long-tail incremental learning. Under this extreme distribution, traditional incremental methods struggle to balance learning between new and old cell types, further exacerbating catastrophic forgetting and efficiency bottlenecks. To address this long-tail phenomenon, existing research has often employed data augmentation strategies to alleviate the low-sample dilemma by expanding data sources. However, in real-world scenarios, single-cell data typically originate from the open world, and its long-tail distribution is more consistent with biological and natural laws. Therefore, it is necessary to explore more practical strategies to address the negative impact of long-tail distributions. Summary of the Invention

[0004] The purpose of the present invention is to address the above-mentioned deficiencies in the prior art and provide a single-cell incremental annotation method based on distribution and expression-aware revisiting to solve the problem that the existing traditional incremental annotation methods are difficult to balance the learning of new and old cell types, further exacerbating the problems of catastrophic forgetting and efficiency bottlenecks.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is: A single-cell incremental annotation method based on distribution and expression-aware revisiting, comprising the following steps: S1. Input the single-cell expression profile sample to be processed into the diffusion model based on distribution-aware conditions to generate gene expression profile data of old cell samples that have participated in the diffusion model training; S2, integrating the gene expression profile data of the old cell sample with the gene expression profile data of the collected new cell sample; S3. The integrated data is input into two different feature extractors in the expression-aware knowledge distillation model to perform multi-perspective gene expression alignment and output the corresponding cell type annotation results.

[0006] Furthermore, S1 specifically includes the following sub-steps: S11. Represent the single-cell expression profile sample to be processed as a true expression profile in matrix form ; S12, the true expression spectrum As input, it is fed into the sparse feature extraction module to extract sparse masks describing key gene expression regions; S13, Conditional probability distribution based on diffusion model, for true expression profile Perform noise addition operation; S14: Take the noise addition result as input and repeat S12-S13 until the final sparse mask is extracted. and the noise expression spectrum corresponding to the final noise addition operation result ; S15, noise expression spectrum As the starting input of diffusion inversion, and through multi-step reverse denoising operation, the generated expression spectrum is gradually restored .

[0007] Furthermore, S14 specifically includes: at each time step of the diffusion model , the input data is constrained by a dynamic sparse mask, which is specifically expressed as: in, Where, Expressed as the conditional probability distribution of the diffusion model; Represents the time step Dynamic sparse mask of; Express endowment The probability weights make Noise distribution with sparse distribution guidance; represents a normal distribution; Represents the time step the degree of retention of original data; The covariance of the noise is a diagonal matrix; is the Sigmoid function; and is an adjustable parameter, Genes in dynamic sampling data The expression level of is a constant.

[0008] Furthermore, in S15, the conditional probability distribution of reverse denoising is expressed as: Where, represents the conditional probability distribution of reverse denoising; and They represent the direction and amplitude of denoising respectively.

[0009] Furthermore, in S15, a sparsity constraint loss is introduced in each step of diffusion inversion recovery. , to keep the generated expression profile and the true expression profile with a consistent sparse structure; Among them, the sparsity constraint loss Expressed as: Where, The time steps Noise expression spectrum after denoising and noise addition ; and is the weight coefficient of sparse area and non-sparse area; is an indicator function, which takes the value 1 when the condition is met, otherwise it is 0; for quantile threshold of ; It is a measure of the error between the generated expression profile data and the true expression profile data.

[0010] Furthermore, S3 specifically includes the following sub-steps: S31. Copy the integrated data into two copies and record them as basic data respectively. With new class data ; S32, the basic class data Input the basic class feature extractor in the knowledge distillation model and convert the new class data Input into the new class feature extractor in the knowledge distillation model and output the basic class high-level features respectively and new class advanced features At the same time, multiple levels of basic class attention maps are extracted from the multi-scale attention modules in the basic class feature extractor and the new class feature extractor. and multiple levels of new class attention maps ; S33, using minimized attention alignment loss , for the basic class attention map and multiple levels of new class attention maps Align the corresponding attention distributions between them; S34, using minimized representation alignment loss , for advanced features of the base class and new class advanced features Perform structural alignment in feature space; S35, will satisfy the minimum representation alignment loss Advanced features of the base class Input into the basic annotator to get the corresponding basic annotation prediction ; Add new class advanced features Input into the new class annotator to get the new class annotation prediction ; S36. Using category prediction consistency loss , for basic annotation prediction and new class annotation prediction Align the annotation distribution.

[0011] Furthermore, in S32, the base class attention graphs of multiple levels in the multi-scale attention module and multiple levels of new class attention maps The calculation expression is: Where, An attention map representing each attention layer; are linear mappings of the input data, representing query, key, and value matrices respectively; d is the dimension of the key; and i is the index.

[0012] Furthermore, in S33, the attention alignment loss is minimized Expressed as: In S34, the characterization alignment loss is minimized Expressed as: In S36, the category prediction consistency loss Expressed as: Based on minimizing attention alignment loss , minimize the representation alignment loss and category prediction consistency loss , perform multi-perspective gene expression revisit alignment: Where, is the dimension of the attention map matrix; is the attention map matrix output by the base class feature extractor; The attention map matrix output by the feature extractor for the new class; is the feature matrix output by the basic class feature extractor; is the feature matrix output by the new class feature extractor; Logits are output by the base class annotator; Logits are output by the new class annotator; , , represents the alignment weight of the regulatory feature and the gene map; Represents the multi-view gene expression revisit alignment loss function.

[0013] Furthermore, fuzzy guidance constraints are introduced into the base class feature extractor to suppress the false recall of the base class feature extractor on unseen cell types, while guiding the new class feature extractor to enhance its adaptability to new types; Among them, the fuzzy guidance constraint is expressed as: Where, represents the fuzzy guided loss function; is the cardinality of the incremental sample set; and is the weight parameter; is the preset minimum false recall threshold; is the Sigmoid activation function; Representative The predicted logit for each sample.

[0014] Further, according to 、 and Compute the summed training loss for the incremental class annotations: in, Where, represents the cross entropy loss function; It is the cross entropy loss function operation; Indicates data splicing operation The formed joint input; Cell type information.

[0015] The single-cell incremental annotation method based on distribution and expression-aware revisiting provided by the present invention has the following beneficial effects: 1. This invention innovatively designs two perceptual paradigms for revisit learning: the conditional diffusion generation paradigm of distribution perception and the fuzzy incremental guidance paradigm of expression perception. By deeply integrating the above-mentioned perceptual paradigms with the core modules, data replay and continuous learning can be efficiently realized, thereby significantly improving the performance and robustness of incremental annotation. Moreover, this invention can be seamlessly integrated into other conventional annotation models, thereby giving full play to its wide applicability and advantages in continuous learning.

[0016] 2. The incremental annotation process of the present invention first uses the conditional diffusion model, leveraging its distribution-aware capabilities to generate replay data of cell samples from the old task in a nearly lossless manner. This replay data is then jointly trained with samples of the new cell type. The jointly trained samples are input into the incremental knowledge distillation architecture, which uses expression awareness to achieve multi-perspective gene expression re-alignment and incorporates fuzzy guided constraints to eliminate the interference of false recall. This design of the present invention effectively solves the core problem of incremental annotation of long-tailed, high-dimensional, sparse single-cell data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Comparison between conventional cell annotation and incremental cell annotation methods; Figure 1 (a) is an independently trained annotation model used to learn unseen cell types; Figure 1 (b) is a pre-trained annotation model used for transfer learning of new cell types; Figure 1 (c) in FIG. 1 is the incremental cell annotation framework proposed in Example 1 of the present invention.

[0018] Figure 2 This is the single cell class incremental annotation process in Example 1; Figure 2 (a) is a flowchart of the process of generating old cell gene expression profiles by distribution-aware conditional diffusion; Figure 2 (b) is a flowchart of the incremental annotation process of single-cell long-tail classes with expression-aware knowledge distillation and fuzzy guidance.

[0019] Figure 3 This is the forward denoising and reverse denoising process for generating single-cell expression profiles using sparse distribution-aware conditional diffusion in Example 1.

[0020] Figure 4 This is the multi-perspective gene expression-aware revisit alignment process under the knowledge distillation architecture in Example 1, including multi-scale attention feature map alignment, high-level feature alignment, and annotation type representation alignment.

[0021] Figure 5 Schematic diagram of fuzzy guidance constraints in Example 1.

[0022] Figure 6Hyperparameter analysis for mitigating forgetting in gene expression-aware alignment on the Vento dataset in Example 2.

[0023] Figure 7 This is the result of the distribution-aware multi-scale gene expression review in Example 2. Due to the large number of genes, the gene expression vector is reshaped into a rectangular form in the figure, where: Figure 7 (a) is the expression follow-up result of PMSC cells in the He dataset; Figure 7 (b) is the expression review result of NK cells in the Madissoon dataset; Figure 7 (c) is the expression follow-up result of Endothelium cells in the Stewart dataset; Figure 7 (d) in the figure is the expression follow-up result of dS1 cells in the Vento dataset.

[0024] Figure 8 This is to evaluate the effect of scLTCIA distribution awareness on the gene expression levels generated by the He, Madissoon, Stewart and Vento datasets in Example 2.

[0025] Figure 9 These are the three key visualization results on the Zheng68k dataset in Example 2; Figure 9 (a) shows the long-tail distribution of experimental settings and cell types, reflecting the extremely unbalanced label distribution characteristics of the dataset in the last incremental stage; Figure 9 (b) shows the visualization of the multi-perspective alignment based on gene expression perception. It compares the changes in the feature distribution of the Base model and the Novel model on the new and old cell types before and after alignment in the second incremental phase, verifying the effectiveness of the feature space alignment strategy. Figure 9 (c) in the figure compares the differences in gene expression distribution between the generated samples and the real samples for the extremely rare cell type CD34+, further verifying the accuracy and biological consistency of the expression recall.

[0026] Figure 10 This is a flowchart of the single-cell incremental annotation method based on distribution and expression-aware revisiting according to Example 1 of the present invention. DETAILED DESCRIPTION

[0027] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0028] Example 1 This embodiment provides a single-cell incremental annotation method based on distribution and expression-aware revisiting. This embodiment first designs a distribution-aware generation framework to accurately revisit old data in the model memory and efficiently replay single-cell data with high-dimensional sparse characteristics for incremental annotation. At the same time, an expression-aware knowledge distillation architecture is designed. By aligning gene representations from multiple perspectives, it significantly improves the perception and representation capabilities of specific expressions of new types of small samples on the basis of strengthening the memory retention of mainstream categories. In addition, fuzzy incremental guidance constraints are introduced to effectively eliminate false memories of new types of cells in the incremental process, thereby reducing the negative impact of long-tail distribution on incremental annotation; The corresponding framework of the method in this embodiment is the single cell incremental annotation framework scLTCIA, reference Figure 1 (c) Figure 2 and Figure 10 , which specifically includes the following: S1. Input the single-cell expression profile sample to be processed into the diffusion model based on distribution-aware conditions to generate gene expression profile data of old cell samples that have participated in the diffusion model training; like Figure 2 As shown in (a), this embodiment uses the pre-trained distribution-aware diffusion model based on the type and number information of single cells to generate gene expression profile data of old cell samples that have participated in training. Subsequently, the generated gene expression profile data is integrated with the gene expression profile data of the newly collected new cell samples and input together. Figure 2 (b) The incremental learning framework shown.

[0029] Generative replay, with its advantage of avoiding large-scale storage consumption, has become an effective strategy to resolve catastrophic forgetting in incremental learning. However, the characteristics of high-dimensional sparse expression profile distribution in single-cell incremental annotation seriously restrict the quality of generative replay. To this end, this embodiment proposes a distribution-aware expression profile generation paradigm, referring to Figure 3 , which specifically includes the following steps: S11. Extracting real single-cell expression profile data; Represent the single-cell expression profile sample to be processed as a true expression profile in matrix form , which has high-dimensional and sparse characteristics and serves as the input data of the model.

[0030] S12, extracting sparse expression structure information; The true expression spectrum As input, it is fed into the sparse feature extraction module to extract sparse masks used to describe key gene expression regions. , to capture the effective information regions in the expression profile; S13, performing a noise addition operation of the diffusion model; Conditional probability distribution based on diffusion model , for the true expression spectrum Perform noise addition operation (noise disturbance processing) to generate the noise addition operation result of the first step , to simulate the gradual perturbation process of data in the latent space; S14, repeating the sparse mask extraction and noise addition process; The result of the noise addition operation is used as the new input to replace the true expression spectrum input in S12 , and repeat S12~S13, that is, the intermediate noise addition operation results generated in each step (such as , Until ) Repeated input sparse extraction module Extract its sparse mask , and through the conditional probability Continue adding noise until the final sparse mask is extracted and the noise expression spectrum corresponding to the final noise addition operation result .

[0031] More specifically, in order to preserve the sparse statistical characteristics of real data, this embodiment introduces a dynamic sparse mask into the diffusion model. This dynamic sparse mask maintains the sparse areas and enhances the noise in the non-sparse areas in the forward diffusion operation. In the reverse operation, it ensures that the sparse distribution of the denoised data is close to the real distribution, thus achieving both generation quality and biological consistency.

[0032] Among them, dynamic sparse mask generation is sampled from the real expression spectrum And the data after the noise operation at each subsequent time step The sparsity statistics of , specifically, express Step gene The mask value of is determined by the sparse distribution of gene expression at that position. If the expression level is very low in the sampled data, Approaches 0; on the contrary, if the gene The higher the expression level, the Approaching 1, the process can be expressed as: At each time step of the diffusion model , the input data is constrained by a dynamic sparse mask, which is specifically expressed as: Where, Expressed as the conditional probability distribution of the diffusion model; Represents the time step Dynamic sparse mask of; Express endowment The probability weights make Noise distribution with sparse distribution guidance; represents a normal distribution; Represents the time step the degree of retention of original data; The covariance of the noise is a diagonal matrix (i.e., each dimension is independent and identically distributed); is the Sigmoid function; and is an adjustable parameter, Genes in dynamic sampling data The expression level of is a constant.

[0033] S15, initializing the reverse diffusion process; Noise expression spectrum As the starting input of diffusion inversion, it enters the reverse generation process constructed by Diffusion U-Net, and through multi-step reverse denoising operations, the generated expression spectrum is gradually restored. , which is used to gradually restore the biological expression profile structure from pure noise.

[0034] Specifically, this embodiment is for the real expression spectrum After normalization and sparse mask processing, sparse pattern learning is optimized. Cell type information With sample identifier Joint encoding into conditional embeddings via pre-trained Text2vec model , and linearly mapped to the embedding space , as the initial condition , and embed standard Gaussian noise Afterwards as initial input.

[0035] In the reverse reasoning stage, the model is trained from the noise expression spectrum Starting from, through multiple denoising steps, The expression spectrum is gradually restored and generated. At each step, the diffusion model predicts the expression spectrum distribution at the current moment, that is, the target data is generated by step-by-step denoising - generating expression spectrum ; To perceive the sparse characteristics of gene expression data, a dynamic sparse mask is introduced , guides the update of specific regions in the denoising process. Specifically, for each time step and genes , sparse mask It is used to dynamically adjust the amplitude and area of ​​denoising to ensure that the generated data retains the sparse distribution characteristics consistent with the real data; the conditional probability distribution of reverse denoising is expressed as: Where, represents the conditional probability distribution of reverse denoising; and They represent the direction and amplitude of denoising respectively.

[0036] In order to minimize the difference between the generated data and the real data, this embodiment introduces a sparse regularization loss for structural constraint. The quantile threshold method is used to dynamically distinguish sparse areas from non-sparse areas, and their errors are calculated separately. Specifically, the sparsity constraint loss is introduced in each step of the diffusion inversion recovery. , in order to keep the generated expression spectrum and the true expression spectrum with the same sparse structure, in order to keep the generated expression spectrum with the same sparse structure as the true expression spectrum, and improve the expression ability of the model in high-dimensional sparse scenarios; sparsity constraint loss Expressed as: Where, The time steps Noise expression spectrum after denoising and noise addition ; and is the weight coefficient of sparse area and non-sparse area; is an indicator function, which takes the value 1 when the condition is met, otherwise it is 0; for The quantile threshold (such as the 50% quantile) is used to dynamically distinguish sparse and non-sparse areas; To generate the error metric between the expression profile data and the true expression profile data, it is expressed as: Where, express and The squared Euclidean distance between them is ; this constraint ensures the global and local sparsity consistency between the generated data and the real data.

[0037] This example introduces a cell type recognition module for semantic constraints, and inputs the intermediate expression profiles during the diffusion process into the cell type identifier. , predict the cell type, and reversely optimize the generation model (diffusion model) based on the classification loss feedback to ensure that the generated spectrum is biologically reasonable and discernible; Based on the above process, the final output of this embodiment is It meets the requirements of sparse expression features and cell type semantic information at the same time, and can be used for downstream single-cell analysis tasks such as cell annotation and lineage inference.

[0038] S2, integrating the gene expression profile data of the old cell sample with the gene expression profile data of the collected new cell sample; Specifically, the old data generated by the diffusion model With new data As input in knowledge distillation, through data splicing operation Integrate to form joint input .

[0039] S3. Knowledge distillation, as a classic regularization strategy, has demonstrated excellent performance in continuous learning. However, traditional methods primarily focus on feature representation learning and struggle to effectively address the incremental annotation challenges posed by the specific gene expression and long-tail distribution in single-cell data. To this end, this example proposes an innovative multi-perspective gene attention mechanism and fuzzy guidance constraints to construct an expression-aware incremental annotation paradigm. Specifically, this embodiment inputs the integrated data into two different feature extractors in the expression-aware knowledge distillation model to perform multi-perspective gene expression alignment and output the corresponding cell type annotation results. Figure 4 , which specifically includes the following steps: S31, constructing expression spectrum data of multi-view input: the integrated data ( ) is copied into two copies and recorded as basic class data With new class data ; S32, the basic class data Input the basic class feature extractor in the knowledge distillation model and convert the new class data Input into the new class feature extractor in the knowledge distillation model and output the basic class high-level features respectively and new class advanced features ; At the same time, both the base class feature extractor and the new class feature extractor are composed of multiple attention layers (structures with queries, keys, and values). In the multi-scale attention module, multiple levels of base class attention maps are obtained respectively. and multiple levels of new class attention maps , used to capture gene expression patterns at different scales; Specifically, the base class attention graphs of multiple levels in the multi-scale attention module and multiple levels of new class attention maps The calculation expression is: Where, Represents the various levels of downsampling, An attention map representing each attention layer; are linear mappings of the input data, representing query, key, and value matrices respectively; d is the dimension of the key; i is the index.

[0040] S33, perform multi-scale attention feature alignment; Adopting the minimum attention alignment loss , for the basic class attention map and multiple levels of new class attention maps Align the corresponding attention distributions between the base class and the new class to achieve distribution consistency in attention response, improving cross-class expression perception capabilities; Minimize attention alignment loss Expressed as: S34, output high-level feature representation and perform representation alignment; Minimize the representation alignment loss , for advanced features of the base class and new class advanced features Structural alignment in feature space is beneficial for feature migration of new classes; Minimize representation alignment loss Expressed as: S35, enter the annotator to predict cell types; will satisfy the minimum representation alignment loss Advanced features of the base class Input into the basic annotator to obtain the corresponding basic annotation prediction and the new class high-level features Input into the new class annotator to get the new class annotation prediction ; S36, performing type prediction alignment; Using category prediction consistency loss , for basic annotation prediction and new class annotation prediction Align the annotation distribution of the dataset to ensure that new classes can obtain stable and reliable classification decision boundaries even without a large number of labeled samples; Category prediction consistency loss Expressed as: After the basic annotation model was trained with the old type, it accumulated rich and detailed gene expression data. and cell type characterization , can be used as a base model to guide the representation alignment of the new model under the knowledge distillation architecture. In addition, in the incremental annotation scenario, the immature new model revisits the gene expression "world view" of the base model and dynamically aligns the receptive field , in order to accurately capture the hierarchical information of specific gene expression in single cells, the alignment constraints of knowledge distillation are formally defined as: Where, is the dimension of the attention map matrix; is the attention map matrix output by the base class feature extractor; The attention map matrix output by the feature extractor for the new class; is the feature matrix output by the basic class feature extractor; is the feature matrix output by the new class feature extractor; Logits are output by the base class annotator; Logits are output by the new class annotator; , , represents the alignment weight of the regulatory feature and the gene map; Represents the multi-view gene expression revisit alignment loss function.

[0041] This embodiment introduces fuzzy-guided long-tail incremental annotation into the basic model (diffusion model); As a preferred embodiment of this invention, refer to Figure 6 , in the distillation architecture, the predictive representation of the base model The dimension corresponds to the number of categories of basic cell types. Although the basic model can effectively revisit the representation of the basic types during the replay of the class increment, its prediction of new types often leads to misclassification due to the probability peak distribution of the basic types. In order to improve the incremental annotation performance, a fuzzy guidance constraint is proposed such as Figure 5 As shown, this method can suppress the false recall of the base model as a teacher on unseen cell types, while effectively guiding the new model to enhance its adaptability to new types, avoid forgetting, and improve its learning ability for long-tail data; Among them, the fuzzy guidance constraint is expressed as: Where, represents the fuzzy guided loss function; is the cardinality of the incremental sample set; and is the weight parameter; is the preset minimum false recall threshold; is the Sigmoid activation function; Representative The predicted logit of samples. In the fuzzy bootstrap constraint formula, the first term exceeds the square penalty constraint The second term provides stable gradient guidance by smoothing logarithmic penalties; by constraining the maximum value of the logit predicted by the basic model to not exceed , they can be classified as unknown types, thus guiding the new model to perceive new type-specific genes in incremental tasks. This method not only enhances the deep perception of new type-specific genes, but also significantly improves the performance of incremental annotation of single-cell long-tail classes.

[0042] Incremental annotation reasoning process; The proposed method is a two-stage incremental annotation framework: first, data replay is achieved through a sparse-aware conditional diffusion model, and then long-tail classes are incrementally annotated using gene-aware and fuzzy guidance. With new data As input in knowledge distillation, through data splicing operation Forming joint input Then, a multi-perspective gene expression revisit alignment is performed. Secondly, the basic model (diffusion model) introduces a fuzzy guidance mechanism to suppress False memories, when guiding the new model, use the cross entropy loss function Calculating losses , significantly improving the ability to perceive genes specifically expressed in single cells under long-tail distribution, effectively alleviating the forgetting problem and achieving continuous and accurate annotation of rare cell types.

[0043] This example mines gene relationships at different scales, integrates low-level local gene expression with high-level global features, and constructs the attention graph of each layer of the output extractor. , , and the cell type representation of the annotator output logit ,in Represents a comment operation.

[0044] Based on this, the total training loss for class incremental annotation in this embodiment is: in, Where, represents the cross entropy loss function; It is the cross entropy loss function operation; Indicates data splicing operation The formed joint input; Cell type information.

[0045] Example 2 This example verifies the performance and effect of the method of Example 1 through experiments, which specifically include the following contents: Experimental setup Benchmark datasets; This example analyzes five single-cell datasets: He, containing 14 cell types and 15,680 cells. Madissoon, containing 25 cell types and 57,020 cells. Stewart, containing 43 cell types and 26,628 cells. Vento, containing 32 cell types and 64,734 cells. Zheng68k, containing 11 cell types and 65,943 cells. The first four datasets are split into five stages of incremental tasks. The specific incremental tasks for each dataset in each stage are shown in the formula As shown, the number of cell types in the dataset is divided into the first stage categories and subsequent 4 incremental phases The Zheng68k dataset, due to its significant long-tail distribution, is divided into three incremental phases for model comparison and validation, computational efficiency evaluation, and visualization analysis under more stringent comparison conditions. For details on the specific divisions of these datasets, see the dataset splits listed in Table 1.

[0046] Baseline comparison method; This example selects a single-cell type annotation baseline method that is easy to modularize and expand, and embeds it into the single-cell incremental annotation framework scLTCIA proposed in Example 1 as the core feature extraction module: CIForm: This method is based on the Transformer architecture. Its sub-vector self-attention mechanism has a good modular design, which makes it easy to split the model into an encoder (for feature extraction) and a classification head (for cell type prediction), so that it can be flexibly embedded in the incremental training process of the continuous learning framework.

[0047] scTransSort: This method uses a Transformer model to extract features from disordered input data, effectively alleviating the sparsity of single-cell data without manual annotation or reliance on external reference data, demonstrating excellent cell type recognition performance. Its end-to-end feature modeling capabilities make it a strong candidate for high-quality feature extraction modules in continuous learning frameworks.

[0048] scPred: This method uses singular value decomposition (SVD) for feature dimensionality reduction and screening, and combines it with support vector machines (SVM) to build a high-precision classification model. It performs well in cell annotation tasks and has good generalization ability, making it suitable as a core feature extraction module in a continuous learning framework.

[0049] Furthermore, to further validate the comparative effectiveness of our method against representative algorithms in the general field of continuous learning, we selected two classic continuous learning methods widely used in image classification tasks for a more rigorous evaluation on the Zheng68k dataset. During this evaluation, the scRNA-seq data was reshaped (a tensor operation that changes the shape of data) into an "image format" suitable for convolution operations, enabling these methods to run on our data, thus enabling a fair comparison. iCaRL: This method innovatively combines a nearest neighbor mean classification strategy, a sample selection mechanism based on herding (a method mentioned in the iCaRL paper), and a feature learning method that combines sample replay and knowledge distillation. It can simultaneously optimize feature representation and classifiers in a quasi-incremental learning scenario, effectively mitigating catastrophic forgetting and improving overall performance.

[0050] DER: This method combines sample replay, knowledge distillation, and regularization constraints to optimize network predictions (logits) rather than relying directly on true labels. This approach is more in line with the general continuous learning (GCL) paradigm, has better model calibration, and demonstrates leading performance on multiple benchmarks.

[0051] Evaluation indicators; Classification accuracy is used to evaluate the performance of baseline models in incremental annotation tasks. Specifically, three core indicators are used: B represents the annotation accuracy of known cell types (base), N represents the annotation accuracy of newly introduced cell types (new), and B+N represents the overall accuracy for all cell types. In addition, the resource overhead of different baseline methods is compared and analyzed in terms of training time, GPU memory usage, and model parameter count, and is reported in minutes (mins), gigabytes (G), and millions of parameters (M). In the distribution-aware gene expression replay experiment, two evaluation indicators, MAE (mean absolute error) and SSIM (structural similarity index), are further introduced to measure the average error between the generated replay data and the real gene expression data.

[0052] Specific implementation details; The evaluation experiment is designed into two levels: conventional benchmark test and enhanced benchmark test.

[0053] In the conventional benchmark test, four public datasets (He, Madissoon, Stewart and Vento) are divided into five task stages, and two task construction strategies are used for evaluation: (1) Ordered task: the long-tail class samples are divided in order from large to small to simulate an idealized incremental learning scenario; (2) Shuffled task: the class order is randomly shuffled while keeping the overall class distribution unchanged to simulate the random, unstructured long-tail distribution that may occur in real biological sequencing, thereby evaluating the robustness and generalization ability of the model in a more complex environment.

[0054] Using the two aforementioned partitioning strategies, we integrated three mainstream single-cell type annotation methods (scPred, CIForm, and scTransSort) into the proposed incremental annotation framework, scLTCIA, to systematically evaluate their performance on the long-tail incremental annotation task. To ensure a fair comparison, all models were trained with a uniform initial learning rate of 1e-4 for the first three incremental phases and 1e-2 for the last two phases. The maximum number of training epochs was 100, and an early stopping mechanism based on the validation set loss was employed, with a tolerance for no improvement of up to 15 epochs.

[0055] In the enhanced benchmark, the large-scale Zheng68k dataset was selected and divided into three task phases, with the number of class samples allocated in descending order: the first phase contains 5 classes, and the second and third phases each contain 3 classes, simulating a sparser and more challenging incremental learning scenario. In this setting, the performance of three types of methods was compared: (1) the performance of three single-cell annotation methods, scPred, CIForm, and scTransSort, integrated in an incremental framework; (2) the direct training results of CIForm, which performs relatively well in a non-incremental setting; and (3) the adaptability of representative general continuous learning methods (iCaRL and DER). In addition to accuracy, the methods were further compared in terms of training time, GPU memory usage, and model parameter consumption, comprehensively evaluating the trade-offs between performance and resource efficiency of different methods.

[0056] Conventional benchmark experiment results Mixed order long tail incremental annotation; The experimental results are shown in Table 2. The scLTCIA of the present invention outperformed other baseline models in the mixed-order long-tail incremental task on all datasets. Compared with other methods, scLTCIA showed stronger stability in the long-tail incremental annotation process, and its accuracy of new and old types fluctuated significantly less at each stage. For example, in the final stage of the Madissoon and Vento datasets, scLTCIA's learning accuracy for new cell types exceeded the suboptimal methods by 7.86% and 6.68%, respectively, highlighting its significant advantage in learning new knowledge. In terms of anti-forgetting ability, scLTCIA achieved an accuracy of 88.51% for old types in the final stage of the He dataset, which was 1.95% higher than the suboptimal method scTransSort, further demonstrating the significant advantage of the framework of the present invention in maintaining the integrity of old knowledge. These results show that scLTCIA can not only efficiently learn new types in long-tail distribution data with biological randomness, but also effectively alleviate the problem of catastrophic forgetting, demonstrating its strong adaptability and robustness in complex biological scenarios.

[0057] Sequential long-tail incremental annotation The experimental results are shown in Table 3. In mixed-order annotation tasks, tail categories (with very few samples) are often overlooked because they have a minimal impact on overall accuracy. However, in sequential long-tail incremental tasks, by learning in descending order based on the number of samples, the model's incremental performance on tail categories can be more effectively evaluated, revealing its ability to handle rare-sample types. For example, in the final stage of the Madissoon dataset, scLTCIA achieved an annotation accuracy of 85% for new types with very few samples, significantly higher than CIForm's 55.08% and scPred's 40.53%. Furthermore, on the Stewart dataset, scLTCIA consistently maintained an annotation accuracy of over 92% for tail categories, and the suboptimal method scTransSort, which extends the framework of our invention, approached scLTCIA's performance across all metrics. These results demonstrate that scLTCIA exhibits excellent scalability and extensibility in sequential long-tail incremental annotation scenarios.

[0058] Ablation studies The results in Table 4 show that the scLTCIA framework has a significant advantage in incremental annotation of a very small number of samples in the tail class, especially in the final stage of the Madissoon dataset, where it exhibits high plasticity for a very small number of cell types. Through ablation experiments, we quantitatively evaluate the contribution of the three key constraints of scLTCIA to incremental annotation on the Madissoon dataset. As shown in Table 4, compared with using only compared to, and The plasticity of incremental annotation is significantly improved, thanks to the knowledge distillation architecture that enhances the specific attention to new types through gene expression awareness. In addition, the introduction of The fuzzy guided incremental constraint significantly improves the annotation accuracy of new types at the tail of the long-tail distribution in the final stage, indicating that it has a significant effect on the uncertainty increment effect of very few samples.

[0059] gene expression-aware alignment analysis; The expression-aware knowledge distillation system is key to maintaining the memory of old cell revisits. Therefore, for the best-performing Vento dataset, we analyzed the hyperparameter sensitivity of the two alignment operations involved in the expression-aware paradigm, multi-scale gene feature map alignment and cell type representation alignment. The experimental results are shown in Figure 2. Figure 6 In the nine experimental settings, by comparing the basic accuracy evaluation indicators of the four incremental stages, gene expression perception shows the best effect in preventing forgetting when and .

[0060] Multi-scale diffusion revisit Figure 7 The effectiveness of the distributed-aware diffusion model proposed in this paper in revisiting cell expression patterns in historical memory is demonstrated. In the four datasets of He, Madissoon, Stewart, and Vento, four types of cells, PMSC, NK, Endothelium, and dS1, were selected respectively. At three different gene dimension scales, scLTCIA almost perfectly reconstructed the cell expression characteristics in historical memory. It is worth noting that it can be observed from the figure that at all scales, the gene expression data showed high sparsity, with a large number of zero values. However, even in such an extremely sparse situation, the model can still maintain extremely high revisit accuracy, which fully demonstrates the robustness and effectiveness of the proposed sparse distribution-aware module.

[0061] Revisit gene expression visualization; T-SNE dimensionality reduction was used to retrospectively visualize the gene expression of the four datasets He, Madissoon, Stewart, and Vento. The gene regularized expression distribution of cells generated by the scLTCIA framework after the second incremental annotation session was compared with the real data. Figure 8As shown in the figure, except for the Madissoon dataset, the samples generated by the latent space in the other datasets are highly consistent with the true distribution after dimensionality reduction, showing good distribution fidelity. In the Madissoon dataset, the generated samples show a distribution pattern that is roughly mirror-symmetrical with the real data. This deviation may be due to slight differences between the generated data and the true distribution in high-dimensional space, especially slight shifts in the sparsity structure. Considering the high sensitivity of T-SNE to local structural changes, these high-dimensional micro-differences may be amplified during the dimensionality reduction process. Overall, the results show that scLTCIA can effectively model high-dimensional sparse expression distributions, and has a strong distribution perception ability when generating data revisits, which is consistent with the results of the previous article. Figure 7 The return visit accuracy shown in confirms each other.

[0062] Enhanced benchmark experiments Additional baseline comparisons; In order to verify the results on a more demanding dataset, we used Zheng68k, a dataset widely used as a benchmark with an extremely obvious long-tail distribution, for testing. The long-tail distribution of cell types in the Zheng68k dataset is shown in the figure below. Figure 9 As shown in (a). The dataset is divided into three incremental sessions, with five initial classes in the first round and three incremental classes in the second and final rounds. CIForm, scPred, and scTransSort are used as baseline feature extraction modules and integrated into our scLTCIA framework to compare the performance of incremental annotation. At the same time, CIForm, the baseline method with the best generalization performance for unseen cells in the final round of the three cell annotation models, was selected and directly trained and tested separately to evaluate non-incremental and incremental methods. To increase the theoretical support for continuous learning in the experiment, two comparison baselines, iCaRL and DER, which are widely acclaimed traditional continuous learning methods in the image field, were introduced for evaluation.

[0063] As can be seen from Table 5, the framework of the present invention shows the best performance in terms of overall incremental annotation accuracy. It is almost on par with the DER method in the traditional continuous learning field in terms of accuracy and computational overhead. At the same time, from the comparison between the non-incremental annotation CIForm-DA method and the incremental CIForm-CIA method, it can be seen that the incremental implementation method performs better in the final B+N overall accuracy. Figure 9The data distribution in (a) also shows that the three cell types used in the Zheng68k dataset in the last incremental session only account for 3.2% of the total number of cells, demonstrating a very extreme long-tail trend. Therefore, it can be concluded that the CIForm-DA baseline method, which performs direct supervised training, still has room for improvement in annotating rare cells. At the same time, CIForm-CIA leverages the fuzzy guidance mechanism within our scLTCIA incremental annotation framework to prevent new cells from being mistakenly predicted as numerous common cells, achieving a remarkable accuracy of 56.09% for the annotation of rare cells in this final round.

[0064] Looking at the overall computational overhead, the scLTCIA framework uses comparable GPU memory to traditional continuous learning frameworks. While slightly more computationally expensive than directly training a cell annotation model, the improved annotation accuracy for rare cells achieved using incremental training, coupled with good control over training time and overhead, is commendable. Overall, the scLTCIA framework is an excellent architecture that balances annotation performance with low computational overhead.

[0065] Alignment visualization of expression distribution; In addition, to verify the multi-perspective alignment capability of the student model to the teacher model in the knowledge distillation architecture of gene expression perception, Figure 8 (b) visualizes the feature distributions of the old and new classes in the second session of the Zheng68k dataset, showing how the base model and the new model represent the same cell types before and after alignment. As shown, before alignment, the distribution of feature points generated by the base model (red) and the new model (blue) differ significantly. The embedding distributions of cells of the old category (circles) and cells of the new category (squares) are inconsistent in the two models, exhibiting a clear separation. However, after applying the distillation alignment strategy, the feature points generated by the two groups of models almost completely overlap, with red and blue points highly overlapping, indicating that the student model has been successfully aligned to the feature space of the teacher model, achieving consistency in the distribution of old and new cell types. This phenomenon intuitively demonstrates the effectiveness of the proposed gene expression-aware multi-perspective distillation alignment strategy in incremental learning, providing solid support for generalization of rare cell types.

[0066] Rare cell expression return visualization; This example further selected the tenth category of CD34+ cells, which only contained 188 samples, as an extremely rare cell type in the Zheng68k dataset. A sparse distribution-aware expression revisit generation strategy was used to compare and analyze the differences in gene expression between the generated samples and the real samples. The results are as follows: Figure 8 (c) shows the regions where gene expression between the generated cells and real cells differ significantly, as indicated by green boxes. The results show that these regions generally capture the significant sparse expression trends in the scRNA-seq data, but some gaps remain in reproducing the expression levels of some specific loci, requiring further optimization of simulation accuracy. This phenomenon is speculated to be due to the extremely low overall proportion of CD34+ cells in Zheng68k, at only 0.29%. Therefore, even with such a small sample size, achieving a preliminary reconstruction of gene expression distribution characteristics is a significant challenge, demonstrating the potential of distribution-aware revisit strategies. The proposed fuzzy-guided incremental strategy also strives to mitigate the impact of revisit accuracy issues caused by long-tail distributions on incremental annotation performance, and will continue to improve its generation accuracy in extremely low-resource scenarios.

[0067] Although the specific embodiments of the invention are described in detail in conjunction with the accompanying drawings, this should not be construed as limiting the scope of protection of this patent. Within the scope described by the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.

Claims

1. A single-cell incremental annotation method based on distribution and expression-aware revisiting, characterized by: The following steps are involved: S1. Input the single-cell expression profile sample to be processed into the diffusion model based on distribution-aware conditions to generate gene expression profile data of old cell samples that have participated in the diffusion model training; S2, integrating the gene expression profile data of the old cell sample with the gene expression profile data of the collected new cell sample; S3. The integrated data is input into two different feature extractors in the expression-aware knowledge distillation model to perform multi-perspective gene expression alignment and output the corresponding cell type annotation results.

2. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 1 is characterized in that The S1 specifically includes the following sub-steps: S11. Represent the single-cell expression profile sample to be processed as a true expression profile in matrix form ; S12, the true expression spectrum As input, it is fed into the sparse feature extraction module to extract sparse masks describing key gene expression regions; S13, Conditional probability distribution based on diffusion model, for true expression profile Perform noise addition operation; S14: Take the noise addition result as input and repeat S12-S13 until the final sparse mask is extracted. and the noise expression spectrum corresponding to the final noise addition operation result ; S15. Noise expression spectrum As the starting input of diffusion inversion, and through multi-step reverse denoising operation, the generated expression spectrum is gradually restored .

3. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 2 is characterized in that The S14 specifically includes: at each time step of the diffusion model , the input data is constrained by a dynamic sparse mask, which is specifically expressed as: in, Where, Expressed as the conditional probability distribution of the diffusion model; Represents the time step Dynamic sparse mask of; Express endowment The probability weights make Noise distribution with sparse distribution guidance; represents a normal distribution; Represents the time step the degree of retention of original data; The covariance of the noise is a diagonal matrix; is the Sigmoid function; and is an adjustable parameter, Genes in dynamic sampling data The expression level of is a constant.

4. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 3 is characterized in that In S15, the conditional probability distribution of reverse denoising is expressed as: Where, represents the conditional probability distribution of reverse denoising; and They represent the direction and amplitude of denoising respectively.

5. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 3 is characterized in that In S15, a sparsity constraint loss is introduced in each step of diffusion inversion recovery. , to keep the generated expression profile and the true expression profile with a consistent sparse structure; Among them, the sparsity constraint loss Expressed as: Where, The time steps Noise expression spectrum after denoising and noise addition ; and is the weight coefficient of sparse area and non-sparse area; is an indicator function, which takes the value 1 when the condition is met, otherwise it is 0; for quantile threshold of ; It is a measure of the error between the generated expression profile data and the true expression profile data.

6. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 1 is characterized in that The S3 specifically includes the following sub-steps: S31. Copy the integrated data into two copies and record them as basic data respectively. With new class data ; S32, the basic class data Input the basic class feature extractor in the knowledge distillation model and convert the new class data Input into the new class feature extractor in the knowledge distillation model and output the basic class high-level features respectively and new class advanced features At the same time, multiple levels of basic class attention maps are extracted from the multi-scale attention modules in the basic class feature extractor and the new class feature extractor. and multiple levels of new class attention maps ; S33, using minimized attention alignment loss , for the basic class attention map and multiple levels of new class attention maps Align the corresponding attention distributions between them; S34, using minimized representation alignment loss , for advanced features of the base class and new class advanced features Perform structural alignment in feature space; S35, will satisfy the minimum representation alignment loss Advanced features of the base class Input into the basic annotator to get the corresponding basic annotation prediction ; Add new class advanced features Input into the new class annotator to get the new class annotation prediction ; S36. Using category prediction consistency loss , for basic annotation prediction and new class annotation prediction Align the annotation distribution.

7. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 6 is characterized in that: In S32, the base class attention graphs of multiple levels in the multi-scale attention module and multiple levels of new class attention maps The calculation expression is: Where, An attention map representing each attention layer; are linear mappings of the input data, representing query, key, and value matrices respectively; d is the dimension of the key; i is the index.

8. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 6 is characterized in that: In S33, minimize the attention alignment loss Expressed as: In S34, the characterization alignment loss is minimized Expressed as: In S36, the category prediction consistency loss Expressed as: Based on minimizing attention alignment loss , minimize the representation alignment loss and category prediction consistency loss , perform multi-perspective gene expression revisit alignment: Where, is the dimension of the attention map matrix; is the attention map matrix output by the base class feature extractor; The attention map matrix output by the feature extractor for the new class; is the feature matrix output by the basic class feature extractor; is the feature matrix output by the new class feature extractor; Logits are output by the base class annotator; Logits are output by the new class annotator; , , represents the alignment weight of the regulatory feature and the gene map; Represents the multi-view gene expression revisit alignment loss function.

9. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 6, characterized in that: Introducing fuzzy guidance constraints into the base class feature extractor to suppress false recall of the base class feature extractor on unseen cell types, while guiding the new class feature extractor to enhance its adaptability to new types; Among them, the fuzzy guidance constraint is expressed as: Where, represents the fuzzy guided loss function; is the cardinality of the incremental sample set; and is the weight parameter; is the preset minimum false recall threshold; is the Sigmoid activation function; Representative The predicted logit for each sample.

10. The single-cell incremental annotation method based on distribution and expression-aware revisiting according to claim 6, characterized in that: according to 、 and Compute the summed training loss for the class incremental annotations: in, Where, represents the cross entropy loss function; It is the cross entropy loss function operation; Indicates data splicing operation The formed joint input; Cell type information.

Citation Information

Patent Citations

  • Single-cell multi-omics data integration method and system based on graph contrast learning

    CN118571328A

  • Single cell subtype sample generation method, system, equipment and medium

    CN119252345A

  • Methods and systems for annotating genomic data

    US20240112813A1

  • Genome-wide prediction method based on deep learning by using genome-wide data and bioinformatics features

    US20250104813A1