Bulk-to-single-cell association method based on long-range interaction theory
By employing a method based on long-range interaction theory and utilizing bulk transcriptome and single-cell data, multi-scale feature vectors and gradient direction vectors are extracted, solving the interference problem in single-cell analysis, achieving accurate identification of associated cells, and improving the sparsity and analysis efficiency of single-cell RNA expression data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGZHOU INST OF TECH
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-31
AI Technical Summary
Existing single-cell analysis methods are affected by interference signals, resulting in serious information mis-extraction. They also lack sufficient sample sets and the time-consuming and laborious process of labeling samples, leading to difficulties in computer extraction and a high error rate.
Using a method based on long-range interaction theory, bulk transcriptome and single-cell data are acquired, multi-scale feature vectors are extracted using a Transformer encoder, gradient direction vectors and long-range interaction forces are calculated, an associated cell model is established, and associated cells are determined using a minimization function and energy calculation.
It accurately identifies associated cells, improves the sparsity of single-cell RNA expression data, and enhances the accuracy and efficiency of single-cell analysis, making it widely applicable in the field of biological detection.
Smart Images

Figure CN122493942A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cell association technology, and in particular to a bulk and single-cell association method based on long-range interaction theory. Background Technology
[0002] The need for timely single-cell data analysis in conjunction with bulk transcriptome data makes it crucial to accurately obtain information from bulk transcriptomes, which has significant biological implications and application value for precision cancer treatment. However, single-cell sequencing data contains a large amount of interference, which is mixed with the real signal, making computer extraction difficult and resulting in a high error rate.
[0003] Existing technologies related to this invention include template matching, knowledge-driven methods, and deep learning methods, which mainly utilize features in the data, such as expression intensity features and latent features, to distinguish them from other interferences.
[0004] Traditional technologies commonly suffer from problems such as incorrect extraction of interference signals, lack of sufficient sample sets in deep learning methods, and time-consuming and laborious sample labeling. Existing methods for single-cell analysis using data also suffer from severe interference and incorrect information extraction. Summary of the Invention
[0005] To address the shortcomings of existing methods, this invention solves the problems of severe interference and mis-extraction of information when using data for single-cell analysis.
[0006] The technical solution adopted in this invention is: a method for bulk and single-cell association based on long-range interaction theory, comprising the following steps: Step 1: Obtain bulk transcriptome and single-cell data; In a preferred embodiment of the present invention, the bulk transcriptome includes the TCGA database.
[0007] In a preferred embodiment of the present invention, the single-cell data includes the PanglaoDB database.
[0008] Step 2: Extract multi-scale feature vectors from bulk transcriptomes and single cells; In a preferred embodiment of the present invention, the multi-scale feature vector extractor includes a Transformer encoder.
[0009] Step 3: Calculate the gradient direction vector of the bulk transcriptome and the principal direction of single-cell features using multi-scale feature vectors; Step 4: Calculate the long-range interaction force between the gradient direction vector of the bulk transcriptome and the principal direction of a single-cell feature; change the number of gene expression segments to generate new states, and calculate the energy in the new states. When the energy in the new states meets the conditions, identify the associated cells. In a preferred embodiment of the present invention, step four specifically includes: Step 41: Introduce lossy free energy , ψ 0 represents lossless free energy; Energy change factor; set up , Indicates energy. For stress tensor, ; It is an elastic small strain tensor. It is a symmetric gradient operator. u For the transfer field; The correlation between single cells is equivalent to minimizing a function; In a preferred embodiment of the present invention, the formula for the minimized function is: ; in, For bounded open sets; K For the boundary; u 0 represents the cellular primitive field; and The coefficient is used to adjust the contribution of the corresponding item based on engineering experience.
[0010] Step 42: Gene expression Divide into a finite number of segments: , ,in The coordinates of the segment center; Let these be characteristic quantities; then the gene expression minimization function is obtained. ; Step 43, Calculation Neighborhood relations of elements R ,when R When =1, class , For Duan Ou Neng ; Step 44: Calculate the classification The mean square of gene expression levels Let the distance between non-associated cell points and typical cell points be... W ;calculate energy ;Pick These are candidate associated cells.
[0011] As a preferred embodiment of the present invention, energy The formula is: ; The first item is a certain feature. ej and non-differential expression features e ne Distance sum; the second term is a certain feature. e j Differential expression features e de Distance and sum; The difference coefficient is denoted as .
[0012] As a preferred embodiment of the present invention, a bulk and single-cell association system based on long-range interaction theory includes: a memory for storing instructions executable by a processor; and a processor for executing the instructions to implement a bulk and single-cell association method based on long-range interaction theory.
[0013] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code implements a bulk and single-cell association method based on long-range interaction theory when executed by a processor.
[0014] The beneficial effects of this invention are: 1. This invention utilizes long-range interaction theory to accurately locate associated cells, improving data sparsity caused by low capture in single-cell sequencing and mitigating the problem of false zero expression in single-cell RNA expression data; 2. This invention utilizes single-cell biomarker discovery and bulk verification of its prognostic / diagnostic efficacy; 3. The method of this invention can be used for the identification of associated cells and is widely used in the field of biological detection. Attached Figure Description
[0015] Figure 1 This is a flowchart of the bulk and single-cell association method based on long-range interaction theory of the present invention; Figure 2 This is a schematic diagram of the associated cell extraction results of the present invention. Detailed Implementation
[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments. The drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0017] like Figure 1 As shown, a method for linking bulk and single cells based on long-range interaction theory includes the following steps: Step 1: Obtain bulk transcriptome and single-cell data; Bulk transcriptomes can utilize databases such as TCGA, which focuses on cancer research and contains a vast amount of sequencing data and clinical information on tumors and normal tissues; they can also be accessed through the GDC Data Portal, where expression matrices can be filtered and downloaded by cancer type. Cell data can be obtained from databases such as PanglaoDB, which is a database focused on cell types and contains a large amount of data from mice and humans.
[0018] Transcriptome expression levels correspond to "stress" and are related to the "strain" (state) of each cell; Step 2: Extract multi-scale feature vectors from bulk transcriptomes and single cells; Multi-scale feature vectors are extracted from bulk and single-cell transcriptome data using an encoder. The encoder can be a Transformer encoder or other encoders used for feature extraction; Step 3: Calculate the gradients of the bulk transcriptome and single cells using multi-scale feature vectors; The gradient calculation formula is:
[0019] The direction of the feature gradient is calculated using the gradient formula: θ
[0020] The gradient direction vector of the bulk transcriptome is obtained by utilizing the characteristic gradient direction of the bulk transcriptome. By statistically analyzing the histograms of each vector direction using the gradient directions of single-cell features, the peak value of the histogram is determined as the main direction of the single-cell feature.
[0021] Step 4: Calculate the long-range interaction force between the gradient direction vector of the bulk transcriptome and the principal direction of a single-cell feature; change the number of gene expression segments to generate new states, and calculate the energy in the new states. When the energy in the new states meets the conditions, identify the associated cells. Otherwise, recalculate the long-range interaction forces; Step four specifically includes: The long-range interaction model was established to simulate the force behavior of different states of microscopic cells. Gene expression levels are subject to various disturbances, resulting in various errors. The "stress" (such as the pixel value) at any point is related not only to the "strain" at that point, but also to the state of all points; Step 41: The disturbance reflects an energy jump, introducing lossy free energy. ψ With lossless free energy ψ 0 is represented as: ;in, Energy change factor; use Indicates energy. For the stress tensor, we have: (1) in, ; It is an elastic small strain tensor. It is a symmetric gradient operator. u For the transfer field; That is, the gradient direction vector corresponding to the bulk transcriptome; Therefore, associating a specific single cell is equivalent to minimizing the following function: (2) in, For bounded open sets; K For the boundary; u 0 represents the cellular primitive field; and The coefficient is used to adjust the contribution of the corresponding item based on engineering experience; Step 42: Divide gene expression into a finite number of segments: Each segment is represented as: The number of gene expression segments is The number of; among them, The coordinates of the segment center; For characteristic quantities, such as gene expression, then: (3) Therefore, it is equivalent to minimizing the following function: (4) Step 43, Calculation The neighborhood relationship is considered to have a critical length of , distance greater than The inter-point influence is negligible. The formula for characterizing the influence length of long-range interaction effects, and the neighborhood relationship, is: (5) Only when R When =1, that is, two categories , There is mutual influence, which is a segmental even energy, and the formula is: (6) Step 44: For a certain category To determine which surrounding cells belong to Then those further away certainly do not belong to... The criterion is the standard deviation of gene expression levels in the two cells, and the formula is: (7) in, , Representing cells respectively x and y The i Each feature is expressed. n For the number of cells, l It is the characteristic number. This refers to the differentiation based on single-cell characteristics.
[0022] get Set the largest point as the background (bg); Gene overexpression and complete non-expression in cells necessarily Maximum, let the distance between the non-associated cell point (background, bg) and the typical cell point (obj) be . W That is, differentiated expression ( de ), which is different from indifferent expression ( ne The distance is W ,but: (8) In the formula, if we take The boundary dividing the expression space into differential and non-differential expression regions is the sum of vector distances; the first term represents a certain feature. e j and non-differential expression features e ne Distance sum; the second term is a certain feature. e j Differential expression features e de Distance and sum; The difference coefficient is denoted as .
[0023] Take the smallest The corresponding cells are candidate associated cells, i.e., those that satisfy... .
[0024] In this embodiment, when calculating equation (8), the difference in expression of adjacent features is taken. To further improve computational efficiency, the simulation point process does not use a uniform distribution to extract each candidate. Instead, it takes the value with the higher probability by referring to the difference between the gradient characteristics of the candidate and the local gradient.
[0025] The above steps are used to obtain associated cells, and the bulk transcriptome is linked to single cells.
[0026] like Figure 2The diagram shows the association roadmap between the bulk transcriptome and single cells. Red cells represent associated cells, while other cells represent cells with different differential expressions.
[0027] The method of this invention can accurately locate cells related to a pre-queried single cell in the tissue to be tested, and can be widely used in the field of biological detection.
[0028] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.
Claims
1. A bulk-to-single-cell association method based on long-range interaction theory, characterized in that, Includes the following steps: Step 1: Obtain bulk transcriptome and single-cell data; Step 2: Extract multi-scale feature vectors from bulk transcriptomes and single cells; Step 3: Calculate the gradient direction vector of the bulk transcriptome and the principal direction of single-cell features using multi-scale feature vectors; Step 4: Calculate the long-range interaction force between the gradient direction vector of the bulk transcriptome and the principal direction of a single-cell feature; change the number of gene expression segments to generate new states, and calculate the energy in the new states. When the energy in the new states meets the conditions, identify the associated cells.
2. The bulk and single-cell association method based on long-range interaction theory according to claim 1, characterized in that, Step four specifically includes: Step 41: Introduce lossy free energy , ψ 0 represents lossless free energy; Set as the energy change factor; , Indicates energy. For stress tensor, ; It is an elastic small strain tensor. It is a symmetric gradient operator. u For the transition field; the associated single cell is equivalent to a minimization function; Step 42: Gene expression Divide into a finite number of segments: ; , The coordinates of the segment center; Let these be characteristic quantities; then the gene expression minimization function is obtained. ; Step 43, Calculation Neighborhood relations of elements R ,when R When =1, class , For Duan Ou Neng ; Step 44: Calculate the classification The mean square of gene expression levels Let the distance between non-associated cell points be... W ;calculate energy ;Pick These are candidate associated cells.
3. The bulk and single-cell association method based on long-range interaction theory according to claim 2, characterized in that, energy The formula is: ; The first item is a certain feature. e j and non-differential expression features e ne Distance sum; the second term is a certain feature. e j Differential expression features e de Distance and; The coefficient of variation is denoted as .
4. The bulk and single-cell association method based on long-range interaction theory according to claim 2, characterized in that, The formula for minimizing the function is: ; in, For bounded open sets; K For the boundary; u 0 represents the cellular primitive field; and The coefficient is used to adjust the contribution of the corresponding item based on engineering experience.
5. The bulk and single-cell association method based on long-range interaction theory according to claim 1, characterized in that, Bulk transcriptomes include the TCGA database.
6. The bulk and single-cell association method based on long-range interaction theory according to claim 1, characterized in that, Single-cell data includes the PanglaoDB database.
7. The bulk and single-cell association method based on long-range interaction theory according to claim 1, characterized in that, Multi-scale feature vector extractors include the Transformer encoder.
8. A bulk and single-cell association system based on long-range interaction theory, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing instructions to implement the bulk and single-cell association method based on long-range interaction theory as described in any one of claims 1-7.
9. A computer-readable medium storing computer program code, characterized in that, The computer program code, when executed by a processor, implements the bulk and single-cell association method based on long-range interaction theory as described in any one of claims 1-7.