Difference quantification comparison method for adaptive immune system, and use thereof

By obtaining B cell and T cell sequences and utilizing the IMGT database and distance calculation methods, the differences in the immune system are quantified, solving the problem that existing technologies cannot quantify the adaptive immune system and achieving greater accuracy in early disease detection and treatment assessment.

WO2025232927A1PCT designated stage Publication Date: 2025-11-13NANJING UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/095859
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2025-05-19
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Current technologies cannot effectively quantify differences in the adaptive immune system, nor can they accurately assess and evaluate immunity in the early stages of disease.

Method used

By obtaining B cell BCR and T cell TCR sequences, using the IMGT database for sequence annotation, constructing a 3D distribution map, and calculating Wasserstein distance and Chebyshev distance, the differences in the immune system were quantified, and cluster analysis was performed using visualization dimensionality reduction tools.

Benefits of technology

It enables quantitative comparison and cluster visualization of the immune system status, allowing for early disease detection, evaluation of treatment effectiveness, and application in the auxiliary diagnosis and drug screening of various diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025095859_13112025_PF_FP_ABST
    Figure CN2025095859_13112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a difference quantification comparison method for an adaptive immune system, and the use thereof. The method comprises obtaining at least one of a BCR heavy chain sequence of a B cell, a BCR light chain sequence of the B cell and a TCR β chain sequence of a T cell of a sample to undergo comparison, and sequencing same; comparing the determined sequence with a gene sequence in the IMGT database to obtain sequence annotation information of the corresponding sequence, and further constructing a 3D graph displaying the adaptive immune condition of the sequence; then by means of using a minimum transformation cost method, quantifying an immune state difference between samples to be tested and between a sample to be tested and the database that has undergone the test. The method for quantifying immune differences can be used for evaluating the immune state of biological samples, and can also be used for evaluating influences of various therapy and intervention methods on the immune system, so as to judge the effects of the therapies and interventions, including but not limited to the drug efficacy, the vaccine efficacy, etc.
Need to check novelty before this filing date? Find Prior Art

Description

A comparative method for differential quantification of adaptive immune system and its application Technical Field

[0001] This invention belongs to the field of immune evaluation technology, specifically relating to a method for quantitative comparison of differences in the adaptive immune system and its application. Background Technology

[0002] With the development of sequencing technology, read length, read depth, and phoned quality score have gradually improved, making large-scale sequencing studies of B lymphocyte and T lymphocyte diversity more reliable. Proteins on the surface of B cells that bind foreign antigens are called B cell receptors (BCRs), and proteins on the surface of T cells that bind foreign antigens are called T cell receptors (TCRs). During the differentiation of B and T lymphocytes from hematopoietic stem cells, VDJ and VJ gene rearrangements occur. After activation, B cells also produce somatic hyper-Mutation (SHM), resulting in approximately 10^11 functional BCRs and 10^7 functional TCRs.

[0003] Adaptive immunity refers to the acquired immunity that develops after birth in response to antigen stimulation. Because adaptive immunity can recognize differences between different antigens, it exhibits different characteristics in different diseases; it is also called specific immunity. The immune repertoire refers to the collection of all functional BCRs and TCRs in the body at a specific time. Identical clusters at the cellular, gene, and complementarity-determining regions (CDR) levels are called clones, and their clonal diversity is called repertoire diversity. Immune diversity is not static and changes and shifts according to an individual's immune status: under abnormal states of antigen stimulation such as infection, tumors, and vaccination, some B and T cells proliferate, altering immune repertoire diversity; as an individual ages, the hematopoietic function of hematopoietic stem cells declines, reducing the types of B and T cells, thus altering the diversity of immune repertoire types; after the effects of chemotherapy, radiotherapy, and other drugs, immune cell regulation and death also alter immune repertoire diversity. By detecting the diversity of the immune repertoire, we can understand the state of the immune system, which has wide application value in detecting the occurrence and development of diseases, evaluating the effectiveness of vaccines and other drugs, and assessing immunity.

[0004] Currently, methods for analyzing immune diversity generally include the following: a. D50 method, which compares the frequency positions corresponding to the 50% of the cumulative frequency curve to determine diversity; b. Abundance method, which estimates diversity by detecting the abundance of each clone and the total number of clones; c. Frequency method, which calculates diversity by statistically analyzing the relative frequencies of each clone; d. Entropy calculation method, which estimates sample diversity by calculating the Shannon Index, Simpson Index, Hill Diversity, etc.; e. Experimental detection method, which detects clonal diversity through experimental methods such as qPCR; f. Machine learning method, which uses methods such as constructing neural networks to cluster clonal data and compare differences in immune diversity.

[0005] Traditional methods can only be used for pairwise comparisons of differences and cannot be quantified. Early disease diagnosis relies on machine learning, which depends on the training dataset and cannot provide specific quantitative indicators. Summary of the Invention

[0006] Purpose of the invention: The technical problem to be solved by the present invention is to provide a method for quantitative comparison of differences in the adaptive immune system and its application, which addresses the shortcomings of the prior art.

[0007] To address the aforementioned technical problems, this invention discloses a method for quantitatively comparing differences in the adaptive immune system, comprising the following steps:

[0008] (1) Obtain a certain amount of biological samples to compare immune diversity differences, obtain any one of the following sequences: B cell BCR heavy chain sequence, BCR light chain sequence, or T cell TCR β chain sequence, and sequence them. The sequenced sequences are compared with gene sequences in the IMGT database to obtain corresponding sequence annotation information, which includes sequence mutation information. The purpose of sequence annotation is to correlate base sequences with the functional regions of translated proteins, and to compare with unmutated sequences to obtain the location and bases of mutations. Protein spatial structure and functional regions do not vary depending on the sequencing instrument. The IMGT database is the most commonly used and authoritative database for labeling the spatial structure of BCR and TCR proteins. Common sequence annotation information includes, but is not limited to: barcodes (UMIs) or read counts or quantities, names of aligned variable gene fragments, names of aligned diversity or joining gene fragments, and complete nucleotide sequences.

[0009] (2) Each difference in base pair sites obtained by comparing with the sequence in the database is recorded as a mutation. Sequences with the same BCR annotation information and mutation or sequences with the same TCR annotation information are regarded as the same clone. The percentage of the number of sequences contained in each clone to the total number of sequencing sequences is recorded as the percentage of sequences contained in the clone. For the sequencing sequence in each biological sample, at the clonal level, the number of mutations in the clone is used as the x-axis, the type of clone with the number of mutations is used as the y-axis, and the percentage of sequences contained in the clone is used as the Z-axis. Each sequence corresponds to a point in 3D space. All sequences in each biological sample form a 3D distribution map in the 3D space, which represents the immune status of the sample.

[0010] (3) Quantify the difference in immune diversity between biological sample A and biological sample B into the minimum transformation cost required to transform the 3D distribution map formed by biological sample A to the 3D distribution map formed by biological sample B.

[0011] In one embodiment, the minimum transformation cost is obtained by: sorting the dimensionality-reduced projection of the 3D distribution map of the corresponding sample onto the YZ-axis plane to obtain the frequency distribution function curve of the "clone-sequence number percentage" of the corresponding sample; and then calculating the minimum transformation cost required to transform the frequency distribution function curve P1 of biological sample A to the frequency distribution function curve P2 of biological sample B using the Wasserstein distance.

[0012] Preferably, when using Wasserstein distance to calculate the minimum transformation cost required to transform the frequency distribution function curve P1 of biological sample A to the frequency distribution function curve P2 of biological sample B, the weight of each frequency distribution function curve on the X-axis is K+N. M 2 , where N M K represents the number of mutations in the clone, and K is the background weight. To apply this model to early-stage, mutation-free BCR and TCRβ (TCRβ does not mutate, only B cell BCR does), we add a background weight to avoid a value of 0. The biological meaning of the background weight is the difficulty required to maintain normal cellular activities other than the immune process; this value should be less than the difficulty of the immune process, and ranges from 0 to 1. When the value is 0.1, the immune system generates cells with a mutation number N. M The difficulty is proportional to 0.1 + N M 2 .

[0013] Organisms possess both cellular immunity and humoral immunity. Humoral immunity includes various types of antibodies, and B cell BCRs are responsible for secreting antibodies to clear pathogens free outside the cell. For pathogens that have already invaded the cell, antibodies cannot clear them; in this case, T cell TCRs are responsible for recognizing and clearing them. Following the calculation method for differences between antibody libraries of the same type between two individuals, as described above, a corresponding distance can be calculated for each antibody type. Antibody types mainly refer to cellular immunity and humoral immunity, including antibody subtypes responsible for mucosal secretion, antibody subtypes responsible for allergies, etc., all of which are B cell BCR heavy and light chains or TCR sequences. For the overall difference between two organisms, it is necessary to integrate the distances between various B cell antibody types and T cell distances. We use the Chebyshev distance, taking the maximum value among multiple distances as the distance between the two individuals. For example, if two people have mucosal antibody data, we can calculate the distance A using the method above. For the same two people, we can calculate the distance B using the same method for blood antibodies. Similarly, we can calculate the distance C for T cells and the distances D, E, and F for other types. Therefore, for these two people, considering mucosal immunity, blood immunity, and other types, the distance is the maximum value among AF.

[0014] For example, the Wasserstein distance can be calculated using the Python function scipy. Specifically, it can be calculated using stats.wasserstein_distance.

[0015] For example, based on the annotation information, two "clone-sequence number percentage" curves are obtained. The first curve is arranged from high to low according to clone size. The clone size A is 0.3, 0.2, 0.2, 0.1, 0.1, 0.1, and the number of mutations A1 of the clones is 2, 2, 1, 1, 0, 0. The background weight is 0.1, and the weight C calculated using mutations is 4.1, 4.1, 1.1, 1.1, 0.1, 0.1.

[0016] The second curve is arranged from highest to lowest clone size. The clone size B is 0.5, 0.2, 0.1, 0.1, 0.05, 0.05, and the number of mutations B1 of the clones is 5, 3, 2, 1, 0, 0. The weights D calculated using mutations are 25.1, 9.1, 4.1, 1.1, 0.1, 0.1.

[0017] The Wasserstein distance can be calculated using the scipy package in Python. The specific code is as follows:

[0018] Scipy.stats.wasserstein_distance([A],[B],[C],[D])

[0019] ABCD represents the list of numbers above, with commas separating the numbers within ABCD, and the entire list enclosed in square brackets. That is: `Scipy.stats.wasserstein_distance([0.3, 0.2, 0.2, 0.1, 0.1, 0.1],[0.5, 0.2, 0.1, 0.1, 0.05, 0.05],[4.1, 4.1, 1.1, 1.1, 0.1, 0.1],[25.1, 9.1, 4.1, 1.1, 0.1, 0.1])`. The software output is the Wasserstein distance between the two curves.

[0020] Preferably, in step (1), the number of sequence alignments and annotations of the biological sample to be tested is not less than 10,000. According to the results of batch experiments, if the number of alignments and annotations is greater than 10,000, the correct clustering results can be obtained. The more data, the more accurate the results. The number of sequence alignments and annotations depends on the strength of the disease's stimulation of the immune system. The stronger the stimulation, the fewer sequence alignments are needed to reflect the differences. The number of biological samples depends on experimental efficiency and consumption. If there is no loss, a biological sample of 10,000 cells is sufficient to describe the immune status of cancer samples, healthy samples, adults, and children.

[0021] In practice, for acute illnesses, including those caused by sudden onset where the immune system is still weak, larger samples can be prioritized for assessment. For common illnesses, however, due to cost and procedural convenience, the biological information contained in a 5mL peripheral blood sample is usually sufficient for analysis. Furthermore, a suitable sample size threshold can be determined by comparing the results of large and small sample sizes.

[0022] The method described in this application has been validated. For non-sudden acute illnesses, the results from 10,000 sequences typically tend to be consistent with those from a large sample; the larger the sample size, the smaller the error. This method can use the results from a small sample to represent the results from a large sample, i.e., the macroscopic immune status, and holds promise for "cancer detection with a drop of blood."

[0023] Based on the aforementioned immune distance quantization, this application further proposes an immune status clustering visualization method. By utilizing the aforementioned immune distance quantization method, the N immune statuses of N samples to be tested are visualized. There are N*N minimum transformation costs among the N samples, corresponding to N*N distance matrices. Then, a visualization dimensionality reduction tool is used to convert the matrix into an image, and the immune statuses to be tested are clustered and classified based on the distance between points.

[0024] The visualization dimensionality reduction tool can be any one of tSNE, PCA, or Umap.

[0025] After calculating the minimum transformation cost between similar data, Chebyshev distance is used to calculate the differences between individuals, a distance matrix is ​​constructed, and tsne is used to plot the graph.

[0026] More specifically, the Silhouette Score (SS) or Davies-Bouldin Index (DBI) is used to evaluate clustering performance. SS values ​​range from -1 to 1; values ​​closer to -1 indicate more misclassification, values ​​closer to 0 indicate more uniform clustering with less clear classification, and values ​​closer to 1 indicate better classification. DBI values ​​range from 0 to positive infinity; values ​​closer to 0 indicate more accurate classification.

[0027] This invention also proposes the application of the above-mentioned method for quantifying differences in immune diversity in the evaluation of immune characteristics of pan-diseases, including the following steps:

[0028] (1) Establish a library of frequency distribution function curves of immune status in pan-diseases, where the pan-diseases are single primary diseases;

[0029] (2) By comparing the minimum conversion cost of the test sample with the frequency distribution function of the immune status in the frequency distribution function curve library, the smaller the cost, the greater the similarity between the immune status of the test sample and the corresponding disease.

[0030] The term "pan-disease" here means no specific disease type and applies to all disease types, such as common cancers, including the top ten cancers in the "China Cancer Report," including but not limited to lung cancer, liver cancer, stomach cancer, colorectal cancer, esophageal cancer, pancreatic cancer, breast cancer, brain tumors, leukemia, and lymphoma, as well as multiple myeloma, febrile infections, and Kawasaki disease. Evaluation of the immune characteristics of samples can be used to assess the differentiation of disease type, disease course, pathogen type, immune development process, and febrile infections. The term "single primary disease" means a disease that occurs independently without any obvious other diseases acting as inducing or promoting factors.

[0031] This invention also proposes the application of the above-mentioned quantitative method for immune diversity differences in evaluating the effectiveness of pan-disease treatment methods and interventional methods.

[0032] Specifically, the effectiveness of treatment or intervention is evaluated by comparing the immune status of the test subjects after treatment or intervention with that of corresponding healthy samples; the smaller the difference, the better the effect.

[0033] This invention also proposes the application of the above method in pan-disease drug screening. The treatment or intervention effect is evaluated by comparing the differences in immune status before and after medication with corresponding healthy samples; the smaller the difference, the better the effect.

[0034] Beneficial Effects: The quantitative immune differential method proposed in this application can effectively assess the impact of various treatments and interventions on the immune system, and thereby determine the effectiveness of these treatments and interventions, including but not limited to drug efficacy and vaccine efficacy. The model is not a traditional empirical probability model comparing pairwise differences, but rather a quantitative model capable of simultaneously distinguishing and determining multiple health conditions. Each health condition has unique characteristics, making it a versatile auxiliary clinical diagnostic tool for various diseases, including cancer and fever. Attached Figure Description

[0035] Figure 1 shows a 3D space constructed using the number of mutations in a clone as the x-axis, the clone size with that number of mutations as the y-axis, and the percentage of sequences contained within that clone (proliferation) as the Z-axis. Each sequencing sequence occupies an independent point in the coordinate space based on its annotation information and number of mutations.

[0036] Figure 2 shows the frequency distribution function curve of "clone-sequence number percentage" after the dimensionality reduction projection of the 3D image onto the YZ axis plane is sorted.

[0037] Figure 3 is a schematic diagram showing that a clone with a fast amplification rate is equivalent to a clone that is more likely to be assigned to each new cell produced.

[0038] Figure 4 shows that the B cell distribution generated by the preference attachment model has scale-free network characteristics;

[0039] Figure 5 shows that the T cell distribution generated by the preference dependency model has scale-free network characteristics;

[0040] Figure 6 shows that the measured B cell distribution is consistent with the fitted data.

[0041] Figure 7 shows that the measured T cell distribution is consistent with the fitted data.

[0042] Figure 8 shows that as the immune response progresses and the total number of cells increases, the correlation between the number of mutations in a clone and the number of clone species exhibits an inverted bell-shaped single-peak distribution, with the peak gradually shifting to the right.

[0043] Figure 9 shows the distribution of B cell clonal mutation number and clonal frequency 3 days after mouse immunization with VLP;

[0044] Figure 10 shows the distribution of B cell clonal mutation number and clonal frequency in mice 28 days after immunization with VLP.

[0045] Figure 11 shows the distribution of cell clonal mutation number and clonal frequency in healthy child B;

[0046] Figure 12 shows the distribution of B cell clonal mutation number and clonal frequency in healthy adults.

[0047] Figure 13 shows the distribution of B-cell clonal mutation number and clonal frequency in myeloma patients;

[0048] Figure 14 shows the trend of the almost linear increase in the average number of mutations in clones as the total number of cells increases during the immunization process.

[0049] Figure 15 shows the calculation of the distance between any two samples. There are N*N distances between N samples, which can be used to construct an N*N distance matrix. Then, visualization software can be used to turn the matrix into an image.

[0050] Figure 16 shows the clustering results of the overall distance between individuals calculated after infection with fungal cell wall (CAWS) and phage coat protein (VLP). The results show that there are significant differences in the activation patterns of the immune system caused by different sources of infection.

[0051] Figure 17 shows the clustering results after calculating the overall distance between individuals in mice immunized with viral capsid proteins at different time points; the immune system exhibits different characteristics at different immunization time points, and the model is accurate enough to distinguish them.

[0052] Figure 18 shows the results of the blind clustering test. The natural grouping of myeloma (MM) patients and healthy controls is consistent with the clinical grouping. Since the results do not use the unique characteristics of myeloma cancer, they can be used as a pan-cancer early screening method.

[0053] Figure 19 shows the clustering results of infectious fever and vasculitis fever;

[0054] Figure 20 shows the results of simultaneous differentiation and judgment for multiple health conditions;

[0055] Figure 21 shows the clustering results for the healthy children and healthy adults groups;

[0056] Figure 22 shows the clustering results for healthy individuals, lung cancer, esophageal cancer, myeloma, and colorectal cancer;

[0057] Figure 23 shows the effect of sample size on the stability of sample results when using TRB sequences;

[0058] Figure 24 shows the effect of sample size on the stability of sample results when using IgM sequences;

[0059] Figure 25 shows the effect of sample size on the stability of sample results when using IgG sequences.

[0060] Figure 26 shows the results of differentiating the immune status of healthy children, healthy adults, infected individuals, Kawasaki disease patients, and multicellular myeloma patients after fixing the sample size of all samples to 10,000 random sequences. Detailed Implementation

[0061] In the very early stages of disease, before any obvious symptoms appear, the immune system is already activated to fight the abnormality. Even after symptoms subside and the individual recovers, the activation of the immune system continues for a period due to immune memory. The immune system has an immune surveillance function, responsible for eliminating abnormal cells; in the very early stages of cancer, this immune surveillance function is bound to be abnormal.

[0062] Therefore, modeling changes in the normal immune system, rather than detecting the disease itself, can enable the early detection of abnormalities in a wide range of diseases, and make disease screening, recurrence diagnosis, assessment of vaccine effectiveness, and optimization of treatment plans more sensitive and accurate.

[0063] This application proposes a method for quantifying the differences in different immune states.

[0064] Specifically, the detection data required by this method can be obtained through known experimental and sequencing methods. After cleaning, deduplication and error correction, and comparison with gene sequences in the IMGT database, the obtained B cell BCR heavy chain, or BCR light chain gene, or T cell TCRβ chain gene sequence containing sequence annotation and mutation information is obtained.

[0065] The input BCR and TCR raw sequences are quality controlled, requiring sufficient sequence length, at least 250 bp, and containing the complete CDR3 gene fragment and partial V and J gene fragments on both sides.

[0066] BCR differential base pairs obtained by sequence alignment with the database are identified, with each difference corresponding to a mutation. B cells with the same BCR annotation information and mutation, or T cells with the same TCR annotation information, are considered as the same clone. The percentage of sequences contained within all clones out of the total number of sequences sequenced is calculated to obtain a frequency distribution function curve of "clone-sequence number percentage". According to existing research, this curve needs to have the power-law distribution property of a scale-free network; otherwise, there will be an error between the measured data and the actual immune status of the individual, and the calculation results can only represent the sample, not the macroscopic individual, resulting in inaccurate differential calculation results.

[0067] After verifying that the data met the standards, we began to quantify the differences between different immune states.

[0068] Specifically, each base pair difference obtained by comparing with sequences in the database is recorded as a mutation. B cells with the same BCR annotation information and mutation, or T cells with the same TCR annotation information, are considered as the same clone. The percentage of sequences contained within a clone relative to the total number of sequences sequenced is recorded as the percentage of sequences contained within that clone. For each sequence in each biological sample, at the clonal level, the number of mutations (Mutations) of that clone is used. The x-axis is the number of mutations, the y-axis is the clone species with that mutation number, and the z-axis is the percentage of sequences contained within the clone. Each sequence corresponds to a point in 3D space. All sequences in each biological sample form a 3D distribution map in this 3D space. The dimensionality reduction projection of this 3D distribution map onto the YZ-axis plane, after sorting, becomes the frequency distribution function curve of "clone-sequence number percentage". The frequency distribution function curve of "clone-sequence number percentage" represents the immune status of the corresponding sample, as shown in Figure (1). All sequences form a 3D distribution map in this 3D space. The dimensionality reduction projection of this 3D distribution map onto the YZ-axis plane, after sorting, becomes the frequency distribution function curve of "clone-sequence number percentage", as shown in Figure 2.

[0069] The amount of reaction process required for the immune system to generate a high-mutation clone is different from that required to generate a low-mutation clone. Therefore, in order to numerically calculate the state of the immune system, it is necessary to quantify the relationship between the number of mutations and the reaction process.

[0070] One sequencing sequence corresponds to one RNA sequence, which, at the extraction rate of conventional RNA extraction, is approximately equivalent to one cell. Therefore, we simulate the cellular-level immune response. At the onset of the immune response, each clone contains only one cell. Each clone expands at different rates based on its reactivity to the disease, and there is a certain probability that mutations will occur and new clones will be generated. We use a complex community network computational method to simulate this process. Clones with faster expansion rates are equivalent to being more likely to be assigned to each new cell generated (as shown in Figure 3). The number of times this new cell attachment process is repeated is equivalent to the total number of cell expansions. This cyclical repetition of the new cell attachment process simulates the expansion of immune cells and the continuous immune process.

[0071] This preference-dependent model can also produce scale-free network power-law characteristics (Figures 4-7) that are identical to those reported in the literature (e.g., Das, J., & Jayaprakash, C. (Eds.). (2019). Systems Immunology: An Introduction to Modeling Methods for Scientists (1st ed.). CRC Press. https: / / doi.org / 10.1201 / 9781315119847). Figures 4 and 5 show the immune system simulated in a computer based on the probabilistic model in Figure 3, simulating individual immune cells with mutations. Some identical cells are grouped into one clone, and the results are statistically analyzed. Figures 6 and 7 show the results of BCR heavy chain and TCRβ chain measurements on peripheral blood B and T cells from healthy adults. The horizontal axis represents the number of antibody mutations, and the vertical axis represents the percentage of clones with that mutation. One curve represents the result of a single simulated immune response, either from a blood sample or a spleen sample. The results showed that the distribution and variation of the number of clone species under different mutation numbers were consistent with the actual measurements (Figure 8-13). Therefore, we believe that the model can represent the immune response process.

[0072] In this model, as the immune process progresses, the total number of cells increases, and the average number of mutations in clones increases almost linearly (Figure 14). A mechanical model is introduced here, treating the immune process as a stimulating force applied to the immune system, and the average number of mutations as a deformation of the immune system. Thus, the immune system generates cells with a mutation number N. M The difficulty of the BCR sequence is proportional to N. M 2This can be analogized to a spring; as the tension and deformation of a spring increase linearly, the elastic potential energy is proportional to the square of the deformation. To apply this model to early-stage, unmutated BCR and TCRβ (TCRβ does not mutate, only B-cell BCR does), we add a background weight of 0.1 to avoid generating a value of 0, i.e., the immune system generates cells with the mutation number N. M The difficulty is proportional to 0.1 + N M 2 In practical applications, the background weight value ranges from 0 to 1 and can be adjusted according to the specific detection results.

[0073] We define the difference between two immune systems, A and B, as the minimum transformation cost required to transition from state A to state B. As in the embodiments below, Wasserstein distance is used to calculate the difference between two 3D distribution graphs. We reduce the dimensionality to a 2D "clone-sequence number percentage" frequency distribution function curve and then use Wasserstein distance to calculate the difference between the two curves. The Wasserstein distance can be calculated using the scipy package in Python. In the embodiments below, the weight formula for each clonal transformation is 0.1 + N. M 2 N M The number of mutations in the corresponding clone.

[0074] Using this distance calculation method, the distance between any two samples can be calculated. There are N*N distances between N samples, which can be used to construct an N*N distance matrix. This matrix can then be visualized using visualization software (Figure 15). tSNE is a commonly used dimensionality reduction clustering visualization method in biology. In the final result image, the x and y coordinates of points are meaningless; only the distances between points are meaningful. The distance data comes from the distance matrix. Since this method is not based on traditional empirical rules derived from data summaries, and the calculation process does not use clinical information, it is essentially a blind test. Only in the final clustering result image are color-coded clinical groups to verify the correctness of the results.

[0075] There are various types of antibodies in organisms. Following the calculation method described above for the differences between antibody libraries of the same type between two individuals, a corresponding distance can be calculated for each antibody type. For the overall difference between two organisms, it is necessary to integrate the distances between various B cell antibody types and T cell distances. We use the Chebyshev distance, taking the maximum value among multiple distances as the distance between the two individuals.

[0076] The quantitative immune difference method proposed in this application can be used to effectively assess the impact of various treatments and interventions on the immune system, and to determine the effectiveness of these treatments and interventions, including but not limited to drug efficacy and vaccine efficacy.

[0077] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0078] Example 1: Blind clustering of immune characteristics of mice immunized with different immune sources.

[0079] Immunization experiments were conducted on C57BL / 6 inbred mice raised in an SPF environment. Spleens from healthy mice, spleens immunized with fungal cell wall (CAWS) via intraperitoneal injection, and spleens immunized with bacteriophage virus capsid (VLP) via intraperitoneal injection were collected. Grouping information is shown in Table 1. The BCR heavy chain and TCRβ sequences of B and T cells in these tissues were measured to verify scale-free network properties. After meeting quality control standards, blind classification was performed using the above method. The final clustering results were color-coded according to the grouping information, as shown in Table 1. Specifically, spleen samples from C57BL6 mice immunized with fungal cell wall (CAWS) and viral capsid (VLP) were measured, identifying antibody subtypes as IgG and IgM, and gene types as BCR heavy chain and TCR beta chain. Due to individual sample differences, the amount of data used ranged from 10,000 to 140,000 sequences after quality control standards were met. After calculating the minimum transition cost between similar data (i.e., calculating the distance based on the results of BCR heavy chain IgG sequence, IgM sequence, and TCR beta sequence), Chebyshev distance was used to calculate the differences between individual mice, a distance matrix was constructed, and tsne was used to plot the distance matrix.

[0080] The results are shown in Figure 16. Each point in the figure represents a sample (blood or spleen). The horizontal and vertical axes are meaningless, and the distance between points represents their similarity, obtained from the visualization of the distance matrix above. The classification results are consistent with the experimental groupings. Different infectious agents induce different immune system activation patterns, and the model's accuracy can distinguish these differences. The SS value is 0.54, and the DBI value is 1.01. VLP is the empty vector of a viral particle vaccine. This method can observe the differences in immune activation induced by different vaccine types.

[0081] Example 2: Clustering of immune characteristics of mice with the same immune source but different immune stages.

[0082] Spleens of mice immunized with phage capsid protein (VLP) for 3, 7, 14, and 28 days were collected. The BCR heavy chain and TCRβ sequences of their B and T cells were measured to verify that they have scale-free network properties. After meeting the quality control standards, blind classification was performed using the above method. In the final clustering results, the classification results icons were colored according to the grouping information.

[0083] Table 1

[0084] The results are shown in Figure 17, and the classification results are consistent with the experimental groupings. The immune system exhibits different characteristics at different immunization time points, and the model's accuracy is sufficient to distinguish them, with an SS value of 0.20 and a DBI value of 5.4.

[0085] Example 3 differentiates the immune characteristics of people in different immune states.

[0086] Peripheral blood samples were collected from healthy adults, healthy children, patients with infectious fever, patients with immune vasculitis fever (Kawasaki disease, KD), and patients with myeloma (MM). Sample information is shown in Table 2.

[0087] Table 2

[0088] Human samples were uniformly peripheral blood samples, except for myeloma patients, whose samples were taken from the bone marrow lesions. Antibody subtypes were measured as IgG, IgM, and IgA, and gene types as BCR heavy chain and TCR beta chain. The data volume was 50,000-100,000 sequences after quality control. After calculating the minimum conversion cost between similar data, Chebyshev distance was used to calculate inter-individual differences, a distance matrix was constructed, and TSNE was used for plotting.

[0089] The BCR heavy chain and TCRβ sequences of B and T cells were measured, and blind classification was performed using the above method. Furthermore, by integrating the distances between various antibody types of B cells and T cells, we used the Chebyshev distance and took the maximum value of the multiple distances as the distance between two individuals.

[0090] In the final clustering results, the classification results icons are colored according to the grouping information. The classification results are consistent with the experimental groups. Cancer immune status differs from healthy status. Since the results do not use the unique features of myeloma, they can be used as a general cancer early screening method (Figure 18, SS: 0.35, DBI: 0.88). Because early and mild Kawasaki disease symptoms only include fever, it is easily misdiagnosed. This model can distinguish between infectious fever and vasculitis fever, which can help clinical diagnosis of etiology and reduce misdiagnosis. For example, after quantifying the distance between samples, the average distance between the test sample and each known point in the infectious and vasculitis groups determines which category it belongs to (Figure 19, SS: 0.2, DBI: 0.54). This model is not a traditional model that compares differences pairwise, but can simultaneously distinguish and determine multiple health conditions. Each health condition has unique characteristics (Figure 20, SS: 0.06, DBI: 1.84), making it a general auxiliary clinical diagnostic tool for multiple diseases, including cancer and fever.

[0091] Example 4: Study on immune development characteristics.

[0092] For the clustering of healthy children and healthy adults, the blinded clustering also showed clear clustering (Figure 21, SS: 0.3, DBI: 1.05). Therefore, this method can be used to compare the differences in immune normality over time at different age stages.

[0093] Example 5

[0094] Classify healthy individuals, lung cancer (LC), esophageal cancer (EC), myeloma (MM), and colorectal cancer (CRC).

[0095] Data were collected from 6 healthy individuals, 6 patients with lung cancer (LC), 6 patients with esophageal cancer (EC), 6 patients with myeloma (MM), and 18 patients with colorectal cancer (CRC). 5 ml of peripheral blood was collected from each group, and IgM, IgG, IgA, and TRB gene RNA from B and T immune cells in the peripheral blood was obtained. RNA was extracted, reverse transcribed, amplified by PCR, library constructed, and sequenced using high-throughput sequencing. Clustering was performed with unknown data labels, and then labels were restored to calculate the clustering effect. The clustering results are shown in Figure 22. The silhouette coefficient (SS) was 0.21, and the DBI was 0.92, indicating good clustering performance. This method can still distinguish between healthy individuals and mixed samples of lung cancer, esophageal cancer, myeloma, and colorectal cancer in a blind test, demonstrating its applicability to these cancers.

[0096] Example 6: The effect of sample size on the stability of results.

[0097] This application also explores the impact of the sample size on the stability of the results.

[0098] Figures 23-25 ​​show that as the sample size increases, the error between large and small samples decreases, and the smaller sample becomes more representative of the large sample data. Figures 23-25 ​​show the results of comparisons using TRB, IgM, and IgG sequences, respectively. Each curve in the figures represents a sample, and both the horizontal and vertical axes are logarithmic. The horizontal axis represents the sample size, and the vertical axis represents the immune distance (i.e., error) between the large and small samples. For example, when the horizontal axis is 25000, the corresponding point signifies: 25000 sequences are randomly selected from a large sample (100,000 sequences), and these 25000 sequences are considered a new small sample. The minimum conversion cost method of this application is used to calculate the immune difference between this small sample and the large sample. As can be seen from the figure, as the sample size increases, the error decreases, and the sample becomes closer to the large sample. It is reasonable to infer that when the data volume exceeds 100,000, the fluctuation will continue to decrease along the curve trend. Within a certain reasonable threshold range, the data of a small sample can be expanded into the data results of a large sample. For example, the result of a 1ml blood sample can represent 5L of blood, which can further reflect the macroscopic immune status. This provides theoretical proof for the blood drop cancer detection, and it is expected that macroscopic immune results can be obtained with a small sample size in the future.

[0099] To further illustrate how to find the minimum sample size, this application uses the immune results of myeloma patients (MM), healthy adults (HC), and Kawasaki disease patients (KD) as examples. As shown in Figures 23-25, the horizontal dashed lines represent the average immune distances between the HC and MM groups, and between the KD and MM groups, respectively. When the data volume exceeds the intersection of the horizontal dashed line and the diagonal line, the error is less than the average distance, and clustering begins to appear. Therefore, for the points between the two horizontal dashed lines, with the corresponding sample size on the x-axis, HC and MM should be distinguishable, but KD and MM cannot be distinguished. Theoretically, when the data volume exceeds the intersection of the horizontal dashed line and the diagonal line, distinction can be achieved. Therefore, all data samples were reduced to 10,000 for clustering, resulting in Figure 26. Figure 26 shows the differentiation of immune status in healthy children, healthy adults, infected individuals, Kawasaki disease patients, and multicellular myeloma patients using 10,000 sequences. The results show that clustering with 10,000 sequences can distinguish HC and MM, but cannot distinguish KD and MM. Since Kawasaki disease (KD) is an acute, self-limiting vasculitis, a larger sample size is preferable for acute cases to minimize error. As shown in Example 3, the results were based on over 50,000 sequences, all of which demonstrated good differentiation. This data further validates the conclusion that "as the sample size increases, the error decreases, and the sample becomes increasingly representative of large samples." The detection curve in the figure is nearly linear; for every tenfold increase in sequencing data volume, the accuracy improves by one decimal place.

[0100] In summary, the method of this invention can infer the overall macroscopic immune status of the human body from a very small amount of data. This method can quantify and differentiate between various diseases, effectively distinguishing between tumors, fever, autoimmune diseases, and other conditions, enabling early screening for diseases at a very early stage, thus achieving early detection and early treatment. It has significant value for both basic research and clinical applications.

[0101] This invention provides a quantitative approach and method for assessing immune diversity differences. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.

Claims

1. A method for quantitatively comparing differences in the adaptive immune system, characterized in that, Includes the following steps: (1) Take a certain amount of biological samples with differences in immune diversity to be compared, obtain at least one of the B cell BCR heavy chain sequence, BCR light chain sequence or T cell TCR β chain sequence, amplify and construct a library and sequence, and obtain the sequence annotation information of the corresponding sequence by comparing the sequence with the gene sequence in the IMGT database. The sequence annotation information includes sequence mutation information. (2) Each difference in base pair sites obtained by comparing with the sequence in the database is recorded as a mutation. Sequences with the same BCR annotation information and mutation or sequences with the same TCR annotation information are regarded as the same clone. The percentage of the number of sequences contained in each clone to the total number of sequences is recorded as the percentage of sequences contained in the clone. For the sequencing sequence in each biological sample, at the clonal level, the number of mutations in the clone is used as the x-axis, the clone type with the number of mutations is used as the y-axis, and the percentage of sequences contained in the clone is used as the Z-axis. Each sequence corresponds to a point in 3D space, and all sequences in each biological sample form a 3D distribution map in this 3D space. (3) Quantify the difference in immune diversity between biological sample A and biological sample B into the minimum transformation cost required to transform the 3D distribution map formed by biological sample A to the 3D distribution map formed by biological sample B.

2. The quantitative comparison method according to claim 1, characterized in that, The minimum transformation cost is obtained by the following method: after sorting the dimensionality-reduced projection of the 3D distribution map of the corresponding sample onto the YZ axis plane, the frequency distribution function curve of the "clone-sequence number percentage" of the corresponding sample is obtained. Then, the minimum transformation cost required to transform the frequency distribution function curve P1 of biological sample A to the frequency distribution function curve P2 of biological sample B is calculated by using Wasserstein distance.

3. The quantitative comparison method according to claim 2, characterized in that, The weight of each frequency distribution function curve on the X-axis is K+N. M 2 , where N M This represents the number of mutations in the clone, and K represents the background weight, with a value ranging from 0 to 1.

4. The quantitative comparison method according to claim 1, characterized in that, In step (1), the number of sequences after alignment and annotation of the biological sample to be tested is no less than 10,000.

5. The quantitative comparison method according to claim 1, characterized in that, The minimum cost is the maximum value among the minimum transformation costs obtained based on at least one of the B cell BCR heavy chain sequence, the BCR light chain sequence, or the T cell TCR β chain sequence.

6. A method for visualizing immune status clustering results, characterized in that, The quantitative comparison method described in any one of claims 1-5 is used to visualize the N immune states of N samples to be tested. There are N*N minimum transformation costs among the N samples. An N*N matrix is ​​constructed, and then a visualization dimensionality reduction tool is used to turn the matrix into an image. The immune states to be tested are clustered and classified according to the distance between the points.

7. The method according to claim 6, characterized in that, The visualization dimensionality reduction tool can be any one of tSNE, PCA, or Umap.

8. The application of the quantitative comparison method according to any one of claims 1-5 in the evaluation of immune characteristics of pan-diseases, characterized in that, Includes the following steps: (1) Establish a library of frequency distribution function curves of immune status in pan-diseases, where the pan-diseases are single primary diseases; (2) By comparing the minimum conversion cost of the test sample with the frequency distribution function of the immune status in the frequency distribution function curve library, the smaller the cost, the greater the similarity between the immune status of the test sample and the corresponding disease.

9. The application of the method according to any one of claims 1-5 in evaluating the effectiveness of pan-disease treatment methods, drug treatment effects, and interventional methods, wherein the treatment or interventional effect is evaluated by comparing the immune status of the test subject after treatment or intervention with the difference before treatment or corresponding healthy samples, wherein... The greater the difference between before and after treatment, or the smaller the difference between after treatment and a healthy sample, the better the effect.

10. The application of the method according to any one of claims 1-6 in pan-disease drug screening.

Citation Information

Patent Citations

  • Method for analyzing lymphocytes or plasma cells in sample by applying immune repertoire sequencing method, application of method and kit of method

    CN112852936A

  • TCR group library framework for diagnosis of multiple diseases

    CN117693793A

  • Adaptive immune system difference quantitative comparison method and application thereof

    CN118441038A

  • Measurement and Comparison of Immune Diversity by High-Throughput Sequencing

    US20140235478A1

  • Immune repertoire patterns

    US20210265008A1