Method for evaluating immune state and immune age of human body
By using multi-dimensional feature set information based on T-cell receptor immune repertoire and multi-task deep learning models, the accuracy problem of immune status and immune age assessment in existing technologies has been solved, enabling accurate assessment and early warning of human immune status and immune age.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAFEI IMMUNOSCIENCE (GUANGDONG) CO LTD
- Filing Date
- 2025-11-19
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have low accuracy in assessing human immune status and immune age, making it difficult to fully capture the dynamic functional state of the immune system, and disease states can interfere with the prediction results.
By employing multi-dimensional feature set information based on the T-cell receptor immune repertoire and combining it with a multi-task deep learning model, an evaluation result set information is generated, including immune health status classification and immune age prediction.
It enables accurate assessment of human immune status and immune age, improves prediction accuracy and robustness in diseased individuals, can identify sub-health states and provide early warning of immune dysregulation, and enhances model performance in situations where healthy samples are scarce.
Smart Images

Figure CN121834643A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biological data processing technology, and more specifically, to a method for assessing human immune status and immune age. Background Technology
[0002] The human immune system is the core of resisting disease and maintaining health. Its functional state is not a simple binary division of "health" or "disease," but a complex, continuous, and dynamic spectrum. Among them, "immunosenescence" is a key physiological process in which the function of the immune system gradually declines with age, and is closely related to increased risk of infection, weakened vaccine response, and increased incidence of cancer and autoimmune diseases.
[0003] Currently, flow cytometry is mainly used to analyze the proportion of immune cell subsets or to detect the levels of specific cytokines and antibodies (such as IgE). However, since it can only provide a static and local snapshot, it is difficult to fully capture the overall and dynamic functional state of the immune system, resulting in low accuracy, which needs further improvement. Summary of the Invention
[0004] Based on this, embodiments of this application provide a method for assessing human immune status and immune age to address the problem of low accuracy in the prior art.
[0005] In a first aspect, embodiments of this application provide a method for assessing human immune status and immune age, the method comprising: Based on a pre-defined T-cell receptor immune repertoire, multi-dimensional feature set information is obtained; The multi-dimensional feature set information is input into a preset multi-task deep learning model to generate an evaluation result set information, wherein the evaluation result set information includes immune health status classification information and immune age prediction information.
[0006] Compared with the prior art, the beneficial effects are as follows: The method for assessing human immune status and immune age provided in this application embodiment allows the terminal device to first quickly obtain multi-dimensional feature set information based on a preset T-cell receptor immune group library, and then input the multi-dimensional feature set information into a preset multi-task deep learning model to accurately generate assessment result set information, thereby accurately assessing human immune status and immune age, and to a certain extent solving the problem of low accuracy in the current system.
[0007] Secondly, embodiments of this application provide a system for assessing human immune status and immune age, the system comprising: Multi-dimensional feature set information acquisition module: used to acquire multi-dimensional feature set information based on a preset T cell receptor immune library; Evaluation result set information generation module: used to input the multi-dimensional feature set information into a preset multi-task deep learning model to generate evaluation result set information, wherein the evaluation result set information includes immune health status classification information and immune age prediction information.
[0008] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect above.
[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.
[0010] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0012] Figure 1 This is a flowchart illustrating a method for assessing human immune status and immune age according to an embodiment of this application; Figure 2 This is a flowchart illustrating the process before step S100 in a method for assessing human immune status and immune age provided in an embodiment of this application. Figure 3 This is a flowchart illustrating step S200 in a method for assessing human immune status and immune age provided in an embodiment of this application. Figure 4 This is a flowchart illustrating step S210 in a method for assessing human immune status and immune age provided in an embodiment of this application. Figure 5 This is a flowchart illustrating step S220 in a method for assessing human immune status and immune age provided in an embodiment of this application. Figure 6 This is a flowchart illustrating step S230 in a method for assessing human immune status and immune age provided in an embodiment of this application. Figure 7 This is a block diagram of a system for assessing human immune status and immune age provided in one embodiment of this application; Figure 8 This is a schematic diagram of a terminal device provided in an embodiment of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] In the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0015] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0016] Currently, among the existing technologies for predicting immune age, the most representative biological age prediction model is the "epigenetic clock" (e.g., Horvath clock) based on DNA methylation levels. This method predicts biological age by detecting the methylation status of specific CpG sites. Its disadvantages are: (1) Insufficient immune specificity: The epigenetic clock reflects the average methylation level of cells throughout the body, which is not specific to the immune system. Its prediction results are easily affected by non-immune tissues or disease states, and cannot accurately characterize "immune age"; (2) Inability to distinguish between aging and disease signals: When an individual suffers from chronic inflammatory diseases (such as allergic rhinitis), the inflammatory state will interfere with the methylation pattern, causing the predicted age to deviate from the true physiological aging trajectory, and cannot effectively remove the confusion of aging signals caused by diseases; (3) High cost of dynamic monitoring: DNA methylation detection is expensive and not convenient for frequent dynamic monitoring. Another technical approach is to use flow cytometry to detect surface markers of lymphocytes such as T cells and NK cells (e.g., CD28, CD45RA, etc.) and construct an immune age prediction model based on the proportion of these cell subpopulations. Its disadvantages are: (1) Low feature dimension and limited information: Flow cytometry can only capture a limited number of cell surface proteins (usually a dozen or so), and cannot reflect the deep information of the TCR repertoire at the molecular level, such as clonal diversity, physicochemical properties of CDR3 sequence, VJ pairing preference, etc., which leads to limited accuracy and information richness of the model; (2) There is also a signal confusion problem: Similar to the epigenetic clock, the disease state will significantly change the proportion of lymphocyte subpopulations, making it difficult for the model to distinguish between temporary immune disorders caused by disease and true, age-related immune aging.
[0017] Currently, in existing technologies for judging immune health status, machine learning (such as support vector machines and random forests) is commonly used to diagnose or classify specific diseases (such as cancer and autoimmune diseases). The disadvantages are: (1) Task isolation and neglect of intrinsic correlation: These models are usually designed as a single "classification task", completely separating immune health status judgment from immune age prediction, and ignoring the intrinsic biological correlation between TCR features and the two dimensions of "functional state" and "time accumulation (aging)". For example, the hydrophobicity of CDR3 may be related to specific antigen response (disease state) or may change systematically with age; (2) Model is susceptible to interference and has poor robustness: When trying to predict age with a single task model, the model cannot identify and exclude the "contamination" of TCR features by disease state, resulting in huge deviations in the age prediction of diseased individuals, and vice versa; (3) Low data utilization efficiency: For the task of "immune age prediction", strictly defined healthy samples are usually very scarce and costly to obtain. Single-task models cannot utilize a large number of existing disease-labeled samples to help learn general, age-related TCR feature patterns, resulting in poor model performance (i.e., underfitting) when data is limited.
[0018] To illustrate the technical solution described in this application, specific embodiments are provided below.
[0019] Please see Figure 1 , Figure 1 This is a flowchart illustrating the method for assessing human immune status and immune age provided in this application embodiment. In this embodiment, the method for assessing human immune status and immune age is executed by a terminal device. It is understood that the types of terminal devices include, but are not limited to, mobile phones, tablets, laptops, Ultra-Mobile Personal Computers (UMPCs), netbooks, Personal Digital Assistants (PDAs), etc. This application embodiment does not impose any restrictions on the specific type of terminal device.
[0020] Please see Figure 1 The method for assessing human immune status and immune age provided in this application includes, but is not limited to, the following steps: In S100, multi-dimensional feature set information is obtained based on a preset T-cell receptor immune repertoire.
[0021] Specifically, the terminal device can first obtain multi-dimensional feature set information based on the preset T-cell receptor immune repertoire. The multi-dimensional feature set information can be a 48-dimensional feature vector, which can represent a panoramic view of the TCR immune repertoire of a sample.
[0022] In some possible implementations, the multi-dimensional feature set information can include three major categories with a total of 48 features, namely, diversity index set information, gene usage preference set information, and CDR3 physicochemical property set information. Among them, the diversity index set information is used to quantify the uniformity and breadth of the distribution of target clone information. The diversity index set information can have 11 dimensions and can include unique clone number information (i.e., richness), Shannon entropy information, Gini coefficient information, clonality index information, and Rényi spectrum information.
[0023] Among them, gene usage preference set information is used to capture bias in gene fragment usage. Gene usage preference set information can be 7-dimensional in total, including V gene entropy information, J gene entropy information, VJ pairing Jensen-Shannon divergence information (i.e., JSD), high-frequency V / J gene information, and high-frequency V / J gene frequency information.
[0024] The CDR3 physicochemical property set information is used to describe the physicochemical properties of key regions of the T cell receptor. The CDR3 physicochemical property set information can be 30-dimensional and includes CDR3 length distribution information (i.e., 5-25), average hydrophobicity information, average net charge information, aromatic amino acid ratio information, and average molecular weight information.
[0025] In some possible implementations, to transform raw TCR sequencing data into structured, numerical feature vectors that are beneficial for subsequent multi-task deep learning models, please refer to [link to relevant documentation]. Figure 2 Before step S100, the method further includes, but is not limited to, the following steps: In S101, raw sequencing data of the T cell receptor were obtained.
[0026] Specifically, the terminal device can first obtain the TCR sequence file through high-throughput sequencing to acquire the raw sequencing data of the T cell receptor, where the raw sequencing data of the T cell receptor is used to describe the raw sequencing data of the TCR.
[0027] In S102, the raw sequencing data of T cell receptors is processed to generate recombinant sequencing data of T cell receptors.
[0028] Specifically, after the terminal device acquires the raw sequencing data of the T cell receptor, the terminal device can perform quality control processing and sequence assembly processing on the raw sequencing data to generate recombinant sequencing data of the T cell receptor. The recombinant sequencing data of the T cell receptor can be the raw sequencing data of the T cell receptor after sequence assembly processing.
[0029] In S103, based on a pre-defined annotation tool, the T-cell receptor recombinant sequencing data is annotated to determine key T-cell receptor region information.
[0030] Specifically, after the terminal device generates T-cell receptor recombinant sequencing data, the terminal device can annotate the V gene fragment, D gene fragment, and J gene fragment in the T-cell receptor recombinant sequencing data based on a preset annotation tool to determine the key region information of the T-cell receptor. The annotation tool can be the IMGT tool or the V-QUEST tool; the key region information of the T-cell receptor can be the CDR3 region.
[0031] In S104, target clone information is determined based on key regions of multiple T cell receptors with the same amino acid sequence.
[0032] Specifically, after the terminal device determines the key region information of the T cell receptor, the terminal device can determine the target clone information based on multiple key region information of the T cell receptor with the same amino acid sequence, thereby defining the CDR3 region with the same amino acid sequence as a clone.
[0033] In S105, the cloning frequency information is determined based on the target cloning information.
[0034] Specifically, after the terminal device determines the target clone information, it can determine the clone frequency information based on the target clone information, thereby calculating the frequency of each clone in the sample.
[0035] In S106, a T-cell receptor immune repertoire was constructed based on T-cell receptor recombinant sequencing data and cloning frequency information.
[0036] Specifically, after the terminal device determines the cloning frequency information, it can construct a T-cell receptor immune library based on the T-cell receptor recombinant sequencing data and cloning frequency information, which is beneficial for the subsequent systematic extraction of 48-dimensional features based on cloning frequency and sequence information.
[0037] In S200, multi-dimensional feature set information is input into a pre-defined multi-task deep learning model to generate evaluation result set information.
[0038] Specifically, after the terminal device acquires multi-dimensional feature set information, it can input the multi-dimensional feature set information into a preset multi-task deep learning model to accurately generate evaluation result set information, thereby achieving accurate assessment of human immune status and immune age. The evaluation result set information includes immune health status classification information and immune age prediction information.
[0039] Specifically, the multi-task deep learning model includes a shared encoder and at least two task-specific decoders, one of which is a classification decoder for determining immune health status, and the other is a regression decoder for predicting immune age.
[0040] In some possible implementations, during the training phase of the multi-task deep learning model, the loss function used by the aforementioned multi-task deep learning model can be: , In the formula, This indicates the total loss information. This represents the preset first weight value; The Huber loss value represents the prediction of immune age, which is used for robust outliers; This represents the preset second weight value. and All of these can be determined using a grid search; The cross-entropy loss value represents the health status classification.
[0041] Meanwhile, during the training phase of the multi-task deep learning model, the optimization strategy adopted by the aforementioned multi-task deep learning model can be combined with the AdamW optimizer, and through dynamic adjustment of the learning rate and gradient pruning, the stability and efficiency of the training process can be ensured.
[0042] For some possible implementations, please refer to [link to relevant documentation] for accurate generation of evaluation result set information. Figure 3 Step S200 includes, but is not limited to, the following steps: In S210, multi-dimensional feature set information is input into the shared encoder to generate T cell receptor shared feature information.
[0043] Specifically, the terminal device can input multi-dimensional feature set information into the shared encoder to generate shared feature information of T cell receptors. The shared encoder includes a first fully connected layer, a second fully connected layer, a first residual block, a second residual block, and a multi-head self-attention layer.
[0044] In some possible implementations, to achieve the generation of shared T cell receptor signature information, please refer to [link to relevant documentation]. Figure 4 Step S210 includes, but is not limited to, the following steps: In S211, the multi-dimensional feature set information is input into the first fully connected layer to generate the first processed data information, so as to map the multi-dimensional feature set information from 48 dimensions to 512 dimensions.
[0045] Specifically, the terminal device can input multi-dimensional feature set information into the first fully connected layer to generate first processed data information, so as to map the multi-dimensional feature set information from 48 dimensions to 512 dimensions. After data processing in the first fully connected layer, batch normalization and SiLU activation function can be applied to accelerate convergence and enhance nonlinear expression capability.
[0046] In S212, the first processed data information is input to the second fully connected layer to generate the second processed data information, so as to map the multi-dimensional feature set information from 512 dimensions to 256 dimensions.
[0047] Specifically, after the terminal device generates the first processed data information, the terminal device can input the first processed data information into the second fully connected layer to generate the second processed data information, so as to map the multi-dimensional feature set information from 512 dimensions to 256 dimensions, thereby realizing the gradual mapping of features to a higher-dimensional space; after data processing in the second fully connected layer, batch normalization and SiLU activation function can also be applied to achieve the same effect.
[0048] In S213, the second processed data information is input into the first residual block to generate the third processed data information, so as to reduce the multi-dimensional feature set information from 256 dimensions to 128 dimensions.
[0049] Specifically, after the terminal device generates the second processed data information, it can input the second processed data information into the first residual block to generate the third processed data information, thereby reducing the multi-dimensional feature set information from 256 dimensions to 128 dimensions. The residual connection can effectively alleviate the gradient vanishing problem in deep networks. Both the first residual block and the subsequent second residual block contain two fully connected layers. After the data processing of the first residual block, batch normalization, SiLU activation, and Dropout regularization can be applied.
[0050] In S214, the third processing data information is input into the second residual block to generate the fourth processing data information, so as to reduce the multi-dimensional feature set information from 128 dimensions to 64 dimensions.
[0051] Specifically, after the terminal device generates the third processed data information, the terminal device can input the third processed data information into the second residual block to generate the fourth processed data information, so as to reduce the multi-dimensional feature set information from 128 dimensions to 64 dimensions, thereby achieving stepwise dimensionality reduction. After the data is processed in the second residual block, batch normalization, SiLU activation and Dropout regularization can also be used.
[0052] In S215, the fourth processing data information is input into the multi-head self-attention layer to generate T cell receptor shared feature information.
[0053] Specifically, after the terminal device generates the fourth processing data information, it can input the fourth processing data information into the multi-head self-attention layer to generate T cell receptor shared feature information. This enables the learning and extraction of a more abstract 64-dimensional shared feature representation that is valuable for both tasks from the 48-dimensional input features. The multi-head self-attention layer can be a 4-head Transformer self-attention layer. The multi-head self-attention layer can dynamically calculate the importance weights between different TCR features, thereby capturing the complex, high-order nonlinear interactions between features, which is difficult to achieve with traditional machine learning models.
[0054] It should be noted that shared encoders force multi-task deep learning models to find common features in disease and aging signals. For example, certain CDR3 length features may be abnormal in disease states and also change systematically with age. This shared representation learning is the foundation for achieving signal decoupling.
[0055] In S220, the shared feature information of T cell receptors is input into the classification decoder to generate immune health status classification information.
[0056] Specifically, after the terminal device generates T-cell receptor shared feature information, the terminal device can input the T-cell receptor shared feature information into the classification decoder to generate immune health status classification information. The classification decoder includes a fully connected layer and a classification layer.
[0057] In some possible implementations, for generating immune health status classification information, please refer to [link / reference]. Figure 5 Step S220 includes, but is not limited to, the following steps: In S221, the shared feature information of T cell receptors is input into the fully connected layer to generate the fifth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions.
[0058] Specifically, the terminal device can input T-cell receptor shared feature information into a fully connected layer to generate fifth-stage processing data, thereby reducing the multi-dimensional feature set information from 64 dimensions to 32 dimensions. This achieves the reception of 64-dimensional shared features and their dimensionality reduction to 32 dimensions through a fully connected layer. After data processing in this fully connected layer, batch normalization, SiLU activation, and Dropout regularization can then be applied.
[0059] In S222, based on the preset Softmax activation function, the fifth processed data information is input into the classification layer to generate immune health status classification information.
[0060] Specifically, after the terminal device generates the fifth processed data information, it can input the fifth processed data information into the classification layer based on a preset Softmax activation function to generate immune health status classification information. This reduces the multi-dimensional feature set information from 32 dimensions to 3 dimensions, thereby enabling the output of the probability of belonging to three categories: healthy, sub-healthy, and disease through a 3-node output layer and the Softmax activation function. The immune health status classification information includes health category probability information, sub-healthy category probability information, and disease category probability information. The health category probability information describes the probability of the healthy category, the sub-healthy category probability information describes the probability of the sub-healthy category, and the disease category probability information describes the probability of the disease category.
[0061] In S230, T-cell receptor shared feature information is input into the regression decoder to generate immune age prediction information.
[0062] Specifically, the terminal device can input the shared feature information of T cell receptors into the regression decoder to generate immune age prediction information. The regression decoder includes a fully connected layer and a regression layer. The regression decoder can be symmetrical to the classification decoder in terms of network structure, that is, they can perform data processing tasks simultaneously.
[0063] In some possible implementations, for generating immune age prediction information, please refer to [link / reference needed]. Figure 6 Step S230 includes, but is not limited to, the following steps: In S231, the shared feature information of T cell receptors is input into the fully connected layer to generate the sixth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions.
[0064] Specifically, the terminal device can input the shared feature information of T cell receptors into the fully connected layer to generate sixth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions. After data processing in the fully connected layer, batch normalization, SiLU activation and Dropout regularization can also be used.
[0065] In S232, based on a preset linear activation function, the sixth processed data information is input into the regression layer to generate immune age prediction information.
[0066] Specifically, after the terminal device generates the sixth processed data information, the terminal device can input the sixth processed data information into the regression layer based on the preset linear activation function to generate immune age prediction information, so as to reduce the multi-dimensional feature set information from 32 dimensions to 1 dimension, thereby realizing the use of the linear activation function in the final output layer to output a continuous immune age scalar value.
[0067] It should be noted that this application can obtain a comprehensive immune profile of an individual through a single inference, including their current immune status and the "physiological age" of their immune system, providing an unprecedented tool for precision health management. The multi-task deep learning model of this application achieved a health status classification accuracy of 93.46% and an average absolute error of 3.39 years for immune age prediction on an independent test set. Through accurate identification of sub-healthy individuals and the discovery of accelerated immune age in AR patients, the multi-task deep learning model of this application demonstrates robustness in complex real-world scenarios. This application can objectively quantify the "sub-healthy" immune transition state, enabling the prediction of disease occurrence. This application provides early warning of immune dysregulation, significantly advancing the intervention window. When predicting the age of an AR patient, the classification decoder "captures" features related to allergies and inflammation, while the regression decoder provides an immune age based on "purer" aging-related features, effectively avoiding misclassification of inflammation as aging. The classification task utilizes a large number of sub-health / disease samples. The general knowledge about the overall structure of the TCR repertoire learned from these samples is transferred to the regression task through a shared encoder, greatly assisting and improving the performance of age prediction models trained on limited healthy samples. This application can automatically discover and prioritize key feature combinations. For example, it may learn that "when the frequency of VJ pairing with high JSD and a CDR3 length of 25 decreases simultaneously, it strongly indicates a sub-health state," a high-order nonlinear relationship that traditional models struggle to capture. By analyzing attention weights, this application can trace back which input features contribute most to the final decision, providing clues for biologists to understand the mechanisms of immune aging. The fusion of three major categories of features—diversity, gene usage, and physicochemical properties—ensures the comprehensiveness of the model's input information and avoids model bias caused by missing features. For example, using diversity metrics alone may miss specific VJ pairing patterns in AR, while comprehensive features enable the model to capture these subtle but crucial signals.
[0068] In some possible implementations, the shared encoder is not limited to fully connected layers and Transformers. Other deep learning modules can be replaced or combined, such as: (1) Convolutional Neural Networks (CNNs): If the feature vectors are arranged in a specific way, a one-dimensional CNN can be used to extract local patterns; (2) Graph Neural Networks (GNNs): If the TCR clones and their relationships are used to construct a graph, GNNs can be used for learning; (3) Autoencoders (VAEs): VAEs can be used to reduce the dimensionality and denoise the features before the latent variables are input into the task decoder; (4) Attention mechanism variants: Other attention mechanisms such as gated attention and external memory-enhanced attention can be used to replace the standard Transformer self-attention; (5) Decoder variants: The decoder can be deeper or wider, and attention mechanisms can be introduced to perform secondary weighting on the shared features. In terms of feature engineering extensions, more types of features can be introduced to extend the feature dimensions, such as the following features: (1) k-mer frequencies of CDR3 sequences; (2) Embedded features extracted from the original CDR3 sequences based on deep learning (such as LSTM); (3) Matching information from public antigen-specific databases.
[0069] In some possible implementations, the multi-task deep learning model of this application can be extended to multimodal learning, so that other omics data can be integrated by simply adding an additional branch to the shared encoder part, such as: (1) epigenetic data (e.g. DNA methylation level); (2) cytokine profile data; (3) microbiome data, so that these different modalities of data can be fused after the shared encoder, which is expected to build a more powerful multi-omics immune assessment model.
[0070] In some possible implementations, in terms of task definition expansion, it can be fine-grained classification, that is, the classification task is not limited to three categories, but can be extended to more disease types, such as subdividing "disease" into autoimmune diseases, chronic infections, cancer, etc.; or it can be adding a prediction task, that is, a third decoder can be easily added to the framework to predict other related indicators, such as: (1) the strength of response to a specific vaccine; (2) the response probability of immune checkpoint inhibitor treatment; (3) the recurrence risk of a specific disease (such as cancer).
[0071] In some possible implementations, the training strategy of the multi-task deep learning model can be optimized, for example: (1) In terms of the loss function: a dynamic weight adjustment strategy can be adopted to automatically adjust the weights according to the learning progress of the two tasks. and (2) In terms of data augmentation: In addition to SMOTE, Generative Adversarial Networks (GANs) can be used to generate more realistic minority class samples.
[0072] In some possible implementations, this application uses "immune age difference" as a biomarker to evaluate whether probiotics, drugs or lifestyle interventions effectively delay immune aging.
[0073] The implementation principle of the method for assessing human immune status and immune age in this application embodiment is as follows: The terminal device can first quickly acquire multi-dimensional feature set information based on a preset T-cell receptor immune repertoire, and then input the multi-dimensional feature set information into a preset multi-task deep learning model to accurately generate evaluation result set information. This creatively applies the multi-task learning (MTL) paradigm to TCR immune repertoire data analysis, solving the two fundamentally related technical problems of "confusion between disease and aging signals" and "scarcity of healthy samples" in this specific field. By forcing the shared encoder to learn a more general TCR feature representation useful for both tasks, interference signals useful only for a single task are naturally filtered out. By using signal decoupling as the model's built-in optimization objective, and then combining it with the joint optimization of multi-task losses, the model is actively guided to discover: which feature changes are mainly related to disease status (captured by the classification decoder), and which feature drifts are mainly related to physiological aging (captured by the regression decoder). The TCR library structure learned from patient data can be transferred to better understand the aging process in healthy individuals. This effectively breaks down the isolation between health and disease data, greatly alleviates the bottleneck of scarce healthy samples, improves data utilization efficiency, and makes it possible to build a high-precision age prediction model with limited healthy samples. At the same time, by dynamically calculating the correlation weights between different features through the Transformer multi-head self-attention mechanism, the model can capture complex feature combination patterns such as "a high Gini coefficient and a low frequency of CDR3 length 25 strongly indicate sub-health", thereby improving the model's discriminative ability and interpretability.
[0074] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0075] Embodiments of this application also provide a system for assessing human immune status and immune age. For ease of explanation, only the parts relevant to this application are shown, such as... Figure 7 As shown, the system 70 includes: Multi-dimensional feature set information acquisition module 71: used to acquire multi-dimensional feature set information based on a preset T cell receptor immune library; Evaluation result set information generation module 72: used to input multi-dimensional feature set information into a preset multi-task deep learning model to generate evaluation result set information, wherein the evaluation result set information includes immune health status classification information and immune age prediction information.
[0076] Optionally, the multi-task deep learning model includes a shared encoder and at least two task-specific decoders, one of which is a classification decoder for determining immune health status, and the other is a regression decoder for predicting immune age. The evaluation result set information generation module 72 includes: T-cell receptor shared feature information generation submodule: used to input multi-dimensional feature set information into the shared encoder to generate T-cell receptor shared feature information; Immune health status classification information generation submodule: used to input T cell receptor shared feature information into the classification decoder to generate immune health status classification information; The immune age prediction information generation submodule is used to input T cell receptor shared feature information into the regression decoder to generate immune age prediction information.
[0077] Optionally, the system 70 also includes: T-cell receptor raw sequencing data acquisition module: used to acquire T-cell receptor raw sequencing data; T-cell receptor recombinant sequencing data generation module: used to perform sequence assembly processing on raw T-cell receptor sequencing data to generate T-cell receptor recombinant sequencing data; T-cell receptor key region information determination module: used to annotate T-cell receptor recombinant sequencing data based on preset annotation tools to determine key region information of T-cell receptor; Target clone information determination module: used to determine target clone information based on key regions of multiple T cell receptors with the same amino acid sequence; Cloning frequency information determination module: used to determine cloning frequency information based on target cloning information; T-cell receptor immune repertoire construction module: used to construct T-cell receptor immune repertoire based on T-cell receptor recombinant sequencing data and cloning frequency information.
[0078] Optionally, the multi-dimensional feature set information includes diversity index set information, gene usage preference set information, and CDR3 physicochemical property set information. The diversity index set information is used to quantify the uniformity and breadth of the distribution of target clone information. The diversity index set information includes unique clone number information, Shannon entropy information, Gini coefficient information, clonality index information, and Rényi spectrum information. The gene usage preference set information is used to capture the bias in gene fragment usage. The gene usage preference set information includes V gene entropy information, J gene entropy information, VJ pairing Jensen-Shannon divergence information, high-frequency V / J gene information, and high-frequency V / J gene frequency information. The CDR3 physicochemical property set information is used to describe the physicochemical properties of key regions of the T cell receptor. The CDR3 physicochemical property set information includes CDR3 length distribution information, average hydrophobicity information, average net charge information, aromatic amino acid ratio information, and average molecular weight information.
[0079] Optionally, the shared encoder includes a first fully connected layer, a second fully connected layer, a first residual block, a second residual block, and a multi-head self-attention layer; the classification decoder includes a fully connected layer and a classification layer; and the regression decoder includes a fully connected layer and a regression layer. The aforementioned T-cell receptor shared feature information generation submodule includes: First processing data information generation unit: used to input multi-dimensional feature set information into the first fully connected layer to generate first processing data information, so as to map the multi-dimensional feature set information from 48 dimensions to 512 dimensions; The second processing data information generation unit is used to input the first processing data information into the second fully connected layer to generate the second processing data information, so as to map the multi-dimensional feature set information from 512 dimensions to 256 dimensions. The third processing data information generation unit is used to input the second processing data information into the first residual block and generate the third processing data information to reduce the multi-dimensional feature set information from 256 dimensions to 128 dimensions. The fourth processing data information generation unit is used to input the third processing data information into the second residual block to generate the fourth processing data information, so as to reduce the multi-dimensional feature set information from 128 dimensions to 64 dimensions. T-cell receptor shared feature information generation unit: used to input the fourth-processed data information into the multi-head self-attention layer to generate T-cell receptor shared feature information; Accordingly, the above-mentioned immune health status classification information generation submodule includes: The fifth processing data information generation unit is used to input the shared feature information of T cell receptors into the fully connected layer to generate the fifth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions; Immune health status classification information generation unit: It is used to input the fifth processed data information into the classification layer based on the preset Softmax activation function to generate immune health status classification information, wherein the immune health status classification information includes health category probability information, sub-health category probability information and disease category probability information; Accordingly, the above-mentioned immune age prediction information generation submodule includes: The sixth processing data information generation unit is used to input the shared feature information of T cell receptors into the fully connected layer to generate the sixth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions. Immune age prediction information generation unit: Based on a preset linear activation function, it inputs the sixth-processed data information into the regression layer to generate immune age prediction information.
[0080] Optionally, the loss function for the above multi-task deep learning model is: , In the formula, For total loss information, The preset first weight value, This represents Huber's loss value. The preset second weight value, This represents the cross-entropy loss value.
[0081] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0082] This application also provides a terminal device, such as... Figure 8 As shown, the terminal device 80 of this embodiment includes: a processor 81, a memory 82, and a computer program 83 stored in the memory 82 and executable on the processor 81. When the processor 81 executes the computer program 83, it implements the steps described in the method embodiment for assessing human immune status and immune age, for example... Figure 1 Steps S100 to S200 are shown; or, when processor 81 executes computer program 83, it implements the functions of each module in the above-described device, for example... Figure 7 The functions of modules 71 and 72 shown.
[0083] The terminal device 80 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device, and includes, but is not limited to, a processor 81 and a memory 82. Those skilled in the art will understand that... Figure 8This is merely an example of terminal device 80 and does not constitute a limitation on terminal device 80. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device 80 may also include input / output devices, network access devices, buses, etc.
[0084] The processor 81 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.; the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0085] The memory 82 can be an internal storage unit of the terminal device 80, such as a hard disk or memory of the terminal device 80. The memory 82 can also be an external storage device of the terminal device 80, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 80. Furthermore, the memory 82 can include both internal storage units and external storage devices of the terminal device 80. The memory 82 can also store computer program 83 and other programs and data required by the terminal device 80. The memory 82 can also be used to temporarily store data that has been output or will be output.
[0086] One embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0087] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the methods, principles and structures of this application should be covered within the scope of protection of this application.
Claims
1. A method of assessing the immune status and immune age of a human body, characterized in that, The method includes: Based on a pre-defined T-cell receptor immune repertoire, multi-dimensional feature set information is obtained; The multi-dimensional feature set information is input into a preset multi-task deep learning model to generate an evaluation result set information, wherein the evaluation result set information includes immune health status classification information and immune age prediction information.
2. The method of claim 1, wherein, The multi-task deep learning model includes a shared encoder and at least two task-specific decoders, one of which is a classification decoder for determining immune health status, and the other is a regression decoder for predicting immune age. The step of inputting the multi-dimensional feature set information into the preset multi-task deep learning model to generate an evaluation result set includes: The multi-dimensional feature set information is input into the shared encoder to generate T cell receptor shared feature information. The shared feature information of the T cell receptors is input into the classification decoder to generate the immune health status classification information; The shared feature information of the T cell receptors is input into the regression decoder to generate the immune age prediction information.
3. The method of claim 1, wherein, Before acquiring multi-dimensional feature set information based on a preset T-cell receptor immune repertoire, the method further includes: Obtain raw sequencing data of T cell receptors; The raw sequencing data of the T cell receptor are processed by sequence assembly to generate recombinant sequencing data of the T cell receptor; Based on a preset annotation tool, the T cell receptor recombinant sequencing data is annotated to determine key T cell receptor region information. Based on the key region information of the T cell receptor with multiple identical amino acid sequences, the target clone information was determined; Based on the target clone information, determine the cloning frequency information; Based on the T cell receptor recombinant sequencing data and cloning frequency information, a T cell receptor immune repertoire was constructed.
4. The method of claim 1, wherein, The multi-dimensional feature set information includes diversity index set information, gene usage preference set information, and CDR3 physicochemical property set information. The diversity index set information is used to quantify the uniformity and breadth of the distribution of target clone information. The diversity index set information includes unique clone number information, Shannon entropy information, Gini coefficient information, clonality index information, and Rényi spectrum information. The gene usage preference set information is used to capture bias in gene fragment usage. The gene usage preference set information includes V gene entropy information, J gene entropy information, VJ pairing Jensen-Shannon divergence information, high-frequency V / J gene information, and high-frequency V / J gene frequency information. The CDR3 physicochemical property set information is used to describe the physicochemical properties of the key regions of the T cell receptor. The CDR3 physicochemical property set information includes CDR3 length distribution information, average hydrophobicity information, average net charge information, aromatic amino acid ratio information, and average molecular weight information.
5. The method of claim 2, wherein, The shared encoder includes a first fully connected layer, a second fully connected layer, a first residual block, a second residual block, and a multi-head self-attention layer; the classification decoder includes a fully connected layer and a classification layer; and the regression decoder includes a fully connected layer and a regression layer. The step of inputting the multi-dimensional feature set information into the shared encoder to generate T cell receptor shared feature information includes: The multi-dimensional feature set information is input into the first fully connected layer to generate the first processed data information, so as to map the multi-dimensional feature set information from 48 dimensions to 512 dimensions; The first processed data information is input into the second fully connected layer to generate the second processed data information, so as to map the multi-dimensional feature set information from 512 dimensions to 256 dimensions; The second processed data information is input into the first residual block to generate the third processed data information, so as to reduce the multi-dimensional feature set information from 256 dimensions to 128 dimensions; The third processed data information is input into the second residual block to generate the fourth processed data information, so as to reduce the multi-dimensional feature set information from 128 dimensions to 64 dimensions; The fourth processed data information is input into the multi-head self-attention layer to generate T cell receptor shared feature information; Accordingly, the step of inputting the T cell receptor shared feature information into the classification decoder to generate the immune health status classification information includes: The shared feature information of the T cell receptor is input into the fully connected layer to generate the fifth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions; Based on the preset Softmax activation function, the fifth processed data information is input into the classification layer to generate the immune health status classification information, wherein the immune health status classification information includes health category probability information, sub-health category probability information and disease category probability information; Accordingly, the step of inputting the T cell receptor shared feature information into the regression decoder to generate the immune age prediction information includes: The shared feature information of the T cell receptor is input into the fully connected layer to generate the sixth processing data information, so as to reduce the multi-dimensional feature set information from 64 dimensions to 32 dimensions; Based on a preset linear activation function, the sixth processed data information is input into the regression layer to generate the immune age prediction information.
6. The method of claim 1, wherein, The loss function of the multi-task deep learning model is: , In the formula, is total loss information, is a preset first weight value, is a Huber loss value, is a preset second weight value, is a cross-entropy loss value.
7. A system for assessing immune status and immune age of a human body, characterized in that, The system includes: Multi-dimensional feature set information acquisition module: used to acquire multi-dimensional feature set information based on a preset T cell receptor immune library; Evaluation result set information generation module: used to input the multi-dimensional feature set information into a preset multi-task deep learning model to generate evaluation result set information, wherein the evaluation result set information includes immune health status classification information and immune age prediction information.
8. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.