Immune state evaluation method and system based on plasma proteomics
Through the Meta-GCN model and multiple machine learning algorithms, immune-related proteins in plasma proteins are identified, and the problem of incomplete immune status assessment in the prior art is solved, and efficient and accurate immune status assessment is achieved.
Patent Information
- Application Number
- CN202510332485.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, immune status assessment relies on white blood cell counts routinely detected, making it difficult to provide comprehensive immune information. In addition, traditional machine learning methods require large-scale bioinformatic data to label large-scale bioinformatic data, which is unrealistic.
The Meta-GCN model was used to identify immune-related proteins in plasma proteins, combining meta-learning and multiple machine learning algorithms, and through a small amount of known immune-related proteins and interaction data, more immune-related proteins and predict immune status scores.
It improves the accuracy and efficiency of immune status assessment, reduces the need for labeled data, and can comprehensively evaluate systemic immune status. External verification also proves the stability and accuracy of the method.
Smart Images

Figure CN120296552A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the application of artificial intelligence models in the field of immunology, and specifically relates to a method and system for evaluating immune status based on plasma proteomics. Background Art
[0002] In the prior art, the evaluation of immune status mainly relies on immune cell data in blood routine tests, such as white blood cell count. However, this method can only reflect limited immune information and it is difficult to provide accurate immune data for identifying immune abnormalities through this immune status evaluation. In contrast, plasma protein data provides a more comprehensive perspective for evaluating the status of the whole body immune system. However, there is a large amount of information involved in plasma proteomics. How to select and choose which plasma proteins as immune-related proteins for machine learning to obtain a reliable immune status evaluation model is a technical difficulty. In addition, traditional machine learning methods often require a large amount of labeled data when dealing with large-scale biological information data, which is unrealistic in practical applications. Therefore, developing a new method that can accurately evaluate an individual's immune status using plasma protein data and has less demand for labeled data has important clinical and scientific research value. Summary of the Invention
[0003] In a first aspect, in order to solve at least one of the above technical problems, the present invention provides a method for evaluating immune status based on plasma proteomics, which is implemented by one or more processors and includes:
[0004] Using a Meta-GCN model to identify immune-related proteins in plasma proteins. The identification process includes taking interaction data with medium confidence among several interaction data between several plasma proteins of healthy people to form a topological structure diagram. The nodes of the topological structure diagram are plasma proteins, and the nodes form an identity matrix in the form of one-hot encoding. The edges of the topological structure diagram are the confidence levels of pairwise protein interactions, and the edges form an adjacency matrix. Inputting the identity matrix and the adjacency matrix into the Meta-GCN model for meta-learning. Each time of meta-learning uses the known immune-related proteins in plasma proteins as positive labels, and randomly selects 2-4 times the number of positive labels of proteins from the remaining plasma proteins as negative labels for learning, and uses the predicted probability that the plasma protein is an immune-related protein as the output. The meta-learning process is repeated multiple times until the prediction results of all plasma proteins in the topological structure diagram converge or the meta-learning reaches the maximum specified number of rounds. Taking the known immune-related proteins and the plasma proteins with predicted probabilities higher than the threshold as the immune-related proteins identified by the Meta-GCN model. There are m immune-related proteins identified by the Meta-GCN model;
[0005] Predict using the immune status score prediction model. The prediction process includes training the prediction model with the immune-related proteins identified by using n of the described Meta-GCN models, where 81 ≤ n ≤ m. Using the expression level characteristic data of the immune-related proteins identified by the n Meta-GCN models of a group of healthy people to be tested as input features, and using the initial immune status score corresponding to the age of this group of healthy people to be tested as a label for model training. The trained prediction model can output its immune status score according to the expression level characteristic data of the immune-related proteins identified by the n Meta-GCN models of the input person to be tested. The immune status score prediction model is selected from one or more of Lasso, LightGBM, XGBoost, and random forest. When there are multiple models, the immune status score is the average of the scores given by multiple models respectively. The calculation formula for the initial immune status score is:
[0006] y = -6.58e -7 x 3 +9.662e -5 x 2 -5.468e -3 +0.6177, where x is the age and y is the initial immune status score.
[0007] In some embodiments, the expression level characteristic data of the immune-related proteins identified by the Meta-GCN model refers to the result after the protein relative fluorescence unit (RFU) is logarithmically transformed by 10. For example, the plasma protein expression levels of healthy people are obtained through the SOMAscan platform, which corresponds to protein relative fluorescence unit data, that is, the protein relative fluorescence unit data can be obtained through the SOMAscan platform.
[0008] In some embodiments, the interaction data among the several plasma proteins is obtained by inputting the names of the several plasma proteins into the STRING database for analysis; the medium confidence level is 0.4; the threshold is 0.95.
[0009] In some embodiments, the maximum specified number of rounds is 50 ± 5 training cycles, with 20 ± 2 iterations per training cycle and 20 ± 2 random node selections per iteration; the learning rate of meta-learning is 0.0001, or 0.0005, or 0.00001; each random node selection includes 10 ± 2 negative-label proteins and 5 ± 1 positive-label proteins; the 10 ± 2 negative-label proteins are generated by randomly selecting with replacement from 100 ± 10 negative-label proteins, and the 5 ± 1 positive-label proteins are generated by randomly selecting with replacement from the known immune-related proteins; the 100 ± 10 negative-label proteins are generated by randomly selecting from plasma proteins excluding the known immune-related proteins; the number of known immune-related proteins is 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40, and 30, 31, 32, 33, 34, or 35 of them are selected from IFNB1, IL1A, IL23R, TGFB1, C2, C3, C5, C6, C7, C9, CCL13, CCL2, CCL20, CCL28, CXCL1, CXCL5, CXCL6, CXCL8, IFNA2, IFNG, IL10, IL13, IL17A, IL1B, IL2, IL22, IL34, IL4, IL5, IL6, IL9, TNF, CXCL9, GDF15, and CSF1; the number of plasma proteins of healthy individuals is 1000 - 1305, or 1100 - 1300, or 1150 - 1250, or 1180 - 1240, or 1190 - 1230, or 1195 - 1220, or 1195 - 1215, or 1200 - 1210, or 1205 - 1210, or 1206, or 1207, or 1208, or 1209.
[0010] In some embodiments, the predicted probability of each protein is the median of 100 ± 5 predicted results converging or 100 ± 5 predicted probabilities corresponding to the maximum specified number of rounds of meta-learning; the y value is mapped to the range of 60 - 95 scores and used as a label for training the prediction model.
[0011] In some embodiments, the graph convolutional network (GCN) aggregates neighborhood information through two convolutional layers, gradually extracting features from the topological structure graph. The adjacency matrix is input into the Meta-GCN model for meta-learning in the data form obtained after normalization by FirstOrderGCN. The dimension of the identity matrix is the number of proteins × the number of proteins. The role of this identity matrix is to provide an initial feature representation for each protein node, where the feature of each node is a one-hot encoding, that is, there is only one 1 in the corresponding row of the node, and the rest are 0. The one-hot encoding can provide a basic node discrimination ability for the model, enabling the Meta-GCN model to learn the relationships and feature propagation between nodes in subsequent graph convolution operations.
[0012] In a second aspect, the present invention also provides another method for evaluating immune status based on plasma proteomics, which is implemented by one or more processors and includes:
[0013] Predict using an immune status score prediction model. The prediction process includes training the prediction model using the immune-related proteins identified by n Meta-GCN models, where 81 ≤ n ≤ 309; using the expression feature data of the immune-related proteins identified by the n Meta-GCN models of a group of healthy people to be tested as input features, and using the initial immune status score corresponding to the age of this group of healthy people to be tested as a label for model training. The trained prediction model can output its immune status score according to the expression feature data of the immune-related proteins identified by the n Meta-GCN models of the input person to be tested; the immune status score prediction model is selected from one or more of Lasso, LightGBM, XGBoost, and random forest. When there are multiple models, the immune status score is the average of the scores of multiple models; the calculation formula for the initial immune status score is:
[0014] y = -6.58e -7 x 3 + 9.662e -5 x 2 - 5.468e -3 + 0.6177, where x is the age and y is the initial immune status score;
[0015] The immune-related proteins identified by the Meta-GCN model are shown in Table 1:
[0016] Table 1. Immune-related proteins identified by the Meta-GCN model
[0017]
[0018]
[0019]
[0020]
[0021]
[0022]
[0023]
[0024] In some embodiments, the expression quantity feature data of the immune-related proteins obtained by the Meta-GCN model recognition refers to the result after the relative fluorescence units of the proteins are logarithmically transformed by 10. The relative fluorescence unit data of the proteins can be obtained through the SOMAscan platform.
[0025] In a third aspect, the present invention also discloses a computer program including instructions, which when executed by one or more processors of a computing system, cause the computing system to execute the immune status assessment method based on plasma proteomics according to any one of the first aspect or the second aspect.
[0026] The present invention also discloses a computing system configured to execute the immune status assessment method based on plasma proteomics according to any one of the first aspect or the second aspect.
[0027] The present invention also discloses a computer-readable storage medium storing instructions that can be executed by one or more processors of a computing system to implement the immune status assessment method based on plasma proteomics according to any one of the first aspect or the second aspect.
[0028] Compared with the prior art, all or some embodiments of the present invention have the following beneficial effects:
[0029] 1. Improve efficiency and accuracy: The present invention greatly shortens the time for identifying immune-related proteins through meta-learning and related ingenious designs, and can quickly identify more immune-related proteins, thereby improving the accuracy of immune status assessment.
[0030] 2. Reduce costs: By reducing the need for a large amount of labeled data, the present invention effectively reduces the cost of data labeling, making the immune status assessment more economical and efficient. This is mainly reflected in that only more than 30 (35 in the embodiment) known immune-related proteins are used as positive labels, combined with interaction data, and the Meta-GCN model is used to accurately find a large number of immune-related proteins (309 in the embodiment), laying a good foundation for accurately predicting the immune status score subsequently.
[0031] 3. Comprehensive evaluation: The present invention first proposes to comprehensively evaluate the systemic immune status using plasma protein data. This innovative method can more accurately reflect an individual's immune status.
[0032] 4. External validation: The present invention verified the relevant methods with an external validation set, fully demonstrating that the method has good stability and accuracy in immune status evaluation. Data validation found that among 309 immune-related proteins, taking 81 kinds can achieve relatively high accuracy.
[0033] The following will further illustrate the concept, specific structure and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. Brief description of the drawings
[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 It is a corresponding relationship diagram between age and immune status score.
[0036] Figure 2 It is a flowchart of immune status evaluation.
[0037] Figure 3 It is a display of the prediction results of multiple machine learning models.
[0038] Figure 4 It is an analysis result diagram of the immune status of two external data sets. Mann-Whitney U test analysis, "ns" indicates p>0.05; "*" indicates p≤0.05; "**" indicates p≤0.01; "***" indicates p≤0.001; "****" indicates p≤0.0001.
[0039] Figure 5 It is a result comparison between an embodiment of the present application and a comparative example. Detailed implementation manners
[0040] For the convenience of those skilled in the art to understand, some terms appearing in this article are explained and described.
[0041] In this article, the singular forms of "a", "an" and "the" also include their plural forms unless the context otherwise indicates. Therefore, for example, "the known immune-related proteins" can be understood to include multiple known immune-related proteins.
[0042] In this text, unless otherwise specified, the terms "comprising", "including", or "containing" mean including the listed technical features, but do not exclude the inclusion of other technical features.
[0043] In this text, a "processor" is a component such as a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU) that can perform corresponding data processing as long as it meets the requirements of a device such as a computer.
[0044] The main technical solutions of the present invention are outlined as follows:
[0045] 1. Search relevant literature for known immune-related protein names as training data for exploring immune-related proteins. There are 35 known immune-related proteins found in the literature: IFNB1, IL1A, IL23R, TGFB1, C2, C3, C5, C6, C7, C9, CCL13, CCL2, CCL20, CCL28, CXCL1, CXCL5, CXCL6, CXCL8, IFNA2, IFNG, IL10, IL13, IL17A, IL1B, IL2, IL22, IL34, IL4, IL5, IL6, IL9, TNF, CXCL9, GDF15, and CSF1.
[0046] Collect the expression data of 1,305 plasma proteins from 171 healthy individuals as training data for immune status assessment.
[0047] 2. Use a small number of known immune-related protein names and combine a meta-learning graph convolutional network (Meta-GCN) with a protein interaction network to identify more key proteins highly related to the immune status. The input of the Meta-GCN model is the plasma protein names (converted to one-hot encoding) and their interaction relationships (interaction data) in the protein interaction network, corresponding to the "points" (nodes) and "edges" of the graph convolutional network respectively. The output of the Meta-GCN model is the probability that a certain plasma protein is an immune-related protein, and this probability value is used to determine whether it is a final immune-related protein.
[0048] Based on the above scheme, the SOMAscan platform was used to obtain the plasma protein expression data of healthy individuals, measuring 1,305 different plasma proteins and recording their respective relative fluorescence units (RFUs). These RFU values were processed by a logarithm base 10 transformation. Using the Meta-GCN model developed by the present invention, a total of 309 proteins highly related to immunity were identified and used as the input of the immune status assessment model.
[0049] 3. The present invention uses a variety of advanced machine learning models, taking the expression levels of the immune-related proteins identified in the above steps as input features, and the output is the immune status score.
[0050] According to the experiments of the present invention, among the currently 309 immune-related proteins, the computer model system must include at least more than 81 protein features to train and construct an effective and reliable immune status evaluation model.
[0051] When constructing the immune status score prediction model, a variety of machine learning algorithms are used, including Lasso regression, support vector machine (SVM), LightGBM, random forest, XGBoost, and decision tree. By comparing the performance of these models in terms of indicators such as mean absolute error (MAE), mean squared error (MSE), root mean squared error (RMSE), coefficient of determination R 2 and the Pearson correlation coefficient between the predicted value and the true value, etc., to determine the optimal model. The output result of this model is the immune status score, which is calculated based on the previous research of Jiyunuo on 19,105 healthy individuals. The formula is:
[0052] y = -6.58e -7 x 3 + 9.662e -5 x 2 - 5.468e -3 + 0.6177, where x is the age and y is the initial immune status score. It should be noted that according to the previous research statistics, the immune status score corresponding to each age is the initial immune status score. After training the model based on the expression data of immune-related proteins and the initial immune status score of healthy people, the final output based on the model is the final immune status score. The research objects are all healthy people, and it is stipulated that their immune status scores are above 60 points, and the scores in the range of 0 - 1 in the original research are mapped to the range of 60 - 95 points for easy understanding and use. Figure 1 The corresponding relationship curve between the original score and age, and the corresponding relationship curve between the transformed score and age are shown. Table 2 shows the comparison table of the transformed age and immune status score. Finally, by calculating the average value of the immune status scores predicted by the 4 types of machine learning models with the best performance of various evaluation indicators, it is used as the final immune status score of an individual.
[0053] Table 2. Reference table of actual age and immune status score
[0054]
[0055]
[0056] 4. Validate the model using an external dataset
[0057] To evaluate the effectiveness and reliability of the model, the model was validated using two external datasets. Both datasets contain data of healthy individuals and COVID-19 patients (infected with the novel coronavirus). The reliability of the immune status score of the model was verified by the correlation between the scores given by the model to them and their age, as well as the recovery status.
[0058] The solution of the present invention will be further explained below in conjunction with embodiments. Those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the embodiments regarding specific techniques or conditions, the techniques or conditions described in the literature in this field are followed.
[0059] The immune status assessment method based on plasma proteomics according to an embodiment of the present invention is analyzed and evaluated according to the following steps, and its process is as Figure 2 shown.
[0060] Step 1: Plasma proteome data of 171 healthy individuals were collected from the literature. Each healthy individual contains 1305 plasma proteins. The data processing in the article includes using the SOMAscan platform to collect the plasma protein expression levels of healthy people, and then performing a logarithm base 10 conversion on the protein relative fluorescence units (RFU).
[0061] Step 2: The Meta-Graph Convolutional Network (Meta-GCN) is a method that combines graph convolutional networks (GCNs) and meta-learning techniques, and the corresponding model is the Meta-GCN model. The Meta-GCN model constructed in the present invention aims to promote the prediction of immune-related plasma proteins through the protein-protein interaction network, and includes three key parts: adjacency matrix construction, GCN training, and meta-learning optimization. These three parts are interrelated, interact with each other, and are closely related. They will be elaborated below respectively, but when explaining one part, it often involves other parts, especially the GCN training and meta-learning optimization are intertwined.
[0062] In the study, 1305 plasma protein names were input into the STRING database (https: / / cn.string-db.org / ) for analysis, and 82,828 interaction data (interaction relationships) among 1207 plasma proteins were obtained. These interaction data were selected based on a confidence threshold of 0.4 (medium confidence), constituting the topological structure of a huge graph. The Meta-GCN model has two inputs, one is the feature matrix, and the other is the adjacency matrix. The two matrices represent the representation of plasma protein names and their relationships with other proteins in the protein interaction network, respectively. The feature matrix represents the nodes of the graph, and the adjacency matrix represents the edges of the graph, jointly representing a topological graph. The researchers used the data in the "combined_score" column (representing the confidence of the interaction between a protein and another protein) in the interaction data to construct the adjacency matrix of protein interactions. This "combined_score" incorporates multiple factors, such as chromosomal proximity, gene fusion, phylogenetic co-occurrence, homology, co-expression, experimental verification, database annotation, and text mining.
[0063] In this embodiment, the graph convolutional network (GCN) aggregates neighborhood information through two convolutional layers and gradually extracts features from the graph. The adjacency matrix is normalized by FirstOrderGCN. Since proteins are represented only by names and lack actual features, the identity matrix (with dimensions of the number of proteins × the number of proteins. The role of this identity matrix is to provide an initial feature representation for each protein node, where the feature of each node is a one-hot encoding, that is, only one 1 exists in the row corresponding to the node, and the rest are 0s. The one-hot encoding can provide a basic node discrimination ability for the model, enabling the model to learn the relationships and feature propagation among nodes in subsequent graph convolution operations.) is used as the feature matrix. In semi-supervised learning, a small number of labeled nodes (35 proteins known to be immune-related found in the literature are labeled as 1) and a large number of unlabeled nodes (other plasma proteins after deducting the aforementioned 35 known immune-related proteins from 1207 plasma proteins) are used for training. Through literature mining, the 35 immune-related proteins listed above are designated as positive labels, and 100 are randomly selected from the remaining plasma proteins as negative labels. The 1207 proteins and interaction data are used for training and prediction of the Meta-GCN model.
[0064] Meta - learning can enhance the adaptability of a model in different tasks or environments. Especially in the few - shot learning scenario, it improves the model performance by acquiring experience from a small number of samples. The meta - learning framework makes the model quickly adapt to new tasks by optimizing the model parameter initialization, and its process is divided into two stages: model training and testing. In the training stage, several nodes of each category (two categories, the class label of immune - related proteins is 1, and the class label of non - immune - related is 0) are randomly selected from the training set to form a support set, and the remaining nodes form a query set. Each meta - learning task is constructed in this way, and this process is repeated 20 times (i.e., 20 iterations below) to generate 20 tasks.
[0065] When using the Meta - GCN model to predict immune - related proteins in plasma, the model parameters are set as 50 training epochs, 20 iterations per training epoch, and the learning rate is 0.0001. During the training process, 35 immune - related proteins are designated as positive samples, and 100 randomly selected proteins are used as negative samples; the remaining proteins are classified into other categories. In each iteration of each training epoch, 20 random node (protein) selections are made. Each random selection includes 10 negative samples (randomly selected with replacement from the aforementioned 100 randomly selected protein negative samples) and 5 positive samples (randomly selected with replacement from the 35 immune - related proteins) to ensure that all known positive samples are covered. This method helps to fully represent positive and negative samples during the training process. In the testing stage, the pre - trained model is used to predict all proteins (nodes). The 100 proteins with the lowest prediction probabilities are relabeled as negative samples for retraining until the prediction results converge (i.e., the loss values of the model on the training set and the validation set tend to be stable and no longer change significantly) or reach the maximum specified number of rounds, completing a full training and prediction process. The maximum specified number of rounds refers to 50 training epochs, 20 iterations per training epoch, and 20 random node selections per iteration.
[0066] To improve the reliability of the results, the above - mentioned entire process (result convergence or reaching the maximum specified number of rounds) is repeated 100 times, and the median prediction probability of each protein is recorded as the final immune - related probability.
[0067] Step 3: Combine the aforementioned 35 known immune - related proteins and the proteins with a prediction probability exceeding 0.95 obtained by the above Meta - GCN model prediction, resulting in a total of 309 proteins, as shown in Table 3, which are used as the names of the input features for the subsequent machine - learning model.
[0068] Table 3. Immune - related proteins identified by the Meta - GCN model and their prediction probabilities
[0069]
[0070]
[0071] Step 4: Each sample uses the expression level of the immune-related protein (i.e., the value after the protein RFU in Step 1 is logarithmically transformed by 10) as the input feature, and the immune status score corresponding to the age of each sample (Table 2) is used as the label for model training. There are a total of 171 samples, and the training set and the test set are divided at a ratio of 8:2. The results of the test set are shown in Figure 3 In order to evaluate the prediction performance of the six algorithms, the Pearson correlation coefficient is used to measure the linear relationship between the predicted value and the assumed true value (Table 2), and multiple indicators such as the mean absolute error, mean square error, root mean square error, and coefficient of determination are calculated. Among them, the smaller the values of the mean absolute error, mean square error, and root mean square error, the better, and the larger the coefficient of determination (ranging from 0 to 1), the better, indicating that the model has a better fitting degree to the data. The results show (Appendix Figure 4 and Table 4) that considering the five evaluation indicators comprehensively, the Pearson correlation coefficients between the prediction results of Lasso, LightGBM, XGBoost, and random forest and the assumed true value all exceed 0.9, and the three error indicators are much smaller than those of decision tree and SVM, and the coefficient of determination is also much larger than that of decision tree and SVM. This indicates that the prediction performance of these algorithms in this study is significantly better than that of decision tree and SVM. Therefore, finally, the immune status scores predicted by the four models of Lasso, LightGBM, XGBoost, and random forest are selected, and their average value is taken as the final immune status score of the individual, and the trained model is saved for subsequent direct verification on other data sets.
[0072] Table 4 Results of various evaluation indicators of the test set
[0073] Pearson coefficient Mean absolute error Mean squared error Root mean squared error Coefficient of determination Decision tree 0.88 2.98 15.90 3.99 0.76 Lasso 0.90 2.82 12.95 3.60 0.81 LightGBM 0.90 2.83 12.93 3.60 0.81 Random forest 0.91 2.65 12.19 3.49 0.82 Support vector machine 0.85 4.60 47.32 6.88 0.30 XGBoost 0.92 2.48 10.30 3.21 0.85
[0074] Step 5: In order to evaluate the effectiveness and reliability of the model, two external data sets were used to verify the model.
[0075] First, for the collected COVID-19 dataset 1, data on 803 plasma proteins in each sample were measured using the TMTpro 16-plex platform. The individual samples included three groups: the healthy group (n = 35, age range 25 - 64 years), the acute infection group (n = 26, age range 25 - 67 years), and the post-acute group (n = 32, age range 19 - 69 years). During the model training process, since the dataset was sourced externally and did not cover all the aforementioned 309 immune-related proteins, the intersection of the immune-related proteins (the 309 obtained from the above model) and all the measured proteins (803) in the external dataset was first screened out, and 81 common features (81 proteins) were determined as the model input. Subsequently, the model was retrained using the healthy person dataset corresponding to these 81 common features, and the model was saved after training. If the validation set covered all the aforementioned 309 immune-related proteins, there was no need to retrain the model.
[0076] Next, the saved model was loaded to predict the immune status scores of the samples in COVID-19 dataset 1. Finally, the average of the prediction results of the four models was taken as the final immune status score of each sample in the independent validation set. To verify the effectiveness of the model in scoring the immune status of the healthy population, first, a correlation analysis was performed between the immune status scores and age of the healthy population in COVID-19 dataset 1. The results are as Figure 4 shown in a of []. The correlation coefficient was r = -0.444, and the p-value was less than 0.01, indicating a significant negative correlation between the immune status score and age, which proved the effectiveness of the model in evaluating the immune status of the healthy population.
[0077] In COVID-19 dataset 2, blood samples were collected within 24 hours (day 0) and on day 14 (negative nucleic acid test) after confirmed COVID-19 infection for analysis, including the NPX values of 1459 proteins that passed quality control. NPX is a relative protein quantification unit on a log2 scale. All treatments started on the day of diagnosis, and all patients had negative PCR test results on day 14. COVID-19 dataset 2 consisted of individuals with mild to moderate symptoms due to COVID-19, and none of these patients required hospitalization. During the model training process, first, the intersection of the immune-related proteins (the 309 obtained from the above model) and all the measured proteins in the external dataset was extracted, and 185 common features (185 proteins) were determined as the model input.
[0078] Next, use these 185 common features and the healthy person training set to retrain the model, and save the model after training is completed. Subsequently, load the saved model to predict the immune status scores of the samples in COVID-19 dataset 2, and finally take the average of the prediction results of the four models as the final immune status scores of each sample in the independent validation set. Calculate the immune status scores of the population at the initial stage of infection (day 0) and the recovery period (day 14). Day 0 and day 14 represent day 0 and day 14 of infection of the same batch of 50 samples respectively, and there is a one-to-one correspondence between them. Further analysis found that there was a significant negative correlation between the immune status score on day 14 and age, with a correlation coefficient of r = -0.608 and a p-value less than 0.0001, which was statistically significant, as shown in Figure 4 panel b. This finding was consistent with the negative correlation between the immune status score and age in the healthy population in COVID-19 dataset 1, further verifying the stability of the immune status assessment model.
[0079] To evaluate the differences in different immune status scores, the application of this model in the immune status assessment of COVID-19 datasets was further verified. In COVID-19 dataset 1, the immune status scores of three groups, namely healthy people, patients in the acute infection period, and patients in the late acute infection period, were compared. In the Mann-Whitney U test, as shown in Figure 4 panel c, although there was no statistically significant difference between the acute infection period and the late acute infection period, the median value of the immune status score in the late acute infection group was slightly higher than that in the acute infection group. Compared with the healthy control group, the immune status score of the healthy group was significantly higher than that of the late acute infection group (p < 0.01), and the difference between the acute infection group and the healthy group was even more significant (p < 0.001). In COVID-19 dataset 2, the immune status scores of the population at the initial stage of infection (day 0) and the recovery period (day 14) were calculated. Through the Mann-Whitney U test analysis, the results were as shown in Figure 4 panel d. The immune status score of the samples on day 14 was significantly higher than that on day 0 (p < 0.0001), indicating that the immune status had been significantly improved after 14 days of recovery.
[0080] Comparative analysis of routine blood immune status assessment methods
[0081] The method for evaluating the immune status of blood routine can be seen in the specific implementation manner of CN202310969500.5. To compare the effects of the two methods in an independent test set, a blood routine dataset of a group of COVID-19 patients was used to test the accuracy of the method for evaluating the immune status of blood routine. All the subjects in the blood routine dataset of COVID-19 patients used in the study on the method for evaluating the immune status of blood routine were COVID-19 patients, and a total of 961 samples were included. The age range of the subjects in the dataset was from 23 years old to 89 years old. In terms of gender distribution, there were 573 male subjects and 388 female subjects. Each sample represented the relevant information recorded for a subject for the first time after the onset of the disease, which was consistent with the status of the samples in the two COVID-19 data in the present invention. They were all in the early stage of COVID-19 infection, and the disease states were highly consistent.
[0082] The results of the method for evaluating the immune status of blood routine are as Figure 5 shown in a of. The correlation between the immune status score and age was not significant (p > 0.05), indicating that the blood routine method could not reliably evaluate the immune status of COVID-19 patients. However, the research results of the present invention are as Figure 5 shown in b and c of. They were the correlations between the immune status scores and age of the patients in the acute infection period in COVID-19 dataset 1 and the patients within 24 hours after positive nucleic acid test in dataset 2, respectively, and both showed a significant negative correlation (p < 0.05), indicating that the method of the present invention could reliably evaluate the immune status of COVID-19 patients. To make the immune status scores of the blood routine patent method better comparable with those of the plasma protein patent method, the immune status scores of the plasma protein method were normalized from 60 - 95 points to between 0 and 1 in this comparative experiment.
[0083] It can be seen from this that the present invention can not only evaluate the immune status of healthy people, but also reliably evaluate the immune status of disease patients. However, the applicable range of the method for evaluating the immune status of blood routine is limited to healthy people. When using blood routine to evaluate the immune status, there are fewer blood routine test items, and it is greatly affected by the abnormal status of the tested subjects. Therefore, it is limited to evaluating the immune status of healthy people. The present invention uses hundreds of immune-related plasma proteins to evaluate the immune status. The influence of individual test items on the results is small, and the developed artificial intelligence algorithm has strong robustness.
[0084] In summary, the computerized immune status assessment system constructed by the present invention can stably quantify the immune status of an individual. The system can not only accurately output the specific score of an individual's immune status, but also clarify the individual's immune level position in the peer group according to the preset age-immune status score mapping relationship. Moreover, by comparing the immune status score data at different infection stages, the system can effectively monitor the change trend of the immune status before and after an individual's infection, and intuitively present the process of the immune status score gradually rising as the body recovers from the infection, providing strong technical support for the dynamic monitoring of the immune status and health management.
[0085] Some embodiments include one or more processors of one or more computing devices (such as a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)), wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the execution of any of the methods described herein. Some embodiments also include one or more non-transitory computer-readable storage media that store computer instructions executable by the one or more processors to execute any of the methods described herein.
[0086] In some embodiments, the immune status assessment method based on plasma proteomics is a method for non-diagnostic and therapeutic purposes.
[0087] Specific examples are used in this article to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. It is possible to make changes and improvements to the present invention without exceeding the concept and scope defined by the claims. In summary, the content of the embodiments of this specification should not be construed as a limitation to the present invention.
Claims
1. An immune status assessment method based on plasma proteomics, characterized in that, The method is implemented by one or more processors and includes: Using the Meta-GCN model to identify immune-related proteins in plasma proteins. The identification process includes taking interaction data with medium confidence from several interaction data among several plasma proteins of healthy people to form a topological structure diagram. The nodes of the topological structure diagram are plasma proteins, and the nodes form an identity matrix in the form of one-hot encoding. The edges of the topological structure diagram are the confidence levels of pairwise protein interactions, and the edges form an adjacency matrix. Inputting the identity matrix and the adjacency matrix into the Meta-GCN model for meta-learning. Each meta-learning uses the known immune-related proteins in plasma proteins as positive labels, and randomly selects proteins with 2 to 4 times the number of positive labels from the remaining plasma proteins as negative labels for learning, and takes the predicted probability that the plasma protein is an immune-related protein as the output. The meta-learning process is repeated multiple times until the prediction results of all plasma proteins in the topological structure diagram converge or the meta-learning reaches the maximum specified number of rounds. Taking the known immune-related proteins and the plasma proteins with predicted probabilities higher than the threshold as the immune-related proteins identified by the Meta-GCN model. There are m immune-related proteins identified by the Meta-GCN model; Using an immune status score prediction model for prediction. The prediction process includes training the prediction model using n immune-related proteins identified by the Meta-GCN model, where 81 ≤ n ≤ m. Using the expression level feature data of the n immune-related proteins identified by the Meta-GCN model of a group of healthy people to be tested as input features, and using the initial immune status score corresponding to the age of this group of healthy people to be tested as a label for model training. The trained prediction model can output its immune status score according to the expression level feature data of the n immune-related proteins identified by the Meta-GCN model of the input person to be tested. The immune status score prediction model is selected from one or more of Lasso, LightGBM, XGBoost, and random forest. When there are multiple models, the immune status score is the average of the scores of multiple models. The calculation formula for the initial immune status score is: y = -6.58e -7 x 3 + 9.662e -5 x 2 - 5.468e -3 + 0.6177, where x is the age and y is the initial immune status score.
2. The method according to claim 1, wherein The expression level feature data of the immune-related proteins identified by the Meta-GCN model refers to the result after the protein relative fluorescence unit is logarithmically transformed by base 10.
3. The method according to claim 1, characterized in that The several interaction data among the several plasma proteins are obtained by inputting the names of the several plasma proteins into the STRING database for analysis; the medium confidence is 0.4; the adjacency matrix is input into the Meta-GCN model for meta-learning in the data form obtained after being normalized by FirstOrderGCN; The threshold is 0.
95.
4. The method according to claim 1, characterized in that, The maximum specified number of rounds is 50 ± 5 training cycles, with 20 ± 2 iterations per training cycle and 20 ± 2 random node selections per iteration; the learning rate of meta-learning is 0.0001, or 0.0005, or 0.00001; each random node selection includes 10 ± 2 negative-label proteins and 5 ± 1 positive-label proteins; the 10 ± 2 negative-label proteins are randomly selected with replacement from 100 ± 10 negative-label proteins, and the 5 ± 1 positive-label proteins are randomly selected with replacement from the known immune-related proteins; the 100 ± 10 negative-label proteins are randomly selected from plasma proteins excluding the known immune-related proteins; the number of known immune-related proteins is 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40, and 30, 31, 32, 33, 34, or 35 of them are selected from IFNB1, IL1A, IL23R, TGFB1, C2, C3, C5, C6, C7, C9, CCL13, CCL2, CCL20, CCL28, CXCL1, CXCL5, CXCL6, CXCL8, IFNA2, IFNG, IL10, IL13, IL17A, IL1B, IL2, IL22, IL34, IL4, IL5, IL6, IL9, TNF, CXCL9, GDF15, and CSF1; the number of plasma proteins of healthy individuals is 1000 - 1305, or 1100 - 1300, or 1150 - 1250, or 1180 - 1240, or 1190 - 1230, or 1195 - 1220, or 1195 - 1215, or 1200 - 1210, or 1205 - 1210, or 1206, or 1207, or 1208, or 1209.
5. The method according to claim 1, characterized in that The predicted probability of each protein is the median of 100 ± 5 prediction probabilities corresponding to the convergence of 100 ± 5 prediction results or the maximum specified number of rounds reached by meta-learning; the y value is mapped to the range of 60 - 95 points and used as a label for the training of the prediction model.
6. An immune status assessment method based on plasma proteomics, characterized in that, The method is implemented by one or more processors and includes: Predict using an immune status score prediction model. The prediction process includes training the prediction model using immune-related proteins identified by n Meta-GCN models, where 81 ≤ n ≤ 309. Using the expression level feature data of the immune-related proteins identified by the n Meta-GCN models of a group of healthy persons to be tested as input features, and using the initial immune status score corresponding to the age of this group of healthy persons to be tested as a label for model training. The trained prediction model can output its immune status score according to the expression level feature data of the immune-related proteins identified by the n Meta-GCN models of the person to be tested. The immune status score prediction model is selected from one or more of Lasso, LightGBM, XGBoost, and random forest. When there are multiple models, the immune status score is the average of the scores given by multiple models. The formula for calculating the initial immune status score is: y = -6.58e -7 x 3 + 9.662e -5 x 2 - 5.468e -3 + 0.6177, where x is the age and y is the initial immune status score; The immune-related proteins identified by the Meta-GCN model are shown in Table 1: Table 1. Immune-related proteins identified by the Meta-GCN model 7. The method according to claim 6, wherein The expression level feature data of the immune-related proteins identified by the Meta-GCN model refers to the result after the protein relative fluorescence unit is logarithmically transformed by base 10.
8. A computer program comprising instructions that, when executed by one or more processors of a computing system, cause the computing system to perform the method according to any one of claims 1 to 7.
9. A computing system configured to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing instructions that can be executed by one or more processors of a computing system to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and system for determining immune state of object based on blood routine data
CN116978566A