Proteomics-based individual future health state prediction method and device and medium
By building a health assessment network based on twin networks and multi-layer perceptrons, combined with proteomic data, the problems of lack of longitudinal data and single outcome prediction in individual future health status prediction are solved, and multi-dimensional, full-view assessment and risk prediction of individual future health status are achieved.
Patent Information
- Application Number
- CN202311503648.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-11-13
AI Technical Summary
The prior art has the limitations of longitudinal data and single outcome prediction in the prediction of individual future health status, and it is impossible to evaluate the individual's future health level in a panoramic manner.
By building a shared network based on the twin network framework and a backbone network of multi-layer perceptron method, combined with proteomic data, a health assessment network is built to achieve multi-dimensional and full-view assessment of individuals' future health status.
The risk prediction of a total of 45 key health outcomes in an individual's future is achieved, providing a comprehensive and meticulous assessment of individual's future health status, and overcoming the limitations of lack of longitudinal data and single outcome prediction.
Smart Images

Figure CN119993470A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and biomedical engineering, and in particular to a method, device and medium for predicting the future health status of an individual based on proteomics. Background Art
[0002] Early health level prediction is of great significance for individual health management and treatment. By determining the individual's potential disease risk and health assessment, it is helpful to take targeted intervention treatments for the individual early and provide the necessary time window to maximize the protection and improvement of the individual's health level. The current technical routes and methods have the following limitations: First, most studies are based on cross-sectional data, and the lack of longitudinal data greatly limits the application of research and development models in real application scenarios; in addition, an individual's future illness and death events are important evaluation indicators of their future health status. However, many current technologies only focus on predicting a single outcome (such as death, specific diseases, aging age, cognitive or motor function decline, etc.), and cannot comprehensively evaluate future health levels.
[0003] Proteins are important functional molecules in cells. As the final product of gene expression, proteomics data can directly reflect the functions of cells or tissues, reveal the internal mechanisms of biological systems, and better reflect the actual status of individuals. Previous studies have determined that abnormal changes in the expression levels of certain specific proteins can be associated with the occurrence and progression of diseases such as cancer, cardiovascular disease, and diabetes, and can serve as potential drug targets and biomarkers. Therefore, the development of a health assessment model based on proteomics data has important clinical significance and application prospects. However, given the complexity of proteomics data and the massive growth in data volume, there is no research on proteomics data for predicting health assessment models with multiple diseases and death as outcomes. Summary of the invention
[0004] The purpose of the present invention is to provide a method, device and medium for predicting the future health status of an individual based on proteomics.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] A method for predicting the future health status of an individual based on proteomics, comprising the following steps:
[0007] Step 1) obtaining proteomics data and preprocessing;
[0008] Step 2) Shared network construction: Construct a neural network to predict individual past and future comorbidity based on the twin network framework;
[0009] Step 3) Backbone network construction: construct a health-specific outcome prediction neural network based on the multi-layer perceptron method;
[0010] Step 4) The health assessment network integrates the shared network and the backbone network architecture, and extracts features to fuse and fine-tune them in the latent space. Through further training, the fine-tuned network parameters are updated to output the future risk assessment probabilities of multiple health-specific outcomes.
[0011] The input of the shared network is the vectorized value of the preprocessed plasma protein level, and the output is the number of past disease categories and the number of future disease categories. The method for determining the number of past disease categories and the number of future disease categories is: the time point for collecting plasma protein is set as the baseline time, and the statistical number of various diseases and death categories in the individual electronic medical record information is determined to be divided into before-baseline and after-baseline according to the baseline time point.
[0012] The shared network includes two branch networks with exactly the same architecture. Each encoding branch network includes four fully connected layers and an output layer, wherein each fully connected layer contains 512, 256, 128 and 64 nodes respectively. The fully connected layers are connected through the ReLu activation function, and the activation function between the last fully connected layer and the output layer adopts a linear function.
[0013] The input of the backbone network is the vectorized value of the preprocessed plasma protein level, and the output is the health assessment information after the baseline plasma protein collection time point, which covers the future risk probability estimation of multiple health-specific outcomes.
[0014] The encoder of the backbone network includes four fully connected layers and an output layer, wherein each fully connected layer contains 512, 256, 128 and 64 nodes respectively, and the fully connected layers are connected through the ReLu activation function. The activation function between the last fully connected layer and the output layer adopts the Sigmoid function.
[0015] The input of the health assessment network is the vector value of plasma protein level, and the output is the same as the backbone network, which is the risk probability estimation of multiple future health-specific outcomes.
[0016] The health assessment network fuses the processed protein features in the two branch encoders of past and future comorbidity outcomes in the shared network with the protein features extracted from the encoding part of the backbone network.
[0017] The health assessment network merges the 64 pre-trained weights of the last fully connected layer of the two encoder branches of the shared network with the 64 pre-trained weights of the last fully connected layer of the backbone network in series; the merged output is sequentially input into the fully connected layers with 128 and 64 nodes respectively and then output through the output layer, wherein the series merged output and the fully connected layers, and the fully connected layers are connected using the ReLu activation function, and the activation function between the last fully connected layer and the output layer uses the Sigmoid function.
[0018] A device for predicting individual future health status based on proteomics comprises a memory, a processor, and a program stored in the memory, wherein the processor implements the method described above when executing the program.
[0019] A storage medium stores a program, which implements the method described above when executed.
[0020] Compared with the prior art, the present invention has the following beneficial effects:
[0021] (1) Previous proteomics-based prediction studies were mainly based on cross-sectional data. The lack of longitudinal clinical data greatly limited the application of the research and development model in real application scenarios. The present invention uses the data of participants' health outcome tracking and follow-up for more than 14 years as input. This longitudinal cohort data enables the model to simulate real application scenarios when it is constructed. The protein and clinical indicators obtained at the baseline are used to build a prediction model to evaluate future outcomes.
[0022] (2) An individual's future illness and death events are important evaluation indicators of his or her future health status. However, most current models only focus on a single outcome. The present invention classifies and organizes the electronic medical records of participants in primary care records, hospital hospitalization records, and death registries during the follow-up period. The electronic medical records cover a variety of clinical health outcomes, which enables the present invention to systematically evaluate future health from a multi-dimensional and full perspective.
[0023] (3) The present invention achieves an estimation of the individual's overall disease burden and a risk prediction of a total of 45 key health outcomes in the future by constructing two important components: a shared network and a health-specific outcome network, thereby achieving a comprehensive and detailed assessment of the individual's future health status. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the method of the present invention;
[0025] Figure 2 Evaluate network model structure for health;
[0026] Figure 3Forest plots showing the predictive performance of different feature sets for 45 health-specific outcomes. DETAILED DESCRIPTION
[0027] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0028] Example 1
[0029] The present invention processes and handles the proteomics data of 1461 plasma proteins, and constructs a future health assessment model based on a deep learning method, wherein the health assessment model integrates two parts: a twin sharing network and a specific outcome prediction network; the health assessment covers 45 longitudinal future health-specific outcomes, including all-cause mortality, 4 disease-specific deaths, 26 emerging diseases and 14 major disease categories (including infectious, blood, endocrine, mental, nervous, circulatory, respiratory, digestive, skin, musculoskeletal and genitourinary system diseases and cancer).
[0030] Specifically, Figure 1 As shown, the following steps are included:
[0031] Step 1) Obtain proteomics data and perform preprocessing.
[0032] Proteomics data acquisition and processing: Blood samples were collected in EDTA (9 ml) vacuum containers and divided into 850 μl EDTA plasma, buffy coat and red blood cell aliquots. The plasma samples to be analyzed were stored in a refrigerator at -80 °C, and the proximity extension assay was combined with next-generation sequencing to measure 1461 unique proteins in parallel. The proteomics data contained proteins measured on four panels of cardiac metabolism, inflammation, neurology and oncology proteins. Protein expression measurement data were obtained based on the standardized measurement process of OLINK in Sweden.
[0033] The proteomic expression data (Normalized Protein eXpression, NPX) preprocessing for modeling used a normalization process as follows: the counts of each sample and each assay were divided by the counts of the extended control, and the ratio was further logarithmically transformed. The intra- and inter-plate differences were minimized by considering the median of the normalized counts of the extended control, the batch-specific NPX median, and the differences in the assay-specific NPX median for each batch.
[0034] Acquisition of health outcomes: The outcomes of the training data are divided into three parts. The two coding branches in the shared network correspond to the number of previous disease categories and the number of future disease categories. The number of health categories is defined as 14 health categories (values from 0 to 14), including infectious diseases, cancer, blood immunity, endocrine, mental behavior, nervous system, eye, ear, circulation, respiratory, digestive, skin, musculoskeletal, and urogenital); the past and future are defined as before and after the individual collects plasma proteomics data. Secondly, the outcomes in the backbone network are defined as the 45 outcomes corresponding to the ICD-10 code (including 14 disease categories, 26 specific diseases, all-cause mortality, and 4 specific deaths).
[0035] Step 2) Shared network construction: Construct a neural network to predict individual past and future comorbidities based on the twin network framework.
[0036] The input of the shared network is 1461 preprocessed vectorized values of plasma protein levels, and the output (model training outcome) is the number of previous disease categories and the number of future disease categories. The method for determining the number of previous disease categories and the number of future disease categories is as follows: the time point for collecting plasma protein is set as the baseline time, and the 14 disease and death categories in the individual electronic medical record information (including infectious diseases, cancer, blood immunity, endocrine, mental behavior, nervous system, eye, ear, circulation, respiratory, digestive, skin, musculoskeletal, urogenital and all-cause death) are determined and divided into the statistical number (0-14) before and after the baseline time point.
[0037] like Figure 2 As shown in the figure, the shared network contains two branch networks with exactly the same architecture. Each encoding branch network includes four fully connected layers and an output layer. Each fully connected layer contains 512, 256, 128 and 64 nodes respectively. The fully connected layers are connected through the ReLu activation function. The activation function between the last fully connected layer and the output layer adopts a linear function.
[0038] Shared network training: First, the encoders for past illness and future illness were trained separately. Then, the weights of the two pre-trained encoders were transferred to the two branches of the shared network and the weights were fine-tuned and updated. The Adam optimizer was used for both separate training and combined training, and the learning rate was set to 1×10 -5 , the training batch size is 128, the number of iterations is 1000, and in order to reduce the risk of overfitting, the iteration node for early termination of training is defined as the validation set loss function has not decreased for 25 consecutive iterations.
[0039] Step 3) Backbone network construction: Construct a health-specific outcome prediction neural network based on the multi-layer perceptron method.
[0040] The input of the backbone network is 1461 preprocessed vectorized values of plasma protein levels, and the output (model training outcome) is the health assessment information after the baseline plasma protein collection time point, which covers 45 health-specific outcomes, including all-cause mortality, 4 disease-specific deaths, 26 emerging diseases and 14 major disease categories; the output information is the risk probability estimate of 45 specific health outcomes.
[0041] like Figure 2 As shown in the figure, the encoder of the backbone network includes four fully connected layers and an output layer, where each fully connected layer contains 512, 256, 128 and 64 nodes respectively. The fully connected layers are connected through the ReLu activation function, and the activation function between the last fully connected layer and the output layer adopts the Sigmoid function.
[0042] Backbone network training: Adam optimizer is used, and the learning rate is set to 1×10 -5 , the training batch size is 128, the number of iterations is 1000, and in order to reduce the risk of overfitting, the iteration node for early termination of training is defined as the validation set loss function has not decreased for 25 consecutive iterations.
[0043] Step 4) The health assessment network integrates the shared network and the backbone network architecture, and extracts features to fuse and fine-tune them in the latent space. Through further training, the fine-tuned network parameters are updated to output the future risk assessment probabilities of multiple health-specific outcomes.
[0044] The health assessment network includes a shared network and a backbone network, so the same 1461 vector values of plasma protein levels will be input into the health assessment network as two input information at the same time, and the output is the same as the backbone network, which is the risk probability estimate of 45 future health-specific outcomes.
[0045] The health assessment network fuses the processed protein features in the two branch encoders of past and future comorbidity outcomes in the shared network with the protein features extracted from the encoding part of the trunk network. Specifically, Figure 2 As shown, the 64 pre-trained weights of the last fully connected layer of the two encoder branches of the shared network and the 64 pre-trained weights of the last fully connected layer of the backbone network are merged in series (i.e., 64+64+64=192 neural node weight values); the merged output is sequentially input into the fully connected layers with 128 and 64 nodes respectively and then output through the output layer, wherein the serial merged output and the fully connected layers, and the fully connected layers are connected with each other using the ReLu activation function, and the activation function between the last fully connected layer and the output layer uses the Sigmoid function.
[0046] Health assessment model training: The shared network and the backbone network have obtained the optimal weight parameters through pre-training. After the health assessment network is built, all pre-training parameters are first frozen during training, and only the last 128 and 64 nodes of the fully connected layer are trained; after the initial training is completed, all the backbone network weight parameters are unfrozen and fine-tuned together. Note that the shared network weight parameters remain frozen in this step. The Adam optimizer is used for training, and the learning rate is set to 1×10 -5 , the training batch size is 128, the number of iterations is 1000, and the iteration node for early termination of training is defined as the validation set loss function has not decreased for 25 consecutive iterations.
[0047] Figure 3 The forest plot shows the prediction performance of different feature sets for 45 health-specific outcomes. Prediction performance of proteomics data (leftmost column) and prediction performance of proteomics data + clinical indicators (rightmost column). To compare the predictive performance of proteomics data, different feature sets were compared, including age + gender, blood indicators (25 types), and clinical indicators (54 types, including demographic information, blood indicators, lifestyle, medical history, medication history, etc.). In each comparison, the vertical line represents the prediction performance of proteomics data alone, the dot represents the prediction performance of the respective feature set alone, and the triangle represents the prediction performance of the respective feature set + proteomics data.
[0048] This embodiment uses the Cox proportional hazard regression model for the predicted risk probability output by the health assessment network model, and calculates the Hazard ratio (HR) and the significant p value. Table 1 reports the HR values of the output risk probability under the correction of different covariate feature sets, and it can be seen that the corresponding p values of the HR values are statistically significant.
[0049]
[0050]
[0051]
[0052] Example 2
[0053] This embodiment provides a device for predicting the future health status of an individual based on proteomics, including a memory, a processor, and a program stored in the memory, and the processor implements the method described above when executing the program.
[0054] Example 3
[0055] This embodiment provides a storage medium on which a program is stored. When the program is executed, the method described above is implemented.
[0056] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0057] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.
Claims
1. A method for predicting individual future health status based on proteomics, characterized in that: The following steps are involved: Step 1) obtaining proteomics data and preprocessing; Step 2) Shared network construction: Construct a neural network to predict individual past and future comorbidity based on the twin network framework; Step 3) Backbone network construction: construct a health-specific outcome prediction neural network based on the multi-layer perceptron method; Step 4) The health assessment network integrates the shared network and the backbone network architecture, and extracts features to fuse and fine-tune them in the latent space. Through further training, the fine-tuned network parameters are updated to output the future risk assessment probabilities of multiple health-specific outcomes.
2. A method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The input of the shared network is the vectorized value of the preprocessed plasma protein level, and the output is the number of past disease categories and the number of future disease categories. The method for determining the number of past disease categories and the number of future disease categories is: the time point for collecting plasma protein is set as the baseline time, and the statistical number of various diseases and death categories in the individual electronic medical record information is determined to be divided into before-baseline and after-baseline according to the baseline time point.
3. A method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The shared network includes two branch networks with exactly the same architecture. Each encoding branch network includes four fully connected layers and an output layer, wherein each fully connected layer contains 512, 256, 128 and 64 nodes respectively. The fully connected layers are connected through the ReLu activation function, and the activation function between the last fully connected layer and the output layer adopts a linear function.
4. A method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The input of the backbone network is the vectorized value of the preprocessed plasma protein level, and the output is the health assessment information after the baseline plasma protein collection time point, which covers the future risk probability estimation of multiple health-specific outcomes.
5. The method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The encoder of the backbone network includes four fully connected layers and an output layer, wherein each fully connected layer contains 512, 256, 128 and 64 nodes respectively, and the fully connected layers are connected through the ReLu activation function. The activation function between the last fully connected layer and the output layer adopts the Sigmoid function.
6. A method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The input of the health assessment network is the vector value of plasma protein level, and the output is the same as the backbone network, which is the risk probability estimation of multiple future health-specific outcomes.
7. The method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The health assessment network fuses the processed protein features in the two branch encoders of past and future comorbidity outcomes in the shared network with the protein features extracted from the encoding part of the backbone network.
8. The method for predicting individual future health status based on proteomics according to claim 1, characterized in that: The health assessment network merges the 64 pre-trained weights of the last fully connected layer of the two encoder branches of the shared network with the 64 pre-trained weights of the last fully connected layer of the backbone network in series; the merged output is sequentially input into the fully connected layers with 128 and 64 nodes respectively and then output through the output layer, wherein the series merged output and the fully connected layers, and the fully connected layers are connected using the ReLu activation function, and the activation function between the last fully connected layer and the output layer uses the Sigmoid function.
9. A device for predicting individual future health status based on proteomics, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Multiplex PCR primer sets for identification of fowl adenovirus serotype and uses thereof
KR102596123B1
System and method for providing personalized health data
US20200185073A1
Clinical omics data processing method and apparatus based on graph neural network, device and medium
US20230028046A1
Cited By
Health data prediction method and device based on protein expression data, equipment and storage medium
CN120808895A