Apparatus for acute altitude stress (AMS) susceptibility grouping
By receiving and analyzing data on specific protein expression levels in subject samples, using machine learning models to perform susceptibility grouping for acute altitude sickness, solving the problem of difficulty in distinguishing and preventing AMS in the prior art, and achieving accurate grouping and prevention efficiency of susceptible populations.
Patent Information
- Application Number
- CN202510055508.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The prior art is difficult to effectively distinguish and prevent acute altitude sickness (AMS), and the lack of reliable objective distinction indicators has increased the difficulty of prevention and treatment.
By receiving data on six protein expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTNS246 and VPS4A S95 in the subject sample, input the acute altitude sickness susceptibility grouping model, and using machine learning technology (such as the random forest model) for training and prediction, output whether the subject belongs to the acute altitude sickness susceptibility population.
Accurate grouping of people susceptible to acute altitude sickness has been achieved, which has reduced the workload of medical staff and improved the prevention and treatment efficiency of AMS.
Smart Images

Figure CN120015309A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of medical care informatics, and in particular relates to a device for acute mountain sickness (AMS) susceptibility grouping. Background Art
[0002] Acute mountain sickness (AMS), also known as acute mountain sickness, generally refers to a series of clinical syndromes such as dyspnea, nausea, dizziness, palpitations, and sleep disorders caused by the human body's inability to adapt to low air pressure and low oxygen within a few hours to a few days when one quickly enters a plateau above 2,500 meters above sea level from a plain or from a plateau to a higher altitude area, causing compensatory dysfunction.
[0003] Every year, millions of people need to quickly enter the plateau for work, tourism, sports, rescue and other reasons. Low pressure and hypoxia are obvious environmental characteristics that threaten the health of people in plateau areas. When plain people who have not adapted to the environment rush to the plateau above 2,500 meters, AMS often occurs, mainly headache, accompanied by clinical symptoms such as dizziness, insomnia, and fatigue. Although most AMS cases are not fatal, they have an acute onset and progress rapidly. If not intervened in time, it may develop into severe altitude diseases such as high-altitude pulmonary edema and high-altitude cerebral edema, and even death. AMS is the most common altitude disease to date. Epidemiological data show that the incidence of AMS is about 9% to 75% after plain people rush into plateau areas. With the increase in altitude, the incidence and mortality of AMS also increase, seriously affecting the physical health and work ability of people who quickly enter the plateau. The risk of AMS depends on personal susceptibility, altitude and climbing speed. At present, there is no reliable objective differentiation indicator for AMS, which to a certain extent increases the difficulty of AMS prevention and treatment. Summary of the invention
[0004] One of the technical problems to be solved by the present invention is how to group the susceptibility of people to acute mountain sickness. The technical problems to be solved are not limited to the technical subjects described, and those skilled in the art can clearly understand other technical subjects not mentioned in this article through the following description.
[0005] To solve the above technical problem, the present invention first provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps:
[0006] S1. Data reception: Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTNS246 and VPS4A S95, in the samples of the subjects;
[0007] S2. Data processing: inputting the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model;
[0008] S3. Output result: output whether the subject belongs to the group susceptible to acute mountain sickness through the acute mountain sickness susceptibility grouping model.
[0009] The plateau acclimatizer may be a person who has no symptoms after rapid exposure to plateau.
[0010] The six proteins (RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95) are all erythrocyte membrane proteins. Among them:
[0011] The GenBank accession number of RAB2B protein is: 84932, and the UniProt annotation number is: Q8WUD1.
[0012] The GenBank accession number of VPS4A protein is: 27183, and the UniProt annotation number is: Q9UN37.
[0013] The GenBank accession number of SPTB protein is: 6710, and the UniProt annotation number is: P11277.
[0014] The GenBank accession number of ANP32B protein is: 10541, and the UniProt annotation number is: Q92688.
[0015] The GenBank accession number of DMTN protein is: 2039, and the UniProt annotation number is: Q08495.
[0016] The GenBank accession number of VPS4A protein is: 27183, and the UniProt annotation number is: Q9UN37.
[0017] SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95 all indicated phosphorylated proteins.
[0018] SPTB T1387 indicates SPTB protein phosphorylated at threonine (T) 1387.
[0019] ANP32B T244 indicates ANP32B protein phosphorylated at threonine (T) 244.
[0020] DMTN S246 indicates DMTN protein phosphorylated at serine (S) 246.
[0021] VPS4A S95 indicates VPS4A protein phosphorylated at serine (S) 95.
[0022] Furthermore, the subject sample may be a peripheral blood red blood cell membrane sample of the subject.
[0023] Furthermore, the acute mountain sickness susceptibility clustering model can be selected from at least one of a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model and a hidden Markov model.
[0024] Furthermore, the acute mountain sickness susceptibility clustering model may be a random forest model.
[0025] Furthermore, the hyperparameters of the random forest model may be as follows: the number of random forest trees may be 500.
[0026] The present invention also provides a device for acute altitude sickness susceptibility grouping, the device comprising a data receiving module, a data processing module and a result output module;
[0027] The data receiving module can be used to receive data on the expression levels of six proteins, namely, RAB2B, VPS4A, SPTB T1387, ANP32BT244, DMTN S246 and VPS4A S95, in the subject sample;
[0028] The data processing module can be used to input the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model;
[0029] The result output module can be used to output whether the subject belongs to the acute mountain sickness susceptible population through the acute mountain sickness susceptibility grouping model.
[0030] The data processing module can analyze or interpret the data based on statistical methods (such as mean, median, mode, standard deviation, variance, percentage, etc.), machine learning methods, or deep learning methods (such as convolutional neural networks, recurrent neural networks, graph neural networks, graph convolutional networks, etc.).
[0031] Furthermore, the data processing module includes a machine learning model submodule, the input of the machine learning model is the data of the expression levels of the six proteins described in this article, and the output is whether the person belongs to the group susceptible to acute mountain sickness.
[0032] The data processing module can be used to perform the following operations: input the data received by the data receiving module into the machine learning model, obtain an output value through the machine learning model, and thus provide an indication of whether the subject belongs to a group of people susceptible to acute mountain sickness.
[0033] The output value can be a class label of the sample, such as {0, 1} or {positive, negative}, etc. The output value can be mapped to a numerical value by mapping "positive" to 1 and mapping "negative" to 0. The output value can also be a probability value or a real value. When the output value is a probability value or a real value, the probability value or the real value can be compared with a preset threshold value, thereby providing an indication of whether the subject belongs to a group of people susceptible to acute mountain sickness, for example, when the probability value or the real value is greater than or equal to the preset threshold value, it indicates that the acute mountain sickness is positive (belongs to the group of people susceptible to acute mountain sickness), and when the probability value or the real value is less than the preset threshold value, it indicates that the acute mountain sickness is negative (does not belong to the group of people susceptible to acute mountain sickness).
[0034] The data processing module may further include a training submodule for training the machine learning model. The machine learning model may be trained based on a labeled data set to optimize its accuracy and / or reliability in clustering people susceptible to acute mountain sickness.
[0035] The data processing module may also include a performance evaluation submodule, which can be used to evaluate the accuracy and / or reliability of the machine learning model, such as drawing an ROC curve and calculating the AUC value of the machine learning model. The AUC value is used as an indicator to measure model performance, and the higher the value, the better the performance of the model. The classification threshold of the machine learning model can also be determined by the performance evaluation submodule. The classification threshold can be used to indicate the decision boundary of the model when distinguishing different categories.
[0036] The data processing module may also include a model optimization submodule, which can be used to optimize the performance of the machine learning model by adjusting hyperparameters and / or using different training strategies.
[0037] The machine learning model submodule may be a pre-trained machine learning model submodule.
[0038] The machine learning model submodule may also be a trained machine learning model submodule obtained after training and evaluation through a training submodule, a performance evaluation submodule, and a model optimization submodule.
[0039] The result output module can output the calculation result of the trained machine learning model as whether the person belongs to the group susceptible to acute mountain sickness.
[0040] Furthermore, the machine learning model can be selected from a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model and a hidden Markov model.
[0041] Furthermore, the machine learning model may be a random forest model.
[0042] The present invention also provides a method for constructing an acute mountain sickness susceptibility clustering model, the method comprising the following steps: using data on the expression levels of six proteins, namely, RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95, in samples of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model.
[0043] In the above method, the method may include the following steps:
[0044] A1) Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95, in samples from known patients with acute mountain sickness and known people adapted to the plateau environment;
[0045] A2) using the data as input features of a machine learning model to train and evaluate the machine learning model;
[0046] A3) obtaining a trained machine learning model according to the evaluation results, i.e., the acute mountain sickness susceptibility clustering model.
[0047] In the case where the input features are clear, the training method of the machine learning model is known to those skilled in the art. For example, after the expression levels of the six proteins described in the present invention are used as the input features of the machine learning model, the model can be trained according to the input data samples (labeled data sets), and the model parameters can be iteratively updated using optimization algorithms such as gradient descent method, Newton iteration method, particle swarm optimization algorithm, artificial fish swarm algorithm, artificial bee colony algorithm, and firefly algorithm to minimize the loss function.
[0048] During the model training process, the model can be evaluated to determine the performance of the model. Evaluation indicators can usually include ROC curve (AUC value), accuracy, precision, recall, F1 score, etc. For example, ROC curve and AUC value are two core indicators used to evaluate the performance of classification models. When evaluating the model based on the ROC curve (AUC value), the closer the ROC curve is to the upper left corner, the better the performance of the model. The point (0,1) in the upper left corner indicates that the model can classify perfectly. AUC is the area under the ROC curve, which is a comprehensive indicator used to quantify the performance of the model. The value range of AUC is between 0 and 1. The closer to 1, the better the model performance is, indicating that the model has a stronger ability to distinguish samples. AUC is not affected by the classification threshold and provides an overall performance evaluation.
[0049] Furthermore, based on the evaluation results of the model, the model can be optimized, for example, by adjusting hyperparameters and / or using different training strategies to optimize the model, so as to further improve the performance of the model. Cross-validation, grid search method and / or Bayesian optimization algorithm can be used for model selection and hyperparameter adjustment.
[0050] The trained machine learning model can be used as an acute mountain sickness susceptibility clustering model for prediction or classification tasks of practical problems.
[0051] In the above method, the data described in A1) can be obtained by protein immunoblotting (Western Blot), immunofluorescence, radioimmunoassay (RIA), immunoprecipitation, immunohistochemistry (IHC), enzyme-linked immunosorbent assay (ELISA), double antibody sandwich ELISA, fluorescence-linked immunosorbent assay (FLISA), enzyme immunoassay (EIA), flow cytometry (FACS), polyacrylamide gel electrophoresis, capillary electrophoresis, near infrared spectroscopy, immunochemiluminescence, colloidal gold immunoassay (GIC), colloidal gold immunochromatography (GICA), fluorescence immunochromatography, surface plasmon resonance technology. , immuno-PCR technology, biotin-avidin technology, liquid chromatography, high performance liquid chromatography, solid phase chromatography, mass spectrometry, liquid chromatography-mass spectrometry (LC-MS), high performance liquid chromatography-mass spectrometry (HPLC-MS), liquid phase secondary mass spectrometry (LC-MS / MS), two-dimensional electrophoresis-mass spectrometry (2DE-MS), two-dimensional liquid chromatography-mass spectrometry (2D-LC-MS), isotope coded affinity tag technology (ICAT), sequence-eluted time of flight mass spectrometry (SELDI-TOF-MS) or protein chip technology detection, but not limited to these.
[0052] In the above method, the machine learning model can be selected from a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model and a hidden Markov model.
[0053] In the above method, the machine learning model may be a random forest model.
[0054] The output value of the random forest model described in this article can be a class label (which class the sample belongs to) or a probability value.
[0055] For class labels: In the random forest model, each sample is assigned a predicted probability of belonging to a certain category. These predicted probabilities can be converted into specific class labels through the classification threshold. The classification threshold is not fixed and can be adjusted according to actual needs. For example, when the sensitivity or specificity of the model is more concerned, the sensitivity and specificity can be weighed by adjusting the classification threshold. Lowering the classification threshold can improve the sensitivity of the model, and increasing the classification threshold can improve the accuracy and specificity of the model. When the category distribution is uneven, the classification threshold needs to be adjusted to improve the performance of the model. The optimal classification threshold can also be determined using methods such as cross-validation and grid search. The threshold with a higher AUC value can also be selected as the optimal classification threshold by drawing the ROC curve and calculating the AUC value. Usually, the classification threshold can be set to the default value (0.5), which means that if the predicted probability is greater than 0.5, the sample is classified as a positive class, otherwise it is classified as a negative class.
[0056] The subjects with a probability greater than or equal to 0.5 were more susceptible to acute mountain sickness than those with a probability less than 0.5.
[0057] For probability values: In the random forest model, each decision tree predicts the sample and gives the probability value of belonging to each category. Then, these probability values are averaged (or other aggregation methods are used) to obtain the final category probability distribution. These probability values can provide more model information and can be used to build more complex decision systems and evaluate model performance.
[0058] The hyperparameters of the random forest model described in this article can be: the number of random forest trees is 500.
[0059] Those skilled in the art will understand that the choice of hyperparameters is not unique, and the purpose of determining hyperparameters is to help the model train for good performance on the data set. As long as the model can predict accurately and does not have underfitting or overfitting, any combination of hyperparameters is acceptable.
[0060] The application method of the acute mountain sickness susceptibility grouping model obtained by any of the methods of the present invention may include the following steps:
[0061] (1) Receive data on the expression levels of six proteins, namely, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95, in samples from the test subjects;
[0062] (2) inputting the data into an acute mountain sickness susceptibility clustering model obtained by any of the methods described in the present invention (for example, a trained random forest model constructed by the method described in the present invention), and obtaining the clustering results of the test subjects through the acute mountain sickness susceptibility clustering model.
[0063] Furthermore, it is unknown whether the test subject suffers from acute altitude sickness.
[0064] Furthermore, the subject to be tested may be an asymptomatic healthy subject or a suspected AMS patient, for example, a suspected AMS patient having at least one of the following symptoms: headache, fatigue, weakness, dizziness, vertigo, sleep disorder, insomnia, loss of appetite, dyspnea, palpitations, nausea, vomiting, etc.
[0065] Furthermore, the subject to be tested may be a person located in a plateau at an altitude of 2000 or 2500 meters or more.
[0066] The present invention also provides a method for grouping people according to their susceptibility to acute mountain sickness, the method comprising the following steps:
[0067] S1. Data reception: Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTNS246 and VPS4A S95, in the samples of the subjects;
[0068] S2. Data processing: inputting the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model;
[0069] S3. Output result: output whether the subject belongs to the group susceptible to acute mountain sickness through the acute mountain sickness susceptibility grouping model.
[0070] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the steps of any method described in this article (a method for constructing an acute mountain sickness susceptibility grouping model or a method for grouping a population for acute mountain sickness susceptibility).
[0071] Furthermore, the computer-readable storage medium may also execute the step of outputting the acute mountain sickness susceptibility grouping results.
[0072] The sample described herein may be derived from a blood sample (including plasma, serum and whole blood).
[0073] Furthermore, the sample may be a peripheral blood red blood cell membrane sample.
[0074] The six proteins RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 described herein can be used to determine whether a subject belongs to a susceptible population for acute mountain sickness (AMS). The expression levels of the six proteins can be detected, and the detection results can be compared with a preset threshold (or reference value) that can effectively distinguish AMS patients from non-AMS subjects to perform grouping. The comparison covers any type of comparison of the expression levels of the six proteins with a preset threshold (or reference value). For example, the comparison can be performed by mathematically correlating the expression levels of the six proteins (processing or calculation based on a statistical analysis method) to obtain a score or index, and comparing the score or index with a preset threshold (or reference value) to obtain a classification result. The comparison can also be performed by using a machine learning model (such as a random forest model, a support vector machine model, etc.) to input the expression levels of the six proteins into a model as an independent variable, and then outputting the model to obtain a classification result. It can be understood by those skilled in the art that the comparison can be performed manually or computer-assisted. For example, the value to be compared (such as the value output by the model) can be compared with a preset threshold (or reference value) manually, or the comparison can be automatically performed by a computer program executing the algorithm. The computer program executing the comparison can provide the clustering results in an appropriate output format.
[0075] The present invention performed proteomic and phosphoproteomic analysis on the red blood cell membranes of AMS patients and nAMS volunteers in the plateau area. By characterizing the protein and phosphorylated protein profiles of the subjects, the relationship between the red blood cell membrane and AMS was analyzed and mined. In addition, the present invention analyzed the association between red blood cell membrane proteins and phosphorylated proteins, studied the correlation between key differential molecules in the RBC membrane (red blood cell membrane) and the severity of AMS, and established an AMS clustering model based on machine learning.
[0076] The present invention screened and obtained six proteins, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4AS95, for the first time, and constructed an acute mountain sickness susceptibility clustering model based on this. At present, the Lake Louise Acute Mountain Sickness Scoring System (LLSS) is mainly relied on to score the subjective symptoms of patients at home and abroad, and questionnaires need to be filled in, the work efficiency is low, and the workload is large. The present invention establishes a clustering model based on machine learning, and verifies its specificity and reliability, which can accurately distinguish AMS patients and non-AMS subjects, can realize rapid clustering of acute mountain sickness, reduce the workload of medical staff, and is suitable for wide promotion and application.
[0077] Definition of terms
[0078] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Meanwhile, in order to better understand the present invention, the definitions and explanations of the relevant terms are provided below.
[0079] The term "device" generally refers to a system comprising modules or units that are operably connected to each other to enable computing. One or more modules or units may be implemented by software (such as computer programs, web applications, mobile applications, and standalone applications, etc.), hardware (such as processors or memories, etc.), or a combination thereof. The device encompasses various forms, such as server computers, desktop computers, laptop computers, notebook computers, small notebook computers, netbook computers, netpad computers, set-top computers, streaming media devices, handheld computers, mobile smartphones, tablet computers, and personal digital assistants, but is not limited thereto.
[0080] The term "data receiving module" generally refers to a functional module for obtaining relevant data. In the present invention, the data receiving module may refer to a functional module for obtaining the 6 protein expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4AS95. The data receiving module may also include reagents and / or instruments for detecting the 6 protein expression levels to perform the step of detecting the 6 protein expression levels in the sample to be tested. The step of detecting the 6 protein expression levels in the sample to be tested may also be performed by an independent device (such as a test kit, a mass spectrometer, a chromatograph, a flow cytometer, etc.), and is not included in the data receiving module.
[0081] The term "data processing module" generally refers to a functional module responsible for data processing (including preprocessing of data, etc.). In the present invention, the data processing module can be used to process and / or analyze the six protein expression level data of RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the sample to be tested, and achieve the purpose of acute mountain sickness susceptibility grouping based on the analysis and processing results. The data processing module can process and analyze the data based on the machine learning model.
[0082] The term "result output module" generally refers to a functional module for displaying result information to a user. In the present invention, the result output module may refer to a functional module for displaying the results of acute mountain sickness grouping. The result output module may include a display, a projector, a monitor, a printer, an audio output device, etc. but is not limited thereto.
[0083] The term "machine learning model" generally refers to a computational model or algorithm. A machine learning model can learn and analyze input data to perform tasks such as prediction, classification, recognition, or decision-making. The learning process of a machine learning model can be based on statistical principles and data pattern recognition, using training data sets to adjust model parameters and optimize the model to improve its prediction or reasoning ability. It can predict or classify new data by learning the mapping relationship between features and labels in the training data. In machine learning, "label" generally refers to the true category or target value of a data sample, which is the "reference answer" when training a machine learning model. In classification tasks, labels can indicate which category a data sample belongs to. In regression tasks, labels can be continuous target values or real target values to be predicted. "Feature" generally refers to an attribute or characteristic that describes a data sample. In machine learning, features are input variables used to train a model, which can be any type of data, such as numeric or categorical. "Training" generally refers to the process of model learning, which uses labeled training data to adjust the parameters of the model so that the model can predict labels as accurately as possible. During the training process, the model can continuously optimize its own parameters to approximate the real data distribution. The machine learning model can adopt various algorithms and techniques to be trained and optimized by supervised learning, unsupervised learning or reinforcement learning. Exemplary machine learning models may include random forest models, support vector machine models, linear discriminant analysis models, recursive feature elimination models, logistic regression models, neural network models, CART models, flextree models, LART models, MART models, k-nearest neighbor classification models, cluster analysis models, principal component analysis models, Bayesian classification models and hidden Markov models, etc. In one or more embodiments of the present invention, a random forest model is used to establish a clustering model for acute mountain sickness susceptibility.
[0084] The term "random forest model" usually refers to an ensemble learning model built using a bagging algorithm with multiple decision trees as the basic learners. For classification problems, the category of the output result is determined by the mode (the majority result) of the output categories of all decision trees, that is, the final output result is determined by voting by all single decision tree models, and the output is the class label of the sample (which category the sample belongs to). For regression problems, the output result can be the average of the decision tree results, and its output is a real number (such as a probability value).
[0085] The term "computer-readable storage medium" generally refers to any medium that can store data and can be read or accessed by a computer, which medium may include a computer program. Computer-readable storage media may include volatile and non-volatile media and removable and non-removable media. As an example and not limitation, computer-readable storage media may include ROM (read-only memory), PROM, EPROM, EEPROM, FLASH-EPROM, CD-ROM, DVD-ROM, RAM (random access memory), SRAM, DRAM, SDRAM, RDRAM, DDR SDRAM, GDDR SDRAM, hard disk drive (HDD), solid state drive (SSD), optical disk drive (such as CD, DVD drive, etc.), flash memory (such as USB), disk drive, tape drive, distributed computing system (such as cloud storage, cloud disk, etc.).
[0086] The terms "device" and "system" generally have the same meaning and are used interchangeably.
[0087] The terms "module" and "unit" generally have the same meaning and are used interchangeably.
[0088] The term "receiver operating characteristic curve (ROC curve)" usually refers to a curve drawn according to a series of different binary classification methods, with the true positive rate as the ordinate and the false positive rate as the abscissa. AUC (area under the curve) represents the area enclosed by the ROC curve and the horizontal axis. The ROC curve reflects the relationship between sensitivity and specificity. The principle of making the ROC curve is to set a number of different critical values for the continuous variable, calculate the corresponding sensitivity and specificity at each critical value, and then draw a curve with sensitivity as the ordinate and 1-specificity as the abscissa. Since the ROC curve is composed of multiple critical values representing their respective sensitivity and specificity, the ROC curve can be used to select the best limit value for a method. The closer the ROC curve is to the upper left corner, the higher the test sensitivity, the lower the false positive rate, and the better the performance of the method. It can be seen that the point on the ROC curve closest to the upper left corner has the largest sum of sensitivity and specificity. This point or the value corresponding to its adjacent point is often used as a reference value (also called a threshold).
[0089] The term "threshold" generally refers to a dividing value used as a reference to obtain information about a measurement result and / or to classify a measurement result. For example, a threshold is set when defining a dividing line between two subsets in a population. Values equal to or above the threshold level define one subset of the population, and values below the threshold level define another subset of the population. The threshold can be determined based on one or more control samples or the entire control sample population. The threshold can be determined before, while, or after performing the measurement of interest.
[0090] The term "reference value" generally refers to a reference level that can be obtained in advance from a subject (including a healthy subject and a patient), and an appropriate reference level can be measured and selected according to techniques known to those skilled in the art.
[0091] The term "comprising" is not intended to be limiting, but is intended to be inclusive and means that there may be other elements besides the listed elements, which can be interpreted as "including but not limited to". The term "comprising" also encompasses the terms "consisting of" and "consisting essentially of". The terms "comprising" and "including" are used interchangeably herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 The exploration process of erythrocyte membrane proteome and phosphoproteome in AMS.
[0093] Figure 2 Detailed demographic characteristics of the participants.
[0094] Figure 3 Characteristic of dysregulated erythrocyte membrane proteins in the AMS group.
[0095] Figure 4 Characteristic of phosphorylated proteins / peptides on erythrocyte membranes.
[0096] Figure 5 These are the characteristics of dysregulated erythrocyte membrane phosphorylated proteins in the AMS group.
[0097] Figure 6 Kinase prediction and motif analysis for dysregulated phosphorylated proteins.
[0098] Figure 7 Figure 3 Correlation analysis between AMS severity and proteomic and phosphoproteomic characteristics.
[0099] Figure 8 Sort the importance of variables in the AMS random forest clustering model.
[0100] Fig. 9 This is the classification effect of the random forest clustering model on the validation data set.
[0101] Fig.10 The present invention is a flow chart for the acute mountain sickness susceptibility grouping.
[0102] Fig.11 The figure is a schematic diagram of the device for grouping the susceptibility to acute mountain sickness in the present invention. DETAILED DESCRIPTION
[0103] The present invention is further described in detail below in conjunction with specific embodiments, and the examples provided are only for illustrating the present invention, rather than for limiting the scope of the present invention. The examples provided below can be used as a guide for further improvements by those of ordinary skill in the art, and do not constitute a limitation of the present invention in any way.
[0104] The experimental methods in the following examples, unless otherwise specified, are all conventional methods, and are performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials, reagents, etc. used in the following examples, unless otherwise specified, can all be obtained from commercial channels.
[0105] Example 1: Screening of erythrocyte membrane proteins and phosphorylated proteins
[0106] Red blood cells are the most abundant cell type in the human body and are responsible for transporting O2 and CO2. Red blood cell membrane protein homeostasis is closely related to hypoxia, which is closely related to membrane deformability, glycolysis, release and transport of oxygen molecules, and regulation of iron ion balance. Under hypoxic conditions, red blood cells need to be able to pass through narrow blood vessels and deliver oxygen to hypoxic tissues. At this time, the deformability of red blood cells is particularly important. Studies have shown that hypoxia can affect the deformability of red blood cells by affecting the cytoskeleton, NO and ATP production, ion homeostasis, and regulation of cell membrane protein composition. In addition, mature red blood cells do not have obvious organelle structures, and pentose phosphate and glycolysis are the main energy supplies. By sensing changes in oxygen content in the circulatory system, red blood cells cause a significant increase in glycolysis under hypoxic conditions, promote the production of ATP and 2,3-BPG (2,3-diphosphoglycerate), thereby promoting the release of oxygen and alleviating hypoxia in peripheral tissues. Acute hypoxia also leads to a significant reduction in red blood cell cytoskeletal proteins and glucose transporters: spectrin, ankyrin, band 3, band 4.1, band 4.2, and GAPDH. In addition, in acute hypoxia, due to insufficient ATP supply, the activity of ion transporters on the red blood cell membrane decreases significantly, suggesting that ion homeostasis in red blood cells is affected. Currently, although studies have explored the effects of hypoxia on the composition and function of red blood cells, there is still a lack of systematic research to explore whether red blood cells in AMS patients show specific changes or more significant adjustments, and whether these changes are directly related to the severity of AMS.
[0107] Protein phosphorylation is one of the important post-translational modifications of organisms and an effective regulatory method for the body to quickly respond to changes in the external environment. Phosphorylation is involved in the regulation of almost all life activities (such as cell proliferation, signal transduction, metabolism, etc.). It is reported that phosphorylation is also involved in the regulation of enzyme activity in erythrocytes and the control of erythrocyte deformability. Proteomics and phosphoproteomics are suitable for studying the pathogenesis of AMS from the perspective of erythrocytes, because there is only regulation of existing proteins in erythrocytes, but no regulation of newly synthesized proteins.
[0108] The present invention performed proteomics and phosphoproteomics analysis on the erythrocyte membranes of AMS patients and nAMS volunteers in plateau areas, and screened erythrocyte membrane proteins and phosphorylated proteins for acute mountain sickness (AMS) clustering. The specific method is as follows:
[0109] 1. Screening sample information
[0110] Clinical demographic information and study design: According to the Lake Louise Scoring System (LLSS), a group of volunteers were recruited (see Table 1, training set sample), including AMS patients (n = 5) and nAMS subjects (n = 5) ( Figure 2 A). Obtain detailed clinical characteristics, including age, blood oxygen saturation (SaO2), blood routine indexes, coagulation indexes ( Figure 1 A). Compared with the nAMS group, the SaO2 in the AMS group was significantly lower ( Figure 2 B), while there was no significant change in heart rate ( Figure 2 C). In addition, the AMS group also showed abnormal coagulation indexes ( Figure 2 D), such as PT increase and PTA decrease, indicating that the coagulation function of the AMS group may be weakened. However, there was no significant change in platelet count and other related indicators ( Figure 2 E. Figure 2 F). Accordingly, indicators such as red blood cell count did not show abnormalities. Then, we extracted red blood cell membrane proteins and mapped the red blood cell membrane proteome and phosphorylated proteome of this cohort ( Figure 1 B).
[0111] The criteria for the Lake Louise Scoring System (LLSS) are as follows: assess the presence or absence of four symptoms: headache symptoms, gastrointestinal symptoms, fatigue and / or weakness symptoms, and dizziness and / or vertigo symptoms. Each of the four symptoms is scored from 0 to 3 points depending on its severity. 0 points indicates no symptoms; 1 point indicates mild (mild) symptoms; 2 points indicate moderate symptoms; and 3 points indicate severe (severe) symptoms. Gastrointestinal symptoms include appetite, nausea and / or vomiting, and the scores are as follows: 0 points indicate no symptoms and good appetite; 1 point indicates loss of appetite or mild nausea; 2 points indicate moderate nausea or vomiting; and 3 points indicate severe nausea and vomiting. When the sum of the four symptom scores is greater than or equal to 4, it can be judged that the patient has acute mountain sickness and is evaluated as an AMS patient; when it is less than or equal to 3, it can be judged that the patient does not have acute mountain sickness and is evaluated as a non-AMS subject (nAMS subject).
[0112] Table 1. Information of 10 screening samples
[0113]
[0114] 2. Sample preparation and protein detection methods
[0115] Separation of erythrocyte membranes from peripheral blood samples: The peripheral blood samples of the subjects were allowed to stand for 30 minutes and then centrifuged at 3000 rpm for 15 minutes. The upper plasma, the middle white membrane, and part of the upper red blood cells in the lower layer were discarded. 300 μL of erythrocyte sediment was mixed with hypotonic Tris-HCl buffer (volume ratio 1:30, pH = 8.0, containing 0.01M Tris-HCl, 2mM sodium orthovanadate, 1mM phenylmethylsulfonyl fluoride and Roche protease inhibitors) and incubated at 4°C for 1 hour to extract erythrocyte membranes. Then, centrifuge at the highest speed at 4°C for 15 minutes to separate the erythrocyte membranes, and wash the membrane sediment with 1 mL of hypotonic Tris-HCl buffer, and repeat the operation until a white precipitate is obtained. SDC lysis buffer (containing 4% weight / volume sodium deoxycholate, pH=8.5, 100 mM Tris-HCl) precooled to 4°C was added to the erythrocyte membrane pellet, and immediately denatured at 95°C for 5 minutes after mixing, and then sonicated in a water bath at maximum power for 10 minutes at 4°C. After centrifugation, the protein concentration was determined by the BCA method.
[0116] Enrichment and detection of phosphorylated peptides: Take 300 μg of protein and reduce and alkylate it at 37°C for 15 minutes with a 1:10 volume of reduction / alkylation buffer (containing 100 mM TCEP-HCl and 400 mM chloroacetamide). Add 3 μg of trypsin, mix thoroughly and digest overnight. The next day, the digestion reaction was terminated with 2% trifluoroacetic acid (TFA). Mix the samples and centrifuge at high speed to remove SDC. The SDC precipitate was washed three times with 100 microliters of 0.5% TFA. Combine all supernatants and purify with a C18 column. Finally, take 1 μ of the peptide mixture for liquid chromatography-tandem mass spectrometry (LC-MS / MS) analysis, and the rest is used for phosphorylated peptide enrichment. Take 200 μg of peptides for phosphorylated peptide enrichment (Thermo Fisher, High-Select TM Fe-NTA Phosphopeptide Enrichment Kit). The enriched phosphopeptides were purified using a homemade C18 column and subjected to LC-MS / MS analysis.
[0117] 3. Characteristics of erythrocyte membrane proteome in AMS patients
[0118] A total of 2383 proteins were identified from the 10 RBC membrane samples, of which 329 proteins were uniquely identified by the AMS group, which was higher than that of the nAMS group ( Figure 3 A). After normalizing the protein quantitative information, we found that the red blood cell membrane proteome characteristics of the two groups were significantly different through UMAP analysis ( Figure 3 B). The results of differential analysis showed that there were 189 differentially expressed proteins (DEPs) between the AMS group and the nAMS group, of which 69 DEPs were down-regulated and 120 DEPs were up-regulated ( Figure 3C). KEGG pathway enrichment analysis showed that 69 downregulated membrane proteins were highly related to ferroptosis, fatty acid metabolism, endocytosis, mineral absorption and tight junction ( Figure 3 D). In addition, we analyzed the KEGG pathways enriched for upregulated membrane proteins. These DEPs were mainly involved in complement and coagulation cascades, TCA cycle, peroxisome, regulation of actin cytoskeleton, and focal adhesion ( Figure 3 E). Finally, we use the heat map ( Figure 3 F) shows Figure 3 D and Figure 3 DEPs enriched in energy metabolism-related pathways in E.
[0119] 4. Phosphoproteomic characteristics of erythrocyte membrane in plateau
[0120] Using phosphoproteomics, 1076 phosphorylated proteins were identified in all erythrocyte membrane samples, of which 922 phosphorylated proteins were commonly found in both AMS and nAMS, 130 phosphorylated proteins were uniquely identified in AMS, and 24 phosphorylated proteins were uniquely identified in nAMS ( Figure 4 A). Based on the probability of phosphorylation modification sites > 0.8, a total of 2913 reliable phosphorylated peptides were screened. Among them, 2417 phosphopeptides were commonly found in AMS and nAMS, and 387 phosphopeptides were uniquely identified in AMS, which was more than that in nAMS ( Figure 4 B).
[0121] We then reviewed the basic information of this dataset. The quantitative information of phosphorylated peptides spans 6 orders of magnitude ( Figure 4 C). In addition, statistical analysis of the phosphorylation modification site types showed that the proportion of serine (S) modification in AMS and nAMS was much greater than that of T and Y modification: S was 84.77% (AMS) and 84.32% (nAMS); T was 13.80% (AMS) and 14.29% (nAMS), and Y was 1.43% (AMS) and 1.39% (nAMS) ( Figure 4 D). About 46% of the phosphorylated proteins have two or more phosphorylation modification sites. Interestingly, the top seven are SPTA1, ADD2, ANK1, SPTB, EPB41, DMTN, and ADD1, all of which are associated with spectrin and actin. Motif analysis of T and S phosphorylated peptides showed that the xxxTPPxxx and xxxRxDSxxx conserved motifs were significantly enriched in T and S modified sequences, respectively ( Figure 4 E), suggesting that specific kinases in erythrocytes may contribute to this event.
[0122] 5. Dysregulation of erythrocyte membrane phosphorylation in AMS patients
[0123] After normalizing the quantitative information of phosphorylated peptides, we found through UMAP analysis that the phosphorylated peptide features of AMS and nAMS are different ( Figure 4 F). Screening of differentially expressed phosphorylated peptides (DEPPs) based on fold change and p-value. The results are shown in Figure 5 As shown in A. The phosphorylated peptides represented by blue were significantly downregulated (n=48), and the phosphorylated peptides represented by red were significantly upregulated (n=108). Further analysis revealed that the change trend of multiple DEPPs in the same protein was consistent, among which the DEPPs in EEPD1, AGO2, SPTB and ANP32B were significantly decreased, while the DEPPs in other proteins were significantly upregulated ( Figure 5 B). The possible kinases of DEPPs were predicted using kinase-substrate annotation information. The results showed that the kinase activity of CSNK2A1 was reduced, while the activities of other kinases were upregulated. In particular, the kinase activity of PRKCG was significantly upregulated ( Figure 6 A). We then performed motif enrichment analysis on S-modified DEPPs and found that the xxxSPxxx motif was significantly enriched ( Figure 6 B).
[0124] 6. DEPP-matched proteins for functional enrichment analysis
[0125] Downregulated phosphorylated proteins were significantly enriched in necroptosis, calcium reabsorption, mineral absorption, and glycolysis / gluconeogenesis ( Figure 5 C). The upregulated proteins were mainly enriched in platelet activation, actin cytoskeleton regulation, ECM-receptor interaction, and other pathways related to intercellular communication and cell morphology maintenance ( Figure 5 D). In addition, we found that the downregulated phosphorylated proteins were mainly located in the cell cortex and cortical cytoskeleton and had functions in spectrin binding, structural composition of the cytoskeleton, and P-type ion transporter activity. These downregulated phosphoproteins were mainly involved in the regulation of actin filament depolymerization and the regulation of cell shape ( Figure 5 E). The upregulated phosphorylated proteins were enriched in membrane rafts, secretory granule membranes, and glycoprotein complexes, which also suggested that RBC membrane proteins had a good extraction efficiency. These upregulated phosphorylated proteins were involved in regulating Fc receptor-mediated stimulation signaling pathways, coagulation, and cellular metal ion homeostasis. We also noted that the upregulated phosphorylated proteins had functions in cadherin binding, cAMP-dependent protein kinase regulation activity, and structural composition of the cytoskeleton ( Figure 5 F).
[0126] 7. Correlation between membrane phosphopeptides / proteins and AMS and construction of classification model
[0127] Correlation analysis was performed on erythrocyte membrane phosphopeptides / proteins and AMS respectively. We used the Lake Louise Subjective Scoring System (LLSS) as the standard for judging the severity of AMS. First, it was found that the severity of AMS was positively correlated with PT and PTR, and negatively correlated with PTA. Then DEPPs with a correlation coefficient greater than 0.6 were selected. Among them, 11 DEPPs, ANP32B T244, NAP1L1 T62, ADD3S42, VPS4A S95, GAB1 S419, ANK1 S1756, CC2D1A T204, BTK S55, SPTB T1387, DMTNS264, and USP6NL S415, were negatively correlated with the severity of AMS, and FLNA S2152 and TUBA4A S439 were positively correlated with the severity of AMS ( Figure 7 A). For DEPs, the results of correlation analysis showed that 9 DEPs (RAB2B, STX6, SLC30A7, VPS4A, CCNY, BMP2K, DHX9, BROX, and DLG1) were negatively correlated with AMS, while GC and HP1BP3 were positively correlated with AMS ( Figure 7 B).
[0128] Subsequently, we investigated whether the 13 DEPPs and 11 DEPs that were highly correlated with AMS severity could be used to distinguish AMS from nAMS. To optimize the classification model, the importance of these 24 variables was calculated by random forest analysis ( Figure 8 A-8B), and determined that 6 variables constitute the best model based on the distribution of error rate with the number of variables. Then combined with the tree with the smallest error rate ( Figure 8 C) and the importance analysis of DEPP and DEP, we selected a combination of 6 variables ( Figure 7 C-7D, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, VPS4A S95) were used to build the final machine learning model.
[0129] Example 2: Construction of acute mountain sickness (AMS) susceptibility clustering model
[0130] In order to achieve the highest clustering efficiency with the least combination of variables, the expression levels of the six variables (RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95) screened in Example 1 and the scores of the Lake Louise subjective scoring system were used as parameter sample data, and a calculation model was obtained by calculation using a machine learning method. The method for establishing an AMS susceptibility clustering model may include the following steps:
[0131] S1) receiving the expression levels of six proteins, namely, RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95, from samples of known patients with acute mountain sickness and known healthy people to obtain the original data set;
[0132] S2) selecting a machine learning model (e.g., random forest model, support vector machine model, linear discriminant analysis model, recursive feature elimination model, logistic regression model, neural network model, CART model, flextree model, LART model, MART model, k nearest neighbor classification model, cluster analysis model, principal component analysis model, Bayesian classification model, hidden Markov model and combinations thereof, etc.);
[0133] S3) using the original data set as an input feature of a machine learning model, training and evaluating the machine learning model based on the original data set, and obtaining an AMS susceptibility clustering model according to the evaluation results.
[0134] This example uses a random forest algorithm to construct an AMS susceptibility clustering model, using the expression levels of the six proteins RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95 in the sample to be tested as the input of the random forest model, and whether the sample belongs to the group susceptible to acute mountain sickness as the output of the random forest model. The hyperparameters of the random forest model are screened and determined using a 5-fold cross-validation combined with a grid search method, and the hyperparameters are as follows:
[0135] Number of random forest trees (n_estimators): 500.
[0136] The random forest algorithm uses the randomForest package, and all programs are executed in the R environment. For parameter values that are not explicitly stated, the default values of the algorithm package are used.
[0137] The model classification threshold is calculated according to the ROC curve, thereby constructing an AMS susceptibility clustering model. By setting a reasonable classification threshold, the output results of the random forest model can be converted into specific clustering results. The classification threshold of the random forest model in this embodiment is 0.5.
[0138] Specifically, when the output result of the model is higher than the classification threshold, the sample to be tested is judged to be AMS positive (belonging to the acute mountain sickness susceptible population); when the output result is lower than the classification threshold, the sample to be tested is judged to be AMS negative (not belonging to the acute mountain sickness susceptible population). Those skilled in the art can set the size of the threshold according to actual needs, and the present invention is not limited here. In addition, the hyperparameter is also non-restrictive and can be adjusted and optimized as needed.
[0139] The model building results are as follows Figure 7 As shown in D, the constructed random forest model achieved a classification effect of area under the ROC curve of 1 in the training set. At the same time, the model was applied to evaluate the 10 participants in this study and clearly distinguished AMS patients from nAMS subjects ( Figure 7 C). These results fully demonstrate the potential clinical application value of this model.
[0140] Example 3: Verification of the clinical application value of the model
[0141] 1. Verify sample information
[0142] Six validation samples different from the screening samples in Example 1 were selected, see Table 2.
[0143] Table 2. Information of 6 verification samples
[0144] ID N10 N12 N9 Y1 Y3 Y4 Data classification Validation set Validation set Validation set Validation set Validation set Validation set Type nMs nMs nMs AMS AMS AMS Symptom score 1 1 1 6 4 5 Blood oxygen saturation 84 87 90 87 84 84 age 22 25 24 22 23 22 gender male male male male male male Heart rate 100 90 96 89 112 80 PT 9.4 9.5 9.7 15.1 11 10.2 PTR 0.79 0.8 0.82 1.28 0.93 0.87 PT-INR 0.79 0.8 0.81 1.29 0.93 0.86 FIB 1.98 1.44 1.96 2.7 2.84 3.18 PTA 133.1 131 126.9 58.8 103.6 116 PLT 306 168 216 229 184 284 PDW 15 17 20 20 21 14 RBC 5.6 5.7 5.2 5.3 5.6 5.6 HCT 51 51 49 52 51 51 HGB 167 175 155 172 162 162
[0145] 2. Sample preparation and protein detection
[0146] Same as Example 1.
[0147] The test results were input into the random forest model constructed in Example 2 for verification. The verification results are as follows: Fig. 9 As shown, the AUC value of the random forest model reached 0.889 ( Fig. 9 A), with a specificity of 66.7%. This indicates that the random forest model characterized by the expression levels of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246, and VPS4A S95 has a good ability to distinguish acute mountain sickness ( Fig. 9 B), has high clinical application value.
[0148] The present invention relates to a method for susceptibility grouping of acute mountain sickness, the flow chart of which is as follows: Fig.10 As shown, including:
[0149] S1. Data reception: Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTNS246 and VPS4A S95, in the samples of the subjects;
[0150] S2. Data processing: inputting the data of the expression levels of the six proteins into the acute mountain sickness susceptibility clustering model; the model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model;
[0151] S3. Output result: output whether the subject belongs to the group susceptible to acute mountain sickness through the acute mountain sickness susceptibility grouping model.
[0152] Accordingly, the present invention provides a device for grouping acute mountain sickness susceptibility, such as Fig.11 As shown, it includes: a data receiving module, a data processing module and a result output module, wherein the data receiving module is used to receive the data of the expression levels of the six proteins RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 in the peripheral blood red blood cells of the subject; the data processing module is used to input the data into the acute mountain reaction susceptibility clustering model (which can be a model pre-trained based on the previous training set, or a model obtained after training with a continuously increasing new training set); the acute mountain reaction susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain reaction patients and known healthy people as training samples to train the acute mountain reaction susceptibility clustering model; the result output module is used to output whether the subject belongs to the acute mountain reaction susceptibility population through the acute mountain reaction susceptibility clustering model. Among them, the data processing module includes a pre-trained machine learning model submodule and / or a trained machine learning model submodule operably connected to the training submodule, the performance evaluation submodule, and the model optimization submodule.
[0153] The present invention has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present invention, and without the need to carry out unnecessary experimental conditions, the present invention can be implemented in a wide range under equivalent parameters, concentrations and conditions. Although the present invention provides specific embodiments, it should be understood that the present invention can be further improved. In a word, according to the principles of the present invention, the application is intended to include any changes, uses or improvements to the present invention, including departure from the disclosed scope in the application, and changes made with conventional techniques known in the art.
Claims
1. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the following steps: S1. Data reception: Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95, in the samples of the subjects; S2. Data processing: inputting the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model; S3. Output result: output whether the subject belongs to the group susceptible to acute mountain sickness through the acute mountain sickness susceptibility grouping model.
2. The computer device according to claim 1, characterized in that The acute mountain sickness susceptibility clustering model is selected from at least one of a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model and a hidden Markov model.
3. The computer device according to claim 2, characterized in that The acute mountain sickness susceptibility clustering model is a random forest model.
4. The computer device according to claim 3, characterized in that The hyperparameters of the random forest model are as follows: The number of random forest trees is 500.
5. A device for grouping people with susceptibility to acute mountain sickness, characterized in that: The device comprises a data receiving module, a data processing module and a result output module; The data receiving module is used to receive data on the expression levels of six proteins, namely, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95, in the subject sample; The data processing module is used to input the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model; The result output module is used to output whether the subject belongs to the acute mountain sickness susceptible population through the acute mountain sickness susceptibility grouping model.
6. The device according to claim 5, characterized in that The data processing module includes a machine learning model submodule, the input of the machine learning model is the data of the expression levels of the six proteins described in claim 1, and the output is whether the person belongs to the group susceptible to acute mountain sickness.
7. A method for constructing a clustering model for acute mountain sickness susceptibility, characterized in that: The method comprises the following steps: using data of expression levels of six proteins, namely, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95, in samples of known acute mountain sickness patients and known plateau environment adaptors as training samples to train an acute mountain sickness susceptibility clustering model.
8. The method according to claim 7, characterized in that The method comprises the following steps: A1) Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95, in samples from known patients with acute mountain sickness and known people who have adapted to the plateau environment; A2) using the data as input features of a machine learning model to train and evaluate the machine learning model; A3) obtaining a trained machine learning model according to the evaluation results, i.e., the acute mountain sickness susceptibility clustering model.
9. A method for grouping people according to their susceptibility to acute mountain sickness, characterized in that: The method comprises the following steps: S1. Data reception: Receive data on the expression levels of six proteins, namely RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95, in the samples of the subjects; S2. Data processing: inputting the data into an acute mountain sickness susceptibility clustering model; the acute mountain sickness susceptibility clustering model uses the data of the expression levels of the six proteins of known acute mountain sickness patients and known plateau environment adaptors as training samples to train the acute mountain sickness susceptibility clustering model; S3. Output result: output whether the subject belongs to the group susceptible to acute mountain sickness through the acute mountain sickness susceptibility grouping model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 7 to 9 are implemented.
Citation Information
Patent Citations
Methods and Compositions for Treating Diseases Associated with Exhausted T Cells
US20210033595A1
Methods and systems for analyzing targetable pathologic processes in covid-19 via gene expression analysis
US20230220470A1
Methods and systems for predicting response to Anti-TNF therapies
US20230282367A1
Systems and Methods for Targeting COVID-19 Therapies
US20250011886A1
Methods for determining the risk of developing liver and lung toxicity
WO2005039588A2