Apparatus for acute mountain sickness (AMS) susceptibility stratification

By receiving and analyzing the expression levels of specific proteins in the red blood cell membranes of subjects, and using machine learning models for susceptibility grouping, the problem of lack of objective indicators in existing technologies has been solved, enabling accurate grouping of susceptibility to acute altitude sickness and reducing workload.

CN120015309BActive Publication Date: 2025-12-12ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510055508.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-12-12
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The lack of reliable objective indicators in current technology to differentiate the susceptibility of people to acute altitude sickness increases the difficulty of prevention and treatment of AMS.

Method used

A computer device is provided that receives data on the expression levels of six proteins—RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95—from a subject's sample, and uses an acute altitude sickness susceptibility clustering model to cluster the subjects. The model can be selected from random forest models, etc., and outputs whether the subjects belong to the susceptible population.

Benefits of technology

It enables accurate grouping of susceptibility to acute altitude sickness, reduces the workload of medical staff, and is suitable for widespread promotion and application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015309B_ABST
    Figure CN120015309B_ABST
Patent Text Reader

Abstract

The application discloses a device for acute mountain sickness (AMS) susceptibility grouping. Specifically disclosed is a computer device, comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to realize the steps of acute mountain sickness grouping by using the data of the expression levels of six proteins, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 on the peripheral blood red cell membranes of a subject. The application establishes an AMS grouping model based on machine learning, and verifies the specificity and reliability thereof, can accurately distinguish AMS patients from non-AMS subjects, can realize the rapid grouping task of acute mountain sickness, reduces the workload of medical personnel, and is suitable for wide promotion and application.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of medical care informatics, and particularly relates to a device for acute mountain sickness (AMS) susceptibility grouping. BACKGROUND

[0002] Acute mountain sickness is also called acute mountain sickness (AMS), which generally refers to a series of clinical syndromes such as dyspnea, nausea, dizziness, palpitation, sleep disorders and the like caused by the body's unadaptation to low pressure and hypoxia within a few hours to a few days when a person rapidly enters a high altitude area above 2500 meters from a plain or enters a higher altitude area from a high altitude area.

[0003] Every year, millions of people need to rapidly enter high altitudes for work, tourism, sports, rescue and the like. Low pressure and hypoxia are obvious environmental characteristics threatening the health of people in high altitude areas. AMS often occurs when a person from a plain who is not adapted to the environment rushes onto a high altitude area above 2500 meters, mainly with headache, accompanied by dizziness, insomnia, fatigue and the like. Although most AMS cases are not fatal, they have a rapid onset and progress. If not intervened in time, they can develop into serious high altitude diseases such as high altitude pulmonary edema, high altitude cerebral edema and even death. AMS is the most common high altitude disease so far. Epidemiological data shows that after people from a plain rapidly enter a high altitude area, the incidence of AMS is about 9% to 75%. With the increase of altitude, the incidence and mortality of AMS also increase, which seriously affects the physical health and working ability of people rapidly entering a high altitude area. The risk of AMS depends on the individual susceptibility, altitude and climbing speed. At present, there is no reliable objective index to distinguish AMS, which to some extent increases the difficulty of AMS prevention and treatment. SUMMARY

[0004] One of the technical problems to be solved by the present application is how to group people according to their susceptibility to acute mountain sickness. The technical problems to be solved are not limited to the technical subject described, and other technical subjects not mentioned herein can be clearly understood by those skilled in the art through the following description.

[0005] To solve the above technical problems, the present application first provides a computer device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the following steps:

[0006] S1, data receiving: receiving data of expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 in a sample of a subject;

[0007] S2, data processing: inputting the data into an acute high altitude reaction susceptibility clustering model; the acute high altitude reaction susceptibility clustering model uses the data of the expression levels of the 6 proteins of known acute high altitude reaction patients and known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility clustering model;

[0008] S3, output result: outputting whether the subject belongs to the acute high altitude reaction susceptible population by the acute high altitude reaction susceptibility clustering model.

[0009] The high altitude environment adaptor can be a person without symptoms after acute high altitude exposure.

[0010] The 6 proteins (RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95) all belong to red blood cell membrane proteins. Among them:

[0011] The GenBank accession number of the RAB2B protein is: 84932, and the UniProt annotation number is: Q8WUD1.

[0012] The GenBank accession number of the VPS4A protein is: 27183, and the UniProt annotation number is: Q9UN37.

[0013] The GenBank accession number of the SPTB protein is: 6710, and the UniProt annotation number is: P11277.

[0014] The GenBank accession number of the ANP32B protein is: 10541, and the UniProt annotation number is: Q92688.

[0015] The GenBank accession number of the DMTN protein is: 2039, and the UniProt annotation number is: Q08495.

[0016] The GenBank accession number of the VPS4A protein is: 27183, and the UniProt annotation number is: Q9UN37.

[0017] SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 all represent phosphorylated proteins.

[0018] SPTB T1387 represents SPTB protein phosphorylated at the 1387th threonine (T).

[0019] ANP32B T244 represents ANP32B protein phosphorylated at the 244th threonine (T).

[0020] DMTN S246 represents DMTN protein phosphorylated at the 246th serine (S).

[0021] VPS4A S95 represents VPS4A protein phosphorylated at the 95th serine (S).

[0022] Further, the sample of the subject can be a sample of peripheral red blood cell membrane of the subject.

[0023] Further, the acute high altitude reaction susceptibility clustering model can be at least one of a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model, and a hidden Markov model.

[0024] Further, the acute high altitude reaction susceptibility clustering model can be a random forest model.

[0025] Further, the hyperparameters of the random forest model can be as follows: the number of random forest trees can be 500.

[0026] The application also provides a device for clustering acute high altitude reaction susceptibility, which comprises a data receiving module, a data processing module, and a result output module.

[0027] The data receiving module can be used to receive data of expression levels of the six proteins RAB2B, VPS4A, SPTB T1387, ANP32BT244, DMTN S246, and VPS4A S95 in a sample of a subject.

[0028] The data processing module can be used to input the data into an acute high altitude reaction susceptibility clustering model; the acute high altitude reaction susceptibility clustering model uses data of expression levels of the six proteins of known acute high altitude reaction patients and known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility clustering model.

[0029] The result output module can be used to output whether the subject belongs to the acute high altitude reaction susceptible population through the acute high altitude reaction susceptibility clustering model.

[0030] The data processing module can analyze or interpret the data based on statistical methods (such as mean, median, mode, standard deviation, variance, percentage, etc.), machine learning-based methods, or deep learning-based methods (such as convolutional neural networks, recurrent neural networks, graph neural networks, graph convolutional networks, etc.).

[0031] Further, the data processing module comprises a machine learning model submodule, the input of the machine learning model is the data of the expression levels of the 6 proteins described herein, and the output is whether it belongs to the susceptible population of acute high altitude reaction.

[0032] The data processing module can be used to perform the following operations: inputting the data received by the data receiving module into the machine learning model, obtaining the output value through the machine learning model, thereby providing an indication of whether the subject belongs to the susceptible population of acute high altitude reaction.

[0033] The output value can be the class label of the sample, such as {0, 1} or {positive, negative}, etc. The output value can be mapped to a numerical value by mapping "positive" to 1 and "negative" to 0. The output value can also be a probability value or a real value. When the output value is a probability value or a real value, the probability value or real value can be compared with a preset threshold, thereby providing an indication of whether the subject belongs to the susceptible population of acute high altitude reaction, for example, when the probability value or real value is greater than or equal to the preset threshold, it indicates acute high altitude reaction positive (belongs to the susceptible population of acute high altitude reaction), and when the probability value or real value is less than the preset threshold, it indicates acute high altitude reaction negative (does not belong to the susceptible population of acute high altitude reaction).

[0034] The data processing module can also comprise a training submodule, which is used to train the machine learning model. The machine learning model can be trained based on a labeled data set to optimize its accuracy and / or reliability in the susceptible population of acute high altitude reaction.

[0035] The data processing module can also comprise a performance evaluation submodule, which can be used to evaluate the accuracy and / or reliability of the machine learning model, such as drawing a ROC curve and calculating the AUC value of the machine learning model. The AUC value is a measure of model performance, and the higher the value, the better the performance of the model. The classification threshold of the machine learning model can also be determined by the performance evaluation submodule. The classification threshold can be used to indicate the decision boundary of the model in distinguishing different categories.

[0036] The data processing module can also comprise a model optimization submodule, which can be used to optimize the performance of the machine learning model by adjusting hyperparameters and / or using different training strategies.

[0037] The machine learning model submodule can be a pre-trained machine learning model submodule.

[0038] The machine learning model submodule can also be a trained machine learning model submodule obtained by training and evaluating the training submodule, performance evaluation submodule, and model optimization submodule.

[0039] The result output module can output the calculation result of the trained machine learning model as whether it belongs to the acute high altitude reaction susceptible population.

[0040] Further, the machine learning model can be selected from a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model, and a hidden Markov model.

[0041] Further, the machine learning model can be a random forest model.

[0042] The present application also provides a method for constructing an acute high altitude reaction susceptibility grouping model, which comprises the following steps: using the data of the expression levels of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the samples of known acute high altitude reaction patients and known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility grouping model.

[0043] In the above method, the method can comprise the following steps:

[0044] A1) receiving the data of the expression levels of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the samples of known acute high altitude reaction patients and known high altitude environment adaptors;

[0045] A2) using the data as input features of a machine learning model, training and evaluating the machine learning model;

[0046] A3) obtaining the trained machine learning model, i.e. the acute high altitude reaction susceptibility grouping model, according to the evaluation result.

[0047] In the case where the input features are clear, the training method of the machine learning model is known to those skilled in the art. For example, after using the expression levels of the six proteins described in the present application as input features of the machine learning model, the model can be trained according to the input data sample (labeled data set), and optimization algorithms such as gradient descent method, Newton iteration method, particle swarm optimization algorithm, artificial fish swarm algorithm, artificial bee colony algorithm and firefly algorithm can be used to iteratively update the model parameters to minimize the loss function.

[0048] During the model training process, the model can be evaluated to determine the performance of the model. Evaluation indicators can generally include ROC curve (AUC value), accuracy, precision, recall, F1 score, etc. For example, the ROC curve and AUC value are two core indicators for evaluating the performance of a classification model. When the model is evaluated based on the ROC curve (AUC value), the closer the ROC curve is to the upper left corner, the better the performance of the model. The point (0, 1) in the upper left corner indicates that the model can perfectly classify. AUC is the area under the ROC curve, which is a comprehensive indicator for quantifying model performance. The AUC value ranges from 0 to 1, and the closer to 1, the better the model performance, indicating that the model has stronger ability to distinguish samples. AUC is not affected by the classification threshold, providing an overall performance evaluation.

[0049] Further, according to the evaluation results of the model, the model can also be optimized, for example, the model can be optimized by adjusting the hyperparameters and / or using different training strategies to further improve the performance of the model. Cross-validation, grid search method and / or Bayesian optimization algorithm can be used for model selection and hyperparameter adjustment.

[0050] The trained machine learning model can be used as an acute high altitude reaction susceptibility clustering model for prediction or classification tasks in practical problems.

[0051] In the above method, the data in A1) can be obtained by Western Blot, immunofluorescence, radioimmunoassay (RIA), co-immunoprecipitation, immunohistochemistry (IHC), enzyme-linked immunosorbent assay (ELISA), double-antibody sandwich ELISA, fluorescence-linked immunosorbent assay (FLISA), enzyme immunoassay (EIA), flow cytometry (FACS), polyacrylamide gel electrophoresis, capillary electrophoresis, near-infrared spectroscopy, immunochemiluminescence, colloidal gold immunotechnology (GIC), colloidal gold immunochromatography technology (GICA), fluorescence immunochromatography technology, surface plasmon resonance technology, immuno-PCR technology, biotin-avidin technology, liquid chromatography, high-performance liquid chromatography, solid-phase chromatography, mass spectrometry, liquid chromatography-mass spectrometry (LC-MS), high-performance liquid chromatography-mass spectrometry (HPLC-MS), liquid chromatography-tandem mass spectrometry (LC-MS / MS), two-dimensional electrophoresis-mass spectrometry (2DE-MS), two-dimensional liquid chromatography-mass spectrometry (2D-LC-MS), isotope-coded affinity tag technology (ICAT), flying mass spectrometry (SELDI-TOF-MS), or protein chip technology, but not limited thereto.

[0052] In the above method, the machine learning model can be selected from a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model, and a hidden Markov model.

[0053] In the above method, the machine learning model can be a random forest model.

[0054] The output value of the random forest model described herein can be a class label (which class the sample belongs to) or a probability value.

[0055] For class labels: In a random forest model, each sample is assigned a predicted probability of belonging to a certain class. These predicted probabilities can be converted into specific class labels by a classification threshold. The classification threshold is not fixed and can be adjusted according to actual needs. For example, when more attention is paid to the sensitivity or specificity of the model, the classification threshold can be adjusted to trade off between sensitivity and specificity, lowering the classification threshold can improve the sensitivity of the model, and raising the classification threshold can improve the accuracy and specificity of the model; when the class distribution is uneven, the classification threshold needs to be adjusted to improve the performance of the model; cross-validation and grid search methods can also be used to determine the optimal classification threshold; the ROC curve can also be drawn and the AUC value can be calculated to select the threshold with a higher AUC value as the best classification threshold. Generally, the classification threshold can be set to the default value (0.5), which means that if the predicted probability is greater than 0.5, the sample is classified as positive, otherwise it is classified as negative.

[0056] Subjects with a probability greater than or equal to 0.5 are more susceptible to acute high altitude reactions than subjects with a probability less than 0.5.

[0057] For probability values: In a random forest model, each decision tree makes a prediction for a sample and gives the probability value of belonging to each class. Then, the probability values are averaged (or other aggregation methods are used) to get the final class probability distribution. These probability values can provide more model information and can be used to build more complex decision systems, evaluate model performance, etc.

[0058] The hyperparameters of the random forest model described herein can be: the number of random forest trees is 500.

[0059] As understood by those skilled in the art, the selection of hyperparameters is not unique, and the purpose of determining hyperparameters is to help the model to train well on the data set. As long as the model can predict accurately and does not appear underfitting and overfitting, any combination of hyperparameters is acceptable.

[0060] The application method of the acute high altitude reaction susceptibility grouping model obtained by any of the methods of the application can include the following steps:

[0061] (1) receiving the expression level data of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95 in the sample of the subject to be tested;

[0062] (2) inputting the data into the acute high altitude reaction susceptibility grouping model obtained by any of the methods of the application (for example, the trained random forest model constructed by the method of the application), and obtaining the grouping result of the subject to be tested by the acute high altitude reaction susceptibility grouping model.

[0063] Further, the subject to be tested is unknown whether or not to have acute high altitude reaction.

[0064] Further, the subject to be tested can be an asymptomatic healthy subject, or a suspected AMS patient, for example, a suspected AMS patient with at least one of the following symptoms: headache, fatigue, weakness, dizziness, dizziness, sleep disorder, insomnia, anorexia, dyspnea, palpitation, nausea, vomiting, etc.

[0065] Further, the subject to be tested can be a person located at an altitude of 2000 or 2500 meters or above.

[0066] The application also provides a method for grouping the susceptibility of a population to acute high altitude reaction, which comprises the following steps:

[0067] S1, data receiving: receiving the expression level data of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95 in the sample of the subject to be tested;

[0068] S2, data processing: inputting the data into the acute high altitude reaction susceptibility grouping model; the acute high altitude reaction susceptibility grouping model uses the expression level data of the six proteins of the known acute high altitude reaction patients and the known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility grouping model;

[0069] S3, output result: outputting whether the subject belongs to the acute high altitude reaction susceptible population by the acute high altitude reaction susceptibility grouping model.

[0070] The application also provides a computer readable storage medium storing a computer program, wherein the computer program, when executed by a processor, can implement the steps of any of the methods described herein (a method for constructing an acute mountain sickness susceptibility grouping model or a method for grouping a population according to acute mountain sickness susceptibility).

[0071] Further, the computer readable storage medium can further execute an output step of the acute mountain sickness susceptibility grouping result.

[0072] The sample described herein can be derived from a blood sample (including plasma, serum and whole blood).

[0073] Further, the sample can be a peripheral red blood cell membrane sample.

[0074] The six proteins RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 described herein can be used to determine whether a subject belongs to an acute mountain sickness (AMS) susceptible population. The grouping can be performed by detecting the expression levels of the six proteins and comparing the detection results with a preset threshold (or reference value) that can effectively distinguish AMS patients and non-AMS subjects. The comparison includes any type of comparison of the expression levels of the six proteins with the preset threshold (or reference value). For example, the comparison can be performed by mathematically correlating (processing or operating based on statistical analysis methods) the expression levels of the six proteins to obtain a score or index, and comparing the score or index with the preset threshold (or reference value) to obtain a classification result. The comparison can also be performed by inputting the expression levels of the six proteins as independent variables into a machine learning model (such as a random forest model, a support vector machine model, etc.), and then outputting the model to obtain a classification result. Those skilled in the art can understand that the comparison can be performed manually or with the aid of a computer. For example, the values to be compared (such as the values output by the model) can be compared with the preset threshold (or reference value) manually, or the comparison can be performed automatically by a computer program executing an algorithm. The computer program performing the comparison can provide the grouping result in a suitable output format.

[0075] The application performs proteomics and phosphoproteomics analysis of red blood cell membranes of AMS patients and nAMS volunteers in plateau areas. By characterizing the protein and phosphoprotein profiles of the subjects, the relationship between red blood cell membranes and AMS is analyzed and mined. In addition, the application analyzes the correlation of red blood cell membrane proteins and phosphoproteins, studies the correlation between key difference molecules in RBC membranes (red blood cell membranes) and the severity of AMS, and establishes an AMS grouping model based on machine learning.

[0076] The present application first screened RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4AS95, and based on this, an acute high altitude reaction susceptibility grouping model was constructed. At present, the main method at home and abroad is to score the subjective symptoms of patients by Lewis Lake acute mountain sickness scoring system (LLSS), which needs to fill in the questionnaire, and the work efficiency is low and the workload is large. The present application establishes a grouping model based on machine learning, and verifies its specificity and reliability, which can accurately distinguish AMS patients and non-AMS subjects, realize the rapid grouping of acute high altitude reaction, reduce the workload of medical personnel, and is suitable for wide promotion and application.

[0077] Definitions of terms

[0078] In the present application, the scientific and technical terms used herein have the meanings commonly understood by a person skilled in the art, unless otherwise specified. In order to better understand the present application, the definitions and explanations of related terms are provided as follows.

[0079] The term "apparatus" generally refers to a system comprising modules or units operably connected to each other to enable computation. It can implement one or more modules or units through software (such as computer programs, web applications, mobile applications, and standalone applications, etc.), hardware (such as processors or memories, etc.), or a combination thereof. The apparatus encompasses various forms, such as, but not limited to, server computers, desktop computers, laptop computers, notebook computers, subnotebook computers, netbook computers, netpad computers, set-top computers, streaming media devices, handheld computers, mobile smartphones, tablet computers, and personal digital assistants, etc.

[0080] The term "data receiving module" generally refers to a functional module for obtaining relevant data. In the present application, the data receiving module can refer to a functional module for obtaining the expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4AS95. The data receiving module can also include reagents and / or instruments for detecting the expression levels of the six proteins to perform the step of detecting the expression levels of the six proteins in the test sample. The step of detecting the expression levels of the six proteins in the test sample can also be performed by an independent apparatus (such as a kit, a mass spectrometer, a chromatograph, a flow cytometer, etc.) and is not included in the data receiving module.

[0081] The term "data processing module" generally refers to a functional module responsible for data processing (including pre-processing of data, etc.). In the present application, the data processing module can be used to process and / or analyze the expression level data of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the sample to be tested, and achieve the purpose of grouping the susceptibility to acute high altitude reaction according to the analysis processing result. The data processing module can process and analyze data based on a machine learning model.

[0082] The term "result output module" generally refers to a functional module for displaying result information to a user. In the present application, the result output module can refer to a functional module for displaying the grouping result of acute high altitude reaction. The result output module can include a display, a projector, a monitor, a printer, an audio output device, etc. but is not limited thereto.

[0083] The term "machine learning model" generally refers to a computational model or algorithm. A machine learning model can learn and analyze input data to perform tasks such as prediction, classification, identification or decision making. The learning process of a machine learning model can be based on statistical principles and data pattern recognition, using a training data set to adjust model parameters and optimize the model to improve its prediction or inference ability. It can predict or classify new data by learning the mapping relationship between features and labels in the training data. In machine learning, "label" generally refers to the true class or target value of a data sample, which is the "reference answer" when training a machine learning model. In a classification task, the label can represent which category the data sample belongs to. In a regression task, the label can be a continuous target value or a real number target value to be predicted. "Feature" generally refers to the attributes or characteristics of a data sample. In machine learning, features are input variables used to train the model, which can be numerical, categorical or any other type of data. "Training" generally refers to the process of model learning, which uses labeled training data to adjust the parameters of the model to enable the model to accurately predict the label to the maximum extent. During the training process, the model can continuously optimize its parameters to approximate the true data distribution. Machine learning models can use various algorithms and techniques to train and optimize through supervised learning, unsupervised learning or reinforcement learning. Exemplary machine learning models can include random forest models, support vector machine models, linear discriminant analysis models, recursive feature elimination models, logistic regression models, neural network models, CART models, flextree models, LART models, MART models, k-nearest neighbor classification models, clustering analysis models, principal component analysis models, Bayesian classification models and hidden Markov models, etc. In one or more embodiments of the present application, a random forest model is used to establish a susceptibility grouping model for acute high altitude reaction.

[0084] The term "random forest model" generally refers to an ensemble learning model that uses bagging algorithm to build on multiple decision trees as base learners. For classification problems, the class of the output result is determined by the mode (the majority result) of the class output by all decision trees, i.e., the final output result is determined by voting of all single decision tree models, and the output is the class label of the sample (which class the sample belongs to). For regression problems, the output result can depend on the average of the results of decision trees, and the output is a real number (such as a probability value).

[0085] The term "computer-readable storage medium" generally refers to any medium that can store data and can be accessed by a computer, which can include a computer program. The computer-readable storage medium can include volatile and non-volatile media and removable and non-removable media. As an example and not a limitation, the computer-readable storage medium can include ROM (Read-Only Memory), PROM, EPROM, EEPROM, FLASH-EPROM, CD-ROM, DVD-ROM, RAM (Random Access Memory), SRAM, DRAM, SDRAM, RDRAM, DDR SDRAM, GDDR SDRAM, hard disk drive (HDD), solid-state drive (SSD), optical disk drive (such as CD, DVD drive, etc.), flash memory (such as USB), magnetic disk drive, tape drive, distributed computing system (such as cloud memory, cloud disk, etc.).

[0086] The terms "device" and "system" generally have the same meaning and can be used interchangeably.

[0087] The terms "module" and "unit" generally have the same meaning and can be used interchangeably.

[0088] The term "receiver operating characteristic curve (ROC curve)" generally refers to a curve plotted with the true positive rate as the ordinate and the false positive rate as the abscissa according to a series of different binary classifications. The AUC (area under the curve) represents the area enclosed by the ROC curve and the horizontal axis. The ROC curve reflects the relationship between sensitivity and specificity. The principle of making a ROC curve is to set a plurality of different critical values for a continuous variable, calculate the corresponding sensitivity and specificity at each critical value, and then plot the curve with sensitivity as the ordinate and 1-specificity as the abscissa. Since the ROC curve is composed of a plurality of critical values representing the sensitivity and specificity of each, the optimal limit value of a method can be selected by means of the ROC curve. The closer the ROC curve is to the upper left corner, the higher the sensitivity of the test and the lower the misjudgment rate, and the better the performance of the method. It can be seen that the point on the ROC curve closest to the upper left corner has the maximum sum of sensitivity and specificity, and the value corresponding to this point or its adjacent point is often used as a reference value (also known as a threshold value).

[0089] The term "threshold value" generally refers to a dividing value used as a reference to obtain information about a measurement result and / or to classify a measurement result. For example, a threshold value is set when defining a dividing line between two subsets in a population. Values equal to or above the threshold level define one subset of the population, and values below the threshold level define another subset of the population. The threshold value can be determined based on one or more control samples or an entire control sample population. The threshold value can be determined before, simultaneously with, or after performing the measurement of interest.

[0090] The term "reference value" generally refers to a reference level that can be obtained in advance from a subject, including a healthy subject and a patient, and the appropriate reference level can be measured and selected according to techniques known to those skilled in the art.

[0091] The term "comprising" is not intended to be limiting, is intended to be inclusive, and means that other elements beyond those listed can be present. The term "comprising" also encompasses the terms "consisting of and "consisting essentially of." The terms "comprising" and "including" are used interchangeably herein. BRIEF DESCRIPTION OF DRAWINGS

[0092] Figure 1 Flow for the exploration of the erythrocyte membrane proteome and phosphoproteome in AMS.

[0093] Figure 2 Detailed demographic characteristics for the participants.

[0094] Figure 3 Characteristics of the dysregulated erythrocyte membrane proteins in the AMS group.

[0095] Figure 4 Characterization of red blood cell membrane phosphoproteins / peptides.

[0096] Figure 5 Characterization of dysregulated red blood cell membrane phosphoproteins in AMS group.

[0097] Figure 6 Kinase prediction and motif analysis of dysregulated phosphoproteins.

[0098] Figure 7 Correlation analysis of AMS severity with proteome and phosphoproteome features.

[0099] Figure 8 Variable importance ranking in AMS random forest clustering model.

[0100] Figure 9 Classification performance of random forest clustering model on validation dataset.

[0101] Figure 10 Flowchart for acute mountain sickness susceptibility clustering in the present invention.

[0102] Figure 11 Schematic diagram of the device for acute mountain sickness susceptibility clustering in the present invention. DETAILED DESCRIPTION

[0103] The present invention will be further described in details with reference to the specific embodiments, and the examples given are only for illustrating the present invention, but not for limiting the scope of the present invention. The examples provided below can be used as a guide for further improvement by those skilled in the art, and do not constitute any limitation on the present invention.

[0104] The experimental methods in the following examples are all routine methods, and are carried out according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained commercially, unless otherwise specified.

[0105] Example 1, Screening of red blood cell membrane proteins and phosphoproteins

[0106] Red blood cells are the most abundant cell type in the human body, responsible for transporting O2 and CO2. Red blood cell membrane protein homeostasis is closely related to hypoxia, and hypoxia is closely related to membrane deformability, glycolysis process, oxygen molecule release and transport, and regulation of iron ion balance. Under hypoxic conditions, red blood cells need to be able to pass through narrow blood vessels and deliver oxygen to hypoxic tissues. At this time, the deformability of red blood cells is particularly important. Studies have shown that hypoxia can affect the deformability of red blood cells by affecting the cytoskeleton, NO and ATP production, ion homeostasis, and regulation of cell membrane protein composition. In addition, mature red blood cells have no obvious organelle structure, and phosphotriose and glycolysis are the main energy supply. By sensing the change of oxygen content in the circulatory system, the glycolysis of red blood cells under hypoxic conditions is significantly increased, promoting the production of ATP and 2,3-BPG (2,3-diphosphoglycerate), thereby promoting oxygen release and relieving hypoxia in peripheral tissues. Acute hypoxia also leads to a significant decrease in red blood cell cytoskeletal proteins and glucose transport proteins: spectrin, ankyrin, band 3, band 4.1, band 4.2, and GAPDH. In addition, due to insufficient ATP supply, the activity of ion transport proteins on the red blood cell membrane is significantly reduced during acute hypoxia, suggesting that ion homeostasis in red blood cells is affected. At present, although there have been studies exploring the effects of hypoxia on the composition and function of red blood cells, there is still a lack of systematic research to explore whether red blood cells in AMS patients exhibit specific changes or more significant adjustments, and whether these changes are directly related to the severity of AMS.

[0107] Protein phosphorylation is one of the important post-translational modifications of organisms and an effective adjustment method for organisms to quickly respond to changes in the external environment. Phosphorylation is involved in the regulation of almost all life activities (such as cell proliferation, signal transduction, metabolism, etc.). It is reported that phosphorylation is also involved in the regulation of enzyme activity in red blood cells and the control of red blood cell deformability. Proteomics and phosphoproteomics are suitable for studying the pathogenesis of AMS from the perspective of red blood cells, because only the regulation of existing proteins in red blood cells, but not the regulation of newly synthesized proteins.

[0108] The present application performs proteomics and phosphoproteomics analysis on the red blood cell membranes of AMS patients and nAMS volunteers in high altitude areas, and screens red blood cell membrane proteins and phosphorylated proteins for acute mountain sickness (AMS) grouping. The specific method is as follows:

[0109] 1. Screening sample information

[0110] Clinical demographic information and study design: According to the Lake Louise Scoring System (LLSS), a group of volunteers (see Table 1, training set samples) was recruited, including AMS patients (n=5) and nAMS subjects (n=5) (Figure 2 A). Obtain detailed clinical characteristics, including age, blood oxygen saturation (SaO2), complete blood count indicators, and coagulation indicators. Figure 1 A). Compared with the nAMS group, the SaO2 in the AMS group was significantly reduced ( Figure 2 B), while the heart rate did not change significantly ( Figure 2 C). In addition, the AMS group also showed abnormalities in coagulation parameters ( Figure 2 D), such as elevated PT and decreased PTA, suggests that coagulation function may be weakened in the AMS group. However, platelet count and other related indicators showed no significant changes. Figure 2 E, Figure 2 F). Correspondingly, no abnormalities were found in indicators such as red blood cell count. Then, we extracted red blood cell membrane proteins and plotted the red blood cell membrane proteome and phosphorylated proteome of this cohort. Figure 1 B).

[0111] The Lake Louise Score System (LLSS) is assessed as follows: Four symptoms are evaluated: headache, gastrointestinal symptoms, fatigue and / or weakness, and dizziness and / or vertigo. Each of the four symptoms is rated from 0 to 3 based on its severity. 0 indicates no symptoms; 1 indicates mild symptoms; 2 indicates moderate symptoms; and 3 indicates severe symptoms. Gastrointestinal symptoms include appetite, nausea, and / or vomiting, rated as follows: 0 indicates no symptoms and good appetite; 1 indicates poor appetite or mild nausea; 2 indicates moderate nausea or vomiting; and 3 indicates severe nausea and vomiting. A total score of 4 or higher indicates acute altitude sickness (AMS) and is assessed as an AMS patient; a total score of 3 or lower indicates no acute altitude sickness and is assessed as a non-AMS subject (nAMS subject).

[0112] Table 1. Information on 10 selected samples

[0113]

[0114] 2. Sample preparation and protein detection methods

[0115] Isolation of red blood cell membrane from peripheral blood samples: The peripheral blood sample of the subject was left to stand for 30 minutes, and then centrifuged at 3000 rpm for 15 minutes. The upper plasma, the intermediate layer of white membrane and the lower part of the upper red blood cells were discarded. 300 microliters of red blood cell precipitate was mixed with hypotonic Tris-HCl buffer (volume ratio 1:30, pH=8.0, containing 0.01M Tris-HCl, 2mM sodium orthovanadate, 1mM benzylsulfonyl fluoride and Roche protease inhibitor), and incubated at 4°C for 1 hour to extract the red blood cell membrane. Then, the red blood cell membrane was separated by centrifugation at the maximum speed for 15 minutes at 4°C, and the membrane precipitate was washed with 1 mL of hypotonic Tris-HCl buffer, and the operation was repeated until a white precipitate was obtained. The red blood cell membrane precipitate was added with pre-cooled SDC lysis buffer (containing 4% weight / volume of deoxycholic acid sodium, pH=8.5, 100mM Tris-HCl) at 4°C, mixed immediately and denatured at 95°C for 5 minutes, and then ultrasonically treated at 4°C for 10 minutes in a water bath at the maximum power. After centrifugation, the protein concentration was determined by the BCA method.

[0116] Enrichment and detection of phosphorylated peptides: 300 μg of protein was reduced and alkylated with 1:10 volume of reducing / alkylating buffer (containing 100 mM TCEP-HCl and 400 mM chloroacetamide) at 37°C for 15 minutes. 3 μg of trypsin was added, and after mixing well, the sample was digested overnight. The next day, the digestion reaction was terminated with 2% trifluoroacetic acid (TFA). The sample was mixed and centrifuged at high speed to remove SDC. The SDC precipitate was washed with 100 microliters of 0.5% TFA for three times. All supernatants were combined and purified with a C18 column. Finally, 1 μg of peptide mixture was taken for liquid chromatography-tandem mass spectrometry (LC-MS / MS) analysis, and the rest was used for phosphorylated peptide enrichment. 200 μg of peptides were taken for phosphorylated peptide enrichment (Thermo, High-Select Phospho- Peptide Enrichment Kit). The enriched phosphorylated peptides were purified with a self-made C18 column and subjected to LC-MS / MS analysis. TM Fe-NTA phospho-peptide enrichment kit). The enriched phosphorylated peptides were purified with a self-made C18 column and subjected to LC-MS / MS analysis.

[0117] 3. Red blood cell membrane proteome characteristics of AMS patients

[0118] A total of 2383 proteins were identified from 10 red blood cell membrane samples, of which 329 proteins were uniquely identified in the AMS group, which was higher than that in the nAMS group ( Figure 3 A) After normalizing the protein quantification information, we found that the red blood cell membrane proteome characteristics of the two groups were significantly different by UMAP analysis ( Figure 3 B) The differential analysis results showed that there were 189 differentially expressed proteins (DEP) between the AMS group and the nAMS group, of which 69 DEP were down-regulated and 120 DEP were up-regulated ( Figure 3C). KEGG pathway enrichment analysis showed that 69 down-regulated membrane proteins were highly related to ferroptosis, fatty acid metabolism, endocytosis, mineral absorption and tight junction Figure 3 D). In addition, we also analyzed the KEGG pathways enriched by up-regulated membrane proteins. These DEPs were mainly involved in complement and coagulation cascade, TCA cycle, peroxisome, regulation of actin cytoskeleton and focal adhesion Figure 3 E). Finally, we showed the DEPs enriched in energy metabolism-related pathways in Figure 3 F) by heat map Figure 3 D and Figure 3 E.

[0119] 4. Phosphoproteomic characterization of red blood cell membranes at high altitude

[0120] Using phosphoproteomic technology, 1076 phosphorylated proteins were identified in all red blood cell membrane samples, of which 922 phosphorylated proteins were universally present in AMS and nAMS, 130 phosphorylated proteins were uniquely identified in AMS, and 24 phosphorylated proteins were uniquely identified in nAMS Figure 4 A). A total of 2913 credible phosphorylated peptides were screened according to the probability of phosphorylation modification sites > 0.8. Among them, 2417 phosphorylated peptides were universally present in AMS and nAMS, and 387 phosphorylated peptides were uniquely identified in AMS, more than in nAMS Figure 4 B).

[0121] We then reviewed the basic information of this dataset. The phosphopeptide quantification information spanned 6 orders of magnitude Figure 4 C). In addition, statistical analysis of the types of phosphorylation modification sites showed that the proportion of serine (S) modification was much higher than that of T and Y modification in AMS and nAMS: S was 84.77% (AMS), 84.32% (nAMS); T was 13.80% (AMS), 14.29% (nAMS), and Y was 1.43% (AMS), 1.39% (nAMS) Figure 4 D). About 46% of the phosphorylated proteins had two or more phosphorylation modification sites. Interestingly, the top 7 were SPTA1, ADD2, ANK1, SPTB, EPB41, DMTN and ADD1, all of which were related to spectrin and actin. Motif analysis of T and S phosphorylated peptides showed that xxxTPPxxx and xxxRxDSxxx conserved motifs were significantly enriched in T and S modification sequences, respectively Figure 4 E), indicating that specific kinases may have contributed to this event in red blood cells.

[0122] 5. Phosphorylation disorder events in red blood cell membranes of AMS patients

[0123] After normalizing the quantitative information of phosphorylated peptides, we found through UMAP analysis that the phosphorylated peptide characteristics of AMS and nAMS are different. Figure 4 F). Differentially expressed phosphorylated peptides (DEPPs) were screened based on fold change and p-value, and the results are as follows: Figure 5 As shown in Figure A, blue represents significantly downregulated phosphorylated peptides (n=48), while red represents significantly upregulated phosphorylated peptides (n=108). Further analysis revealed consistent trends in the changes of multiple DEPPs within the same protein. DEPPs were significantly decreased in EEPD1, AGO2, SPTB, and ANP32B, while they were significantly upregulated in other proteins. Figure 5 B). Kinase-substrate annotation information was used to predict potential kinases in DEPPs. Results showed that CSNK2A1 kinase activity was decreased, while the activities of other kinases were upregulated. In particular, PRKCG kinase activity was significantly upregulated (B). Figure 6 A). Then, we performed motif enrichment analysis on S-modified DEPPs and found significant enrichment of the xxxSPxxx motif (). Figure 6 B).

[0124] 6. DEPP-matched proteins are used for functional enrichment analysis.

[0125] Downregulated phosphorylated proteins were significantly enriched in necroptosis, calcium reabsorption, mineral uptake, and glycolysis / gluconeogenesis. Figure 5 C). Upregulated proteins are mainly enriched in platelet activation, actin cytoskeleton regulation, ECM-receptor interaction, and other pathways related to intercellular communication and cell morphology maintenance. Figure 5 D). Furthermore, we found that the downregulated phosphorylated proteins are mainly located in the cellular cortex and cortical cytoskeleton, and have functions related to spectroscopy binding, cytoskeleton structural composition, and P-type ion transporter activity. These downregulated phosphoproteins are primarily involved in the regulation of actin filament depolymerization and cell shape regulation. Figure 5 E). Upregulated phosphorylated proteins were enriched in membrane rafts, secretory granule membranes, and glycoprotein complexes, suggesting good extraction efficiency of RBC membrane proteins. These upregulated phosphorylated proteins are involved in regulating Fc receptor-mediated stimulus signaling pathways, coagulation, and cellular metal ion homeostasis. We also noted that the upregulated phosphorylated proteins have functions related to cadherin binding, cAMP-dependent protein kinase regulatory activity, and cytoskeleton structural composition. Figure 5 F).

[0126] 7. Correlation and classification model construction of membrane phosphopeptides / proteins with AMS

[0127] Correlation analysis was performed on erythrocyte membrane phosphopeptides / proteins and acute myeloid leukemia (AMS). The Lake Louise Subjective Rating System (LLSS) was used as the standard for assessing AMS severity. First, it was found that AMS severity was positively correlated with PT and PTR, and negatively correlated with PTA. Then, DEPPs with correlation coefficients greater than 0.6 were selected. Among them, 11 DEPPs—ANP32B T244, NAP1L1 T62, ADD3S42, VPS4A S95, GAB1 S419, ANK1 S1756, CC2D1A T204, BTK S55, SPTB T1387, DMTNS264, and USP6NL S415—were negatively correlated with AMS severity, while FLNA S2152 and TUBA4A S439 were positively correlated with AMS severity. Figure 7 A). For DEPs, correlation analysis showed that 9 DEPs (RAB2B, STX6, SLC30A7, VPS4A, CCNY, BMP2K, DHX9, BROX, and DLG1) were negatively correlated with AMS, while GC and HP1BP3 were positively correlated with AMS. Figure 7 B).

[0128] Subsequently, we investigated whether the 13 DEPPs and 11 DEPs highly correlated with AMS severity could be used to distinguish between AMS and nAMS. To optimize the classification model, the importance of these 24 variables was calculated using random forest analysis. Figure 8 A-8B), and based on the distribution of error rate with the number of variables, determined 6 variables as the optimal model composition. Then, combined with the tree with the minimum error rate ( Figure 8 C) and the importance analysis of DEPP and DEP, we selected a combination of 6 variables ( Figure 7 The final machine learning model is constructed using C-7D, RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95.

[0129] Example 2: Construction of a Susceptibility Group Model for Acute Altitude Sickness (AMS)

[0130] To achieve the highest clustering efficiency with the fewest possible variable combinations, the expression levels of the six variables selected in Example 1 (RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95) and the ratings from the Lake Louise Subjective Rating System were used as parameter sample data. A computational model was obtained through machine learning methods. The method for establishing the AMS susceptibility clustering model may include the following steps:

[0131] S1) receiving the expression levels of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the known acute high altitude reaction patient and the known healthy person sample, to obtain an original data set;

[0132] S2) selecting a machine learning model (such as a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model, a hidden Markov model, and combinations thereof).

[0133] S3) taking the original data set as the input features of the machine learning model, training and evaluating the machine learning model based on the original data set, and obtaining an AMS susceptibility clustering model according to the evaluation result.

[0134] In this embodiment, the random forest algorithm is used to construct the AMS susceptibility clustering model, and the expression levels of the six proteins RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in the to-be-tested sample are taken as the input of the random forest model, and whether it belongs to the acute high altitude reaction susceptible population is taken as the output of the random forest model. The hyperparameters of the random forest model are screened and determined by 5-fold cross-validation combined with grid search method, and the hyperparameters are as follows:

[0135] The number of random forest trees (n_estimators) is 500.

[0136] The random forest algorithm uses the randomForest package, and all programs are executed in the R environment. For the parameter values not explicitly stated, the default values of the algorithm package are used.

[0137] The classification threshold of the AMS susceptibility clustering model is calculated according to the ROC curve, so as to construct the AMS susceptibility clustering model. By setting a reasonable classification threshold, the output result of the random forest model can be converted into a specific clustering result. In this embodiment, the classification threshold of the random forest model is 0.5.

[0138] Specifically, when the output result of the model is higher than the classification threshold, the to-be-tested sample is judged as AMS positive (belongs to the acute high altitude reaction susceptible population); when the output result is lower than the classification threshold, the to-be-tested sample is judged as AMS negative (does not belong to the acute high altitude reaction susceptible population). The size of the threshold value can be set according to the actual needs of those skilled in the art, and the present application does not limit it. In addition, the hyperparameters are also non-limiting, and can be adjusted and optimized as needed.

[0139] The model construction results are shown in FIG. 2A. Figure 7 D. The constructed random forest model achieved a classification effect of area under the ROC curve 1 in the training set. At the same time, the model was applied to evaluate 10 participants in this study, and clearly distinguished AMS patients and nAMS subjects (FIG. 2B). Figure 7 C). These results fully demonstrate the potential clinical application value of the model.

[0140] Example 3, verification of the clinical application value of the model

[0141] 1. Verification sample information

[0142] Six verification samples different from the screening samples in Example 1 were selected, as shown in Table 2.

[0143] Table 2, information of six verification samples

[0144] ID N10 N12 N9 Y1 Y3 Y4 Data Classification Validation Set Validation Set Validation Set Validation Set Validation Set Validation Set Type nAMS nAMS nAMS AMS AMS AMS Symptom Score 1 1 1 6 4 5 Oxygen Saturation 84 87 90 87 84 84 Age 22 25 24 22 23 22 Gender Male Male Male Male Male Male Heart Rate 100 90 96 89 112 80 PT 9.4 9.5 9.7 15.1 11 10.2 PTR 0.79 0.8 0.82 1.28 0.93 0.87 PT-INR 0.79 0.8 0.81 1.29 0.93 0.86 FIB 1.98 1.44 1.96 2.7 2.84 3.18 PTA 133.1 131 126.9 58.8 103.6 116 PLT 306 168 216 229 184 284 PDW 15 17 20 20 21 14 RBC 5.6 5.7 5.2 5.3 5.6 5.6 HCT 51 51 49 52 51 51 HGB 167 175 155 172 162 162

[0145] 2. Sample preparation and protein detection

[0146] The same as Example 1.

[0147] The detection results were input into the random forest model constructed in Example 2 for verification. The verification results are shown in FIG. 3A. Figure 9 The AUC value of the random forest model reached 0.889 (FIG. 3B), and the specificity was 66.7%. It is shown that the random forest model characterized by the expression levels of RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTNS246 and VPS4A S95 has good discrimination ability for acute high altitude reaction (FIG. 3A), and has high clinical application value. Figure 9 B). Figure 9

[0148] The method for acute high altitude reaction susceptibility grouping according to the present application has a flow chart as shown in FIG. 4, which comprises the following steps. Figure 10

[0149] S1, data receiving: receiving the data of the expression levels of RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTNS246 and VPS4A S95 in the sample of the subject;

[0150] S2, data processing: inputting the data of the expression levels of the six proteins into the acute high altitude reaction susceptibility grouping model; the model uses the data of the expression levels of the six proteins of known acute high altitude reaction patients and known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility grouping model; ​​

[0151] S3. Output Results: The acute altitude sickness susceptibility cluster model outputs whether the subject belongs to the population susceptible to acute altitude sickness.

[0152] Accordingly, the present invention provides a device for susceptibility grouping for acute altitude sickness, such as... Figure 11 As shown, the system includes a data receiving module, a data processing module, and a result output module. The data receiving module receives data on the expression levels of six proteins—RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246, and VPS4A S95—in the peripheral blood erythrocytes of the subjects. The data processing module inputs the data into an acute altitude sickness susceptibility grouping model (which can be a pre-trained model based on previous training sets or a model obtained by continuously adding new training sets). The acute altitude sickness susceptibility grouping model is trained using data on the expression levels of the six proteins from known acute altitude sickness patients and known healthy individuals as training samples. The result output module outputs whether the subject belongs to the acute altitude sickness susceptibility group through the acute altitude sickness susceptibility grouping model. The data processing module includes a pre-trained machine learning model submodule and / or a trained machine learning model submodule operatively connected to the training submodule, performance evaluation submodule, and model optimization submodule.

[0153] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A computer apparatus comprising a memory, a processor, and a computer program stored on the memory, wherein, The processor executes the computer program to implement the following steps: S1, data receiving: receiving the data of the expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 in the sample of the subject; S2, data processing: inputting the data into the acute high altitude reaction susceptibility clustering model; the acute high altitude reaction susceptibility clustering model takes the data of the expression levels of the six proteins of the known acute high altitude reaction patients and the known high altitude environment adaptors as the training sample, to train the acute high altitude reaction susceptibility clustering model; S3, outputting the result: outputting whether the subject belongs to the acute high altitude reaction susceptible population through the acute high altitude reaction susceptibility clustering model.

2. The computer apparatus of claim 1, wherein, The acute high altitude reaction susceptibility clustering model is at least one selected from the group consisting of a random forest model, a support vector machine model, a linear discriminant analysis model, a recursive feature elimination model, a logistic regression model, a neural network model, a CART model, a flextree model, a LART model, a MART model, a k-nearest neighbor classification model, a cluster analysis model, a principal component analysis model, a Bayesian classification model and a hidden Markov model.

3. The computer apparatus of claim 2, wherein, The acute high altitude reaction susceptibility clustering model is a random forest model.

4. The computer apparatus of claim 3, wherein, The hyperparameters of the random forest model are as follows: the number of random forest trees is 500.

5. A device for acute altitude reaction susceptibility stratification, characterized by The device comprises a data receiving module, a data processing module and a result output module; The data receiving module is used to receive the data of the expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 in the sample of the subject; The data processing module is used to input the data into the acute high altitude reaction susceptibility clustering model; the acute high altitude reaction susceptibility clustering model takes the data of the expression levels of the six proteins of the known acute high altitude reaction patients and the known high altitude environment adaptors as the training sample, to train the acute high altitude reaction susceptibility clustering model; The result output module is used to output whether the subject belongs to the acute high altitude reaction susceptible population through the acute high altitude reaction susceptibility clustering model.

6. The apparatus of claim 5, wherein, The data processing module comprises a machine learning model submodule, the input of the machine learning model is the data of the expression levels of the six proteins in claim 1, and the output is whether it belongs to the acute high altitude reaction susceptible population.

7. A method of constructing a susceptibility sub-population model for acute high altitude reaction, characterized by, The method comprises the following steps: taking the data of the expression levels of RAB2B, VPS4A, SPTB T1387, ANP32B T244, DMTN S246 and VPS4A S95 in the samples of the known acute high altitude reaction patients and the known high altitude environment adaptors as the training sample, to train the acute high altitude reaction susceptibility clustering model.

8. The method of claim 7, wherein, The method comprises the following steps: A1) receiving data of expression levels of RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in samples of known acute high altitude reaction patients and known high altitude environment adaptors; A2) taking the data as input features of a machine learning model, training and evaluating the machine learning model; A3) obtaining the trained machine learning model, i.e. the acute high altitude reaction susceptibility clustering model, according to the evaluation result.

9. A method for stratifying a population for susceptibility to acute high altitude reaction, characterized in that, The method comprises the following steps: S1, data receiving: receiving data of expression levels of RAB2B, VPS4A, SPTBT1387, ANP32B T244, DMTN S246 and VPS4A S95 in samples of subjects; S2, data processing: inputting the data into the acute high altitude reaction susceptibility clustering model; the acute high altitude reaction susceptibility clustering model takes the data of expression levels of the 6 proteins in known acute high altitude reaction patients and known high altitude environment adaptors as training samples to train the acute high altitude reaction susceptibility clustering model; S3, outputting results: outputting whether the subject belongs to the acute high altitude reaction susceptible population by the acute high altitude reaction susceptibility clustering model.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program, when executed by a processor, implements the steps of the method of any one of claims 7-9.