Early high altitude pulmonary edema disease data collection system based on multi-gene mutation characteristics

By using a data collection system based on multi-gene mutation characteristics, combined with LASSO regression and fuzzy logic, key data were screened out, solving the data screening problem in the early disease data collection system for high-altitude pulmonary edema, and achieving accurate prediction and efficient system operation.

CN118899033BActive Publication Date: 2026-04-14GENERAL HOSPITAL OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GENERAL HOSPITAL OF PLA
Filing Date
2024-07-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies in early-stage pulmonary edema data collection systems at high altitudes cannot simultaneously ensure data integrity and the accuracy of user predictions, especially when the database is experiencing peak information levels, making it impossible to filter out key data.

Method used

A data collection system based on multi-gene mutation characteristics is adopted. Through data acquisition, user differentiation, data filtering and optimization matching modules, combined with LASSO regression model and fuzzy logic, the most characteristic historical data and altitude information are selected, and the amount of data is optimized to ensure the accuracy of prediction results and the smoothness of the system.

Benefits of technology

It achieves accurate prediction of early-stage pulmonary edema in high-altitude areas under limited data conditions, reduces system memory pressure, and ensures efficient data collection and accurate prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118899033B_ABST
    Figure CN118899033B_ABST
Patent Text Reader

Abstract

The application discloses a high-altitude pulmonary edema early disease data collection system based on multi-gene mutation characteristics and relates to the field of bioinformatics, and is used for solving the problem of data collection of various users when the database information has reached the peak, that is, all parameters cannot be collected as reference, some data must be screened out, but the prediction result of the user cannot be affected, and the problem of screening the data collection of various users, comprising a data acquisition module, a user distinguishing module, a data screening module and an optimization matching module; the modules are signal-connected, receive user gene data and user average altitude sent by the data acquisition, establish PRS standardization processing, divide user risks, determine the importance of the historical data of multiple categories in predicting various user data through a LASSO regression model according to the occurrence frequency and the historical data uploading speed of the historical data, screen out data with a characteristic coefficient reduced to zero, so that the prediction result deviation is small, and the system runs smoothly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics, and more specifically, to a data collection system for early-stage pulmonary edema in high-altitude areas based on multi-gene mutation characteristics. Background Technology

[0002] High-altitude pulmonary edema (HAPE) is a common acute illness in high-altitude environments, characterized by pulmonary effusion, which can be life-threatening in severe cases. Early detection and intervention are crucial for the prevention and treatment of HAPE. In recent years, advancements in genomics have made it possible to predict individual disease risk through polygenic mutation characteristics.

[0003] The existing technology has the following shortcomings:

[0004] Currently, common prediction systems, such as the SiGCD network system, assess the likelihood of early-stage high-altitude pulmonary edema based on gene prediction. However, in the initial stages of experiments, all parameters need to be comprehensively collected from all subjects to ensure data integrity and analytical accuracy. This strategy is suitable for research phases and scenarios requiring high-precision data collection. But when the database information reaches its peak, it becomes impossible to collect all parameters as a reference; some data must be filtered out without affecting the users' prediction results. Clearly, the selection of data for different user groups is a pressing issue that needs to be addressed. Therefore, this paper proposes a data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics.

[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a data collection system for early-stage pulmonary edema in high-altitude areas based on multi-gene mutation characteristics, which addresses the problems mentioned in the background art by employing different product testing methods.

[0007] To achieve the above objectives, the present invention provides the following technical solution: a data collection system for early-stage pulmonary edema in high-altitude areas based on multi-gene mutation characteristics, comprising a data acquisition module, a user differentiation module, a data filtering module, and an optimization matching module; and signal connections between the modules.

[0008] The data acquisition module is used to collect user genetic data and average user altitude, and send them to the user differentiation module. It also collects the frequency of historical data occurrences and historical data upload speed, and sends them to the data filtering module. The user genetic data refers to the multi-gene risk score.

[0009] The user segmentation module receives user genetic data and average user altitude, establishes PRS standardization processing to make it comparable among different individuals, classifies users by risk based on the results, and sends the user segmentation results to the data filtering module.

[0010] The data filtering module is used to obtain user segmentation results and the most characteristic historical time periods. Based on the user segmentation results, it divides the historical data occurrence frequency and historical data upload speed into multiple categories. The importance of these factors in predicting the early risk of high-altitude pulmonary edema in various types of users is determined by the LASSO regression model. The retained data information is obtained, and the importance of each data point is sent to the optimization matching module.

[0011] The optimization matching module receives the importance of each data point, obtains the data event density and calculates the data volume limit, and uses fuzzy logic to determine the most characteristic historical time period through comprehensive analysis, and then sends it to the data filtering module.

[0012] In a preferred embodiment, user gene data is obtained by calculating PRS; in step A1, after selecting the target gene associated with high-altitude pulmonary edema, the mutation site of each target gene is determined; gene sequencing is performed on each user to obtain the gene fragment containing the target gene in their gene data.

[0013] Step A2 involves obtaining the risk effect value for each mutation site through large-scale population studies; then calculating the contribution of each mutation site to the disease risk, usually expressed as a log-odds ratio.

[0014] Step A3: For each mutation site, calculate the risk score for that individual mutation site based on the user's genotype and corresponding weight, using the following formula:

[0015]

[0016] In the formula, It is the weight of the i-th mutation site. It is the user's genotype at the i-th mutation site, where i is the i-th mutation site;

[0017] Step A4: Sum the single-gene risk scores of all mutation sites to obtain the polygenic risk score (PRS). .

[0018] In a preferred embodiment, the user's average altitude is obtained by averaging the altitude at which the user was located during the measurement. .

[0019] In a preferred embodiment, the comprehensive risk score is a combined assessment indicator that integrates a standardized polygenic risk score and the user's average altitude, and its calculation formula is as follows:

[0020]

[0021] In the formula, For comprehensive risk scoring, as well as The preset scaling factor for the standardized PRS and the average altitude of users, and as well as All are greater than 0; where t represents the t-th user;

[0022] Risk stratification is set based on each user's comprehensive risk score. After sorting the comprehensive risk scores from largest to smallest, the user's risk level is defined.

[0023] In a preferred embodiment, the user segmentation result value is preferentially defined as x, where x represents the x-th user category; the number of occurrences of early-stage pulmonary edema characteristics in users is obtained by using historical data. Historical data upload speeds for various data points were obtained through network monitoring tools. .

[0024] In a preferred embodiment, the frequency of occurrence of historical data for multiple categories and the historical data upload speed are obtained, and their importance in predicting the early risk of high-altitude pulmonary edema for various types of users is determined by a LASSO regression model.

[0025] Step B1: Construct a dataset X containing the historical occurrence frequency and upload speed of each category, and a response variable z. X is set as a feature matrix, including... and ;

[0026] Step B2 involves standardizing the data to eliminate the influence between different units of measurement: In the formula, Let be the mean of the j-th feature. Let be the standard deviation of the j-th feature;

[0027] Step B3: Select the optimal one through cross-validation. value: ,in, This is the cross-validation error; the optimal value is the minimum average validation error during the validation process.

[0028] Step B4, use the optimal Value training for LASSO regression model: In the formula, It is the response variable of the p-th sample. It is the j-th standardized feature variable of the p-th sample. It is the intercept term. It is the regression coefficient of the j-th eigenvector;

[0029] In step B5, the LASSO regression model reduces the coefficients of some features to zero, selecting features with non-zero coefficients as important features, thus obtaining the important feature set F: Data whose feature coefficients have decreased to zero are filtered out.

[0030] In a preferred embodiment, the computational limit data amount is obtained by determining the maximum amount of data that the current system's computing power and storage capacity can handle in a single computation. Where t represents the t-th time period, and the data event density is obtained by measuring the frequency of events consistent with the prediction of high-altitude pulmonary edema in various types of data among various types of users. .

[0031] In a preferred embodiment, the data event density and the computational constraint data volume are defined as input variables, and they are respectively divided into different fuzzy sets;

[0032] Each historical period is defined as an output variable and divided into fuzzy sets.

[0033] Formulate fuzzy rules to describe the impact of data event density and computational constraints on the amount of data in each historical period;

[0034] Fuzzy reasoning is performed based on fuzzy rules to determine each historical period.

[0035] In a preferred embodiment, historical time periods labeled "High" are obtained and their numerical values ​​are collected and sorted in descending order. The first historical time period result is labeled as the most characteristic historical time period and sent to the data filtering module for calculation.

[0036] The technical effects and advantages of this invention are as follows:

[0037] 1. This invention classifies users by comprehensively analyzing user genetic data and average user altitude. Based on the user categories and classification results, it determines the frequency of historical data occurrences and historical data upload speeds for multiple categories. Using a LASSO regression model, it obtains the important features of each data point in each user category and filters out data whose feature coefficients are reduced to zero. This ensures that the prediction results have small deviations while the system runs smoothly.

[0038] 2. This invention, by comprehensively analyzing the density of data events and calculating the limited amount of data, formulates a set of fuzzy rules for fuzzy inference to determine the most characteristic historical time period. At the same time, it divides the historical time span according to the limited amount of data. This not only ensures that the system can collect the target parameters at once, but also optimizes the parameters for verifying the importance of data, making the parameters for calculating the importance of data more accurate and facilitating the relief of system memory pressure. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the modules of the high-altitude pulmonary edema early disease data collection system based on multi-gene mutation characteristics according to the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Based on siGCD: a web server used to collect gene prediction data to assess the potential early-stage disease of high-altitude pulmonary edema. When used in research phases and scenarios requiring high-precision data collection, the data typically collected includes, but is not limited to, genomic data, physiological parameters, and environmental parameters. Genomic data includes the frequency of target gene mutations; physiological parameters include heart rate, blood oxygen saturation, respiratory rate, blood pressure, and body temperature; environmental parameters include altitude, air pressure, ambient temperature, and humidity. However, it is obvious that these parameters cannot be collected or used in all scenarios. The selection of data has not been discussed; it should be evaluated based on the correlation, jumps, and acquisition difficulty of individual parameters. Therefore, a data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics is proposed.

[0042] Example 1: This invention discloses a data collection system for early-stage pulmonary edema in high-altitude areas based on multi-gene mutation characteristics, such as... Figure 1 As shown, it includes a data acquisition module, a user differentiation module, a data filtering module, and an optimized matching module;

[0043] The data acquisition module is used to collect user genetic data and average user altitude, and send them to the user differentiation module. It also collects the number of occurrences of historical data and the historical data upload speed, and sends them to the data filtering module.

[0044] Among them, user genetic data refers to polygenic risk scores, which are used to estimate the overall impact of multiple gene variations in an individual's genome on a certain disease or specific trait. In the risk assessment of high-altitude pulmonary edema, calculating PRS can more accurately predict an individual's genetic risk.

[0045] Step A1: After selecting the target genes associated with high-altitude pulmonary edema, determine the mutation sites of each target gene; perform gene sequencing on each user to obtain the gene fragments containing the target genes in their gene data.

[0046] Step A2 involves obtaining the risk effect value for each mutation site through large-scale population studies; then calculating the contribution of each mutation site to the disease risk, usually expressed as a log-odds ratio.

[0047] Step A3: For each mutation site, calculate the risk score for that individual mutation site based on the user's genotype and corresponding weight, using the following formula:

[0048]

[0049] In the formula, It is the weight of the i-th mutation site. It is the user's genotype at the i-th mutation site, where i is the i-th mutation site;

[0050] Step A4: Sum the single-gene risk scores of all mutation sites to obtain the polygenic risk score (PRS). ;

[0051] The average user altitude refers to the average altitude at which the user is located during the measurement. It's understandable that a user may use the system multiple times, so the average altitude reduces the uncertainty caused by large values ​​and also intuitively expresses the user's altitude over a long period. Its formula is as follows: Where k represents the number of times a user uses the service, k = 1, 2, 3, 4...n; and n represents the total number of users.

[0052] The user segmentation module receives user genetic data and average user altitude, establishes PRS standardization processing to make it comparable among different individuals, classifies users by risk based on the results, and sends the user segmentation results to the data filtering module.

[0053] To ensure comparability of the calculated PRS among different individuals, it is standardized. The standardized PRS, combined with the user's average altitude, is calculated using the following formula:

[0054]

[0055] In the formula, Let PRS be the average value of all individuals in the group. PRS standard deviation in the population;

[0056] The purpose of standardized PRS is to eliminate the incomparability caused by differences in the absolute values ​​of gene scores between individuals. Standardized PRS represents the degree to which an individual's PRS deviates from the population mean (in standard deviations). For example, a standardized PRS of 1 means that the individual's PRS is one standard deviation higher than the population mean.

[0057] The comprehensive risk score is a combined assessment indicator that integrates standardized multigene risk scores and the user's average altitude. Its calculation formula is as follows:

[0058]

[0059] In the formula, For comprehensive risk scoring, as well as The preset scaling factor for the standardized PRS and the average altitude of users, and as well as All are greater than 0; where t represents the t-th user;

[0060] Furthermore, risk stratification is established based on each user's comprehensive risk score, ranking the comprehensive risk scores from highest to lowest. For low risk: Low risk is defined as being below the lower quartile (25%) of the population distribution; high risk is defined as being between the middle two quartiles (25% to 75%) of the population distribution; and high risk is defined as being above the upper quartile (75%) of the population distribution.

[0061] It should be noted that the definition of user risk level can be adjusted according to the actual situation. For example, although this embodiment defines user risk level with three levels of risk, it can actually be divided into more than three levels based on the standardized PRS and the values ​​of the average altitude of users, so as to better filter out the data that needs to be removed from each category of users.

[0062] The data filtering module is used to obtain user segmentation results and the most characteristic historical time periods. Based on the user segmentation results, it divides the historical data occurrence frequency and historical data upload speed into multiple categories. The importance of these factors in predicting the early risk of high-altitude pulmonary edema in various types of users is determined by the LASSO regression model. The retained data information is obtained, and the importance of each data point is sent to the optimization matching module.

[0063] First, define the user segmentation result value as x, where x represents the xth user category. Then, obtain the number of historical data occurrences and the historical data upload speed for x categories corresponding to the user categories.

[0064] The logic for obtaining the frequency of historical data occurrences refers to the frequency of occurrence of each historical data point within the current category, used to predict the early symptoms of high-altitude pulmonary edema in users. Let y be the number of predicted data points, where y represents the y-th predicted data point. The frequency of historical data occurrences for each data point is then obtained using a built-in system counter. ;

[0065] The logic for obtaining historical data upload speed refers to the upload speed of each historical data point used to predict the early symptoms of high-altitude pulmonary edema in the current category. The faster the upload speed, the more attention the system pays to the data size or importance, and the faster the system can collect the data. The historical data upload speed of each data point can be obtained through network monitoring tools. ;

[0066] We obtained the frequency of occurrence and upload speed of historical data for multiple categories, and used the LASSO regression model to determine their importance in predicting the early risk of high-altitude pulmonary edema for various types of users.

[0067] Step B1: Construct a dataset X containing the historical occurrence frequency and upload speed of each category, and a response variable z. X is set as a feature matrix, including... and ;

[0068] Step B2 involves standardizing the data to eliminate the influence between different units of measurement: In the formula, Let be the mean of the j-th feature. Let be the standard deviation of the j-th feature;

[0069] Step B3: Select the optimal one through cross-validation. value: ,in, This is the cross-validation error; the optimal value is the minimum average validation error during the validation process.

[0070] Step B4, use the optimal Value training for LASSO regression model: In the formula, It is the response variable of the p-th sample. It is the j-th standardized feature variable of the p-th sample. It is the intercept term. It is the regression coefficient of the j-th eigenvector;

[0071] In step B5, the LASSO regression model reduces the coefficients of some features to zero, selecting features with non-zero coefficients as important features, thus obtaining the important feature set F: Data whose feature coefficients are reduced to zero are filtered out to ensure accurate prediction results while the system runs smoothly.

[0072] This invention categorizes users by comprehensively analyzing user genetic data and average user altitude. Based on the user categories and the categorization results, it then uses a LASSO regression model to obtain the important features of each data point in each user category, and filters out data whose feature coefficients are reduced to zero. This ensures that the prediction results have a small deviation while the system runs smoothly.

[0073] Example 2

[0074] In Embodiment 1 of this invention, a key example is given illustrating the operational strategy of using LASSO regression model to screen out data whose feature coefficients are reduced to zero after classifying users into categories. However, Embodiment 1 considers all historical data as the basis for evaluation. Obviously, although this operational strategy can ensure the reliability of the data screening operation, in reality, the system cannot take all historical data into account. The most representative historical period should be selected for evaluation to avoid system lag. To address the above issues, Embodiment 2 of this invention further refines the approach.

[0075] The optimization matching module receives the importance of each data point, obtains the data event density and the data volume limit for calculation, and uses fuzzy logic to determine the most characteristic historical time period through comprehensive analysis, and then sends it to the data filtering module.

[0076] The computational limit refers to the maximum amount of data that the current system's computing power and storage capacity can handle in a single computation, in order to avoid system lag due to excessive data processing. The historical time span, or time period t, is divided based on the size of the computational data volume in a single operation, where t represents the t-th time period. The time period t is determined by the computational limit. It changes with the changes;

[0077] It is understandable that the amount of data in a single calculation has a certain historical time period. For example, if the amount of data in a single calculation is 25.6GB, and there are 64,000 data points per day on average, then it contains approximately 256,000 data points, and the historical time span is 4 days.

[0078] Data event density refers to the frequency of events consistent with the prediction of high-altitude pulmonary edema occurring in various types of data across different user groups. Time periods with high event density are generally more representative. The data event density is calculated as the ratio of the number of events consistent with the prediction of high-altitude pulmonary edema occurring within time period t to the time period t itself. ;

[0079] Specifically, fuzzy logic is used to determine each historical period based on the density of data events and the computational constraints on the amount of data.

[0080] For example, "High", "Low", and "Medium" refer to data event density, while "Much", "Moderate", and "Less" refer to computational limits on the amount of data.

[0081] Develop a set of fuzzy rules to describe the impact of different input variables on the output variable. The rules can be defined based on expertise or obtained through data analysis and experimentation. For example:

[0082] The data event density is labeled as X, the calculation limit data volume is labeled as U, and each historical time period is labeled as C_results;

[0083] Then it can be defined as:

[0084] Rule 1: IF (X is High) AND (U is Much) THEN (C_results is High)

[0085] Rule 2: IF (U is Low) AND (U is Less) THEN (C_results is Low) ...

[0087] Specifically, fuzzy reasoning is performed based on fuzzy rules to determine the most characteristic historical period among each historical period;

[0088] It should be noted that the division of fuzzy sets can be adjusted according to the actual situation. For example, although this embodiment uses three fuzzy sets as an example, the data event density and the amount of data to be calculated can actually be divided into more than three sets to facilitate more precise adjustment based on different similarities.

[0089] Furthermore, the judgment of whether the data event density and the computational limit data volume are high, medium, or low can be made by setting thresholds according to the actual situation. For example, when the data event density exceeds 75% of the set threshold, it is marked as "High", and when the computational limit data volume is higher than 70% of the set threshold, it is marked as "Much", etc., which will not be elaborated here.

[0090] Specifically, the historical time periods marked as "High" are obtained and their numerical values ​​are collected and sorted in descending order. The results of the first historical time period are marked as the most characteristic historical time period and brought into the data filtering parameter calculation in Example 1 to calculate the accuracy of data importance.

[0091] This invention, through comprehensive analysis of data event density and computational constraints, formulates a set of fuzzy rules for fuzzy inference to determine the most characteristic historical time periods. At the same time, it divides the historical time span based on computational constraints, which not only ensures that the system can collect the target parameters at once, but also optimizes the parameters for verifying the importance of data, making the parameters for calculating the importance of data more accurate and facilitating the reduction of system memory pressure.

[0092] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0093] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0094] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0096] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0097] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0098] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0099] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0100] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A data collection system for early-stage pulmonary edema in high-altitude areas based on multi-gene mutation characteristics, characterized by: It includes a data acquisition module, a user differentiation module, a data filtering module, and an optimized matching module; signal connections between the modules; The data acquisition module is used to collect user genetic data and average user altitude, and send them to the user differentiation module. It also collects the frequency of historical data occurrences and historical data upload speed, and sends them to the data filtering module. The user genetic data refers to the multi-gene risk score. The user segmentation module receives user genetic data and average user altitude, establishes PRS standardization processing to make it comparable among different individuals, classifies users by risk based on the results, and sends the user segmentation results to the data filtering module. The data filtering module is used to obtain user segmentation results and the most characteristic historical time periods. Based on the user segmentation results, it divides the historical data occurrence frequency and historical data upload speed into multiple categories. The importance of these factors in predicting the early risk of high-altitude pulmonary edema in various types of users is determined by the LASSO regression model. The retained data information is obtained, and the importance of each data point is sent to the optimization matching module. The optimization matching module receives the importance of each data point, obtains the data event density and the data volume limit for calculation, and uses fuzzy logic to determine the most characteristic historical time period through comprehensive analysis, and then sends it to the data filtering module. We obtained the frequency of occurrence and upload speed of historical data for multiple categories, and used the LASSO regression model to determine their importance in predicting the early risk of high-altitude pulmonary edema for various types of users. Step B1: Construct a dataset X containing the historical occurrence frequency and upload speed of each category, and a response variable z. X is set as a feature matrix, including... and ; Step B2 involves standardizing the data to eliminate the influence between different units of measurement: In the formula, Let be the mean of the j-th feature. Let be the standard deviation of the j-th feature; Step B3: Select the optimal one through cross-validation. value: ,in, This is the cross-validation error; the optimal value is the minimum average validation error during the validation process. Step B4, use the optimal Value training for LASSO regression model: In the formula, It is the response variable of the p-th sample. It is the j-th standardized feature variable of the p-th sample. It is the intercept term. It is the regression coefficient of the j-th eigenvector; In step B5, the LASSO regression model reduces the coefficients of some features to zero, selecting features with non-zero coefficients as important features, thus obtaining the important feature set F: Data whose feature coefficients have decreased to zero are filtered out; The computational limit data volume is obtained by determining the maximum amount of data that the current system's computing power and storage capacity can handle in a single computation. Where t represents the t-th time period, and the data event density is obtained by measuring the frequency of events consistent with the prediction of high-altitude pulmonary edema in various types of data among various types of users. .

2. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 1, characterized in that: User gene data is obtained by calculating PRS; in step A1, after selecting the target gene associated with high-altitude pulmonary edema, the mutation site of each target gene is determined; gene sequencing is performed on each user to obtain the gene fragment containing the target gene in their gene data. Step A2 involves obtaining the risk effect value for each mutation site through large-scale population studies; then calculating the contribution of each mutation site to the disease risk, usually expressed as a log-odds ratio. Step A3: For each mutation site, calculate the risk score for that individual mutation site based on the user's genotype and corresponding weight, using the following formula: : In the formula, It is the weight of the i-th mutation site. It is the user's genotype at the i-th mutation site, where i is the i-th mutation site; Step A4: Sum the single-gene risk scores of all mutation sites to obtain the polygenic risk score (PRS). .

3. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 2, characterized in that: The average altitude of the user is obtained by averaging the altitude at which the user was located during the measurement. .

4. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 3, characterized in that: The comprehensive risk score is a combined assessment indicator that integrates standardized multigene risk scores and the user's average altitude. Its calculation formula is as follows: : In the formula, For comprehensive risk scoring, as well as The preset scaling factor for the standardized PRS and the average altitude of users, and as well as All are greater than 0; where t represents the t-th user; Risk stratification is set based on each user's comprehensive risk score. After sorting the comprehensive risk scores from largest to smallest, the user's risk level is defined.

5. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 4, characterized in that: First, define the user segmentation result value as x, where x represents the x-th user category; obtain the result by using historical data to predict the frequency of occurrence of early-stage pulmonary edema characteristics in users. Historical data upload speeds for various data points were obtained through network monitoring tools. .

6. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 1, characterized in that: The data event density and the computational constraint data volume are defined as input variables, and they are divided into different fuzzy sets respectively. Each historical period is defined as an output variable and divided into fuzzy sets. Formulate fuzzy rules to describe the impact of data event density and computational constraints on the amount of data in each historical period; Fuzzy reasoning is performed based on fuzzy rules to determine each historical period.

7. The data collection system for early-stage high-altitude pulmonary edema based on multi-gene mutation characteristics according to claim 6, characterized in that: Obtain the historical time period marked as "High" and collect its numerical values. Sort the results in descending order and mark the first historical time period as the most characteristic historical time period. Send it to the data filtering module for calculation.

Citation Information

Patent Citations

  • Risk event early screening system and method

    CN109978396A

  • Plateau pulmonary edema disease risk self-assessment method and system

    CN117153406A