A cardiovascular and cerebrovascular disease risk prediction system for aerospace occupational groups
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-24
Smart Images

Figure CN122455327A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of physical examination data processing technology, specifically to a cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals. Background Technology
[0002] Cardiovascular and cerebrovascular disease risk prediction involves analyzing an individual's physiological indicators, lifestyle habits, and other data to assess the probability of a cardiovascular or cerebrovascular event occurring within a specific time period. This provides a basis for health management adjustments, enables early warning, and supports the development of targeted intervention measures.
[0003] Risk assessment models based on traditional epidemiological studies are commonly used in related technologies for prediction. These models primarily construct prediction formulas based on conventional factors such as age, blood pressure, blood lipids, and smoking history. However, existing risk prediction models directly substitute static values of physiological indicators such as blood pressure, blood lipids, and heart rate into the prediction formula to calculate the risk probability after collecting these values at a certain time point. This assessment method does not consider the dynamic changes in physiological indicators under different work conditions, leading to discrepancies between the risk assessment results obtained using the same static indicator values and the individual's actual risk level at different work stages. Consequently, the accuracy of predicting cardiovascular and cerebrovascular disease risk is poor. Summary of the Invention
[0004] This application provides a cardiovascular and cerebrovascular disease risk prediction system for the aerospace profession, which can improve the accuracy of cardiovascular and cerebrovascular disease risk prediction, thereby achieving accurate and effective intervention guidance.
[0005] Firstly, this application provides a cardiovascular and cerebrovascular disease risk prediction system for the aerospace profession, the system comprising: The data acquisition module is used to acquire historical longitudinal physiological data and corresponding job task cycle data of the aerospace profession. The job task cycle data is used to mark the work rhythm state of the aerospace profession in different time periods. The first influencing factor processing module is used to perform phased differential analysis on historical longitudinal physiological data based on job task cycle data, and to screen out the first influencing factors whose numerical fluctuation exceeds a preset threshold under different work rhythm states. The second influencing factor processing module is used to input historical longitudinal physiological data into a preset feature association model and filter out the second influencing factors whose association values exceed the preset association threshold. The model building module is used to build a dynamic load response module using the first influencing factor, a static baseline module using the second influencing factor, and to build a risk prediction model based on the dynamic load response module and the static baseline module. The model processing module is used to collect the current physiological indicators and current work rhythm status of the person being tested, and input the current physiological indicators and current work rhythm status into the risk prediction model to obtain the risk prediction probability; The output module generates health recommendations for the individuals being tested based on the predicted risk probability and the specific impact type of the primary influencing factor.
[0006] By adopting the above technical solution, the dynamic load response module constructed using the first influencing factor can reflect the real-time impact of changes in work rhythm on individuals, while the static baseline module constructed using the second influencing factor can assess an individual's baseline risk level. The risk prediction model composed of these two modules achieves a comprehensive assessment of both static baseline risk and dynamic load risk. When predicting risk for the individuals under test, current physiological indicators and current work rhythm status are simultaneously input into the risk prediction model, ensuring that the output risk prediction probability reflects both the individual's baseline health status and the impact of current work status, thereby improving the accuracy of cardiovascular and cerebrovascular disease risk prediction. Based on the accurate risk prediction probability and the specific impact type of the first influencing factor, health recommendations are generated, achieving accurate and effective intervention guidance.
[0007] Secondly, this application provides a computer program product that, when the computer program product is run on the cardiovascular and cerebrovascular disease risk prediction system for the aerospace occupation group, causes the display system to execute the aforementioned system.
[0008] Thirdly, this application provides a computer storage medium that stores multiple instructions adapted for loading and execution by a processor of any of the above-mentioned systems.
[0009] Fourthly, this application provides an electronic device including a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform any of the systems described above.
[0010] In summary, the beneficial effects of the technical solution of this application include: By adopting the above technical solution, the dynamic load response module constructed using the first influencing factor can reflect the real-time impact of changes in work rhythm on individuals, while the static baseline module constructed using the second influencing factor can assess an individual's baseline risk level. The risk prediction model composed of these two modules achieves a comprehensive assessment of both static baseline risk and dynamic load risk. When predicting risk for the individuals under test, current physiological indicators and current work rhythm status are simultaneously input into the risk prediction model, ensuring that the output risk prediction probability reflects both the individual's baseline health status and the impact of current work status, thereby improving the accuracy of cardiovascular and cerebrovascular disease risk prediction. Based on the accurate risk prediction probability and the specific impact type of the first influencing factor, health recommendations are generated, achieving accurate and effective intervention guidance. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating a cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, according to an embodiment of this application. Figure 2 This is a schematic diagram of the structure of a cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, according to an embodiment of this application. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0012] Explanation of reference numerals in the attached drawings: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0014] In the description of the embodiments of this application, words such as "illustrative," "for example," or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "illustrative," "for example," or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of words such as "illustrative," "for example," or "for example" is intended to present the relevant concepts in a specific manner.
[0015] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0016] Please see Figure 1 This document presents a flowchart illustrating a cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, as provided in an embodiment of this application. The system can be implemented using a computer program, a microcontroller, or run on a von Neumann architecture-based cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals. The computer program can be integrated into the application or run as a standalone tool application. The following provides a detailed description of each module of the cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals. Specifically, the system includes a data acquisition module, a first influencing factor processing module, a second influencing factor processing module, a model building module, a model processing module, and an output module. The specific implementation steps of each module are as follows.
[0017] S101: Obtain historical longitudinal physiological data and corresponding job task cycle data of the aerospace occupation group. The job task cycle data is used to mark the work rhythm state of the aerospace occupation group in different time periods. Historical longitudinal physiological data refers to the continuous or periodic collection of physiological health indicators of aerospace professionals over a relatively long period. This data includes, but is not limited to, multi-dimensional physiological parameters such as blood pressure, heart rate, blood lipids, blood glucose, body mass index, electrocardiogram parameters, and vascular elasticity indicators. Each parameter is marked with a specific timestamp, forming a time-series dataset. Job task cycle data represents the time intervals of different work stages experienced by aerospace professionals. It typically includes fields such as task type, start and end times, and work intensity level, reflecting the time nodes of each stage of an aerospace project from feasibility study and development to mission execution. Work rhythm status represents the classification of workload patterns and physiological stress levels exhibited by aerospace professionals at different mission stages. Typical examples include the normal maintenance period and the high-load critical period. The former corresponds to routine work stages such as daily work, theoretical research, and technical management, while the latter corresponds to high-intensity work stages such as key project nodes, mission execution periods, and emergency responses.
[0018] Specifically, the data acquisition process first requires extracting multi-year physical examination data and health monitoring data of the target aerospace occupational group from the occupational health record system. This data needs to be standardized and preprocessed to ensure the comparability of data from different periods and different physical examination institutions. At the same time, task execution records associated with these personnel are extracted from the project management system, including information such as the project tasks participated in, task stages, working hours, and duty records. Then, an association mapping relationship between personnel identification and task records is established, and the physiological data time points of each personnel are accurately aligned with the task stage at that time, forming a longitudinal health dataset with work rhythm labels.
[0019] In some embodiments, historical data acquisition and integration can be achieved through various methods: Optionally, a distributed data acquisition scheme can be adopted. First, qualified personnel health records are exported in batches from the occupational health center databases of various aerospace units via data interfaces, and valid samples with more than three years of employment and a physical examination record completeness of more than 80% are selected. Then, the task participation history of these personnel is extracted from the aerospace project management platform, and the task status within a certain time window before and after each physical examination time point is used as the work rhythm label for that time point using a timestamp matching algorithm. Finally, records with too many outliers and missing values are removed through a data cleaning module to construct a standardized longitudinal dataset. Optionally, a hybrid mode combining active reporting and passive collection can be adopted. Based on the collection of routine physical examination data, participants are required to regularly report simple physiological indicators through a mobile health management application. These high-frequency data and the comprehensive data from regular physical examinations together constitute the longitudinal dataset. At the same time, project progress data and personnel work hour data are extracted from the information management system, and the characteristic patterns of high-load intensive work periods and normal maintenance periods are automatically identified through a machine learning classifier, and the work rhythm status is automatically labeled. It is understood that other methods can also be used to acquire data, which are not limited here.
[0020] S102: Based on job task cycle data, perform phased differential analysis on historical longitudinal physiological data to screen out the first influencing factor whose numerical fluctuation exceeds the preset threshold under different work rhythm states; The phased differential analysis involves dividing historical longitudinal physiological data into different period subsets based on work rhythm states, calculating the statistical characteristics of each physiological indicator within each period subset, and quantifying the significance of differences between different periods using statistical tests. Numerical fluctuation amplitude represents the drastic change of a physiological indicator across different work rhythm states. It is typically quantified by calculating a combined score of distributional variability and mean shift, reflecting the indicator's sensitivity to workload changes. The first influencing factor represents physiological indicators that show significant response characteristics to changes in occupational stress. These indicators reflect an individual's physiological adaptability and stress sensitivity to workload, showing significant increases or abnormal fluctuations during periods of high workload. The preset threshold is a critical value standard used to determine whether the fluctuation amplitude is sufficiently significant. This threshold can be determined based on clinical medical standards, statistical significance levels, or data-driven quantile methods.
[0021] Specifically, the phased differential analysis process first divides each individual's longitudinal physiological data into data segments for different periods based on the work rhythm status labels in the job task cycle data, forming a normal characteristic set and a load characteristic set. Then, for each physiological indicator dimension, descriptive statistics of the indicator in the two characteristic sets are calculated, including mean, median, standard deviation, skewness, kurtosis, etc., and distribution charts are drawn for visual comparison. Next, the distribution variability is calculated to quantify the information difference between the two distributions, and the mean drift is calculated to reflect the degree of deviation of the center position. The corresponding rate of change for the standard deviation is also calculated. Finally, the distribution variability and mean drift are converted into a comprehensive fluctuation index through weighted summation or other synthesis methods. This comprehensive index is compared with a preset threshold. Physiological indicators exceeding the threshold are marked as the first influencing factor. These factors can effectively reflect the dynamic impact of work rhythm on health status.
[0022] S103: Input historical longitudinal physiological data into a preset feature association model to filter out the second influencing factor whose association value exceeds the preset association threshold; Feature association models refer to statistical or machine learning models used to assess the strength of the association between various physiological indicators and the risk of cardiovascular and cerebrovascular diseases. These models can quantify the predictive contribution of each feature to disease outcomes. Association values are numerical measures of the statistical strength or predictive importance of a physiological indicator in relation to the incidence of cardiovascular and cerebrovascular diseases. Different models have different representations of association values, reflecting the indicator's independent predictive ability for disease risk. Secondary influencing factors represent baseline physiological indicators that have a stable association with the long-term risk of cardiovascular and cerebrovascular diseases. These indicators typically reflect an individual's baseline health status and disease susceptibility, and are not significantly affected by short-term fluctuations in work rhythm. Preset association thresholds are critical criteria used to determine whether a physiological indicator has a significant association with disease risk. These thresholds are usually determined based on statistical significance levels, clinical experience, or model performance optimization results.
[0023] Specifically, the first step is to label the historical longitudinal physiological data for outcomes. This involves identifying individuals who experienced cardiovascular events and those who remained healthy during the observation period based on subsequent health records or disease diagnosis information, thus constructing a training dataset with classification or survival analysis labels. Next, all physiological indicators are input as feature variables into a pre-defined feature association model. The model learns the mapping relationship between features and disease outcomes, automatically calculating the association strength index for each feature. After model training, the association values corresponding to each physiological indicator are extracted. These association values reflect the indicator's independent predictive ability for disease risk, controlling for the influence of other confounding factors. Finally, all physiological indicators are sorted according to their association values, and indicators with association values exceeding a pre-defined association threshold are selected as secondary influencing factors. These factors typically have stable predictive capabilities across time periods, are independent of specific work rhythm states, and reflect an individual's long-term health risk baseline.
[0024] S104: Construct a dynamic load response module using the first influencing factor, construct a static baseline module using the second influencing factor, and construct a risk prediction model based on the dynamic load response module and the static baseline module; The dynamic load response module is a model component specifically designed to assess the impact of workload on an individual's physiological state. This module uses the primary influencing factor as input and, combined with current work rhythm information, outputs a risk score reflecting the immediate stress load level. The static baseline module represents a model component specifically designed to assess an individual's baseline health risk level. This module uses the secondary influencing factor as input and outputs a score reflecting the individual's inherent risk baseline. This score is relatively stable and does not fluctuate drastically with short-term changes in work status. The risk prediction model is a comprehensive predictive system that integrates information from both the dynamic load response and static baseline dimensions to comprehensively assess the probability of cardiovascular and cerebrovascular diseases in the aerospace occupational group. This model can distinguish the relative contributions of short-term stress risk and long-term baseline risk.
[0025] Specifically, the model building process first designs the architecture of the dynamic load response module for the first influencing factor. This module needs to be able to handle the interaction effect between features and work rhythm states. Therefore, a model structure containing interaction terms can be adopted, concatenating or interactively calculating the numerical values of the first influencing factor with the encoding of work rhythm states, so that the model can learn the differentiated impact of physiological index changes on risk under different work states. Then, a static baseline module is designed for the second influencing factor. This module is relatively simple and can use linear weighting or nonlinear transformation to map the baseline index to the basic risk score. Next, a fusion mechanism is designed to integrate the outputs of the two sub-modules. Common fusion methods include weighted summation, product form, concatenation followed by processing through a fully connected layer, etc. The fusion parameters need to be learned through training data. Finally, the disease outcome labels in historical data are used to train the entire risk prediction model end-to-end.
[0026] S105: Collect the current physiological indicators and current work rhythm status of the personnel to be tested, and input the current physiological indicators and current work rhythm status into the risk prediction model to obtain the risk prediction probability; The "test subject" refers to an individual member of the aerospace profession who needs to undergo cardiovascular and cerebrovascular disease risk assessment. This individual can be a newly hired employee undergoing their first assessment or an existing employee requiring regular follow-up. Current physiological indicators represent the latest physiological health data collected at the assessment point. This data should cover all physiological parameters corresponding to the first and second influencing factors selected in the preceding steps. Current work rhythm status indicates the work stage classification of the test subject at the assessment point. This status needs to be determined based on information such as the type of task currently being participated in, work intensity, and task progress. Risk prediction probability refers to the likelihood of the test subject experiencing a cardiovascular or cerebrovascular disease event within a specific future time window, typically expressed as a value between zero and one, with the closer the value is to one, the higher the risk.
[0027] Specifically, the risk prediction process first requires collecting the current physiological indicator data of the individuals to be tested. This data can come from regular physical examination reports, real-time monitoring data from wearable devices, or results of specific health examinations. After data collection, it needs to be standardized to ensure that the numerical units and dimensions are consistent with the data format used during model training. Simultaneously, it is necessary to determine the current work rhythm status of the individuals to be tested. This can be obtained by querying information on the projects and tasks the individuals are currently involved in, recent work hours statistics, shift records, etc., or by using an automated system to determine the status according to preset rules. Then, the standardized physiological indicator values and work rhythm status codes are organized according to the input format required by the model to form a feature vector. Next, the inference interface of the pre-trained risk prediction model is called, and the feature vector is input into the model for forward computation. Internally, the model processes the corresponding features through the dynamic load response module and the static baseline module, respectively. Then, the two parts of information are integrated at the fusion layer, and finally, the risk prediction probability is generated at the output layer.
[0028] S106: Generate health recommendations for the individuals to be tested based on the predicted risk probability and the specific impact type of the primary influencing factor.
[0029] Specifically, the health recommendation generation process first requires stratifying the probability of risk prediction. Based on preset risk thresholds, individuals are categorized into four levels: extremely low, low, medium, and high. Different risk levels correspond to intervention strategies of varying intensities and urgency. Next, the model's interpretability analysis module extracts the primary influencing factor contributing most to the current risk prediction, identifying which physiological indicators show significant abnormalities under the current work rhythm and the degree of deviation of these abnormalities from the individual's historical baseline or the normal range of the population. Following this, based on the identified main risk factor types, corresponding intervention measures are matched from a preset health recommendation knowledge base. This knowledge base stores standardized recommendation templates for different physiological indicator abnormalities, including dietary adjustments, exercise prescriptions, stress management, and medication consultations. Simultaneously, considering the individual's current work rhythm, targeted workload management recommendations are proposed, such as increasing rest frequency during high-intensity periods, optimizing work time allocation, and requesting task adjustments. Finally, the risk level assessment results, main risk factor analysis, personalized intervention recommendations, and follow-up monitoring plans are integrated into a structured health recommendation report.
[0030] Based on the above embodiments, as an optional implementation method, the work rhythm state includes at least a normal maintenance period and a high-load intensive period. Regarding the phased differential analysis of historical longitudinal physiological data based on job task cycle data in S102, the method of screening out the first influencing factor whose numerical fluctuation amplitude exceeds the preset threshold under different work rhythm states is described in detail below. The first influencing factor processing module includes a feature set construction unit, a difference calculation unit, a numerical fluctuation calculation unit, and an indicator determination unit. The specific implementation steps of the above units are as follows.
[0031] S201: Based on job task cycle data, historical longitudinal physiological data are divided into time periods according to normal maintenance period and high load tackling period, and normal feature set and load feature set are constructed respectively. The normal characteristic set refers to the set of physiological indicator data extracted from the normal maintenance period, including multiple measurements of various physiological parameters and their statistical characteristics during this period, reflecting the baseline physiological level of the subject under normal working conditions. The load characteristic set refers to the set of physiological indicator data extracted from the high-load intensive period, including multiple measurements of various physiological parameters and their statistical characteristics during this period, reflecting the physiological stress response level of the subject under high-intensity working conditions.
[0032] In practice, the process begins by reading the start and end times of each normal maintenance period and high-intensity intensive period recorded in the job task cycle data. Then, iterates through each record in the historical longitudinal physiological data, determining the corresponding work rhythm stage based on the record's timestamp. All physiological indicator measurements with timestamps falling within the normal maintenance period are categorized into the normal characteristic set, and all physiological indicator measurements with timestamps falling within the high-intensity intensive period are categorized into the load characteristic set. During the construction process, data integrity is maintained, and each physiological indicator is divided into two subsets. Each subset contains all measurement values, measurement times, and measurement contexts for that indicator within the corresponding period.
[0033] S202: For each physiological indicator, calculate the distribution difference and mean shift between the normal characteristic set and the load characteristic set; Distribution variability refers to the degree of difference in the numerical distribution of the same physiological indicator in the normal characteristic set and the load characteristic set, quantifying the statistical distance between the two data sets in terms of distribution shape, dispersion, skewness, etc. Mean shift refers to the difference between the means of the same physiological indicator in the normal characteristic set and the load characteristic set, reflecting the direction and magnitude of the systematic numerical shift of the indicator when the work rhythm changes.
[0034] When calculating the distribution variability, firstly, all values of the indicator in the normal characteristic set and the load characteristic set are extracted. The probability distribution characteristic parameters of each set are calculated, including mean, standard deviation, skewness, and kurtosis. Then, a statistical distribution distance metric is used to calculate the quantitative difference between the two distributions. The larger the quantitative difference value, the more significant the difference in the distribution pattern of the indicator under the two working conditions. When calculating the mean shift, the arithmetic mean of all values of the indicator in the normal characteristic set and the arithmetic mean of all values in the load characteristic set are calculated separately. The difference between the two means yields the mean shift. A positive value indicates that the average level of the indicator increases during high-load periods, while a negative value indicates that the average level decreases. The absolute value reflects the magnitude and intensity of the shift. Through the joint analysis of these two indicators, sensitive indicators that show significant changes under different working conditions can be identified.
[0035] S203: Calculate the amplitude of numerical fluctuation of each physiological indicator under different work rhythm states based on the degree of distribution difference and the mean drift. Among them, the numerical fluctuation amplitude refers to the comprehensive change intensity of physiological indicators when switching between different work rhythm states. By integrating information from two dimensions, namely the distribution difference and the mean drift, the sensitivity of this indicator to changes in workload is quantified.
[0036] Specifically, the calculation integrates two parameters: distribution variability and mean drift. First, the absolute value of the mean drift is taken to eliminate directionality and retain only amplitude information. Then, the normalized distribution variability and the normalized absolute value of the mean drift are weighted and synthesized. The weighting parameters are determined based on the relative importance of the two indicators in reflecting load sensitivity. The weighted synthesis is calculated by multiplying the distribution variability by its weighting coefficient and the absolute value of the mean drift by its weighting coefficient, and then adding the two weighted values to obtain the numerical fluctuation amplitude of the physiological indicator. Indicators with large numerical fluctuation amplitudes indicate significant systematic changes and distribution patterns when the work rhythm changes. These indicators are sensitive to changes in workload and show significant differences under different work conditions. By calculating the numerical fluctuation amplitudes of all physiological indicators, a ranking of the load sensitivity of each indicator is formed.
[0037] Based on the above embodiments, as an optional implementation method, the following provides a detailed description of each subunit of the numerical fluctuation calculation unit. The numerical fluctuation calculation unit includes a fluctuation sensitivity factor calculation subunit, a load response intensity factor calculation subunit, a target fluctuation sensitivity factor determination subunit, and a numerical fluctuation amplitude determination subunit. The specific implementation steps of each of the above subunits are as follows.
[0038] S2031: Analyze the distribution difference, extract the variance change information of physiological indicators during the normal maintenance period and the high-load intensive period, calculate the fluctuation sensitivity factor, and the fluctuation sensitivity factor characterizes the degree of instability of an individual's physiological response to changes in workload. Among them, variance change information refers to the variance values and their changes of physiological indicators in the two stages of normal maintenance and high-load intensive periods, reflecting the differences in the discrete fluctuation characteristics of the indicator under different working conditions. The fluctuation sensitivity factor is a parameter that quantifies the degree of discrete response of an individual's physiological indicator to changes in workload. The larger the value, the more obvious the increased fluctuation instability of the indicator during the high-load period.
[0039] In the specific calculation, the variances of all values of the physiological indicator in the normal characteristic set and the load characteristic set are calculated first. The variance in the normal period reflects the natural fluctuation range of the indicator under normal working conditions, while the variance in the high-load period reflects the fluctuation range of the indicator under high-intensity working conditions. Then, the ratio of the high-load period variance to the normal period variance is calculated to obtain the variance change factor. This factor is standardized and used as a fluctuation sensitivity factor. When the ratio is greater than one, it indicates that the fluctuation amplitude of the indicator increases during the high-load period, i.e., the stability decreases. The larger the ratio, the stronger the impact of the workload on the stability of the indicator, and the higher the sensitivity of the indicator to load changes.
[0040] S2032: Analyze the mean drift, extract the mean deviation information of physiological indicators during the normal maintenance period and the high-load challenge period, calculate the load response intensity factor, and the load response intensity factor characterizes the magnitude of the influence of workload on individual physiological indicators; Among them, mean deviation information refers to the mean level and specific value and direction of the deviation of the physiological indicator in the two stages of normal maintenance and high-load stress period, reflecting the systematic displacement characteristics of the indicator when the workload changes. The load response intensity factor is a parameter that quantifies the magnitude of the deviation of the central position of an individual physiological indicator when the workload changes. The larger the value, the greater the response intensity of the indicator to the influence of the workload.
[0041] In the specific calculation, the absolute value of the mean drift is used as the basic data. This absolute value already reflects the magnitude of the mean shift of the physiological indicator from the normal period to the high-load period. Then, this absolute value is divided by the mean of the indicator in the normal period to obtain the relative shift ratio. This ratio reflects the strength of the shift relative to the baseline level. The relative shift ratio is standardized and used as the load response intensity factor. The load response intensity factor eliminates the influence of different physiological indicator dimensions and baseline levels, making the load response intensity of different indicators comparable. The larger the value, the greater the relative magnitude of the mean drift of the indicator when the workload changes, and the stronger the response to the workload.
[0042] S2033: Pathological direction determination of fluctuation sensitivity factor. When the deviation direction of fluctuation sensitivity factor points to the direction of increased risk of cardiovascular and cerebrovascular diseases, the pathological risk coefficient is applied to amplify the fluctuation sensitivity factor to obtain the target fluctuation sensitivity factor. The fluctuation sensitivity factor refers to the load response intensity factor calculated in the preceding steps, reflecting the degree of mean deviation of physiological indicators when workload changes. Pathological directionality determination refers to the analytical process of judging whether the deviation of physiological indicators is in a direction that increases the risk of cardiovascular and cerebrovascular diseases; for example, deviations such as increased blood pressure, increased blood lipids, and decreased heart rate variability all point to increased risk. The pathological risk coefficient is a weighting coefficient used to amplify the fluctuation sensitivity factor when the deviation direction points to increased disease risk, emphasizing the importance of pathological deviations. The target fluctuation sensitivity factor is the final trend deviation value obtained after pathological directionality determination and amplification.
[0043] Specifically, the direction of the shift is first determined by the sign of the mean shift. Then, the clinical significance of the physiological indicator is used to determine whether the shift indicates an increased risk of cardiovascular and cerebrovascular diseases. For example, a positive shift in systolic blood pressure, diastolic blood pressure, and total cholesterol indicates an increased risk, while a negative shift in heart rate variability indicates an increased risk. When a pathological direction is identified, the load response intensity factor is amplified by multiplying it by a pathological risk coefficient greater than one to obtain the target fluctuation sensitivity factor. When a non-pathological or protective direction is identified, the load response intensity factor is not amplified or is attenuated using a coefficient less than one. This pathologically oriented weighting process gives higher weight to shifts that have a real impact on disease risk, improving the clinical relevance of influencing factor identification.
[0044] S2034: Based on the fluctuation sensitivity factor and the target fluctuation sensitivity factor, a preset comprehensive evaluation function is used to calculate the numerical fluctuation amplitude of each physiological indicator under different work rhythm states.
[0045] Among them, the comprehensive evaluation function refers to the mathematical function used to integrate the volatility sensitivity factor and the target volatility sensitivity factor, and obtain the comprehensive evaluation result through weighted summation or nonlinear combination.
[0046] In practice, the volatility sensitivity factor and the target volatility sensitivity factor are first normalized to unify their numerical ranges to the same interval. Then, weighting coefficients are assigned to the two types of indicators in the comprehensive evaluation function. The volatility dispersion weight reflects the contribution of discrete changes to risk, and the target trend deviation weight reflects the contribution of mean drift to risk. The normalized volatility sensitivity factor is multiplied by its weighting coefficient, and the normalized target volatility sensitivity factor is multiplied by its weighting coefficient. The two weighted values are then added to obtain the numerical volatility amplitude of the physiological indicator. This numerical volatility amplitude comprehensively reflects the combined volatility intensity of the physiological indicator under different work rhythm states, encompassing both discrete and trend changes. A larger value indicates a more significant overall response of the indicator to changes in workload. S204: Physiological indicators whose numerical fluctuation exceeds a preset threshold are identified as the primary influencing factors.
[0047] In practice, the calculated fluctuation range of each physiological indicator is compared with a preset threshold. When the fluctuation range of an indicator exceeds the preset threshold, that indicator is marked as the first influencing factor and added to the influencing factor list. The preset threshold is determined by statistically analyzing the distribution of the fluctuation ranges of all physiological indicators in the training sample, selecting quantile values that can effectively distinguish between load-sensitive and stable indicators as thresholds, or by determining clinically significant fluctuation range critical points through correlation analysis with cardiovascular and cerebrovascular disease risk. The resulting set of first influencing factors includes all physiological indicators that exhibit significant fluctuations when the work rhythm changes.
[0048] Based on the above embodiments, as an optional implementation method, the following provides a detailed description of each unit of the second influencing factor processing module. The second influencing factor processing module includes an attribute label determination unit, an association value calculation unit, and an association value comparison unit. The specific implementation steps of each of the above units are as follows.
[0049] S301: Determine the attribute labels of the initial influencing factors based on historical longitudinal physiological data. The attribute labels are used to indicate the physiological index category to which the initial influencing factors belong. The initial influencing factors refer to all physiological indicators extracted from historical longitudinal physiological data, which have not yet undergone correlation screening with cardiovascular and cerebrovascular diseases, and serve as the original candidate factor set for feature analysis. Attribute labels are classification markers used to identify the physiological indicator category to which the initial influencing factors belong, clarifying which physiological system or measurement type the indicator belongs to. Physiological indicator categories refer to a classification system of physiological indicators based on physiological function systems or clinical measurement types, including categories such as cardiovascular system indicators, metabolic system indicators, inflammatory markers, and electrophysiological indicators.
[0050] In practice, the system iterates through all physiological indicator names contained in the historical longitudinal physiological data, determines the physiological indicator category corresponding to each indicator based on a pre-defined physiological indicator classification knowledge base, and attaches this category information as an attribute label to the corresponding initial influencing factor. The determination of the attribute label is based on the physiological indicator classification standards in the medical field. For example, systolic blood pressure, diastolic blood pressure, and heart rate are classified as cardiovascular system indicators; total cholesterol, triglycerides, and blood glucose are classified as metabolic system indicators; and heart rate variability and electrocardiogram parameters are classified as electrophysiological indicators.
[0051] Before the attribute label determination unit executes step S301, the relevant units for data cleaning are as follows: The data cleaning subunit is used to obtain baseline initial data and remove subjects who have been diagnosed with cardiovascular and cerebrovascular diseases before enrollment, subjects with a history of malignant tumors, and subjects with missing follow-up records. The data association subunit is used to identify the remaining subjects after screening as valid baseline research subjects and associate them with cardiovascular and cerebrovascular disease incidence or mortality data up to a preset deadline to obtain historical longitudinal physiological data.
[0052] Baseline initial data refers to the basic information and initial physiological measurements of all potential study participants collected at the start of the study, including basic information, medical history, initial physical examination data, and other complete profile information. Participants with a pre-existing history of cardiovascular or cerebrovascular disease are those who have been clinically diagnosed with coronary heart disease, stroke, myocardial infarction, or other cardiovascular or cerebrovascular diseases before inclusion in the study cohort. Participants with a history of malignant tumors are those diagnosed with various malignant tumors before the start of the study; these diseases can interfere with the risk assessment of cardiovascular or cerebrovascular diseases. Participants with missing follow-up records are those who failed to complete scheduled follow-up examinations during the study or whose follow-up data is significantly incomplete, resulting in the inability to obtain complete longitudinal observation data. Valid baseline study participants refer to the set of participants who meet the study inclusion criteria after screening according to exclusion criteria; these individuals do not have the target disease at the start of the study and have complete follow-up conditions. The pre-set cutoff date is the end point of the study observation period, used to define the upper limit of the time range for longitudinal follow-up data collection. Cardiovascular and cerebrovascular disease morbidity or mortality outcome data refers to the information on whether a study subject experienced a cardiovascular or cerebrovascular event or died from a cardiovascular or cerebrovascular disease during the observation period, including the specific time of the event and the type of disease.
[0053] In practice, the initial baseline data of all aerospace professionals was first extracted from the health record system. Then, three exclusion criteria were established for a layered screening process. The first layer of screening read the medical history field for each subject, checking for any diagnoses of cardiovascular and cerebrovascular diseases, and excluding subjects whose diagnoses were earlier than the study start date. The second layer of screening read the tumor history field, excluding subjects with diagnoses of malignant tumors. The third layer of screening calculated the completeness of each subject's follow-up records, determining the ratio of actual follow-up visits to planned follow-up visits, and excluding subjects whose ratios were lower than the preset completeness requirement. After this, the remaining subjects were formally confirmed as valid baseline study subjects, and a study cohort roster was established. Then, all cardiovascular and cerebrovascular disease diagnoses and death records for each valid baseline study subject from the study start date to the preset end date were extracted from the disease monitoring system and death registration system, and these outcome data were organized according to subject number and time sequence. Simultaneously, all periodic physical examination data, physiological indicator measurements, and work status records for each subject within the same time period were extracted from the health monitoring system. Outcome data and physiological monitoring data are associated and matched according to object number. For each object, a complete longitudinal data record is formed, which includes baseline characteristics, physiological measurements at multiple time points, work rhythm status annotation, and final disease outcome. The longitudinal data records of all objects are summarized to form a historical longitudinal physiological dataset.
[0054] S302: Based on attribute labels, input historical longitudinal physiological data into a preset feature association model to calculate the association value of each initial influencing factor on the prediction of cardiovascular and cerebrovascular diseases. The correlation value refers to the quantitative parameter output by the feature association model, which represents the correlation strength between an initial influencing factor and the prediction result of cardiovascular and cerebrovascular diseases. The larger the value, the more significant the contribution of the factor to the disease prediction.
[0055] In practice, the numerical sequence of each initial influencing factor in historical longitudinal physiological data, along with its attribute labels, is input into the feature association model. The feature association model is trained on a large-scale historical sample, and internally, statistical or machine learning methods are used to analyze the association patterns between various physiological indicators and cardiovascular and cerebrovascular disease events. For each initial influencing factor, the model calculates its independent and combined contributions to the prediction of cardiovascular and cerebrovascular diseases based on its numerical distribution characteristics in historical data, its temporal evolution trend, its interaction with other factors, and the category information provided by the attribute labels, outputting a standardized association value.
[0056] S303: Compare the correlation value with the preset correlation threshold, and determine the initial influencing factor whose correlation value exceeds the preset correlation threshold as the second influencing factor.
[0057] In practice, the association value corresponding to each initial influencing factor is read one by one and compared with a preset association threshold. When the association value of an initial influencing factor is greater than the preset association threshold, it is determined that the factor has a statistically significant association with cardiovascular and cerebrovascular diseases, and the factor is marked as a second influencing factor and added to the second influencing factor set. When the association value is less than or equal to the preset association threshold, it is determined that the association between the factor and the disease is insufficient to support it as a key predictive feature, and it is excluded from subsequent analysis. The preset association threshold is determined based on cross-validation analysis during the training phase, selecting a critical value that can effectively remove redundant and noisy features while retaining sufficient predictive information. This threshold simultaneously considers statistical significance and clinical practical significance.
[0058] In one specific embodiment, multiple machine learning algorithms were used to comprehensively screen candidate influencing factors. The LASSO-penalized Cox-min model screened for hypertension and age; the Group SCAD-penalized Cox model screened for age, hypertension, gamma-glutamyl transferase (GGT), and diabetes; and the LASSO-penalized Cox-1se model screened for diabetes, hypertension, family history of hypertension, age, TC / HDL-C, waist circumference, systolic blood pressure, and GGT. Combining the results of each model, factors that repeatedly appeared in multiple models were identified as age, hypertension, diabetes, and GGT. After filtering these candidate factors using attribute labels, eight secondary influencing factors were finally determined: diabetes, family history of hypertension, hypertension, TC / HDL-C, age, systolic blood pressure, waist circumference, and GGT. This set of secondary influencing factors includes both traditional cardiovascular risk factors such as age, hypertension, and diabetes, as well as specific factors for young and middle-aged populations such as family history of hypertension, TC / HDL-C, and GGT.
[0059] Based on the above embodiments, as an optional implementation method, the various units of the model building module are described in detail below. The model building module includes a static baseline module building unit and a dynamic load response module building unit. The specific implementation steps of the above units are as follows.
[0060] S401: Use the second influencing factor to train the survival analysis model and construct a static baseline module for outputting individual basic health risk values; The survival analysis model is a statistical model built based on time-to-event data. It is used to analyze the relationship between the time from the observation point to the occurrence of a specific event and influencing factors, and to assess the probability of an individual's risk of experiencing the target event at different time points. The individual baseline health risk score is a quantitative score of the risk of cardiovascular and cerebrovascular diseases calculated based on an individual's baseline physiological characteristics. This score reflects the individual's inherent health risk level under normal working conditions and is not affected by short-term workload fluctuations. The static baseline module is a functional component in the risk prediction system specifically used to assess an individual's inherent health risk. This module calculates risk based on stable physiological characteristics that do not change with work rhythms, providing an individual with a basic risk level determination.
[0061] In practice, the data from all valid baseline subjects in the historical longitudinal physiological data are first feature-extracted according to the second influencing factor to form a training sample set. Each sample includes the subject's second influencing factor value, follow-up duration, and whether a cardiovascular or cerebrovascular event occurred. Then, a survival analysis algorithm suitable for right-censored data processing is selected as the model framework. The sample set is input into the model for parameter learning. The model determines the weight coefficients of each influencing factor and the baseline risk function by fitting the relationship between the second influencing factor and the time of disease event occurrence. After training, the model is validated and calibrated to ensure that the model's risk prediction accuracy meets the preset standard. Finally, the trained survival analysis model is packaged into a static baseline module. This module receives the individual's second influencing factor value as input and calculates and outputs the individual's baseline health risk value, which represents the individual's baseline risk level of cardiovascular and cerebrovascular disease occurrence without considering the impact of workload.
[0062] S402: Calculate the rate of change of physiological indicators of an individual between the normal maintenance period and the high-load stress period using the first influencing factor, and construct a dynamic load response module for outputting the individual stress sensitivity coefficient.
[0063] The "normal maintenance period" refers to the time during which aerospace professionals are in a routine maintenance work state. During this period, the work intensity and pressure are at normal levels, and the individual's physiological state is relatively stable. The physiological indicator change rate is the ratio of the magnitude of change in an individual's physiological indicator between different work rhythm periods to the baseline value, used to quantify the impact of workload changes on physiological state. The individual stress sensitivity coefficient is a parameter that quantifies the intensity of an individual's physiological stress response to workload changes. The larger the coefficient, the more severe the individual's physiological response to workload fluctuations, and the more obvious the dynamic fluctuations in health risk. The dynamic load response module is a functional component in the risk prediction system specifically used to assess the impact of workload changes on individual health risks. This module captures the dynamic changes in individual physiological indicators with work rhythms.
[0064] In practice, the primary influencing factor data for each study subject is first extracted from historical longitudinal physiological data. Based on work rhythm annotation information, the data is segmented into normal maintenance period data and high-intensity stress period data. For each subject's primary influencing factor, the average value of the indicator during the normal maintenance period and the average value during the high-intensity stress period are calculated. Then, the difference between the two average values is calculated, and this difference is divided by the average value during the normal maintenance period to obtain the physiological indicator change rate. The absolute value of the change rate reflects the sensitivity of the indicator to changes in workload. The change rates of all primary influencing factors for an individual are comprehensively processed. Through weighted summarization or cluster analysis, the change rate information of multiple indicators is integrated into a single individual stress sensitivity coefficient. This coefficient comprehensively characterizes the overall response strength of the individual's physiological system to workload fluctuations. The physiological indicator change rate calculation process and the stress sensitivity coefficient generation process are encapsulated into a dynamic load response module. This module receives the primary influencing factor values of an individual in different work rhythm periods as input and outputs the individual's stress sensitivity coefficient.
[0065] Based on the above embodiments, as an optional implementation method, the model building module also includes a risk integration function building unit and a model optimization unit. The specific implementation steps of the above units are as follows.
[0066] S501: Determine the load weighting factor corresponding to the work rhythm state, and construct a risk integration function with the baseline health risk value, stress sensitivity coefficient and load weighting factor as input parameters and the risk of cardiovascular and cerebrovascular diseases as output. The workload weighting factor is a moderating parameter that quantifies the impact of different work rhythm states on the risk of cardiovascular and cerebrovascular diseases. This factor reflects the amplification or mitigation effect of the current workload intensity on an individual's health risk. The risk integration function is a calculation function that integrates multiple risk factors of different natures into a single comprehensive risk score through mathematical relationships. This function establishes a quantitative mapping relationship between input parameters and the risk of cardiovascular and cerebrovascular diseases.
[0067] In practice, firstly, a corresponding load weighting factor is assigned to each work rhythm state based on its category. The load weighting factor for the normal maintenance period is set as the baseline value, while the load weighting factor for the high-load intensive period is set to a value higher than the baseline value to reflect the amplified risk effect of increased work pressure. Then, the structure of the risk integration function is designed. This function receives three types of input parameters: the first is the basic health risk value from the static baseline module; the second is the stress sensitivity coefficient from the dynamic load response module; and the third is the load weighting factor determined based on the current work rhythm state. The risk integration function is designed using an additive or multiplicative combination, with the basic health risk value as the risk base term and the interaction term between the stress sensitivity coefficient and the load weighting factor as the risk adjustment term. Through a reasonable mathematical combination, the function output can simultaneously reflect the individual's inherent risk level and the dynamic impact of workload. The output value of the function is the predicted risk value of cardiovascular and cerebrovascular diseases for an individual under a specific work rhythm state.
[0068] S502: Using records of cardiovascular and cerebrovascular disease events as supervisory labels, the risk prediction model is constructed by optimizing the weight coefficients and load weight factors of each input parameter in the risk integration function through the gradient descent algorithm.
[0069] Among them, the cardiovascular and cerebrovascular disease event record refers to the actual outcome information of cardiovascular and cerebrovascular disease events that occurred in the research subjects, recorded in historical longitudinal physiological data, including detailed records such as whether the disease occurred, the time of onset, and the type of disease. Supervision labels refer to the true labeled values used to guide the learning of model parameters during the training process of a machine learning model; these labels represent the true result category or numerical value of the training samples. Gradient descent algorithm is an optimization algorithm that finds the minimum value of the loss function by iteratively calculating the gradient of the loss function with respect to the model parameters and updating the parameter values in the reverse direction of the gradient.
[0070] In practice, the process begins by extracting records of cardiovascular and cerebrovascular disease events from historical longitudinal physiological data for all research subjects. Subjects experiencing disease events are labeled as positive, and those without are labeled as negative, forming a supervised label sequence. Each subject's baseline health risk value, stress sensitivity coefficient, and corresponding work rhythm status at the corresponding time point are used as input data. These are substituted into a risk integration function to calculate a predicted risk value. The predicted risk value is then compared with the subject's supervised label, and the deviation between the predicted value and the true label is calculated using either cross-entropy loss or mean squared error loss, forming the loss value. Gradient descent is applied to calculate the partial derivatives of the loss value with respect to the weight coefficients and load weight factors of each input parameter in the risk integration function, obtaining the gradient vectors for each parameter. The values of each weight coefficient and load weight factor are updated according to the reverse direction of the gradient and a preset learning rate, gradually reducing the loss value. This process is repeated multiple times, iterating through all training samples in each iteration and continuously updating the parameters until the loss value converges to a minimum or the preset number of iterations is reached. The risk integration function obtained after optimization has the optimal parameter configuration. This function and its parameters are solidified into a risk prediction model. This model receives an individual's current baseline health risk value, stress sensitivity coefficient, and work rhythm status, and outputs a quantitative prediction value of the individual's risk of cardiovascular and cerebrovascular diseases.
[0071] To verify the predictive performance of the risk prediction model in this application, please refer to Table 1. Table 1 shows the performance comparison results between the model constructed in this study and existing mainstream cardiovascular disease risk prediction models at different prediction timeframes. The table horizontally lists five prediction models, including the model constructed in this study, the ischemic cardiovascular disease risk assessment table, the Framingham Heart Study cardiovascular disease risk prediction model, the China-PAR model, and the SCORE2 model. Vertically, the table lists six prediction timeframes from 5 to 10 years. The predictive performance of each model is quantitatively evaluated using its AUC value and 95% confidence interval. A higher AUC value (closer to 1) indicates stronger discriminative ability, and a narrower confidence interval indicates better stability of the prediction results.
[0072]
[0073] Table 1: Comparison of the performance of the risk prediction model and the control model As shown in Table 1, the AUC values of the model constructed in this study remained above 0.890 across all prediction timeframes, reaching a maximum of 0.926 for the 5-year prediction, with a 95% confidence interval of 0.894 to 0.954, demonstrating high predictive accuracy and stability. In contrast, the AUC values of the ischemic cardiovascular disease risk assessment scale ranged from 0.834 to 0.903 across all prediction timeframes, the Framingham model from 0.837 to 0.910, the China-PAR model from 0.833 to 0.901, and the SCORE2 model from 0.831 to 0.898. The AUC values of the model constructed in this study were higher than the other four control models at all six prediction timeframes, with a particularly significant advantage in short- to medium-term predictions at 5, 6, and 7 years. This comparative result shows that the prediction model constructed in this study for the characteristics of the aerospace occupation group, by incorporating the influencing factors of aerospace mission load, effectively improves the prediction accuracy of cardiovascular and cerebrovascular disease risk for this specific group, and overcomes the limitation of general models being insufficiently adaptable to the aerospace occupation group.
[0074] Based on the above embodiments, as an optional implementation method, the various units of the output module are described in detail below. The output module includes a suggestion item acquisition unit and a suggestion item sorting unit. The specific implementation steps of the above units are as follows.
[0075] S601: Based on the specific impact type of the first influencing factor, obtain the atomized suggestion entries corresponding to the specific impact type from the preset occupational health intervention rule base; Specifically, the "specific impact type" refers to the subcategories of the primary influencing factor based on its physiological function and health impact mechanism, used to identify the corresponding health intervention direction. The pre-established occupational health intervention rule base is a structured knowledge base containing various health influencing factors and their corresponding intervention measures. This rule base stores health management recommendations and intervention plans for different impact types. Atomized suggestion entries are indivisible basic health intervention suggestion units formulated for a single specific impact type in the rule base. Each entry contains a clear intervention goal, intervention measure, and implementation method. In practice, firstly, all primary influencing factors identified by the screened personnel are identified, and the attribute information of each primary influencing factor is read. Based on the factor's physiological function characteristics, clinical significance, and health impact pathway, its specific impact type is determined. Then, the pre-established occupational health intervention rule base is accessed. This rule base uses the specific impact type as the index key and stores one or more sets of atomized suggestion entries corresponding to each impact type. Through query operations, using the specific impact type of each primary influencing factor as the search condition, all atomized suggestion entries associated with that impact type are extracted from the rule base. The recommendations in the rule base are based on evidence-based medicine and occupational health management standards. They provide recommendations for controlling blood pressure and blood lipids for cardiovascular-related impacts, dietary and exercise regulation for metabolic-related impacts, and psychological adjustment and work-rest management for stress-related impacts.
[0076] If the specific impact type of the identified primary influencing factor belongs to a metabolic-related indicator, then a low glycemic index diet recommendation and a fragmented exercise recommendation are matched from the rule base; if the specific impact type of the identified primary influencing factor belongs to a hemodynamic indicator, then an anti-sedentary reminder strategy and a breathing regulation recommendation are matched from the rule base.
[0077] S602: Associate the atomic suggestion items with the risk prediction probability of the test subject, and prioritize the atomic suggestion items according to the value of the risk prediction probability from high to low to generate personalized health suggestions for the test subject.
[0078] In practice, each atomized suggestion item is first bidirectionally associated with the primary influencing factor that generated it. Then, this primary influencing factor is associated with the risk prediction probability of the individual being tested. Through this association, each atomized suggestion item carries corresponding risk prediction probability information, reflecting the contribution of the health problem associated with that suggestion item to the overall risk. Next, all the aggregated atomized suggestion items are sorted based on their associated risk prediction probability values. Items with higher risk prediction probability values are placed at the top, and those with lower values at the bottom.
[0079] Optionally, after generating targeted health recommendations, the physiological indicators of the astronauts under test can be continuously tracked over the next work cycle. If the physiological indicators do not show the expected improvement, the weight of the dynamic load response module in the risk prediction model can be increased, and the astronauts under test can be marked as low-intervention response individuals. The recommendation strength can be automatically upgraded when recommendations are generated next time.
[0080] In practice, the process begins by issuing personalized health advice to the astronauts being tested, while simultaneously establishing an intervention effect tracking file for each individual in the health management system. This file records the time of advice issuance, the content of the advice, and baseline physiological index values. The duration of the next work cycle is determined based on the characteristics of the space mission cycle, typically set as a complete work rhythm cycle including at least one routine maintenance period and one high-intensity intensive period. During this time period, regular physical examinations and physiological index measurements are conducted on the personnel being tested according to a preset monitoring frequency. Each measurement collects all index values for the primary and secondary influencing factors, and the measurement data is recorded in the tracking file to form time-series data. At each measurement point, the current measurement value is compared with the baseline measurement value to calculate the magnitude and direction of change for each physiological index. The tracking process continues until the observation period of a complete work cycle is completed. Afterward, the collected physiological index change data is evaluated, and the measurement values of each primary and secondary influencing factor at the end of the cycle are compared with the baseline value to calculate the actual improvement magnitude. This actual improvement magnitude is then compared with a pre-set expected improvement threshold for determination. When the actual improvement of most key indicators is less than the expected improvement threshold or some indicators deteriorate, the physiological indicators of the test subject are determined to have not improved as expected. Under this determination, a dual response mechanism is executed. The first response is adaptive adjustment of model parameters. The parameter configuration of the risk integration function in the risk prediction model is accessed, the weight coefficient corresponding to the dynamic load response module is located, and the value of this weight coefficient is increased by a preset adjustment step. The increased weight coefficient makes the risk prediction model pay more attention to the impact of workload fluctuations when calculating the risk of the test subject, thus improving the sensitivity of its risk assessment. The second response is population stratification identification and intervention strategy enhancement. A low-intervention response group label is added to the test subject in the health management system, and this label is persistently stored in the personnel's health record. When the test subject re-enters the health recommendation generation process, before generating personalized health recommendations, it is first checked whether the person carries the low-intervention response group label. If the label exists, the recommendation strength upgrade mechanism is triggered. On the basis of atomic recommendation items, the frequency requirements for the implementation of intervention measures are increased, the scope of intervention is expanded to cover more lifestyle factors, and the strictness of restrictive recommendations is increased, generating an enhanced version of personalized health recommendations.
[0081] The following is an embodiment of the system of this application. Please refer to it. Figure 2 This illustration shows a schematic diagram of a cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, provided in an exemplary embodiment of this application. The system can be implemented entirely or partially through software, hardware, or a combination of both. The cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals includes: The data acquisition module is used to acquire historical longitudinal physiological data and corresponding job task cycle data of the aerospace profession. The job task cycle data is used to mark the work rhythm state of the aerospace profession in different time periods. The first influencing factor processing module is used to perform phased differential analysis on historical longitudinal physiological data based on job task cycle data, and to screen out the first influencing factors whose numerical fluctuation exceeds a preset threshold under different work rhythm states. The second influencing factor processing module is used to input historical longitudinal physiological data into a preset feature association model and filter out the second influencing factors whose association values exceed the preset association threshold. The model building module is used to build a dynamic load response module using the first influencing factor, a static baseline module using the second influencing factor, and to build a risk prediction model based on the dynamic load response module and the static baseline module. The model processing module is used to collect the current physiological indicators and current work rhythm status of the person being tested, and input the current physiological indicators and current work rhythm status into the risk prediction model to obtain the risk prediction probability; The output module generates health recommendations for the individuals being tested based on the predicted risk probability and the specific impact type of the primary influencing factor.
[0082] This application also provides a computer storage medium that can store multiple instructions. These instructions are adapted for a processor to load and execute the cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, as described in the above embodiments. The specific execution process can be found in the detailed description of the embodiments, which will not be repeated here. Please see Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 300 may include: at least one processor 301, at least one network interface 304, user interface 303, memory 305, and at least one communication bus 302.
[0083] The communication bus 302 is used to enable communication between these components.
[0084] The user interface 303 may include a display screen and a camera.
[0085] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0086] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of digital signal processing, field-programmable gate array, or programmable logic array. The processor 301 may integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.
[0087] The memory 305 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 305 may include a non-transitory computer-readable medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various system embodiments described above, etc.; the data storage area may store data involved in the various system embodiments described above, etc. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. Figure 3 As shown, the memory 305, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a cardiovascular and cerebrovascular disease risk prediction system for the aerospace profession.
[0088] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 301 can be used to call the application program stored in the memory 305 for a cardiovascular and cerebrovascular disease risk prediction system for the aerospace profession. When executed by one or more processors, the electronic device executes one or more systems as described in the above embodiments.
[0089] An electronic device readable storage medium stores instructions that, when executed by one or more processors, cause the electronic device to perform one or more systems as described in the above embodiments.
[0090] It should be noted that, for the sake of simplicity, the aforementioned system embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0091] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0092] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings or direct couplings or communication connections may be through some service interfaces; indirect couplings or communication connections between apparatuses or units may be electrical or other forms.
[0093] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0094] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0095] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the system of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0096] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and practical application disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure.
Claims
1. A cardiovascular and cerebrovascular disease risk prediction system for aerospace professionals, characterized in that, The system includes: The data acquisition module is used to acquire historical longitudinal physiological data and corresponding job task cycle data of the aerospace profession. The job task cycle data is used to mark the work rhythm state of the aerospace profession in different time periods. The first influencing factor processing module is used to perform phased differential analysis on the historical longitudinal physiological data based on the job task cycle data, and to screen out the first influencing factors whose numerical fluctuation amplitude exceeds a preset threshold under different work rhythm states. The second influencing factor processing module is used to input the historical longitudinal physiological data into a preset feature association model and filter out the second influencing factors whose association values exceed the preset association threshold. The model building module is used to build a dynamic load response module using the first influencing factor, a static baseline module using the second influencing factor, and to build a risk prediction model based on the dynamic load response module and the static baseline module. The model processing module is used to collect the current physiological indicators and current work rhythm status of the person to be tested, and input the current physiological indicators and current work rhythm status into the risk prediction model to obtain the risk prediction probability; The output module is used to generate health recommendations for the person to be tested based on the predicted risk probability and the specific impact type of the first influencing factor.
2. The system according to claim 1, characterized in that, The work rhythm state includes at least a normal maintenance period and a high-load intensive period. The first influencing factor processing module includes: The feature set construction unit is used to divide the historical longitudinal physiological data into time periods according to the normal maintenance period and the high-load tackling period based on the job task cycle data, and construct normal feature set and load feature set respectively; The difference calculation unit is used to calculate the distribution difference and mean shift between the normal characteristic set and the load characteristic set for each physiological indicator. The numerical fluctuation amplitude calculation unit is used to calculate the numerical fluctuation amplitude of each of the physiological indicators under different working rhythm states based on the distribution difference degree and the mean drift amount. The indicator determination unit is used to determine the physiological indicators whose numerical fluctuation exceeds a preset threshold as the first influencing factor.
3. The system according to claim 2, characterized in that, The numerical fluctuation calculation unit includes: The fluctuation sensitivity factor calculation subunit is used to analyze the distribution difference, extract the variance change information of physiological indicators during the normal maintenance period and the high-load intensive period, and calculate the fluctuation sensitivity factor. The fluctuation sensitivity factor characterizes the degree of instability of an individual's physiological response to changes in workload. The load response intensity factor calculation subunit is used to analyze the mean drift, extract the mean deviation information of physiological indicators during the normal maintenance period and the high load challenge period, and calculate the load response intensity factor. The load response intensity factor characterizes the magnitude of the influence of workload on individual physiological indicators. The target fluctuation sensitivity factor determination subunit is used to determine the pathological direction of the fluctuation sensitivity factor. When the deviation direction of the fluctuation sensitivity factor points to the direction of increased risk of cardiovascular and cerebrovascular diseases, the pathological risk coefficient is applied to amplify the fluctuation sensitivity factor to obtain the target fluctuation sensitivity factor. The numerical fluctuation amplitude determination subunit is used to calculate the numerical fluctuation amplitude of each physiological indicator under different work rhythm states by using a preset comprehensive evaluation function based on the fluctuation sensitivity factor and the target fluctuation sensitivity factor.
4. The system according to claim 1, characterized in that, The second influencing factor processing module includes: An attribute label determination unit is used to determine the attribute labels of the initial influencing factors based on the historical longitudinal physiological data, wherein the attribute labels are used to indicate the physiological index category to which the initial influencing factors belong; The correlation value calculation unit is used to input the historical longitudinal physiological data into a preset feature correlation model based on the attribute labels, and calculate the correlation value of each of the initial influencing factors on the prediction of cardiovascular and cerebrovascular diseases. The correlation value comparison unit is used to compare the correlation value with a preset correlation threshold, and to determine the initial influencing factor whose correlation value exceeds the preset correlation threshold as the second influencing factor.
5. The system according to claim 4, characterized in that, The attribute label determination unit includes: The data cleaning subunit is used to obtain baseline initial data and remove subjects who have been diagnosed with cardiovascular and cerebrovascular diseases before enrollment, subjects with a history of malignant tumors, and subjects with missing follow-up records. The data association subunit is used to identify the remaining subjects after screening as valid baseline research subjects and associate them with cardiovascular and cerebrovascular disease incidence or mortality data up to a preset deadline to obtain historical longitudinal physiological data.
6. The system according to claim 1, characterized in that, The model building module includes: The static baseline module construction unit is used to train the survival analysis model using the second influencing factor and construct a static baseline module for outputting individual baseline health risk values. The dynamic load response module construction unit is used to calculate the rate of change of physiological indicators of an individual between the normal maintenance period and the high load challenge period using the first influencing factor, and to construct a dynamic load response module for outputting the individual stress sensitivity coefficient.
7. The system according to claim 6, characterized in that, The model building module includes: The risk integration function construction unit is used to determine the load weighting factor corresponding to the work rhythm state, and uses the basic health risk value, the stress sensitivity coefficient and the load weighting factor as input parameters, and the risk of cardiovascular and cerebrovascular disease as output to construct the risk integration function; The model optimization unit is used to optimize the weight coefficients of each input parameter and the load weight factor in the risk integration function by using records of cardiovascular and cerebrovascular disease events as supervision labels and employing a gradient descent algorithm to construct a risk prediction model.
8. The system according to claim 1, characterized in that, The output module includes: The suggestion item acquisition unit is used to acquire atomic suggestion items corresponding to the specific impact type from a preset occupational health intervention rule base according to the specific impact type of the first influencing factor. The suggestion item sorting unit is used to associate the atomic suggestion items with the risk prediction probability of the person to be tested, and to prioritize the atomic suggestion items according to the value of the risk prediction probability from high to low, so as to generate personalized health suggestions for the person to be tested.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading and execution by a processor of the system as described in any one of claims 1 to 8.
10. An electronic device, characterized in that, The device includes a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the system as described in any one of claims 1 to 8.