Heat stroke early diagnosis method combining pattern classification and clustering
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-11
AI Technical Summary
首先,现有方法难以同时兼顾环境热强度的动态变化和个体生理差异
该结合模式分类和聚类的热射病早期诊断方法,通过融合环境热强度数据与个体生理响应数据,采用模式分类与聚类分析结合深度学习建模,实现对高温环境下个体热应激状态的精确评估。与现有技术相比,本发明具有以下有益效果:一方面,本方法能够同时考虑环境动态变化和个体生理差异,实时分析复杂场景下的多源数据,并通过动态调整和历史数据校正机制优化诊断结果,从而提高热射病早期诊断的准确性和可靠性;另一方面,本发明能够生成干预信号序列并建立反馈循环,及时提示高风险个体采取干预措施,实现个性化风险预测和早期干预。综上,本发明能够显著提升热射病早期诊断的科学性和有效性,为高温环境下人群健康提供可靠保障。
Smart Images

Figure CN122552091A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical health and bioinformatics, specifically to an early diagnostic method for heatstroke that combines pattern classification and clustering. Background Technology
[0002] In the field of healthcare, early diagnosis of heatstroke is of great significance, especially in high-temperature environments and extreme climates, where public safety heavily depends on timely identification of heatstroke. However, existing methods for early diagnosis of heatstroke have significant shortcomings, mainly in the following aspects: First, existing methods struggle to simultaneously account for dynamic changes in environmental heat intensity and individual physiological differences. Environmental factors such as temperature, humidity, and radiation significantly influence the human body's heat stress response, while physiological indicators like heart rate, skin temperature, and sweat secretion vary considerably among individuals. This makes diagnostic tools prone to misdiagnosis or missed diagnosis when dealing with diverse populations. Second, existing diagnostic methods lack the ability to effectively analyze and dynamically adjust data under complex scenarios. In high-temperature outdoor work or extreme climate conditions, the human body's response to heat stress exhibits high time-varying and individual variability. Existing tools struggle to adapt to these changes in real time, limiting the reliability and accuracy of diagnostic results. Furthermore, existing methods are insufficient in early intervention and risk prediction. Current diagnostic tools often lack the ability to establish systematic predictive models for high-risk individuals and lack mechanisms for optimization using historical data and feedback information, thus failing to achieve personalized risk assessment and timely intervention for different individuals and environmental conditions. Summary of the Invention
[0003] The purpose of this invention is to provide an early diagnostic method for heatstroke that combines pattern classification and clustering, thereby addressing the problems existing in the prior art.
[0004] To achieve the above objectives, the present invention provides the following technical solution: an early diagnosis method for heatstroke combining pattern classification and clustering, comprising: S1, real-time collection of temperature, humidity, and radiation data in a high-temperature environment via a sensor network, combined with heart rate, skin temperature, and sweat secretion indicators monitored by an individual wearable device, fused into a comprehensive environmental heat intensity data and an individual physiological response dataset to obtain preliminary heat stress indicators; S2, based on the preliminary heat stress indicators, using a support vector machine algorithm combined with cluster analysis to classify the dynamic changes in environmental heat intensity; if the classification result shows that the heat intensity exceeds a preset threshold, then matching the trend of change with the individual physiological response dataset to identify high-risk individuals. S3. Obtain the determined physical differences of high-risk individuals, model and analyze the diversity differences through a deep learning network, generate a personalized heatstroke symptom prediction model, and determine the probability distribution of potential symptoms; S4. Extract the probability distribution from the personalized heatstroke symptom prediction model, and for real-time data input in complex scenarios, if the probability distribution is higher than the warning threshold, activate the dynamic adjustment mechanism to obtain an optimized diagnostic result sequence; S5. Use the obtained optimized diagnostic result sequence to iteratively optimize the core contradictions. If the reliability of the sequence is limited after iteration, incorporate historical diagnostic data for correction to determine the final early diagnosis output of heatstroke.
[0005] Preferably, step S1 includes acquiring and aligning ambient temperature, ambient humidity, ambient radiation, heart rate, skin temperature, and sweat secretion values to obtain a synchronous multi-source data set; constructing a multi-dimensional feature matrix based on the synchronous multi-source data set; reducing the dimensionality of the multi-dimensional feature matrix to obtain comprehensive environmental heat intensity data and individual physiological response datasets; if the comprehensive environmental heat intensity data is greater than a preset heat intensity threshold, classifying and mapping the comprehensive environmental heat intensity data and individual physiological response datasets to obtain preliminary heat stress values.
[0006] Preferably, step S2 includes acquiring preliminary heat stress data, processing the preliminary heat stress data using a support vector machine algorithm combined with cluster analysis to obtain environmental heat intensity classification results; if the environmental heat intensity classification results exceed a preset threshold, analyzing the environmental heat intensity classification results to determine the heat intensity change trend; matching the heat intensity change trend with the individual physiological response dataset to obtain trend matching results; extracting abnormal physiological fluctuation data based on the trend matching results; and determining the physical differences of high-risk individuals based on the abnormal physiological fluctuation data.
[0007] Preferably, step S3 includes acquiring physiological data and reducing the dimensionality of the physiological data to obtain the physical differences of high-risk individuals; extracting abnormal feature values and removing abnormal feature values based on the physical differences of high-risk individuals to obtain a diversity difference feature vector; modeling the diversity difference feature vector to generate a personalized heatstroke symptom prediction model; using the personalized heatstroke symptom prediction model to output a prediction result matrix, and judging the probability distribution of potential symptoms based on the prediction result matrix.
[0008] Preferably, step S4 includes inputting a first real-time vital sign sequence into a personalized heatstroke symptom prediction model and extracting a first symptom probability distribution; if the first symptom probability distribution is higher than a warning threshold, activating a dynamic adjustment mechanism for the first real-time vital sign sequence to obtain a second feature weight allocation matrix; performing calculations on the first symptom probability distribution according to the second feature weight allocation matrix to obtain a second symptom probability distribution; classifying the second symptom probability distribution to determine the risk level; and generating an optimized diagnostic result sequence based on the risk level.
[0009] Preferably, step S5 includes acquiring a first diagnostic result sequence, calculating a sequence reliability metric based on the first diagnostic result sequence, determining whether the sequence reliability metric is restricted, including determining a restricted status identifier if the sequence reliability metric is lower than a preset threshold, obtaining a correction weight matrix by weighting historical diagnostic data according to the restricted status identifier, obtaining a fusion feature set through the correction weight matrix, and determining the final early diagnosis output of heatstroke based on the fusion feature set.
[0010] Preferably, it also includes S6: generating an intervention signal sequence through the final early diagnosis output of heatstroke; if the signal sequence indicates the need for timely intervention, constructing a feedback loop based on the output, and judging the interaction between environmental heat intensity and individual physical condition in the loop, specifically including obtaining physiological characteristic data and heat stress state to generate an early diagnosis output of heatstroke; and generating an intervention signal sequence based on the early diagnosis output of heatstroke.
[0011] Preferably, step S6 further includes constructing a feedback loop if the intervention signal sequence indicates a need for timely intervention; obtaining the environmental heat intensity and the individual's thermoregulation ability in the feedback loop; constructing a risk assessment matrix based on the thermoregulation ability; and determining the interaction between the environmental heat intensity and the individual's physical condition in the feedback loop.
[0012] Preferably, the method further includes S7: extracting interaction data from the feedback loop, reclassifying and processing it using a support vector machine algorithm to obtain refined heat stress indicators, which are used as initialization inputs for subsequent diagnostic cycles. Specifically, this includes obtaining physiological characteristics and environmental loads from historical diagnostic feedback loops, extracting interaction data between physiological characteristics and environmental loads in the time dimension, and performing time-domain feature extraction based on the interaction data to obtain an interaction feature set containing body temperature fluctuation sequences and heart rate variation sequences.
[0013] Preferably, step S7 further includes using a support vector machine algorithm to reclassify the body temperature fluctuation sequence and heart rate variation sequence in the interaction feature set to obtain a classification boundary in a multidimensional space; if the classification boundary meets a preset convergence threshold, redundant data is removed based on the classification boundary to obtain refined features; and the refined heat stress value is determined based on the refined features as the initialization input for subsequent diagnostic cycles.
[0014] As can be seen from the above technical solution, the present invention has the following beneficial effects: This early diagnosis method for heatstroke, combining pattern classification and clustering, integrates environmental heat intensity data with individual physiological response data. It employs pattern classification and clustering analysis combined with deep learning modeling to achieve accurate assessment of individual heat stress status under high-temperature environments. Compared with existing technologies, this invention offers the following advantages: Firstly, it simultaneously considers dynamic environmental changes and individual physiological differences, analyzes multi-source data in complex scenarios in real time, and optimizes diagnostic results through dynamic adjustment and historical data correction mechanisms, thereby improving the accuracy and reliability of early heatstroke diagnosis. Secondly, it generates intervention signal sequences and establishes feedback loops, promptly prompting high-risk individuals to take intervention measures, achieving personalized risk prediction and early intervention. In summary, this invention significantly improves the scientific rigor and effectiveness of early heatstroke diagnosis, providing reliable protection for population health under high-temperature conditions. Attached Figure Description
[0015] Figure 1 This is a flowchart of the early diagnosis method for heatstroke according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] like Figure 1As shown, this invention provides a technical solution: an early diagnosis method for heatstroke combining pattern classification and clustering, including: S1, real-time collection of temperature, humidity, and radiation data in a high-temperature environment via a sensor network, combined with heart rate, skin temperature, and sweat secretion indicators monitored by an individual wearable device, fused into comprehensive environmental heat intensity data and an individual physiological response dataset to obtain preliminary heat stress indicators; S2, based on the preliminary heat stress indicators, using a support vector machine algorithm combined with cluster analysis to classify the dynamic changes in environmental heat intensity. If the classification result shows that the heat intensity exceeds a preset threshold, the change trend is matched with the individual physiological response dataset to determine the physical differences of high-risk individuals; S3, obtaining the determined physical differences of high-risk individuals, using a deep learning network to model and analyze the diversity differences, generating a personalized heatstroke symptom prediction model to determine potential symptoms. S4. Extract the probability distribution from the personalized heatstroke symptom prediction model. For real-time data input in complex scenarios, if the probability distribution is higher than the warning threshold, activate the dynamic adjustment mechanism to obtain an optimized diagnostic result sequence; S5. Use the obtained optimized diagnostic result sequence to iteratively optimize the core contradiction. If the reliability of the sequence is limited after iteration, incorporate historical diagnostic data for correction to determine the final early diagnosis output of heatstroke; S6. Generate an intervention signal sequence based on the final early diagnosis output of heatstroke. If the signal sequence indicates the need for timely intervention, construct a feedback loop based on the output to determine the interaction between environmental heat intensity and individual physical condition in the loop; S7. Extract the interaction data from the feedback loop, reclassify and process it using the support vector machine algorithm to obtain refined heat stress indicators, which are used as the initial input for subsequent diagnostic cycles.
[0018] In the above implementation, a sensor network is used to continuously collect temperature, humidity, and radiation data in a high-temperature environment, while individual wearable devices are used to simultaneously monitor heart rate, skin temperature, and sweat secretion indicators. The collected environmental and physiological data are aligned according to a unified timestamp and then sequentially processed for outlier removal, missing data compensation, smoothing filtering, and normalization to form a comprehensive environmental heat intensity data set and an individual physiological response dataset. The comprehensive environmental heat intensity data characterizes the degree of external heat exposure, while the individual physiological response dataset characterizes the physiological load on the human body under heat exposure conditions. By fusing environmental and physiological data, preliminary heat stress indicators are obtained, enabling the system to incorporate both environmental and individual factors, avoiding biased judgments caused by relying solely on a single temperature or physiological parameter.
[0019] Preliminary heat stress indicators serve as input data for the Support Vector Machine (SVM) algorithm and cluster analysis. The SVM algorithm is used to classify patterns in the dynamic changes of environmental heat intensity, categorizing environmental heat intensity states into low-risk, medium-risk, high-risk, and dangerous states. Cluster analysis is used to classify the differences in physiological responses among individuals under the same environmental heat intensity conditions, classifying individuals into different heat-sensitive types based on the magnitude of heart rate increase, the rate of skin temperature change, the trend of sweat secretion changes, and physiological recovery time. If the classification results show that the heat intensity exceeds a preset threshold, the system matches the trend of environmental heat intensity changes with the individual's physiological response dataset to determine the physical differences in high-risk individuals. These physical differences reflect an individual's tolerance to high-temperature, high-humidity, or strong radiation environments and their tendency to respond abnormally.
[0020] Once the physical differences of high-risk individuals are identified, they are input into a deep learning network for modeling and analysis. The deep learning network jointly learns from changes in environmental heat intensity, physiological responses, and individual physical differences in continuous time series to establish a personalized heatstroke symptom prediction model. This prediction model outputs a probability distribution of potential symptoms based on data trends within the current and historical time windows. This probability distribution represents the probability of an individual experiencing mild heat stress, moderate heat-related discomfort, or early-stage heatstroke. In this way, the system identifies the current risk status while predicting the development direction of potential symptoms based on data trends.
[0021] The system extracts a probability distribution from the personalized heatstroke symptom prediction model and compares this distribution with a warning threshold. In complex scenarios, such as rapid changes in ambient temperature, continuously rising humidity, fluctuations in radiation intensity, or short-term anomalies in individual physiological data, a dynamic adjustment mechanism is activated if the probability distribution exceeds the warning threshold. This mechanism adjusts the sampling frequency, feature weights, judgment threshold, sliding window length, and model confidence correction parameters to adapt the diagnostic results to real-time data changes. After dynamic adjustment, the system generates an optimized diagnostic result sequence, which continuously characterizes changes in an individual's early heatstroke risk over multiple diagnostic periods.
[0022] The optimized diagnostic result sequence is further used for iterative optimization of the core contradictions. These core contradictions manifest as the balance between diagnostic sensitivity and false alarm rate, and between diagnostic real-time performance and diagnostic accuracy. The system assesses the reliability of the current diagnostic sequence by analyzing the consistency of consecutive diagnostic results, probability fluctuation amplitude, data integrity, and model confidence. If the reliability of the iterative sequence is limited, such as due to missing sensor data, frequent fluctuations in model output, inconsistencies in diagnostic results, or probability distributions approaching threshold boundaries, historical diagnostic data is incorporated for correction. Historical diagnostic data includes the individual's past heat exposure records, physiological response baseline, past warning results, post-intervention recovery status, and statistical data of similar populations. After correction using historical data, the system determines the final early diagnosis output for heatstroke.
[0023] The final early diagnosis output for heatstroke is used to generate an intervention signal sequence. This sequence includes on-site alarm signals, mobile terminal notification signals, management platform warning signals, cooling equipment activation signals, and medical assistance signals. If the signal sequence indicates a need for timely intervention, the system constructs a feedback loop based on the diagnostic output. This feedback loop continuously observes the relationship between environmental heat intensity and individual physiological responses before and after intervention, and assesses the interaction between environmental heat intensity and individual constitution. If, after a reduction in environmental heat intensity, the individual's heart rate and skin temperature remain abnormal, the individual is considered to have high heat sensitivity; if the environmental heat intensity is high and the fluctuation range of the individual's physiological response is within a safe range, and the recovery time meets preset recovery conditions, the individual is considered to have short-term heat tolerance.
[0024] Interaction data is extracted from the feedback loop and re-input into a support vector machine algorithm for classification. The reclassified results are used to refine and improve heat stress indicators, ensuring that the initial input for the next diagnostic cycle includes the diagnostic results, intervention feedback, and individual physical changes from the previous cycle. Through this closed-loop processing, this method achieves continuous updates between environmental data, physiological data, model prediction results, intervention feedback data, and subsequent diagnostic inputs, thereby improving the individualization, real-time performance, and reliability of early heatstroke diagnosis.
[0025] S1 includes acquiring and aligning ambient temperature, ambient humidity, ambient radiation, heart rate, skin temperature, and sweat secretion values to obtain a synchronous multi-source data set; constructing a multi-dimensional feature matrix based on the synchronous multi-source data set; reducing the dimensionality of the multi-dimensional feature matrix to obtain comprehensive environmental heat intensity data and individual physiological response datasets; if the comprehensive environmental heat intensity data is greater than a preset heat intensity threshold, classifying and mapping the comprehensive environmental heat intensity data and individual physiological response datasets to obtain preliminary heat stress values.
[0026] In this embodiment, the data acquisition step is used to organize environmental and individual data into data under the same time reference. Environmental data includes ambient temperature, ambient humidity, and ambient radiation values, while individual data includes heart rate characteristics, skin temperature, and sweat secretion values. First, a diagnostic time window is set in the data acquisition gateway. For continuous monitoring of outdoor high-temperature operations, the diagnostic time window is set to 30 seconds; for monitoring at fixed positions, the diagnostic time window is set to 60 seconds. After each diagnostic time window, the data acquisition gateway summarizes all valid sampled data within that time window and generates one set of data to be processed.
[0027] For the same type of data within the same diagnostic time window, the system first counts the number of valid sampling points within that window, then adds up the values of all valid sampling points and divides by the number of valid sampling points to obtain the aligned values within that diagnostic time window. This yields aligned values for ambient temperature, ambient humidity, ambient radiation, heart rate characteristics, skin temperature, and sweat secretion. If a sensor of a certain type does not upload a sampling value at the exact end of the diagnostic time window, the system selects the most recent sampling value before and after that end time and interpolates them according to the sampling time distance; the closer the sampling value is to the end time, the greater its proportion in the interpolation result. Through this processing method, sensor data from different sampling frequencies are unified into the same diagnostic time window, forming a synchronous multi-source data set.
[0028] After the synchronous multi-source data set is formed, the system performs a validity check. Each type of sensor has its detection range pre-written; data exceeding this range is deleted. For data of the same type within two adjacent diagnostic time windows, the system calculates the difference between the two values and compares it to the maximum permissible variation recorded in the sensor's calibration file. If the difference exceeds the maximum permissible variation, the latter data is marked as an abrupt change anomaly. Abrupt change anomalies are not directly included in subsequent calculations; instead, compensation is performed using data from the preceding and following valid diagnostic time windows. During compensation, the system weights the data according to the time distance between the two valid diagnostic time windows and the anomaly window, with data closer to the anomaly window having a larger weighting. If no valid data exists before or after the anomaly window, the diagnostic time window does not output preliminary heat stress values and returns a data missing marker to the acquisition gateway.
[0029] After completing data alignment and validity checks, the system constructs a multidimensional feature matrix based on the synchronized multi-source dataset. The multidimensional feature matrix is arranged chronologically, with each row representing a diagnostic time window and each column representing a feature class. The first column represents ambient temperature, the second ambient humidity, the third ambient radiation, the fourth heart rate, the fifth skin temperature, and the sixth sweat secretion. This arrangement organizes environmental and physiological changes within multiple consecutive diagnostic time windows into structured data, upon which subsequent dimensionality reduction, threshold comparison, and classification mapping are performed.
[0030] Because ambient temperature, humidity, radiation, heart rate, skin temperature, and sweat secretion are measured in different units, the system standardizes each feature column before dimensionality reduction. The historical mean and standard deviation required for standardization are obtained from the calibration samples before system deployment. For each feature column, the system first calculates the average value of the feature from the calibration samples, then calculates the dispersion of the feature relative to the average value, and stores the dispersion as the standard deviation in the parameter table. After real-time data enters the system, the system subtracts the corresponding historical mean from the current feature value and then divides it by the corresponding historical standard deviation to obtain the standardized feature value. Through this process, the six different units of data are transformed to a unified numerical scale, avoiding unreasonable weighting of features with large numerical units in subsequent calculations.
[0031] The comprehensive environmental thermal intensity data is calculated from three standardized features: ambient temperature, ambient humidity, and ambient radiation. Specifically, the system extracts the standardized columns for ambient temperature, ambient humidity, and ambient radiation from the multidimensional feature matrix to form an environmental feature submatrix. Then, the system calculates the covariance between each pair of these three features. When calculating the covariance between two features, the system multiplies the two standardized values within the same diagnostic time window, sums the products over all valid diagnostic time windows, and then divides by the number of valid diagnostic time windows minus one. Following this method, the system obtains a 3-row, 3-column environmental covariance matrix.
[0032] After establishing the environmental covariance matrix, the system calculates the eigenvalues of the matrix and the corresponding projection weights for each eigenvalue. The system selects the projection weight corresponding to the eigenvalue with the largest value as the weight for calculating environmental heat intensity. This projection weight includes temperature weight, humidity weight, and radiation weight. For any diagnostic time window, the system multiplies the standardized environmental temperature value by the temperature weight, the standardized environmental humidity value by the humidity weight, and the standardized environmental radiation value by the radiation weight, then adds the three products to obtain the comprehensive environmental heat intensity data for that diagnostic time window. If the calculated comprehensive environmental heat intensity data decreases with increasing temperature, humidity, and radiation, the system simultaneously reverses the values of the three weights, causing the comprehensive environmental heat intensity data to increase with increasing external heat load. After this processing, environmental temperature, environmental humidity, and environmental radiation are compressed into a single continuous value, which is the comprehensive environmental heat intensity data.
[0033] The individual physiological response dataset is calculated from three standardized features: heart rate, skin temperature, and sweat secretion. The system extracts the standardized columns for heart rate, skin temperature, and sweat secretion from the multidimensional feature matrix to form an individual physiological response feature submatrix. Following the same covariance calculation method as for environmental features, the system obtains a 3x3 physiological covariance matrix. Then, the system calculates the eigenvalues and corresponding projection weights of this physiological covariance matrix and selects the two largest eigenvalues corresponding to two sets of projection weights. For any given diagnostic time window, the system uses the first set of projection weights to weightedly sum the standardized values of heart rate, skin temperature, and sweat secretion to obtain the first physiological response component; then, it uses the second set of projection weights to weightedly sum these values to obtain the second physiological response component. The first physiological response component reflects the main response direction of the human body to heat exposure, while the second physiological response component reflects supplementary changes that differ from the main response direction. The two physiological response components together constitute the individual physiological response dataset.
[0034] The heat intensity threshold is calculated and determined using historical calibration samples before system deployment, and is not manually entered. The historical calibration samples include multiple sets of labeled data, each set containing comprehensive environmental heat intensity data and a corresponding status marker. A status marker of 0 indicates that the sample has not reached the heat intensity limit; a status marker of 1 indicates that the sample has reached the heat intensity limit. The system first arranges the comprehensive environmental heat intensity data from all historical samples in ascending order, and then takes the median value of two adjacent comprehensive environmental heat intensity data as candidate thresholds. For each candidate threshold, the system counts the number of samples in four categories: Category 1: the number of samples that actually reached the limit and are judged as exceeding the limit; Category 2: the number of samples that actually reached the limit but are judged as not exceeding the limit; Category 3: the number of samples that actually did not reach the limit and are judged as not exceeding the limit; Category 4: the number of samples that actually did not reach the limit but are judged as exceeding the limit.
[0035] After statistically analyzing the sample sizes across the four categories, the system calculates the sensitivity and specificity for each candidate threshold. Sensitivity is calculated by dividing the number of samples that actually exceeded the limit and were judged as exceeding it by the total number of samples that actually exceeded the limit. Specificity is calculated by dividing the number of samples that did not actually exceed the limit and were judged as not exceeding it by the total number of samples that did not actually exceed the limit. The system then adds the sensitivity and specificity together and subtracts 1 to obtain the comprehensive discrimination value for that candidate threshold. The candidate threshold with the highest comprehensive discrimination value is determined as the thermal intensity threshold. If two or more candidate thresholds have the same comprehensive discrimination value, the candidate threshold with the smallest value is selected as the thermal intensity threshold, allowing the system to proceed to the subsequent analysis process earlier in the early diagnostic stage. This thermal intensity threshold is stored in conjunction with the sensor combination, application scenario, and monitored population category; when the sensor combination, application scenario, or monitored population category changes, the system re-executes the above threshold calculation process.
[0036] During real-time diagnosis, the system compares the comprehensive environmental heat intensity data of the current diagnostic time window with the heat intensity threshold. When the comprehensive environmental heat intensity data is not greater than the heat intensity threshold, the system records the current window as a normal monitoring state and proceeds to the next diagnostic time window. When the comprehensive environmental heat intensity data is greater than the heat intensity threshold, the system determines that the current external heat load has entered a range that requires judgment in conjunction with individual physiological responses, and inputs the comprehensive environmental heat intensity data, the first physiological response component, and the second physiological response component from the individual physiological response dataset into the classification mapping model.
[0037] The training of the classification mapping model is completed before system deployment. Training samples consist of historical comprehensive environmental heat intensity data, historical individual physiological response data, and manually confirmed heat stress states. During the training phase, the system first assigns weights to the comprehensive environmental heat intensity data, the first physiological response component, and the second physiological response component, and sets one bias parameter. For each input training sample, the system calculates the product of the three input data points and their corresponding weights, then adds these three products to the bias parameter to obtain a discrimination score. The discrimination score is converted into a risk mapping value between 0 and 1 using a logical function. Subsequently, the system compares the risk mapping value with the actual state of the training samples and adjusts the weights and bias parameters based on the comparison error. This process is repeated until the overall error on the training samples reaches the preset convergence condition. After training, the weights of the comprehensive environmental heat intensity data, the physiological response component weights, and the bias parameter are embedded into the classification mapping model, which is directly invoked during the real-time diagnosis phase.
[0038] After triggering classification mapping within the current diagnostic time window, the system calculates preliminary heat stress values according to the trained classification mapping model. The specific process is as follows: first, the comprehensive environmental heat intensity data is multiplied by the corresponding weight; then, the first physiological response component is multiplied by its corresponding weight; then, the second physiological response component is multiplied by its corresponding weight; finally, the above three products are added to the bias parameter to obtain the discrimination score for the current window. The system then inputs the discrimination score into a logic function, converting it into a risk mapping value between 0 and 1, and multiplies this risk mapping value by 100 to obtain a preliminary heat stress value within the range of 0 to 100. The higher the preliminary heat stress value, the higher the degree of heat stress indicated by both the current environmental heat load and the individual's physiological response. This preliminary heat stress value serves as input data for subsequent support vector machine classification and cluster analysis to further identify the physical differences in high-risk individuals.
[0039] S2 includes acquiring preliminary heat stress data, processing the preliminary heat stress data using a support vector machine algorithm combined with cluster analysis to obtain environmental heat intensity classification results; if the environmental heat intensity classification results exceed a preset threshold, analyzing the environmental heat intensity classification results to determine the heat intensity change trend; matching the heat intensity change trend with the individual physiological response dataset to obtain trend matching results; extracting abnormal physiological fluctuation data based on the trend matching results; and determining the physical differences of high-risk individuals based on the abnormal physiological fluctuation data.
[0040] In this embodiment, the individual physiological response dataset includes physiological response components obtained by dimensionality reduction of heart rate features, skin temperature and sweat secretion data, while retaining the original heart rate feature values, skin temperature values and sweat secretion values for subsequent abnormal fluctuation extraction.
[0041] First, preliminary heat stress data are organized according to the time sequence of the diagnostic time window. For each diagnostic time window, one set of data to be classified is generated. This data includes at least the preliminary heat stress value for the current window, the overall environmental heat intensity data for the current window, the increase in overall environmental heat intensity relative to the previous window, the number of increases in overall environmental heat intensity within three consecutive windows, the average of the preliminary heat stress values within five consecutive windows, and individual physiological response data within the same window. The increase in overall environmental heat intensity is calculated by subtracting the overall environmental heat intensity data of the previous window from the current window's data; a result greater than 0 is considered an increase, and a result not greater than 0 is considered no increase. The number of increases within three consecutive windows is obtained by comparing adjacent windows one by one, with a maximum of two. The average of the preliminary heat stress values within five consecutive windows is obtained by adding the five preliminary heat stress values and then dividing by 5. Through this process, the system simultaneously incorporates single-point risk, short-term changes, and continuous changes into the classification input.
[0042] The Support Vector Machine (SVM) algorithm was trained before system deployment. Training samples were derived from historical monitoring data. Each sample set included preliminary heat stress values, comprehensive environmental heat intensity data, the increase in comprehensive environmental heat intensity, the number of consecutive increases, the average value of consecutive windows, and the manually confirmed environmental heat intensity level. The environmental heat intensity level was divided into four levels: Level 1 was low heat intensity, Level 2 was medium heat intensity, Level 3 was high heat intensity, and Level 4 was dangerous heat intensity. Before training, the system standardized each type of input feature by subtracting the historical mean from the current value of the feature and then dividing by the historical standard deviation, thus bringing features with different dimensions to the same numerical scale.
[0043] Support Vector Machines (SVMs) employ a multi-class classification approach to determine four levels of environmental heat intensity. The system pairs the four levels together to train six classifiers. Each classifier is responsible for distinguishing only two levels. During training, the system finds a classification boundary based on labeled samples, ensuring that samples on either side of the boundary belong to different levels as much as possible, while maximizing the distance between the nearest sample and the boundary. After training, each classifier retains the support samples used to determine the classification boundary, their weights, class orientation parameters, and bias parameters. In real-time operation, the data to be classified is sequentially fed into the six classifiers. Each classifier calculates a discrimination score based on the similarity between the data and the support samples. If the discrimination score is greater than 0, one level is output; otherwise, the other level is output. After each of the six classifiers completes its judgment, the system counts the votes for each of the four levels and uses the level with the most votes as the SVM classification result. If two levels have the same number of votes, the system compares the absolute values of the discrimination scores of the corresponding classifiers for these two levels; the level with the larger absolute value is the output level.
[0044] The penalty and kernel parameters in the Support Vector Machine (SVM) are determined through optimization using historical samples. During the training phase, the system sets multiple candidate parameter combinations and divides the historical samples into training and validation samples. Each candidate parameter combination undergoes one training iteration, and the number of correctly classified cases is counted on the validation samples. The parameter combination with the highest number of correct classifications is determined as the final parameter combination. If two parameter combinations have the same number of correct classifications, the combination that minimizes the number of missed classifications for high and dangerous heat intensities is selected. In this way, the parameters of the SVM are derived from the validation results of historical samples, rather than being arbitrarily assigned manually.
[0045] Cluster analysis is used to supplement support vector machines in identifying patterns of thermal intensity changes. The system uses five consecutive diagnostic time windows as one cluster analysis unit, selecting preliminary thermal stress values, comprehensive environmental thermal intensity data, the increase in comprehensive environmental thermal intensity, and the number of increases within that unit as cluster inputs. The number of clusters is determined before system deployment. The system performs trial calculations on historical samples with 2, 3, 4, 5, and 6 clusters respectively. For each cluster size, the system first calculates the distance from each sample to its own cluster center, and then calculates the distance from each sample to the nearest other cluster center. In distance calculation, the difference between each feature in the sample and the corresponding feature at the cluster center is first calculated, then the squares of each difference are summed, and finally the square root of the sum is taken to obtain the distance from the sample to the cluster center. The smaller the distance to the current cluster and the larger the distance to other clusters, the clearer the cluster division. The system calculates the cluster quality evaluation result based on this distance relationship and selects the cluster size corresponding to the highest evaluation result. If the evaluation results of the two clustering numbers are the same, the clustering number with the smaller number should be selected to avoid splitting the same heat intensity change process into too many categories.
[0046] After the clustering model is established, each cluster center corresponds to one type of heat intensity change pattern. Based on the change characteristics of each cluster in historical samples, the system labels the clusters as stable, continuously increasing, rapidly increasing, and fluctuating increasing. During real-time operation, data from the current five consecutive diagnostic time windows are input into the clustering model. The system calculates the distance from the current data to each cluster center and assigns the current data to the cluster center with the smallest distance, thus obtaining the cluster label. This cluster label reflects the current heat intensity change pattern and does not directly replace the support vector machine classification level, but is used to correct the environmental heat intensity classification results.
[0047] After the support vector machine (SVM) classification results and cluster labels are fused, the environmental heat intensity classification result is obtained. The specific fusion process is as follows: the low, medium, high, and hazardous heat intensities output by the SVM are converted into levels 1, 2, 3, and 4, respectively. When the cluster label is stable, the classification level remains unchanged; when the cluster label is continuously increasing, the system checks whether the comprehensive environmental heat intensity in at least 3 out of 5 consecutive windows is higher than the previous window; if so, the classification level is increased by 1 level, up to a maximum of level 4; when the cluster label is rapidly increasing, the system checks whether the increase in comprehensive environmental heat intensity in the current window exceeds the historical safe increase threshold; if so, the classification level is increased by 1 level, up to a maximum of level 4; when the cluster label is fluctuating, the system counts the number of times the classification level exceeds a preset threshold within 5 consecutive windows; if the number reaches 3 times, the classification level is increased by 1 level, up to a maximum of level 4. The level after the above corrections is the environmental heat intensity classification result.
[0048] The preset thresholds used for environmental heat intensity classification results are determined using historical calibration samples before deployment. These historical calibration samples include environmental heat intensity classification levels and corresponding risk states. Risk states are provided by manual review records and are divided into those not in a high-risk state and those in a high-risk state. The system sequentially verifies levels 1, 2, and 3 as candidate thresholds. For each candidate threshold, the system counts the number of samples in four categories: samples that actually entered a high-risk state and whose classification level exceeds the candidate threshold; samples that actually entered a high-risk state but whose classification level does not exceed the candidate threshold; samples that actually did not enter a high-risk state and whose classification level does not exceed the candidate threshold; and samples that actually did not enter a high-risk state but whose classification level exceeds the candidate threshold. Subsequently, the system calculates the sensitivity and specificity of the candidate threshold. Sensitivity is obtained by dividing the number of samples that actually entered a high-risk state and were correctly identified by the total number of samples that actually entered a high-risk state; specificity is obtained by dividing the number of samples that actually did not enter a high-risk state and were correctly identified by the total number of samples that did not actually enter a high-risk state. The system adds the sensitivity and specificity, then subtracts 1 to obtain the comprehensive discriminant value of the candidate threshold. The candidate threshold with the highest comprehensive discriminant value is determined as the preset threshold. If multiple candidate thresholds have the same comprehensive discriminant value, the system selects the candidate threshold with the smallest value to allow for earlier trend analysis in the early diagnosis stage.
[0049] When the environmental heat intensity classification result does not exceed the preset threshold, the system only records the classification level, cluster label, and current time for that window, and then proceeds to the next diagnostic window. When the environmental heat intensity classification result exceeds the preset threshold, the system performs trend analysis on the environmental heat intensity classification result. The trend analysis uses five consecutive diagnostic time windows as the base windows. The system sequentially reads the environmental heat intensity classification level and comprehensive environmental heat intensity data within each of the five windows, first comparing whether each window has increased relative to the previous window, and then counting the number of times the increase judgment is valid within the five windows. If the comprehensive environmental heat intensity data of at least three of the five windows is higher than that of the previous window, and the classification level of the fifth window is higher than that of the first window, then it is determined to be an upward trend in heat intensity. If the classification levels of all five windows reach or exceed the preset threshold, and the comprehensive environmental heat intensity data does not fall below the heat intensity threshold, then it is determined to be a sustained high-level trend. If the difference between the highest and lowest classification levels within the five windows reaches two levels, and there are at least two instances where the data jumps from never exceeding the preset threshold to exceeding the preset threshold, then it is determined to be a trend of increasing fluctuation. The above trend types are the heat intensity change trends.
[0050] The safe increase threshold for heat intensity trends is determined using historical samples. The system extracts the increase in comprehensive environmental heat intensity between consecutive diagnostic time windows from historical samples that have never entered a high-risk state, and arranges them in ascending order of value. The system counts the total number of sorted samples, multiplies the total number of samples by 0.95, and rounds the result up to obtain the target index. Subsequently, the system reads the increase in comprehensive environmental heat intensity corresponding to the target index and uses it as the safe increase threshold. During real-time operation, if the current increase in comprehensive environmental heat intensity is greater than this safe increase threshold, it is determined that the current heat intensity increase has exceeded the main fluctuation range in historical safe samples, and this result is used as one of the criteria for judging a rapid upward trend.
[0051] Once the trend of heat intensity change is determined, the system matches this trend with an individual physiological response dataset. Before matching, the system first determines the standard delay time between environmental change and physiological response. The standard delay time is calculated from historical samples before deployment. In each set of historical samples, the system identifies the point in time when the overall environmental heat intensity begins to rise continuously, and then identifies the point in time when heart rate characteristics, skin temperature, or sweat secretion begin to rise continuously; the time difference between these two points is recorded as one delay sample. The system arranges all delay samples in ascending order and takes the value at the median as the standard delay time. Subsequently, the system calculates the absolute value of the difference between each delay sample and the standard delay time, and then sorts these absolute values again, taking the value at the median as the allowable delay deviation. The standard delay time and the allowable delay deviation together define the time range for trend matching.
[0052] During real-time matching, the system first reads the start and end times corresponding to the heat intensity change trend, then shifts the entire time period forward by a standard delay time to form a physiological response observation period. Within this period, the system reads the individual physiological response dataset, as well as raw heart rate, skin temperature, and sweat secretion values. The system performs directional matching, amplitude matching, and temporal matching respectively. Directional matching is implemented as follows: when the heat intensity trend is upward or persistently high, the system checks if at least two of the heart rate, skin temperature, and sweat secretion values are increasing; an upward trend is determined by the value at the end of the current observation period being greater than the value at the beginning. Amplitude matching is implemented as follows: the system subtracts the individual baseline average value of each physiological indicator from its maximum value within the observation period and compares it with the individual baseline fluctuation threshold for that indicator; indicators exceeding the individual baseline fluctuation threshold are considered amplitude abnormalities. Temporal matching is implemented as follows: the system finds the time point when a physiological indicator first rises for two consecutive windows and determines whether this time point falls within the range formed by adding or subtracting the allowable delay deviation from the standard delay time.
[0053] The individual baseline fluctuation threshold is calculated from the individual's historical data under non-hyperthermal conditions. The system records the heart rate characteristics, skin temperature, and sweat secretion data of each monitored subject in a non-hyperthermal state in the individual baseline table. For each indicator, the system first calculates the average of all baseline samples, then calculates the difference between each baseline sample and the average, then squares each difference, adds them together, divides by the number of baseline samples minus 1, and finally takes the square root of the result to obtain the individual baseline standard deviation. The individual baseline fluctuation threshold is the individual baseline average plus twice the individual baseline standard deviation. During real-time matching, if the current physiological indicator is higher than this individual baseline fluctuation threshold, the system determines that the indicator exceeds the individual baseline fluctuation range.
[0054] After direction matching, amplitude matching, and time matching are completed, the system generates a trend matching result. A strong match is recorded when all three matches are satisfied; a medium match is recorded when two matches are satisfied; and a weak match is recorded when fewer than two matches are satisfied. Strong and medium matches proceed to the abnormal physiological fluctuation data extraction process. Weak matches are recorded as normal responses and are not included in the high-risk individual physical difference feature extraction process. Through this matching method, the system does not simply equate increased environmental heat intensity with high individual risk, but rather requires a correspondence between the environmental trend and the physiological response in terms of direction, amplitude, and time.
[0055] The extraction of abnormal physiological fluctuation data is accomplished using a combination of individual baseline fluctuation thresholds, continuous window conditions, and recovery conditions. The system examines heart rate characteristics, skin temperature, and sweat secretion item by item. If a certain indicator is higher than the corresponding individual baseline fluctuation threshold in two out of three consecutive diagnostic windows, that indicator is extracted as abnormally elevated data. If a certain indicator increases progressively within three consecutive diagnostic windows, and the third window is higher than the individual baseline fluctuation threshold, that indicator is extracted as continuously deteriorating data. If the overall environmental heat intensity data has decreased to below the heat intensity threshold, but a certain physiological indicator remains higher than the individual baseline fluctuation threshold in the subsequent three diagnostic windows, that indicator is extracted as recovering lag data. Abnormally elevated data, continuously deteriorating data, and recovering lag data collectively constitute abnormal physiological fluctuation data.
[0056] A baseline of similar populations is also introduced to identify individual differences. These populations are categorized by age group, work intensity level, type of equipment worn, and monitoring scenario. For each population group, the system calculates the population mean and standard deviation of heart rate characteristics, skin temperature, and sweat secretion at the same environmental heat intensity classification level based on historical samples. The population deviation threshold is calculated as the population mean plus twice the population standard deviation. During real-time operation, if an individual's heart rate characteristics, skin temperature, or sweat secretion exceed the population deviation threshold at the same environmental heat intensity classification level, the system extracts that indicator as population deviation data. This data is used to distinguish between "high environmental intensity" and "abnormal individual response to the same environment."
[0057] High-risk individual physical differences are determined based on abnormal physiological fluctuation data. The system statistically analyzes the types of abnormal indicators, the number of times abnormalities occur, the number of abnormal duration windows, and the recovery status of each individual within the current diagnostic period. If both heart rate and skin temperature show abnormal increases simultaneously, and the number of duration windows reaches three or more, it is extracted as a combined cardiac-thermal sensitivity feature. If sweat secretion first exceeds the individual's baseline fluctuation threshold and then decreases to below the individual's baseline average, while skin temperature remains consistently above the individual's baseline fluctuation threshold, it is extracted as a feature of insufficient heat dissipation response. If, after a decrease in overall environmental heat intensity, heart rate or skin temperature still fails to return to below the individual's baseline fluctuation threshold within three consecutive windows, it is extracted as a feature of delayed recovery. If, within the same environmental heat intensity classification level, any physiological indicator of this individual exceeds the population deviation threshold, it is extracted as a population deviation feature.
[0058] S3 includes acquiring physiological data and reducing the dimensionality of the physiological data to obtain the physical differences of high-risk individuals; extracting and removing abnormal feature values based on the physical differences of high-risk individuals to obtain a diversity feature vector; modeling the diversity feature vector to generate a personalized heatstroke symptom prediction model; using the personalized heatstroke symptom prediction model to output a prediction result matrix, and judging the probability distribution of potential symptoms based on the prediction result matrix.
[0059] In this embodiment, step S3 is used to further organize the high-risk individual-related data obtained in the aforementioned steps into feature inputs that can be incorporated into the prediction model. The system first acquires physiological data for the same monitored individual within a continuous diagnostic time window. This physiological data includes heart rate characteristic values, skin temperature values, sweat secretion values, heart rate change rate, skin temperature change rate, sweat secretion change rate, number of abnormal persistence windows, and recovery time. The heart rate change rate is obtained by subtracting the heart rate characteristic value of the previous window from the current window's heart rate characteristic value, and then dividing by the time interval between the two windows. The skin temperature change rate and sweat secretion change rate are calculated in the same way. The number of abnormal persistence windows is obtained by statistically analyzing the number of diagnostic time windows where the physiological indicators recover to within the individual's baseline fluctuation threshold after the environmental heat intensity drops below the heat intensity threshold.
[0060] Before dimensionality reduction, the system first performs personal baseline correction. The personal baseline is derived from the individual's historical monitoring data under non-high-temperature conditions. The system calculates the personal baseline averages for heart rate characteristics, skin temperature, and sweat secretion. During calculation, all valid values for each indicator in the historical non-high-temperature samples are summed and then divided by the number of valid samples. In real-time calculation, the system subtracts the personal heart rate baseline average from the current window's heart rate characteristic value to obtain the heart rate offset; subtracts the personal skin temperature baseline average from the current window's skin temperature value to obtain the skin temperature offset; and subtracts the personal sweat secretion baseline average from the current window's sweat secretion value to obtain the sweat secretion offset. Through this processing, the system obtains the degree of physiological deviation of the individual relative to their normal state.
[0061] During dimensionality reduction, the system arranges physiological data from multiple consecutive diagnostic time windows into a physiological feature matrix in chronological order. Each row of the matrix corresponds to a diagnostic time window, and each column corresponds to a physiological feature. The system first standardizes the data in each column. The standardization process is as follows: subtract the mean of the feature in historical calibration samples from the current feature value, and then divide by the standard deviation of the feature in historical calibration samples. The standard deviation is calculated by first calculating the difference between each historical sample and the mean, then squaring the differences and summing them, then dividing by the number of historical samples minus 1, and finally taking the square root of the result. After standardization, the system calculates the co-variation relationship between physiological features. This is done by subtracting the mean of the corresponding column from the standardized values of any two features within the same window, multiplying the products over all effective windows, summing the products, and then dividing by the number of effective windows minus 1 to obtain the co-variation value between the two features.
[0062] The system generates projected weighted reassemblies based on the synergistic changes among various features and ranks them according to their contribution to the overall data change. The system selects the top three projected weights with the highest contribution and weights them for the standardized heart rate feature, skin temperature, sweat secretion, rate of change, number of abnormal duration windows, and recovery time, respectively, resulting in three dimensionality-reduced components. The first dimensionality-reduced component represents the overall physiological load, the second represents the change in heat dissipation response, and the third represents the recovery lag state. These three dimensionality-reduced components are combined with the cardiac-thermal joint sensitivity feature, insufficient heat dissipation response feature, recovery lag feature, and population deviation feature obtained in claim 3 to form the high-risk individual physical difference features.
[0063] Subsequently, the system extracts abnormal feature values based on the physical differences of high-risk individuals. These abnormal feature values refer to invalid values caused by sensor detachment, poor contact, communication interruption, or momentary impact; they are not genuine abnormal physiological responses used to identify risk. The judgment boundary for abnormal feature values is determined through historical valid samples. The system arranges the historical valid samples of a particular feature in ascending order of value, taking the 25th and 75th percentiles; then, it subtracts the 25th percentile from the 75th percentile to obtain the interquartile range; the lower boundary is the 25th percentile minus 1.5 times the interquartile range, and the upper boundary is the 75th percentile plus 1.5 times the interquartile range. When the current feature value is below the lower boundary or above the upper boundary, it is marked as a candidate abnormal feature value.
[0064] The system does not directly delete all candidate abnormal feature values, but continues to check their temporal continuity. If a candidate abnormal feature value appears only within one diagnostic time window, and the same type of feature in the preceding and following windows is within the judgment boundary, then the candidate abnormal feature value is considered an isolated interference value and is removed. If the same feature exceeds the judgment boundary for three consecutive diagnostic time windows, and the direction of change is consistent with the direction of change in the overall environmental heat intensity, then the feature value is considered a true physiological abnormal response and is retained. Removed feature values are compensated using adjacent valid feature values: the system reads the most recent valid value before and after the removal window, and assigns weights according to their temporal distance from the removal window; the closer the distance, the higher the proportion of compensation. After completing the removal and compensation, the system obtains a diversity difference feature vector.
[0065] The diversity feature vector is composed in a fixed order, including the overall physiological load component, heat dissipation response component, recovery hysteresis component, cardiothermal joint sensitivity marker, heat dissipation insufficiency marker, recovery hysteresis marker, population deviation marker, number of abnormal duration windows, maximum physiological deviation amplitude, mean recovery time, and response delay time after environmental changes. Labeled features are represented by 0 and 1, with 0 indicating no corresponding feature and 1 indicating its presence. Continuous numerical features are represented using standardized numerical values. One diversity feature vector is generated for each diagnostic time window, and multiple consecutive diagnostic time windows form the model input sequence.
[0066] The personalized heatstroke symptom prediction model is trained using diverse feature vectors. Training samples include historical feature vector sequences and manually reviewed symptom labels. Symptom labels include aggravated heat stress, abnormally high body temperature, persistent abnormal heart rate, decreased heat dissipation capacity, delayed recovery, and early risk of heatstroke. During training, the system inputs feature vectors from multiple consecutive diagnostic time windows into a temporal neural network, and the model outputs predicted values for each symptom category. The system compares the predicted values with the manually reviewed labels and adjusts the model parameters based on the error. Training continues until the validation sample error reaches a convergence condition. The convergence condition is determined through historical training records: the system statistically analyzes the error decrease when the validation error reaches a stable stage during past training, and uses the average error decrease over 10 consecutive training rounds not exceeding this stable decrease as the convergence threshold. Once the model reaches this convergence threshold, training stops, and the model parameters are saved.
[0067] When generating a personalized model, the system incorporates the individual's historical verification data into the general training samples. The weight of the individual samples is higher than that of the general samples, ensuring the model output closely reflects the individual's physiological responses. The general samples continue to participate in training to maintain the model's ability to recognize different heat exposure scenarios. The updated model needs to be validated using the individual's historical verification samples. If the number of missed early-stage heatstroke risks does not increase after the update, and the overall number of correct predictions increases, the updated personalized heatstroke symptom prediction model is saved; otherwise, the previous model continues to be used.
[0068] During model execution, the system outputs a prediction result matrix using a personalized heatstroke symptom prediction model. The rows of the prediction result matrix correspond to potential symptom categories, and the columns correspond to future diagnosis time windows. When the prediction range is three future diagnosis time windows, the prediction result matrix consists of 6 rows and 3 columns. The 6 rows correspond to aggravated heat stress, abnormally high body temperature, persistent abnormal heart rate, decreased heat dissipation capacity, delayed recovery, and early risk of heatstroke, respectively. The 3 columns correspond to the first, second, and third future diagnosis time windows, respectively. Each value in the matrix represents the probability of the corresponding symptom occurring within the corresponding future window.
[0069] The probability distribution calculation process is as follows: The model first outputs the raw score of each symptom category within each future window; the system takes the maximum value among all raw scores of all symptoms within the same future window as the baseline value; then, the raw score of each symptom is subtracted from the baseline value, and an exponential transformation is performed; finally, the exponential transformation result of a certain symptom is divided by the sum of the exponential transformation results of all symptoms within the window to obtain the probability of the symptom occurring within that future window. Through this process, the sum of the probabilities of all symptom categories within the same future window is 1. The system reads the probability change of the same symptom in the next three windows row by row, and reads the probability proportion of each symptom within the same future window column by column, thus obtaining the probability distribution of potential symptoms.
[0070] S4 includes inputting the first real-time vital signs sequence into the personalized heatstroke symptom prediction model and extracting the first symptom probability distribution; if the first symptom probability distribution is higher than the warning threshold, activating the dynamic adjustment mechanism for the first real-time vital signs sequence to obtain the second feature weight allocation matrix; performing calculations on the first symptom probability distribution according to the second feature weight allocation matrix to obtain the second symptom probability distribution; classifying the second symptom probability distribution to determine the risk level, and generating an optimized diagnostic result sequence based on the risk level.
[0071] In this embodiment, the first real-time vital sign sequence consists of vital sign data within a continuous diagnostic time window, including heart rate characteristic values, skin temperature values, sweat secretion values, heart rate change rate, skin temperature change rate, sweat secretion change rate, number of abnormal persistence windows, and recovery time. The system organizes the data in units of diagnostic time windows, with each window forming a set of vital sign vectors. Multiple consecutive windows are arranged according to the acquisition time sequence to obtain the first real-time vital sign sequence. The heart rate change rate is obtained by subtracting the heart rate characteristic value of the previous window from the current window's heart rate characteristic value, and then dividing by the time interval between the two windows; the skin temperature change rate and sweat secretion change rate are obtained in the same way. The number of abnormal persistence windows is obtained by counting the number of windows that continuously exceed the individual's baseline fluctuation threshold; the recovery time is obtained by counting the number of windows during which the vital sign data returns to within the individual's baseline fluctuation threshold after the environmental heat intensity drops below the heat intensity threshold.
[0072] Before the first real-time vital signs sequence is input into the personalized heatstroke symptom prediction model, the system performs data scaling. For heart rate characteristics, skin temperature, and sweat secretion, the system reads the individual's historical data under non-high-temperature conditions and calculates the individual baseline mean and individual baseline standard deviation. The individual baseline mean is obtained by adding the historical valid values and dividing by the number of valid samples; the individual baseline standard deviation is obtained by calculating the difference between each historical valid value and the individual baseline mean, summing the squared differences, dividing by the number of valid samples minus 1, and then taking the square root. Before the real-time vital signs data enters the model, the current value is subtracted from the corresponding individual baseline mean and then divided by the corresponding individual baseline standard deviation to obtain standardized vital signs data. After this processing, the first real-time vital signs sequence is input into the personalized heatstroke symptom prediction model, and the model outputs a first symptom probability distribution. This first symptom probability distribution is arranged according to symptom category and future diagnosis time window. Symptom categories include aggravated heat stress, abnormally high body temperature, persistent abnormal heart rate, decreased heat dissipation capacity, delayed recovery, and early risk of heatstroke.
[0073] The warning threshold is determined by historical validation samples. These samples include symptom probabilities output by the model and risk statuses after manual review. The system sets candidate warning thresholds at intervals of 0.01 between 0 and 1, validating each one individually. For each candidate warning threshold, the system counts four types of quantities: the number of actual risks with a probability higher than the candidate value, the number of actual risks with a probability lower than the candidate value, the number of actual risks without a probability higher than the candidate value, and the number of actual risks without a probability higher than the candidate value. Subsequently, the system calculates sensitivity and specificity. Sensitivity equals the number of actual risks correctly identified divided by the total number of actual risks; specificity equals the number of actual risks without a probability correctly identified divided by the total number of actual risks without a probability. The system adds sensitivity and specificity and subtracts 1 to obtain the comprehensive discriminant value of the candidate warning threshold. The candidate value with the largest comprehensive discriminant value is determined as the warning threshold; if multiple candidate values have the same comprehensive discriminant value, the candidate value with the smallest value is selected, allowing early risks to enter a dynamic adjustment process.
[0074] During real-time operation, the system reads the probability distribution of the first symptom. If the probability of early heatstroke risk is higher than the warning threshold, or if the probability of two or more of the following—whether it is aggravated heat stress, abnormally high body temperature, persistent abnormal heart rate, decreased heat dissipation capacity, or delayed recovery—is higher than the warning threshold, then the dynamic adjustment mechanism is activated. The dynamic adjustment mechanism calculates a second feature weight allocation matrix for the first real-time vital sign sequence. The rows of this matrix correspond to symptom categories, and the columns correspond to vital sign features. The values in the matrix represent the contribution of a particular vital sign feature to the probability adjustment of a specific symptom category.
[0075] The second feature weighting matrix is derived from data reliability, risk contribution, and trend enhancement. Data reliability is calculated using data integrity rate and data stability. Data integrity rate equals the number of valid sampling points in the current window divided by the theoretical number of sampling points. Data stability is determined based on the jumps between adjacent windows. If the difference between adjacent windows does not exceed the maximum allowable change in the device calibration file, the stability is set to the maximum value. If it exceeds the maximum allowable change but is compensated for by adjacent valid data, the stability is reduced. If missing data cannot be compensated for, the stability of this feature in the current window is set to the minimum value. The system multiplies the data integrity rate by the data stability to obtain the data reliability of this feature.
[0076] Risk contribution is calculated using historical validation samples. During the validation phase, the system sequentially masks one vital sign feature, reruns the model after masking, and records the increase in error for early heatstroke risk prediction. The greater the error increase after masking a feature, the greater its contribution to risk assessment. The system sums the error increases for all features and then divides the error increase for a single feature by the total increase to obtain the risk contribution of that feature. Trend enhancement is determined based on changes in vital signs within three consecutive diagnostic windows. If a vital sign increases twice consecutively, and the value in the last window is higher than the individual baseline fluctuation threshold, then that vital sign receives a trend enhancement; otherwise, the baseline trend value is maintained. The individual baseline fluctuation threshold is obtained by adding twice the individual baseline standard deviation to the individual baseline mean.
[0077] During matrix generation, the system first assigns basic weights to each symptom category. For example, abnormally elevated body temperature corresponds to skin temperature values and the rate of change in skin temperature; persistent abnormal heart rate corresponds to heart rate characteristic values and the rate of change in heart rate; decreased heat dissipation capacity corresponds to the rate of change in sweat secretion and persistently elevated skin temperature; delayed recovery corresponds to recovery time and the number of abnormal duration windows; and early risk of heatstroke corresponds to a combination of the above multiple signs. Subsequently, the system multiplies the basic weights by data reliability, risk contribution, and trend enhancement, respectively, to obtain the adjusted weights for the current window. For the same symptom category, the system adds up all the adjusted weights in that row, and then divides each adjusted weight by the sum of the weights in that row, so that the sum of all feature weights corresponding to that symptom category is 1, thus obtaining the second feature weight allocation matrix.
[0078] When calculating the probability distribution of the first symptom based on the second feature weight allocation matrix, the system first calculates the probability adjustment coefficient for each symptom category. Specifically, the system reads the feature weights from the row corresponding to the symptom category, multiplies each feature weight by the standardized value of the corresponding feature in the first real-time vital sign sequence, and then sums the products to obtain the probability adjustment coefficient for that symptom category. The system then multiplies the probability value of the corresponding symptom in the first symptom probability distribution with this probability adjustment coefficient to obtain an intermediate probability value. To ensure a consistent scale among the probabilities of each symptom within the same future diagnostic window, the system sums the intermediate probability values of all symptoms within that future window and then divides each intermediate probability value by the sum to obtain the second symptom probability distribution. This second symptom probability distribution simultaneously reflects the model's original prediction results, real-time vital sign quality, vital sign risk contribution, and short-term trend.
[0079] The risk level is determined based on the probability distribution of the second symptom. The system first calculates the comprehensive risk value for each future diagnosis time window. During calculation, the system multiplies the probability of each symptom by its corresponding weight, and then sums the products. Symptom weights are determined from historical review samples: the system counts the number of times each type of symptom occurs in high-risk samples, and then divides the number of occurrences of a particular symptom by the sum of the number of occurrences of all symptoms to obtain the symptom weight. Early-stage heatstroke risk is a direct target symptom, and its weight is no less than the weight of any other individual symptom.
[0080] The risk level threshold is also determined by historical review samples. The system maps the comprehensive risk value to the manual review level, which includes low risk, medium risk, high risk, and dangerous risk. The system sets candidate cutoff points in intervals of 0.01 between 0 and 1, which are then combined to form the cutoff points between low and medium risk, medium and high risk, and high and dangerous risk. Each set of candidate cutoff points is verified on historical review samples. The system counts the number of correct level judgments and the number of missed high and dangerous risk judgments. The set of candidate cutoff points with the most correct level judgments and the fewest missed high and dangerous risk judgments is determined as the risk level threshold. During real-time judgment, if the comprehensive risk value is below the first cutoff point, low risk is output; if it reaches the first cutoff point but is below the second cutoff point, medium risk is output; if it reaches the second cutoff point but is below the third cutoff point, high risk is output; and if it reaches the third cutoff point or above, dangerous risk is output.
[0081] The optimized diagnostic result sequence consists of risk levels for consecutive future diagnostic time windows. Each diagnostic result includes a window number, a second symptom probability distribution, a comprehensive risk value, a risk level, and the main symptom category that triggered the risk level. The system arranges these diagnostic results in future window order to obtain the optimized diagnostic result sequence. If the risk level of a single window increases but the risk levels of the windows before and after it do not increase, and the data reliability of that window is lower than a preset reliability threshold, the system marks that window as a result to be confirmed. If the risk level increases for two consecutive windows, the system retains the increased risk level. The preset reliability threshold is determined through historical validation samples. The system tests candidate values between 0 and 1 at 0.01 intervals and selects the candidate value that minimizes the number of false triggers due to noise and does not increase the number of missed high-risk diagnoses as the preset reliability threshold.
[0082] S5 includes acquiring a first diagnostic result sequence, calculating a sequence reliability metric based on the first diagnostic result sequence, determining whether the sequence reliability metric is restricted, including identifying a restricted status identifier if the sequence reliability metric is lower than a preset threshold, obtaining a correction weight matrix by weighting historical diagnostic data based on the restricted status identifier, obtaining a fusion feature set through the correction weight matrix, and determining the final early diagnosis output for heatstroke based on the fusion feature set.
[0083] In this embodiment, the first diagnostic result sequence consists of continuous diagnostic results generated in the aforementioned steps. The system reads the risk level, symptom probability distribution, comprehensive risk value, main triggering symptom category, data completeness, and model output confidence information within multiple diagnostic time windows in chronological order to form the first diagnostic result sequence. Each diagnostic time window corresponds to one diagnostic result. Continuous diagnostic results are used to determine whether the current output is stable, whether it is affected by missing data, and whether historical diagnostic data needs to be called for correction.
[0084] The sequence reliability metric is calculated using four results: data integrity, probabilistic stability, risk level consistency, and model confidence. Data integrity is calculated as follows: the number of windows in the first diagnostic result sequence containing complete sign data and complete symptom probability distributions is counted, and then divided by the total number of windows that the sequence should include to obtain the data integrity value. Probabilistic stability is calculated as follows: the system compares the probability values of the same symptom category in each of two adjacent diagnostic time windows, subtracts the probability value of the previous window from the probability value of the later window, and takes the absolute value of the difference; then, the absolute values of the differences for all symptom categories and all adjacent windows are summed, and divided by the number of differences compared to obtain the average probability fluctuation value; since the symptom probability itself is between 0 and 1, the system subtracts this average probability fluctuation value from 1 to obtain the probabilistic stability value. The larger the average probability fluctuation, the lower the probabilistic stability.
[0085] Risk level consistency is calculated based on risk level jumps. The system first assigns low risk, medium risk, high risk, and dangerous risk as 1, 2, 3, and 4 respectively. Then, it compares the risk level values of two adjacent windows. If the risk level value changes, it is recorded as one jump; otherwise, it is not recorded as a jump. The system divides the number of jumps by the number of comparisons between adjacent windows to obtain the jump ratio, and then subtracts the jump ratio from 1 to obtain the risk level consistency value. Model confidence is obtained through symptom probability distribution. The system reads the highest symptom probability value in each diagnosis time window, then adds the highest symptom probability values from multiple consecutive windows and divides them by the number of windows to obtain the model confidence value.
[0086] All four values are between 0 and 1. The system assigns weights to data integrity, probabilistic stability, risk level consistency, and model confidence, and calculates a sequence reliability metric. The weights are determined using historical validation samples before system deployment. Specifically, the system first establishes a historical diagnostic result sequence sample containing manually reviewed conclusions. Then, it removes one of the following: data integrity, probabilistic stability, risk level consistency, and model confidence, respectively, and observes the decrease in the number of correct diagnoses. The greater the decrease in the number of correct diagnoses after removing a particular item, the greater its contribution to reliability. The system adds up the contributions of all four items and then divides each item's contribution by the total contribution to obtain its corresponding weight. In real-time calculation, the system multiplies each of the four reliability values by its corresponding weight and then adds the four products to obtain the sequence reliability metric.
[0087] The preset reliability threshold is determined through historical verification samples. Each sample in the historical verification samples includes a sequence reliability metric and a manually verified sequence status, categorized as reliable or restricted. The system sets values between 0 and 1 as candidate thresholds at 0.01 intervals and verifies them one by one. For each candidate threshold, the system counts the number of samples that are actually restricted but correctly identified as restricted, the number that are actually restricted but correctly identified as reliable, the number that are actually reliable but correctly identified as reliable, and the number that are actually reliable but correctly identified as restricted. Subsequently, the system divides the number of samples that are actually restricted but correctly identified as restricted by the total number of samples that are actually restricted to obtain the restricted identification rate; and divides the number of samples that are actually reliable but correctly identified as reliable by the total number of samples that are actually reliable to obtain the reliable identification rate. The system adds the restricted identification rate and the reliable identification rate and subtracts 1 to obtain the comprehensive discrimination value for that candidate threshold. The candidate threshold with the largest comprehensive discrimination value is determined as the preset reliability threshold; when multiple candidate thresholds have the same comprehensive discrimination value, the candidate threshold with the fewest restricted missed detections is selected.
[0088] During real-time assessment, the system compares the sequence reliability metric of the current first diagnostic result sequence with a preset reliability threshold. When the sequence reliability metric is not lower than the threshold, the system considers the current sequence to be in a reliable state, and the first diagnostic result sequence directly enters the final output confirmation process. When the sequence reliability metric is lower than the threshold, the system determines a restricted state identifier. The restricted state identifier is further generated based on four reliability sub-items: when data integrity is lower than its sub-item threshold, writing data loss is restricted; when probabilistic stability is lower than its sub-item threshold, writing probability fluctuation is restricted; when risk level consistency is lower than its sub-item threshold, writing level jumps are restricted; when model confidence is lower than its sub-item threshold, writing confidence is insufficient and restricted. The above sub-item thresholds are determined using the same historical verification method as the reliability threshold.
[0089] After determining the restricted status indicator, the system retrieves historical diagnostic data. This historical diagnostic data includes data from the same individual, similar population groups, and similar environmental scenarios. Each historical diagnostic data entry includes historical environmental heat intensity data, historical vital sign data, historical symptom probability distribution, historical risk level, historical intervention results, and manual review conclusions. The system prioritizes historical data from the same individual with the same environmental heat intensity level; if the data is insufficient, it supplements it with historical data from the same age group, work intensity, equipment worn, and environmental scenario.
[0090] The calibration weight matrix is generated based on the similarity between historical diagnostic data and the current first diagnostic result sequence. The system calculates environmental similarity, physiological similarity, trend similarity, and conclusion confidence. The calculation process for environmental similarity is as follows: read the current environmental heat intensity data and historical environmental heat intensity data, calculate the absolute value of the difference between the two, and then divide it by the standard deviation of the historical environmental heat intensity data to obtain the standardized environmental difference; then divide by 1 and add the standardized environmental difference to obtain the environmental similarity. Physiological similarity is calculated in the same way. The data involved in the comparison include heart rate characteristics, skin temperature, sweat secretion, recovery time, and the number of abnormal persistence windows; the system calculates the standardized differences for each item, then averages them, and divides by 1 and adds the average difference to obtain the physiological similarity.
[0091] Trend similarity is determined by the direction of risk level changes. The system reads the direction of risk level changes in consecutive windows in the current sequence and the direction of risk level changes in windows of the same length in historical sequences; when both are increasing, both are decreasing, or both are stable, they are considered to be in the same direction; the number of times the direction is consistent is divided by the total number of comparisons to obtain the trend similarity. The reliability of the conclusion is determined based on the historical data source. Data that has been manually reviewed and has complete results after intervention is given the highest reliability; data that has been manually reviewed but has incomplete intervention results has a lower reliability; data that has not been manually reviewed is not included in the correction weight matrix.
[0092] The system multiplies environmental similarity, physiological similarity, trend similarity, and conclusion reliability by their respective contribution weights, then sums the products to obtain the initial correction weight for each historical diagnostic data point. The contribution weight is determined using historical validation samples by sequentially removing environmental similarity, physiological similarity, trend similarity, or conclusion reliability, and then statistically analyzing the decrease in the number of correct diagnoses after correction; the greater the decrease, the higher the corresponding contribution weight. After calculating the initial correction weights, the system sums the initial correction weights of all selected historical data points, then divides the initial correction weight of each historical data point by this sum, ensuring that the sum of the weights of all historical data points equals 1. The normalized weights are arranged according to "historical data entry" and "feature category to be corrected," forming a correction weight matrix.
[0093] When obtaining the fusion feature set through the calibration weight matrix, the system first extracts current features from the current first diagnostic result sequence, including the current symptom probability distribution, current comprehensive risk value, current risk level, current main triggering symptom, and current reliability component. Then, the system reads the corresponding historical features from the selected historical diagnostic data. For each feature category, the system multiplies that feature from each historical data point by the corresponding weight in the calibration weight matrix, and then sums the products to obtain the historical weighted result for that feature category. Next, the system fuses the current features with the historical weighted results. When the current sequence is limited by data missing information, the proportion of the historical weighted result in the fusion is increased; when it is limited by probability fluctuations, the proportion of the historical symptom probability change trend in the fusion is increased; when it is limited by level jumps, the proportion of the historical risk level stability result in the fusion is increased; when it is limited by insufficient confidence, the proportion of the feature corresponding to the manual review conclusion in the fusion is increased. After fusion, a fusion feature set is obtained, which includes the fusion symptom probability distribution, fusion comprehensive risk value, fusion risk level tendency, fusion main triggering symptom, and fusion reliability result.
[0094] The final early diagnosis output for heatstroke is determined based on the fusion feature set. The system first calculates the fusion comprehensive risk value based on the probability distribution of fusion symptoms. During calculation, the fusion probability of each symptom type is multiplied by its corresponding weight, and all products are summed. Symptom weights are determined from historical review samples: the system counts the number of times each symptom type appears in the high-risk samples, then divides the number of occurrences of a particular symptom type by the sum of the occurrences of all symptoms to obtain its weight; early heatstroke risk is a direct target symptom, and its weight is no less than that of other individual symptoms. Subsequently, the system compares the fusion comprehensive risk value with the risk level threshold to determine the final risk level. The risk level threshold is determined through historical review samples. The system tests candidate boundary point combinations between 0 and 1 at 0.01 intervals and selects the boundary point combination with the highest number of correct level judgments and the fewest missed high-risk judgments as the risk level threshold.
[0095] The final early diagnosis output for heatstroke includes the final risk level, the fusion composite risk value, the main triggering symptom category, the restricted status indicator, the number of historical data points involved in correction, the historical correction weighting results, and the output time. When the final risk level is high risk or dangerous risk, the system outputs a timely intervention indicator; when the final risk level is medium risk and the fusion composite risk value increases for two consecutive diagnostic cycles, the system outputs a continuous monitoring indicator. Thus, when the reliability of the first diagnostic result sequence is insufficient, the system performs weighted correction using historical diagnostic data and generates a stable final early diagnosis output for heatstroke.
[0096] S6 includes acquiring physiological characteristic data and heat stress status to generate an early diagnosis output for heatstroke; generating an intervention signal sequence based on the early diagnosis output for heatstroke; constructing a feedback loop if the intervention signal sequence indicates a need for timely intervention; acquiring environmental heat intensity and individual body constitution to obtain thermoregulation ability in the feedback loop; constructing a risk assessment matrix based on thermoregulation ability to determine the interaction between environmental heat intensity and individual body constitution in the feedback loop.
[0097] In this implementation, physiological characteristic data and heat stress status within the current diagnostic period are first acquired. Physiological characteristic data includes heart rate values, skin temperature values, sweat secretion values, heart rate change rate, skin temperature change rate, sweat secretion change rate, number of abnormal duration windows, and recovery time. Heat stress status includes the fused comprehensive risk value, risk level, main triggering symptom category, and restricted status identifier obtained in the preceding steps. The system binds the above data in units of diagnostic time windows, generating one diagnostic input record for each diagnostic time window.
[0098] The early diagnosis output for heatstroke is generated jointly from physiological characteristic data and heat stress status. The system first reads the integrated risk value and risk level, and then reads the main triggering symptom categories. If the risk level is high risk or dangerous risk, an early heatstroke risk identifier is written into the diagnostic output; if the risk level is medium risk, and the integrated risk value increases for two consecutive diagnostic cycles, a continuous observation identifier is written into the diagnostic output; if the risk level is low risk, a routine monitoring identifier is written into the diagnostic output. The diagnostic output also includes the monitored individual's ID, environmental heat intensity data, abnormal physiological indicator names, risk trigger time, and restricted status identifier, so that subsequent intervention signals can be mapped to specific individuals, specific times, and specific triggering causes.
[0099] When generating an intervention signal sequence based on the early diagnosis output of heatstroke, the system converts the risk level into an intervention signal level. Low risk corresponds to a routine recording signal, medium risk to an enhanced monitoring signal, high risk to an on-site alert signal and a management terminal warning signal, and dangerous risk to an immediate intervention signal, a cooling equipment activation signal, and a medical assistance signal. Each intervention signal includes the signal level, trigger time, trigger reason, target personnel number, current risk level, and treatment action. Intervention signals from multiple consecutive diagnostic cycles are arranged chronologically to form an intervention signal sequence.
[0100] Whether an intervention signal sequence indicates a need for timely intervention is determined by an intervention trigger threshold. This threshold is determined from historical review samples. These samples include a fusion of comprehensive risk values, risk levels, manually confirmed intervention needs, and post-intervention results. The system sets candidate intervention thresholds at 0.01 intervals and verifies each one. For each candidate threshold, the system counts the number of samples that actually required intervention and were correctly triggered, the number of samples that actually required intervention but were not triggered, the number of samples that did not require intervention and were not triggered, and the number of samples that did not require intervention but were triggered. The system then divides the number of samples that actually required intervention and were correctly triggered by the total number of samples that actually required intervention to obtain the intervention recognition rate; and divides the number of samples that did not require intervention and were not triggered by the total number of samples that did not require intervention to obtain the non-intervention recognition rate. The system adds the intervention recognition rate and the non-intervention recognition rate and subtracts 1 to obtain the comprehensive discrimination value for the candidate intervention threshold. The candidate intervention threshold with the highest comprehensive discriminant value is determined as the intervention trigger threshold; if multiple candidate intervention thresholds have the same comprehensive discriminant value, the candidate intervention threshold with the fewest actual number of samples that need intervention but have not been triggered is selected.
[0101] During real-time operation, if a high-risk or dangerous risk signal appears in several pre-signal sequences, or if the fused comprehensive risk value exceeds the intervention trigger threshold, the system determines that timely intervention is needed and constructs a feedback loop. The feedback loop includes a pre-intervention data node, an intervention signal node, a post-intervention data node, and a next-cycle diagnostic node. The pre-intervention data node records environmental thermal intensity and physiological characteristic data within one diagnostic time window before the intervention trigger; the intervention signal node records the intervention type and triggering reason; the post-intervention data node records environmental thermal intensity and physiological characteristic data within multiple consecutive diagnostic time windows after the intervention; and the next-cycle diagnostic node records the regenerated risk level and fused comprehensive risk value after the intervention. These nodes are connected in chronological order to form a closed-loop data chain.
[0102] In the feedback loop, the system acquires environmental heat intensity and individual physical condition, and calculates thermoregulation capacity accordingly. Environmental heat intensity includes comprehensive environmental heat intensity data and its changes. The change in comprehensive environmental heat intensity is obtained by subtracting the comprehensive environmental heat intensity data after intervention from the comprehensive environmental heat intensity data before intervention; a result greater than 0 indicates a decrease in environmental heat load. Individual physical condition includes individual baseline heart rate, individual baseline skin temperature, individual baseline sweat secretion, number of abnormal duration windows, recovery time, and physical condition differences among high-risk individuals. Individual baseline data is derived from historical data under non-high-temperature conditions. The individual baseline mean is obtained by summing historical valid values and dividing by the number of valid samples. The individual baseline standard deviation is obtained by calculating the difference between each historical valid value and the individual baseline mean, squared the differences, summed them, divided by the number of valid samples minus 1, and then taking the square root. The individual baseline fluctuation threshold is the individual baseline mean plus twice the individual baseline standard deviation.
[0103] The calculation of thermoregulation capacity is divided into physiological recovery amount, recovery time correction, and environmental heat intensity correction. Physiological recovery amount is calculated separately for heart rate, skin temperature, and sweat secretion. Taking heart rate as an example, the system first calculates the portion of the pre-intervention heart rate characteristic value exceeding the individual's baseline fluctuation threshold, then calculates the portion of the post-intervention heart rate characteristic value exceeding the individual's baseline fluctuation threshold, and subtracts the two to obtain the heart rate recovery amount. Skin temperature recovery amount and sweat secretion recovery amount are obtained in the same way. The system then divides each of the three recovery amounts by its corresponding individual baseline standard deviation to obtain a recovery result on a uniform scale, and adds the three recovery results to obtain the comprehensive physiological recovery result. Subsequently, the system divides the comprehensive physiological recovery result by the recovery time to obtain the recovery result per unit time. If the comprehensive environmental heat intensity decreases significantly but the recovery result per unit time remains low, the thermoregulation capacity is reduced; if the comprehensive environmental heat intensity does not decrease significantly but the recovery result per unit time increases, the thermoregulation capacity is increased.
[0104] Whether the ambient heat intensity has decreased significantly is determined by an environmental degradation threshold. This threshold is calculated from historical effective intervention samples. The system extracts the overall decrease in ambient heat intensity from each effective intervention sample and sorts them in ascending order. The total number of samples is then multiplied by 0.5, and the result is rounded up to obtain a target number. The overall decrease in ambient heat intensity corresponding to this target number is determined as the environmental degradation threshold. During real-time operation, if the current overall decrease in ambient heat intensity exceeds this threshold, a significant decrease in ambient heat intensity is determined.
[0105] When constructing a risk assessment matrix based on thermoregulation capacity, the system uses the environmental heat intensity level as the matrix rows and the individual's physical condition as the matrix columns. The environmental heat intensity level includes low heat load, medium heat load, high heat load, and dangerous heat load, assigned values of 1, 2, 3, and 4, respectively; the individual's physical condition includes normal regulation, delayed regulation, insufficient heat dissipation, and delayed recovery, assigned values of 1, 2, 3, and 4, respectively. The system then generates correction values based on thermoregulation capacity: 0 for thermoregulation capacity above the upper threshold; 1 for thermoregulation capacity between the lower and upper thresholds; and 2 for thermoregulation capacity below the lower threshold. The upper and lower thresholds of thermoregulation ability are determined by historical intervention samples. The system sorts the thermoregulation ability values in the historical samples from smallest to largest, multiplies the total number of samples by 0.33 and rounds up to obtain the lower threshold number, multiplies the total number of samples by 0.67 and rounds up to obtain the upper threshold number, and then reads the thermoregulation ability value of the corresponding number as the lower and upper thresholds of thermoregulation ability.
[0106] The risk assessment value of each unit in the risk assessment matrix is obtained by adding the environmental heat intensity level, the individual's physical condition, and the thermoregulation capacity correction value. The system divides the sum by the maximum allowed sum in the matrix to ensure the risk assessment value falls between 0 and 1. During real-time judgment, the system locates the matrix row based on the current environmental heat intensity level and the matrix column based on the current individual's physical condition, reads the risk assessment value of the corresponding unit, and combines this with the direction of change in the feedback loop to determine the interaction between environmental heat intensity and individual physical condition. If the environmental heat intensity level increases while the thermoregulation capacity decreases, it is determined that the environmental heat intensity has a reinforcing adverse effect on the individual's physical condition; if the thermoregulation capacity remains below the lower threshold after the environmental heat intensity decreases, it is determined that individual physical condition is the main factor contributing to the continued risk; if the environmental heat intensity decreases and the thermoregulation capacity increases above the upper threshold, it is determined that the intervention is effective in reducing the risk.
[0107] S7 includes acquiring physiological characteristics and environmental load from historical diagnostic feedback loops, extracting interaction data between physiological characteristics and environmental load in the time dimension; extracting time-domain features based on the interaction data to obtain an interaction feature set containing body temperature fluctuation sequences and heart rate variability sequences; using a support vector machine algorithm to reclassify the body temperature fluctuation sequences and heart rate variability sequences in the interaction feature set to obtain classification boundaries in multidimensional space; if the classification boundaries meet a preset convergence threshold, redundant data is removed based on the classification boundaries to obtain refined features; and the refined heat stress values are determined based on the refined features as initialization inputs for subsequent diagnostic cycles.
[0108] In this implementation, the system reads closed-loop recorded data from the historical diagnostic feedback loop. Data sources include pre-intervention data, post-intervention data, and data transmitted from the next diagnostic cycle. Physiological characteristics include skin temperature, heart rate, sweat secretion, skin temperature change rate, heart rate change rate, recovery time, and the number of abnormal duration windows. Environmental load includes comprehensive environmental heat intensity data, environmental temperature, environmental humidity, and environmental radiation. The system first sorts the above data by timestamp, then binds it to diagnostic time windows, establishing a correspondence between environmental load and physiological characteristics within the same time window.
[0109] When extracting interactive data over time, the system first calculates the change in environmental load. The calculation process is as follows: subtract the comprehensive environmental heat intensity data of the previous diagnostic time window from the comprehensive environmental heat intensity data of the current diagnostic time window to obtain the change in environmental load. Then, the system calculates the changes in skin temperature and heart rate separately, i.e., subtracting the skin temperature value of the previous window from the skin temperature value of the current window, and subtracting the heart rate characteristic value of the previous window from the heart rate characteristic value of the current window. If skin temperature or heart rate increases synchronously in subsequent windows after an increase in environmental load, it is recorded as a unidirectional response; if skin temperature or heart rate does not decrease within a preset recovery window after a decrease in environmental load, it is recorded as a hysteresis response; if the change in environmental load is at a low level but the physiological change exceeds the individual's baseline fluctuation threshold, it is recorded as a sensitive response. These response records, together with the corresponding time windows, form the interactive data.
[0110] The individual baseline fluctuation threshold is determined using historical data from non-hyperthermic states. The system first calculates the individual's baseline average for skin temperature and heart rate by summing historical valid values from non-hyperthermic states and then dividing by the number of valid samples. Next, the individual baseline standard deviation is calculated by first determining the difference between each historical valid value and the individual baseline average, then squaring these differences, summing them, dividing by the number of valid samples minus 1, and finally taking the square root. The individual baseline fluctuation threshold is the individual baseline average plus twice the individual baseline standard deviation. This threshold is used to determine whether skin temperature and heart rate exceed the individual's normal fluctuation range.
[0111] When extracting temporal features, the system uses multiple consecutive diagnostic time windows as one analysis segment. The body temperature fluctuation sequence consists of the offset of skin temperature relative to the individual's baseline. Within each window, the system subtracts the average baseline skin temperature from the current skin temperature value to obtain the body temperature offset. These offsets are then arranged chronologically to form the body temperature fluctuation sequence. From this sequence, the system extracts the body temperature fluctuation amplitude, the number of consecutive temperature rises, and the body temperature recovery time. The body temperature fluctuation amplitude is obtained by subtracting the lowest body temperature offset from the highest offset within the analysis segment; the number of consecutive temperature rises is obtained by comparing the body temperature offsets of adjacent windows; and the body temperature recovery time is obtained by counting the number of windows during which the body temperature offset exceeds the individual's baseline fluctuation threshold and returns to within that threshold.
[0112] Heart rate variability sequences are generated using the same method. In each window, the system subtracts the individual's baseline heart rate average from the current heart rate feature value to obtain the heart rate offset, which is then arranged chronologically to form the heart rate variability sequence. The system extracts heart rate fluctuation amplitude, number of consecutive heart rate rises, average heart rate change, and heart rate recovery time from the heart rate variability sequence. The average heart rate change is obtained by calculating the absolute value of the difference between the heart rate offsets of two adjacent windows, summing all the absolute values of the differences, and dividing by the number of differences. The body temperature fluctuation sequence, heart rate variability sequence, environmental load change, unidirectional response markers, recovery hysteresis response markers, and sensitive response markers together constitute the interactive feature set.
[0113] Before reclassification using Support Vector Machines (SVM), the system standardizes the interaction feature set. For continuous numerical features, the current value is subtracted from the average of historical labeled samples, and then divided by the standard deviation of historical labeled samples. Labeled features are represented by 0 and 1; 0 indicates no corresponding response, and 1 indicates a corresponding response. SVM training samples come from historical feedback loops. Each sample set includes the interaction feature set and manually verified heat stress states. Heat stress states are categorized as low heat stress, moderate heat stress, high heat stress, and dangerous heat stress. The system uses multiple binary classifiers to distinguish between the four states, with each class classifying two heat stress states. During training, the system searches for boundary parameters that maximize the distance between the two classes of samples on either side of the classification boundary. During real-time reclassification, the interaction feature set is sequentially fed into each binary classifier, and each classifier outputs its judgment result. The system counts the number of votes for each heat stress state, and the state with the highest number of votes is selected as the reclassification result; if the number of votes is the same, the state with the higher absolute value of the discrimination score is chosen.
[0114] The classification boundary in multidimensional space is determined by the boundary parameters obtained after training the support vector machine. When updating the classification boundary, the system records the boundary change between two adjacent training rounds. The calculation process for the boundary change is as follows: read the boundary parameters of the previous and subsequent rounds one by one, calculate the absolute value of the difference between the corresponding parameters, sum all the absolute values of the differences, and finally divide by the number of parameters. The preset convergence threshold is determined by historical training records. The system selects a training phase in historical training where the number of correctly classified validation samples does not increase for five consecutive rounds, and the number of missed classifications for high heat stress and dangerous heat stress does not increase, as the stable phase; calculate the average value of the boundary change in each round within this stable phase, and use this average value as the convergence threshold. During real-time updates, if the boundary change is lower than the convergence threshold for two consecutive rounds, and the number of missed classifications for high heat stress and dangerous heat stress in the validation samples does not increase, then the classification boundary is determined to meet the preset convergence threshold.
[0115] Once the classification boundary meets the convergence threshold, the system removes redundant data based on the classification boundary. Redundant data includes low-contribution features and repetitive features. The process for determining low-contribution features is as follows: the system temporarily removes one feature from the interaction feature set and re-executes support vector machine classification; if removing this feature does not decrease the number of correctly classified validation samples and does not increase the number of missed classifications for high heat stress and dangerous heat stress, then this feature is determined as a low-contribution feature and removed. The process for determining repetitive features is as follows: the system compares the direction of change of two features in historical samples, counts the number of times they simultaneously increase, simultaneously decrease, or simultaneously remain unchanged, and then divides by the total number of comparisons to obtain the repetition ratio. The preset repetition threshold is determined through historical validation samples. The system tests candidate values between 0 and 1 at intervals of 0.01 and selects candidate values that can reduce the number of features without increasing the number of high-risk missed classifications as the repetition threshold. When the repetition ratio is higher than the repetition threshold, the system retains the feature that contributes more to the classification boundary and removes the other feature.
[0116] After removing redundant data, the remaining body temperature fluctuation features, heart rate variability features, environmental load change features, and interaction response markers form refined features. The system determines the refined heat stress value based on the refined features. The specific process is as follows: read the contribution weight of each refined feature in the support vector machine classification boundary, multiply the standardized value of each refined feature by the corresponding contribution weight, and then add all the products to obtain the refined discrimination score. Subsequently, the system converts the refined discrimination score into a value in the range of 0 to 100 as the refined heat stress value; during the conversion, the lowest effective discrimination score in the historical samples corresponds to 0, the highest effective discrimination score corresponds to 100, and when the current discrimination score is between the two, it is converted according to its position between the lowest and highest effective discrimination scores.
[0117] The refined heat stress values are written into the initialization input of subsequent diagnostic cycles. When the next diagnostic cycle starts, the system simultaneously reads the original environmental data, original physiological data, refined characteristics, the interaction results of the previous feedback loop, and the refined heat stress values.
[0118] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for early diagnosis of heatstroke combining pattern classification and clustering, characterized in that: include: S1. Real-time collection of temperature, humidity and radiation data in high-temperature environments through sensor networks, combined with heart rate, skin temperature and sweat secretion indicators monitored by individual wearable devices, to form a comprehensive environmental heat intensity data and individual physiological response dataset, and to obtain preliminary heat stress indicators; S2. Based on the preliminary heat stress index, the dynamic changes in environmental heat intensity are classified using the support vector machine algorithm combined with cluster analysis. If the classification results show that the heat intensity exceeds the preset threshold, the trend of change is matched with the individual physiological response dataset to determine the physical differences of high-risk individuals. S3. Obtain the physical differences of high-risk individuals after identification, model and analyze the diversity differences through deep learning networks, generate a personalized heatstroke symptom prediction model, and determine the probability distribution of potential symptoms. S4. Extract the probability distribution from the personalized heatstroke symptom prediction model. For real-time data input in complex scenarios, if the probability distribution is higher than the warning threshold, activate the dynamic adjustment mechanism to obtain an optimized diagnostic result sequence. S5. Using the obtained optimized diagnostic result sequence, iteratively optimize the core contradiction. If the reliability of the sequence is limited after iteration, historical diagnostic data is incorporated for correction to determine the final early diagnosis output of heatstroke.
2. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that: S1 includes: The ambient temperature, ambient humidity, ambient radiation, heart rate characteristics, skin temperature, and sweat secretion values are acquired and aligned to obtain a synchronized multi-source data set. Construct a multidimensional feature matrix based on a synchronous multi-source dataset; Dimensionality reduction of the multidimensional feature matrix yields comprehensive environmental thermal intensity data and individual physiological response datasets; If the comprehensive environmental heat intensity data is greater than the preset heat intensity threshold, the comprehensive environmental heat intensity data and the individual physiological response dataset are classified and mapped to obtain preliminary heat stress values.
3. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that: S2 includes: Preliminary thermal stress data were obtained, and the preliminary thermal stress data were processed using the support vector machine algorithm combined with cluster analysis to obtain the environmental thermal intensity classification results. If the environmental heat intensity classification result exceeds the preset threshold, the environmental heat intensity classification result will be analyzed to determine the heat intensity change trend. The trend of heat intensity change is matched with individual physiological response datasets to obtain trend matching results; Abnormal physiological fluctuation data are extracted based on trend matching results, and the physical differences of high-risk individuals are determined based on the abnormal physiological fluctuation data.
4. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that: S3 includes: Acquire physiological data and reduce its dimensionality to obtain the physical differences of high-risk individuals; Based on the physical differences of high-risk individuals, abnormal feature values are extracted and removed to obtain a diversity difference feature vector; Model the feature vectors of diversity differences to generate a personalized heatstroke symptom prediction model; A personalized heatstroke symptom prediction model is used to output a prediction result matrix, and the probability distribution of potential symptoms is determined based on the prediction result matrix.
5. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that: S4 includes: Input the first real-time vital signs sequence into the personalized heatstroke symptom prediction model and extract the probability distribution of the first symptom. If the probability distribution of the first symptom is higher than the warning threshold, then the dynamic adjustment mechanism for the first real-time vital signs sequence is activated to obtain the second feature weight allocation matrix. The second symptom probability distribution is obtained by calculating the first symptom probability distribution based on the second feature weight allocation matrix; The probability distribution of the second symptom is classified to determine the risk level, and an optimized diagnostic result sequence is generated based on the risk level.
6. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that: S5 includes: Obtain the first diagnostic result sequence and calculate the sequence reliability measure based on the first diagnostic result sequence; Determine whether the sequence reliability metric is restricted, including identifying a restricted status indicator if the sequence reliability metric is below a preset threshold; The historical diagnostic data are weighted according to the restricted state identifier to obtain the correction weight matrix, and the fusion feature set is obtained through the correction weight matrix. The final early diagnosis output for heatstroke is determined based on the fusion feature set.
7. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 1, characterized in that, It also includes S6, which generates an intervention signal sequence through the final early diagnosis output of heatstroke. If the signal sequence indicates a need for timely intervention, a feedback loop is constructed based on the output to determine the interaction between environmental heat intensity and individual physical condition in the loop. Specifically, this includes: Generate early diagnosis output for heatstroke by acquiring physiological characteristic data and heat stress status; An intervention signal sequence is generated based on the early diagnosis of heatstroke.
8. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 7, characterized in that: S6 further includes: If the intervention signal sequence indicates a need for timely intervention, then a feedback loop is constructed; The feedback loop obtains information about the ambient heat intensity and the individual's physical condition to gain the body's thermoregulation ability. A risk assessment matrix is constructed based on the body's thermoregulation capacity to determine the interaction between environmental heat intensity and individual physical condition in the feedback loop.
9. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 7, characterized in that, It also includes S7, extracting interaction impact data from the feedback loop, reclassifying and processing it using the support vector machine algorithm to obtain refined heat stress indicators, which are used as initial inputs for subsequent diagnostic cycles. Specifically, this includes: Obtain physiological characteristics and environmental load from historical diagnostic feedback loops, and extract the interaction data between physiological characteristics and environmental load over time. Temporal features are extracted from the interactive data to obtain an interactive feature set containing body temperature fluctuation sequences and heart rate variation sequences.
10. The method for early diagnosis of heatstroke combining pattern classification and clustering according to claim 9, characterized in that: The S7 also includes: The support vector machine algorithm is used to reclassify the body temperature fluctuation sequence and heart rate variation sequence in the interaction feature set to obtain the classification boundary in multidimensional space. If the classification boundary meets the preset convergence threshold, redundant data is removed based on the classification boundary to obtain refined features. The heat stress value after refining is determined based on the refining characteristics and used as the initial input for subsequent diagnostic cycles.