Public health early warning system based on big data
By adopting big data processing and nonlinear model optimization techniques in the public health warning system, the threshold is dynamically adjusted, and the limitations of the existing system in handling nonlinear relationships and dynamic environmental changes are solved, and a more sensitive and accurate warning of public health events is achieved.
Patent Information
- Application Number
- CN202510085989.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing public health warning system has limitations in dealing with nonlinear relationships and dynamic environmental changes, resulting in insensitive or accurate responses and insufficient real-time performance, resulting in loss of time and waste of resources.
The public health warning system based on big data is adopted, and the regional health data aggregation module, nonlinear model optimization module, abnormal data analysis module and early warning signal release module are used to process nonlinear relationships in combination with the radial basis function core, and the threshold is dynamically adjusted to achieve accurate detection and early warning of abnormal data points in the health data flow.
It improves the response speed of public health events and the timeliness of preventive measures, reduces the possibility of false positives, ensures the effective operation of the system in various situations, and improves the response speed and accuracy of the model.
Smart Images

Figure CN119943403A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of health management, and in particular to a public health early warning system based on big data. Background Art
[0002] The field of health management technology involves the use of information technology, data analysis and communication technology to optimize the health status of individuals and groups. This field includes the collection, analysis and presentation of health data, systems that support clinical and public health decision-making through data, and the use of health data from various sources, including electronic health records, patient self-reported information, and data collected by real-time health monitoring devices. The goal is to achieve early diagnosis of diseases, continuous monitoring of treatment effects, and early warning of public health events through predictive models and data visualization tools. Health management technology also emphasizes the formulation of personalized medical plans, and improves patient compliance and overall health management efficiency through technical means.
[0003] Among them, the public health early warning system is a system that uses a wide range of data sources and advanced analytical techniques to monitor, predict and notify about public health threats. The system is designed to process health data from multiple channels, such as medical records, drug sales, disease reporting systems and environmental monitoring equipment. By analyzing the data, the system can identify health crises or outbreaks, and warn the public and health agencies in advance so that they can take preventive measures or prepare response measures. The main purpose of such systems is to improve the speed and efficiency of public health responses, reduce the spread and impact of diseases, and protect the health and safety of the community.
[0004] Although the existing public health early warning system covers the collection and analysis of health data, it has obvious limitations in dealing with nonlinear relationships and dynamic environmental changes. Traditional health management technology relies on linear models and static threshold settings, which results in the system's response being not sensitive or accurate enough when faced with a rapidly changing epidemic environment or atypical health data. Without the use of dynamic threshold adjustment, the peak of seasonal influenza is misjudged as an epidemic outbreak, causing unnecessary public panic and waste of resources. Existing technologies are also insufficient in real time, with slow processing and response speeds, which results in the loss of precious time in public health emergencies and thus affects the control and prevention of the disease. Summary of the invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a public health early warning system based on big data.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A public health early warning system based on big data includes:
[0007] The regional health data aggregation module extracts the population epidemic incidence rate and medical facility visit frequency data based on the national health database to generate regional health monitoring data, and based on the regional health monitoring data, evaluates the regional health status trend and obtains health trend indicators;
[0008] The nonlinear model optimization module selects a radial basis function kernel to process nonlinear relationships based on the health trend index, adjusts kernel function parameters and penalty parameters, obtains a health warning model, analyzes abnormal data points in the current health data stream based on the health warning model, and generates anomaly detection results;
[0009] The abnormal data analysis module analyzes abnormal points in the data based on the abnormal detection results, including sudden increases in hospital visit rates and abnormal drug sales, performs time series abnormal point analysis, compares them with seasonal and cyclical changes, and adjusts the threshold to obtain the adjusted threshold, performs abnormal data verification according to the adjusted threshold, and generates an abnormal verification result;
[0010] Based on the abnormal verification results, the early warning signal release module summarizes the historical data analysis and the current abnormal conditions, establishes an early warning signal, releases the early warning information through the operation interface, conveys the early warning information to the relevant health departments and the public, and generates real-time health warnings.
[0011] As a further solution of the present invention, the steps of acquiring the regional health monitoring data are specifically as follows:
[0012] Based on the national health database and extracting the epidemic records of the target area, the population-weighted epidemic incidence rate of the community is calculated in detail using the formula:
[0013]
[0014] The adjusted weighted incidence rate data of the epidemic is obtained, among which p i represents the population of the i-th region, e i represents the original epidemic incidence rate in the i-th region, w i represents the weight factor for regional health resource allocation, R a represents the adjusted epidemic-weighted incidence data;
[0015] Using the adjusted epidemic weighted incidence data, query the medical records of the corresponding medical facilities and calculate the weighted frequency of medical facilities using the formula:
[0016]
[0017] The adjusted visit frequency data is obtained, where v j is the number of visits to the jth medical facility, u j is the weight coefficient adjusted according to the severity of the epidemic, kj is the weight coefficient of the medical facility service quality, F a represents the adjusted visit frequency data;
[0018] The health risk index of the region is calculated by combining the data obtained from the adjusted visit frequency data and the adjusted epidemic weighted incidence data, using the formula:
[0019]
[0020] Get regional health monitoring data, where R a represents the adjusted epidemic weighted incidence data, F a represents the adjusted visit frequency data, H a Represents regional health monitoring data.
[0021] As a further solution of the present invention, the steps of obtaining the health trend indicator are specifically as follows:
[0022] Extract monthly health indicators from the regional health monitoring data, calculate moving averages and smooth seasonal fluctuations using the formula:
[0023]
[0024] Get the weighted monthly average health index, where H k represents the health monitoring data of the kth month, w k represents the activity weight of the kth month, H avg represents the weighted monthly average health index;
[0025] Using the weighted monthly average health index, the trend component is calculated through time series analysis using the formula:
[0026]
[0027] The enhanced trend component data is obtained, where H avg (t+i) is the monthly average health index at time t, d i is the distance weight based on time i, T b represents the enhanced trend component data;
[0028] Based on the enhanced trend component data, the exponential smoothing algorithm is used to predict the health index in the future time period, through the formula:
[0029]
[0030] Generate health trend indicators, where T brepresents the trend component data after strengthening, α is the smoothing coefficient, which controls the prediction reaction speed, Δt is the future time point, τ is the time delay parameter for adjusting the sensitivity of future prediction, and H b Represents the health index predicted for a future period of time.
[0031] As a further solution of the present invention, the steps of obtaining the health warning model are specifically as follows:
[0032] The health trend index is used as input data, and the radial basis function kernel is applied to reconstruct the core architecture of the model, through the formula:
[0033]
[0034] Generate kernel transformation data, where x, x′ represent differentiated data points, γ represents the shape of the control kernel function, β is the bias term added to the kernel function, and K c (x, x′) represents the adjusted kernel space data;
[0035] Based on the kernel conversion data, the adjustment method of the kernel function parameters and the penalty parameters is refined, through the formula:
[0036] C opt =argmin C (∑(y i -y pred,i ) 2 +λC 2 )
[0037] Generate optimized parameter settings, where C is the penalty parameter and y i ,y pred,i They refer to the real-time observation value and the model prediction value respectively, λ is used to adjust the importance of the penalty term, C opt represents the optimized penalty parameter;
[0038] Applying the optimized parameter settings, the formula is:
[0039]
[0040] Construct a health warning model, where M(x) is the health warning model, α i is the coefficient corresponding to the data point, y i is the target value of the data point, K c (x,x i ) is the kernel data after radial basis function processing, and θ is the threshold for adjusting the output result.
[0041] As a further solution of the present invention, the step of obtaining the abnormality detection result is specifically:
[0042] The health warning model is used to analyze each data point in the current health data stream, and the difference between the model's output of the data point and the predetermined threshold is calculated to create an innovative formula:
[0043] R(x)=M(x)-T+σ∥x∥ 2
[0044] Generate model response data points, where M(x) is the health warning model, T is the decision threshold, σ is the sensitivity adjustment coefficient added to the difference, and R(x) represents the adjusted model response;
[0045] The model response data points are evaluated and the adaptive threshold judgment logic is used to determine whether the data points are abnormal, using the formula:
[0046]
[0047] Generate the abnormal state of the data point, where R(x) represents the adjusted model response, δ is the dynamic threshold adjusted based on the statistical characteristics of the data, and D(x) represents the abnormal state of the data point;
[0048] Based on the abnormal status of the data points, the data points marked as abnormal are collected, and the formula is:
[0049] A d ={x|D(x)=1}
[0050] Construct anomaly detection results, where A d is the set of abnormal data points, and D(x) is the abnormal state of the data point.
[0051] As a further solution of the present invention, the step of obtaining the adjusted threshold value is specifically:
[0052] Data points are collected from the anomaly detection results, and the sudden increase in hospital visit rate and the abnormal data of drug sales are quantitatively analyzed, and the standard deviation of the abnormal data points is calculated, through the formula:
[0053]
[0054] Generate standard deviation data points where x i is the value of the outlier data point, is the mean of the data points, and Nf represents the standard deviation calculated from the abnormal data;
[0055] Using the standard deviation data points, combined with time series analysis, seasonal and cyclical adjustments are made to the abnormal data using the formula:
[0056]
[0057] Generate time series adjusted data, where Nf represents the standard deviation calculated from the abnormal data, α t ,ω,φ,β t are the amplitude, frequency, phase and amplitude of the cosine term of the periodic adjustment, respectively, and S(x) represents the data after seasonal and periodic adjustment;
[0058] According to the adjusted data of the time series, the threshold is readjusted and the authenticity of the abnormal data is reflected by the formula:
[0059] Vf=max(S(x))+κ
[0060] Generate adjusted thresholds, where S(x) represents the data after seasonal and periodic adjustment, κ is a small amount used to increase the sensitivity of the threshold, and Vf is the newly set threshold.
[0061] As a further solution of the present invention, the step of obtaining the abnormal verification result is specifically:
[0062] Using the adjusted threshold, the data points in the current data stream are evaluated, the difference between each data point and the threshold is calculated, and the outliers are preliminarily marked, using the formula:
[0063] R i (x) = xV f
[0064] Generates the difference result of preliminary anomaly labeling, where x is the value of the data point and V f is the adjusted threshold, R i (x) is the difference result of the preliminary abnormal marking;
[0065] A time series analysis is performed on the difference results of the preliminary anomaly markings to analyze whether the anomaly data points show deviations from the regular seasonal or cyclical patterns, using the formula:
[0066]
[0067] Generate time series analysis results, where R i (x) is the difference result of the preliminary abnormal marking, c j ,ω,φ j Adjusting the complexity and sensitivity of time series analysis, T a (x) represents the weighted time series anomaly index.
[0068] According to the time series analysis results, the threshold is readjusted and abnormal data is captured and determined, through the formula:
[0069] V n =V f +ζ·max(Ta (x)
[0070] Generate anomaly verification results, where x is the value of the data point, V f is the adjusted threshold, T a (x) represents the weighted time series anomaly index, V n Indicates an abnormal validation result.
[0071] As a further solution of the present invention, the steps of obtaining the real-time health warning are specifically as follows:
[0072] Summarize the data from the abnormal verification results, compare with the historical health data, and evaluate the current abnormal situation, calculate the deviation of the historical data, using the formula:
[0073]
[0074] Generate historical deviation analysis results, where x k Indicates the value of the current data point, h k represents the average value of the same period in history, α k represents the weight factor, and Pm represents the calculated historical deviation result;
[0075] Based on the historical deviation analysis results, combined with public health research, the warning level is determined using the formula:
[0076]
[0077] Generate warning level results, where Pm represents the calculated historical deviation result, θ is the threshold, and Lg is the quantitative warning level;
[0078] Using the warning level results, issue a warning through the operation interface using the formula:
[0079] Pg=ifLg>1then'Publish'else'Monitor'
[0080] Generate real-time health warnings, where Lg is the quantitative warning level and Pg represents the operational decision based on Lg.
[0081] Compared with the prior art, the advantages and positive effects of the present invention are:
[0082] In the present invention, by conducting an in-depth analysis of the national health database, key data such as the incidence of population epidemics and the frequency of visits to medical facilities are extracted to achieve comprehensive utilization of health data, thereby effectively monitoring and predicting regional health trends. The radial basis function kernel is used to process nonlinear relationships, optimize the detection accuracy of abnormal health events, improve the response speed and accuracy of the model, analyze the abnormal data points in the current health data stream, combine time series analysis with seasonal and cyclical changes, so that the early warning system can respond to environmental changes more sensitively, reduce the possibility of false alarms, dynamically adjust the threshold, further ensure the effective operation of the system under various circumstances, integrate historical data analysis with current abnormal conditions, and establish early warning signals that can promptly notify the public and health agencies, greatly improving the response speed of public health events and the timeliness of preventive measures. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 is a system flow chart of the present invention;
[0084] Figure 2 A flow chart of regional health monitoring data in the present invention;
[0085] Figure 3 A flow chart of the health trend indicator in the present invention;
[0086] Figure 4 It is a flow chart of the health early warning model in the present invention;
[0087] Figure 5 It is a flow chart of the abnormal detection results in the present invention;
[0088] Figure 6 A flow chart of the adjusted threshold value in the present invention;
[0089] Figure 7 It is a flow chart of abnormal verification results in the present invention;
[0090] Figure 8 This is a flow chart of the real-time health warning in the present invention. DETAILED DESCRIPTION
[0091] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0092] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.
[0093] Embodiment 1
[0094] See also Figure 1 , a public health early warning system based on big data includes:
[0095] The regional health data aggregation module extracts the population epidemic incidence rate and medical facility visit frequency data based on the national health database, generates regional health monitoring data, and evaluates regional health status trends and obtains health trend indicators based on the regional health monitoring data;
[0096] The nonlinear model optimization module selects the radial basis function kernel to process the nonlinear relationship based on the health trend index, adjusts the kernel function parameters and penalty parameters, obtains the health warning model, analyzes the abnormal data points in the current health data stream based on the health warning model, and generates anomaly detection results;
[0097] The abnormal data analysis module analyzes the abnormal points in the data based on the abnormal detection results, including the sudden increase in hospital visit rate and abnormal drug sales, performs time series abnormal point analysis, and compares it with seasonal and cyclical changes. It also adjusts the threshold and obtains the adjusted threshold. It verifies the abnormal data according to the adjusted threshold and generates the abnormal verification results.
[0098] The warning signal release module is based on the abnormal verification results, summarizes the historical data analysis and the current abnormal conditions, establishes the warning signal, releases the warning information through the operation interface, conveys the warning information to the relevant health departments and the public, and generates real-time health warnings.
[0099] Health trend indicators include the rate of change of the epidemic, the increase or decrease in the number of medical visits by the population, and the health risk rating. The health early warning model includes a nonlinear relationship analysis module, a warning threshold setting module, and a model verification component. The anomaly detection results include the type of abnormal data points detected, the impact range assessment, and the emergency response level. The adjusted thresholds include adjustment parameters based on seasonal factors, adjustment parameters based on cyclical changes, and dynamic anomaly judgment logic. The anomaly verification results include abnormal events that have passed verification, data false alarms that have failed verification, and time series anomaly comparison results. Real-time health warnings include regional warning levels, emergency announcement content, and a list of warning receiving units.
[0100] See also Figure 2 ,The specific steps for obtaining regional health monitoring data are:
[0101] Based on the national health database and extracting the epidemic records of the target area, the population-weighted epidemic incidence rate of the community is calculated in detail using the formula:
[0102]
[0103] The adjusted weighted incidence rate data of the epidemic is obtained, among which p i represents the population of the i-th region, e i represents the original epidemic incidence rate in the i-th region, w i represents the weight factor for regional health resource allocation, R a represents the adjusted epidemic-weighted incidence data;
[0104] Using the adjusted epidemic weighted incidence data, query the medical records of the corresponding medical facilities and calculate the weighted frequency of visits to the medical facilities using the formula:
[0105]
[0106] The adjusted visit frequency data is obtained, where v j is the number of visits to the jth medical facility, u j is the weight coefficient adjusted according to the severity of the epidemic, k j is the weight coefficient of the medical facility service quality, F a represents the adjusted visit frequency data;
[0107] The health risk index of the region is calculated by combining the data obtained from the adjusted visit frequency data and the adjusted epidemic weighted incidence data, using the formula:
[0108]
[0109] Get regional health monitoring data, where R a represents the adjusted epidemic weighted incidence data, Fa represents the adjusted visit frequency data, H a Represents regional health monitoring data.
[0110] The adjusted epidemic weighted incidence data formula is:
[0111]
[0112] p i : The population of the i-th region;
[0113] e i : the raw epidemic incidence in region i, based on the number of reported cases in the past year divided by the population;
[0114] w i :Based on the weight factor of regional health resource allocation, it reflects the resource allocation of the region relative to other regions;
[0115] Suppose there are three regions with data as follows:
[0116] Region 1: p1 = 500,000, e1 = 0.02 (i.e. 2% incidence), w1 = 1.2;
[0117] Region 2: p2=1000000, e2=0.015, w2=1.0;
[0118] Region 3: p3 = 200000, e3 = 0.03, w3 = 0.8;
[0119] Calculate the weighted epidemic separately:
[0120] 500000×0.02×1.2=12000
[0121] 1000000×0.015×1.0=15000
[0122] 200000×0.03×0.8=4800
[0123] Calculate the total weighted epidemic:
[0124]
[0125] Calculate the weighted population:
[0126] 500000×1.2=600000
[0127] 1000000×1.0=1000000
[0128] 200000×0.8=160000
[0129] Calculate the total weighted population:
[0130]
[0131] Calculate R a :
[0132]
[0133] This result R a ≈0.018 means the overall adjusted epidemic incidence is 1.8%.
[0134] Adjusted visit frequency data formula:
[0135]
[0136] v j : Number of visits to the jth medical facility;
[0137] u j : Weight coefficient adjusted according to the severity of the epidemic;
[0138] k j : Weight coefficient of medical facility service quality;
[0139] Assume the following data for three medical facilities:
[0140] Facility 1: v1=10000, u1=1.1, k1=1.5;
[0141] Facility 2: v2 = 20000, u2 = 0.9, k2 = 1.3;
[0142] Facility 3: v3 = 8000, u3 = 1.3, k3 = 1.0;
[0143] Calculate the weighted number of visits separately:
[0144] 10000×1.1×1.5=16500
[0145] 20000×0.9×1.3=23400
[0146] 8000×1.3×1.0=10400
[0147] Calculate the total weighted visits:
[0148]
[0149] Calculate weighted visits:
[0150] 10000×1.5=15000
[0151] 20000×1.3=26000
[0152] 8000×1.0=8000
[0153] Calculate the total weighted visits:
[0154]
[0155] Calculate F a :
[0156]
[0157] This result F a ≈1.027 represents the adjusted mean visit frequency, reflecting an approximately 2.7% increase in the use of medical facilities.
[0158] Regional health monitoring data formula:
[0159]
[0160] R a : Adjusted epidemic weighted incidence data;
[0161] F a : Adjusted visit frequency data;
[0162] H a : Regional health monitoring data, representing a comprehensive health risk index;
[0163] Continue using the results calculated in the previous two steps: R a ≈0.018 (1.8% epidemic-weighted incidence),
[0164] F a ≈1.027 (adjusted visit frequency);
[0165] calculate
[0166]
[0167] calculate
[0168]
[0169] Calculate H a :
[0170]
[0171] This result H a ≈1.027 represents the overall regional health risk index, which assesses the overall health risk level of the region by considering the combined impact of the epidemic and the utilization rate of medical facilities.
[0172] See also Figure 3 , the specific steps for obtaining health trend indicators are:
[0173] Extract monthly health indicators from regional health monitoring data, calculate moving averages and smooth seasonal fluctuations using the formula:
[0174]
[0175] Get the weighted monthly average health index, where H k represents the health monitoring data of the kth month, w k represents the activity weight of the kth month, H avg represents the weighted monthly average health index;
[0176] Using the weighted monthly average health index, the trend component is calculated through time series analysis using the formula:
[0177]
[0178] The enhanced trend component data is obtained, where H avg (t+i) is the monthly average health index at time t, d i is the distance weight based on time i, T b represents the enhanced trend component data;
[0179] Based on the enhanced trend component data, the exponential smoothing algorithm is used to predict the health index in the future time period, through the formula:
[0180]
[0181] Generate health trend indicators, where T b represents the trend component data after strengthening, α is the smoothing coefficient, which controls the prediction reaction speed, Δt is the future time point, τ is the time delay parameter for adjusting the sensitivity of future prediction, and H b Represents the health index predicted for a future period of time.
[0182] Weighted monthly average health index formula:
[0183]
[0184] H k : The health monitoring data for the kth month are directly obtained from the regional health database;
[0185] w k : The activity weight for the kth month, assuming that this is calculated based on the population activity data for that month, such as holidays, flu season, etc.;
[0186] Assume that the monthly health monitoring data and weights are as follows (simplified to 6 months of data): H = [82, 85, 78, 90, 88, 84], w = [1.0, 1.1, 0.9, 1.2, 1.1, 1.0];
[0187] Calculate H avg :
[0188]
[0189] This value of 82.51 represents the half-year average health index after taking into account the monthly activity weights, which can be used to further analyze health trends.
[0190] The enhanced trend component data formula is:
[0191]
[0192] H avg (t+i): monthly average health index for time t;
[0193] d i : Distance weight based on time i, the closer to the current month, the higher the weight;
[0194] Assume that at time t = 4 (i.e. April): H avg =[80, 82, 84, 86, 88, 90, 92] (simplified data), d = [0.5, 0.8, 0.9, 1.0, 0.9, 0.8, 0.5];
[0195] Calculate T b :
[0196]
[0197] The value of 124.41 represents the trend component after weighted average, showing the trend of recent health data.
[0198] The formula for predicting the health index in the future time period is:
[0199]
[0200] T b : Trend component data;
[0201] α: smoothing coefficient, set to 0.3;
[0202] Δt: future time point, set to January;
[0203] τ: time delay parameter, set to 0.5;
[0204] Calculate H b :
[0205]
[0206] This value of 57.62 indicates the predicted health index for the next month, reflecting the expected trend in health status.
[0207] See also Figure 4 ,The specific steps for obtaining the health warning model are:
[0208] Taking the health trend indicator as input data, applying the radial basis function kernel, and reconstructing the core architecture of the model, the formula is:
[0209]
[0210] Generate kernel transformation data, where x, x′ represent differentiated data points, γ represents the shape of the control kernel function, β is the bias term added to the kernel function, and K c (x, x′) represents the adjusted kernel space data;
[0211] Based on the kernel conversion data, the adjustment method of the kernel function parameters and penalty parameters is refined, through the formula:
[0212] C opt =argmin C (∑(y i -y pred,i ) 2 +λC 2 )
[0213] Generate optimized parameter settings, where C is the penalty parameter and y i ,y pred,i They refer to the real-time observation value and the model prediction value respectively, λ is used to adjust the importance of the penalty term, C opt represents the optimized penalty parameter;
[0214] The optimized parameter settings are applied through the formula:
[0215]
[0216] Construct a health warning model, where M(x) is the health warning model, α i is the coefficient corresponding to the data point, y i is the target value of the data point, K c (x,x i ) is the kernel data after radial basis function processing, and θ is the threshold for adjusting the output result.
[0217] Adjusted kernel space data formula:
[0218]
[0219] Assume there are two data points x=1 and x′=3, parameters γ=0.5 and β=0.1, calculate the Euclidean distance between x and x′:
[0220] ||xx′||=|1-3|=2
[0221] Substituting the distance into the formula:
[0222] ∥xx′∥ 2 =2 2 =4
[0223] K c (x,x′)=e -0.5×4 +0.1 = e -2 +0.1≈0.1353+0.1=0.2353
[0224] Here K c (x, x′) = 0.2353 represents the new distance between data points x and x′ after kernel transformation, which is used for further model calculations.
[0225] Optimized penalty parameter formula:
[0226]
[0227] Let y i is the actual observed value [2, 3] and y pred,i is the model prediction value [2.1, 2.9], the penalty parameter C = 1, and the regularization parameter λ = 0.01;
[0228] Calculate the residual sum of squares:
[0229] ∑(y i -y pred,i ) 2 =(2-2.1) 2 +(3-2.9) 2 =0.01+0.01=0.02
[0230] Combined penalty term:
[0231]
[0232] Here C opt = 0.03 represents the balance point between model error and complexity under given λ and C.
[0233] Health early warning model formula:
[0234]
[0235] Assume α = [0.5, 0.5], target value y = [2, 3], kernel conversion data Kc (x,x i ) = [0.2353, 0.2353], threshold θ = 0.05, and n = 2;
[0236] Calculate the model output:
[0237] M(x)=0.5×2×0.2353+0.5×3×0.2353-0.05
[0238] M(x)=0.2353+0.35295-0.05=0.53825
[0239] Here, M(x)=0.53825 represents the output value predicted by the model, reflecting the possibility of health risk.
[0240] See also Figure 5 , the specific steps for obtaining anomaly detection results are:
[0241] Use the health warning model to analyze each data point in the current health data stream, calculate the difference between the model's output of the data point and the predetermined threshold, and innovate the formula:
[0242] R(x)=M(x)-T+σ∥x∥ 2
[0243] Generate model response data points, where M(x) is the health warning model, T is the decision threshold, σ is the sensitivity adjustment coefficient added to the difference, and R(x) represents the adjusted model response;
[0244] Evaluate the model response data points and use the adaptive threshold judgment logic to determine whether the data points are abnormal, using the formula:
[0245]
[0246] Generate the abnormal state of the data point, where R(x) represents the adjusted model response, δ is the dynamic threshold adjusted based on the statistical characteristics of the data, and D(x) represents the abnormal state of the data point;
[0247] Based on the abnormal state of the data point, collect the data points marked as abnormal, through the formula:
[0248] A d ={x|D(x)=1}
[0249] Construct anomaly detection results, where A d is the set of abnormal data points, and D(x) is the abnormal state of the data point.
[0250] The adjusted model response formula is:
[0251] R(x)=M(x)-T+σ∥x∥ 2
[0252] Parameter explanation and calculation method:
[0253] M(x): This is the output of the health warning model for data point x. Assume that M(x) is a function obtained through training data, which is simplified to M(x) = 2x here;
[0254] T: decision threshold, set to 1.5, which is set based on the average response value of historical data;
[0255] σ: sensitivity adjustment coefficient, set to 0.05, which is obtained through the optimization process according to the model performance;
[0256] x: the value of the current data point, assuming x = 2
[0257] Calculate M(x)=2×2=4;
[0258] Calculate ∥x∥ 2 =2 2 =4;
[0259] Substituting the values into the formula we get:
[0260] R(x)=4-1.5+0.05×4=4-1.5+0.2=2.7
[0261] The result R(x)=2.7 represents the adjusted model response and is used for the next step of abnormality determination.
[0262] The abnormal state formula of a data point:
[0263]
[0264] δ: dynamic threshold, set to the mean response plus twice the standard deviation, assumed to be 0.5 here (based on statistical analysis of the data);
[0265] Use the R(x)=2.7 calculated in the previous step;
[0266] Compare R(x) with δ=0.5, because 2.7>0.5;
[0267] D(x)=1, indicating that x=2 is an abnormal data point.
[0268] The set formula for abnormal data points is:
[0269] A d ={x|D(x)=1}
[0270] A d: A collection of abnormal data points, which will collect all points marked as 1 by D(x);
[0271] Suppose there is a data set {1, 2, 3}, and D(x) of each point is calculated as 0, 1, 0 respectively;
[0272] Only x=2 is marked as an anomaly, so A d ={2}.
[0273] See also Figure 6 , the specific steps for obtaining the adjusted threshold are:
[0274] Collect data points from the anomaly detection results, perform quantitative analysis on the sudden increase in hospital visit rate and the abnormal data of drug sales, and calculate the standard deviation of the abnormal data points, using the formula:
[0275]
[0276] Generate standard deviation data points where x i is the value of the outlier data point, is the mean of the data points, and Nf represents the standard deviation calculated from the abnormal data;
[0277] Using standard deviation data points, combined with time series analysis, seasonal and cyclical adjustments are made to the outliers using the formula:
[0278]
[0279] Generate time series adjusted data, where Nf represents the standard deviation calculated from the abnormal data, α t ,ω,φ,β t are the amplitude, frequency, phase and amplitude of the cosine term of the periodic adjustment, respectively, and S(x) represents the data after seasonal and periodic adjustment;
[0280] According to the adjusted data of the time series, the threshold is readjusted and the authenticity of the abnormal data is reflected by the formula:
[0281] Vf=max(S(x))+κ
[0282] Generate adjusted thresholds, where S(x) represents the data after seasonal and periodic adjustment, κ is a small amount used to increase the sensitivity of the threshold, and Vf is the newly set threshold.
[0283] The formula for calculating the standard deviation from abnormal data is:
[0284]
[0285] n: is the number of data points;
[0286] x i : The value of each data point;
[0287] The average of all data points;
[0288] Nf: standard deviation calculated from abnormal data;
[0289] Collect all data points x i ;
[0290] Calculate the average of the data points
[0291] Compute the square of the deviation of each data point from the mean:
[0292] Sum the squares of all differences:
[0293] Calculate the standard deviation Nf:
[0294] Suppose there are 5 data points: 10, 12, 14, 16, 18;
[0295]
[0296] The square of the difference = (10-14) 2 , (12-14) 2 , (14-14) 2 , (16-14) 2 , (18-14) 2 =16,4,0,4,16;
[0297] Total = 40;
[0298] Calculate Nf:
[0299]
[0300] This standard deviation Nf≈2.83 represents the degree to which abnormal data points deviate from the mean.
[0301] The formula for seasonally and periodically adjusted data is:
[0302]
[0303] Nf: standard deviation calculated in the previous step;
[0304] α t ,ω,φ,β t : are the amplitude, frequency, phase and amplitude of the cosine term of the periodic adjustment respectively;
[0305] T: time period;
[0306] Use Nf obtained in the previous step;
[0307] Choose an appropriate period T, amplitude α t and β t , as well as frequency ω and phase φ;
[0308] Calculate the value of the periodic function for each t and accumulate to obtain S(x);
[0309] Assume Nf = 2.83, select T = 1, α1 = 1, β1 = 0.5,
[0310] Calculate S(x):
[0311]
[0312] S(x)=2.83+1×1+0.5×0=3.83
[0313] This S(x)=3.83 represents the data value after seasonal and cyclical adjustments.
[0314] The new threshold formula is:
[0315] Vf=max(S(x))+κ
[0316] S(x): adjusted time series data obtained from the previous step;
[0317] κ: a small amount used to increase the threshold sensitivity;
[0318] Vf: newly set threshold;
[0319] Find the maximum value of S(x);
[0320] Add a small positive value κ to set the new threshold;
[0321] Assume that the maximum value of S(x) is 3.83, κ = 0.2;
[0322] Calculate Vf:
[0323] Vf=3.83+0.2=4.03
[0324] This Vf=4.03 is the new threshold after adjustment to more accurately identify abnormal data points.
[0325] See also Figure 7 , the specific steps for obtaining the abnormal verification results are:
[0326] Using the adjusted threshold, evaluate the data points in the current data stream, calculate the difference between each data point and the threshold, and preliminarily mark the outliers, using the formula:
[0327] R i (x) = xV f
[0328] Generates the difference result of preliminary anomaly labeling, where x is the value of the data point and V f is the adjusted threshold, R i (x) is the difference result of the preliminary abnormal marking;
[0329] Perform time series analysis on the difference results of the preliminary anomaly markings to analyze whether the anomaly data points show deviations from the regular seasonal or cyclical patterns, using the formula:
[0330]
[0331] Generate time series analysis results, where R i (x) is the difference result of the preliminary abnormal marking, c j ,ω,φ j Adjusting the complexity and sensitivity of time series analysis, T a (x) represents the weighted time series anomaly index.
[0332] According to the time series analysis results, the threshold is readjusted and abnormal data is captured and determined by the formula:
[0333] V n =V f +ζ·max(T a (x))
[0334] Generate anomaly verification results, where x is the value of the data point, V f is the adjusted threshold, T a (x) represents the weighted time series anomaly index, V n Indicates an abnormal validation result.
[0335] The difference result formula of the preliminary abnormality mark is:
[0336] R i (x) = xV f
[0337] This formula is used to calculate the difference between each data point x and the adjusted threshold V f The difference, R i The value of (x) will be used to determine
[0338] Check whether the data point is abnormal;
[0339] x: the value of the data point, for example x=105;
[0340] V f : The adjusted threshold is set to V f =100;
[0341] Calculate R i (x):
[0342] R i (105) = 105 - 100 = 5
[0343] This means that the data point x=105 is 5 higher than the threshold, and according to the subsequent judgment criteria, it is an outlier.
[0344] The weighted time series anomaly index formula is:
[0345]
[0346] This formula is used to weight the flagged anomalous data points, incorporating their deviations from seasonal and cyclical patterns;
[0347] R i (x) = 5 (obtained from the previous step);
[0348] c j : Periodicity coefficient, for example c1 = 0.3;
[0349] ω: frequency, set to ω = 1 (every year);
[0350] φ j : Phase, set φ1 = 0;
[0351] m: number of cycles, set to m=1;
[0352] Calculate T a (x):
[0353] T a (105)=5×(1+0.3cos(2π×1×105+0))
[0354] =5×(1+0.3cos(210π))
[0355] =5×(1+0.3×1)
[0356] =5×1.3=6.5
[0357] T a (105)=6.5 indicates that after considering seasonal and cyclical factors, the abnormality level of data point 105 is 6.5.
[0358] Abnormal verification result formula:
[0359] V n =V f +ζ·max(T a (x)
[0360] This formula is used to adjust the threshold to better capture actual abnormal data;
[0361] V f =100(previously set);
[0362] ζ: adjustment coefficient, set to ζ = 0.05;
[0363] max(T a (x)) = 6.5 (obtained from the previous step);
[0364] Calculation process:
[0365] V n =100+0.05×6.5=100+0.325=100.325
[0366] Abnormal verification result V n Adjusted to 100.325 for future anomaly detection.
[0367] See also Figure 8 , the specific steps for obtaining real-time health warnings are:
[0368] Summarize the data from the abnormal verification results, compare with the historical health data, and evaluate the current abnormal situation, calculate the deviation of the historical data, and use the formula:
[0369]
[0370] Generate historical deviation analysis results, where x k Indicates the value of the current data point, h k represents the average value of the same period in history, α k represents the weight factor, and Pm represents the calculated historical deviation result;
[0371] Based on the results of historical deviation analysis and combined with public health research, the warning level is determined using the formula:
[0372]
[0373] Generate warning level results, where Pm represents the calculated historical deviation result, θ is the threshold, and Lg is the quantitative warning level;
[0374] Using the warning level results, issue warnings through the operation interface using the formula:
[0375] Pg=ifLg>1then'Publish'else'Monitor'
[0376] Generate real-time health warnings, where Lg is the quantitative warning level and Pg represents the operational decision based on Lg.
[0377] The calculated historical deviation result formula is:
[0378]
[0379] x k : The value of the current outlier point;
[0380] h k :Historical data for the same period;
[0381] α k : The importance weight of each data point;
[0382] Pm: historical deviation analysis results;
[0383] Data collection: First collect the current outlier data x k and the corresponding historical data h k ;
[0384] Weight assignment: Assign a weight α to each data point based on its importance k ,The basis for weight allocation is the reliability of data and the relevance of health events, etc.;
[0385] Deviation calculation: Calculate the deviation of each data point and adjust the weight. The calculation formula is: The absolute value of the difference indicates the degree of deviation, and the weight index indicates the importance of the deviation;
[0386] Sum calculation: Calculate the sum of all adjusted deviations to obtain the total historical deviation Pm;
[0387] Assume there are three data points, the current value is x = [100, 150, 120], the historical value is h = [90, 160, 115], and the weight is α = [0.5, 1, 0.75];
[0388] Calculate Pm:
[0389] Pm=|100-90| 0.5 +∣150-160∣ 1 +∣120-115∣ 0.75
[0390] Pm=10 0.5 +10 1 +5 0.75
[0391] Pm=3.16+10+3.34
[0392] Pm=16.5
[0393] This value Pm=16.5 represents the comprehensive historical deviation degree.
[0394] Quantitative warning level formula:
[0395]
[0396] Pm: historical deviation analysis results;
[0397] θ: threshold, used to normalize the warning level;
[0398] Lg: Quantitative warning level;
[0399] Set threshold θ: Set a reasonable threshold based on public health standards or historical data;
[0400] Calculate the warning level: Calculate the warning level based on the historical deviation Pm and threshold θ. The level is determined by the ratio Decide and then make adjustments to the upper and lower limits;
[0401] Level range adjustment: Ensure that the warning level is between 1 and 5, and use the min and max functions to adjust;
[0402] Assume Pm = 16.5, threshold θ = 4;
[0403] Calculate Lg:
[0404]
[0405] Lg=min(max(1,5),5)
[0406] Lg=min(5,5)
[0407] Lg=5
[0408] This value Lg=5 represents the highest level of health warning.
[0409] Operation decision formula based on Lg:
[0410] Pg=ifLg>1then'Publish'else'Monitor'
[0411] Lg: calculated warning level;
[0412] Pg: Operation decision based on Lg, indicating whether to issue an early warning;
[0413] Decision condition setting: set a condition. When Lg is greater than 1, it means that an early warning needs to be issued, otherwise continuous monitoring is performed;
[0414] Logical judgment implementation: Apply logical judgment to determine the operation, which is implemented through simple conditional judgment. If Lg>1, execute publishing ('Publish'), otherwise execute monitoring ('Monitor');
[0415] Assume that according to the calculation in the previous step, Lg = 5;
[0416] Calculate Pg:
[0417] Pg=if5>1then'Publish'else'Monitor'
[0418] Pg = 'Publish'
[0419] This operation Pg='Publish' indicates that the warning information will be published according to the high warning level.
[0420] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A public health early warning system based on big data, characterized in that: The system comprises: The regional health data aggregation module extracts the population epidemic incidence rate and medical facility visit frequency data based on the national health database to generate regional health monitoring data, and based on the regional health monitoring data, evaluates the regional health status trend and obtains health trend indicators; The nonlinear model optimization module selects a radial basis function kernel to process nonlinear relationships based on the health trend index, adjusts kernel function parameters and penalty parameters, obtains a health warning model, analyzes abnormal data points in the current health data stream based on the health warning model, and generates anomaly detection results; The abnormal data analysis module analyzes abnormal points in the data based on the abnormal detection results, including sudden increases in hospital visit rates and abnormal drug sales, performs time series abnormal point analysis, compares them with seasonal and cyclical changes, and adjusts the threshold to obtain the adjusted threshold, performs abnormal data verification according to the adjusted threshold, and generates an abnormal verification result; Based on the abnormal verification results, the early warning signal release module summarizes the historical data analysis and the current abnormal conditions, establishes an early warning signal, releases the early warning information through the operation interface, conveys the early warning information to the relevant health departments and the public, and generates real-time health warnings.
2. The public health early warning system based on big data according to claim 1 is characterized in that: The steps for obtaining the regional health monitoring data are specifically as follows: Based on the national health database and extracting the epidemic records of the target area, the population-weighted epidemic incidence rate of the community is calculated in detail using the formula: The adjusted weighted incidence rate data of the epidemic is obtained, among which p i represents the population of the i-th region, e i represents the original epidemic incidence rate in the i-th region, w i represents the weight factor for regional health resource allocation, R a represents the adjusted epidemic-weighted incidence data; Using the adjusted epidemic weighted incidence data, query the medical records of the corresponding medical facilities and calculate the weighted frequency of medical facilities using the formula: The adjusted visit frequency data is obtained, where v j is the number of visits to the jth medical facility, u j is the weight coefficient adjusted according to the severity of the epidemic, k j is the weight coefficient of the medical facility service quality, F a represents the adjusted visit frequency data; The health risk index of the region is calculated by combining the data obtained from the adjusted visit frequency data and the adjusted epidemic weighted incidence data, using the formula: Get regional health monitoring data, where R a represents the adjusted epidemic weighted incidence data, F a represents the adjusted visit frequency data, H a Represents regional health monitoring data.
3. The public health early warning system based on big data according to claim 2 is characterized in that: The steps for obtaining the health trend indicator are specifically as follows: Extract monthly health indicators from the regional health monitoring data, calculate moving averages and smooth seasonal fluctuations using the formula: Get the weighted monthly average health index, where H k represents the health monitoring data of the kth month, w k represents the activity weight of the kth month, H avg represents the weighted monthly average health index; Using the weighted monthly average health index, the trend component is calculated through time series analysis using the formula: The enhanced trend component data is obtained, where H avg (t+i) is the monthly average health index at time t, d i is the distance weight based on time i, T b represents the enhanced trend component data; Based on the enhanced trend component data, the exponential smoothing algorithm is used to predict the health index in the future time period, through the formula: Generate health trend indicators, where T b represents the trend component data after strengthening, α is the smoothing coefficient, which controls the prediction reaction speed, Δt is the future time point, τ is the time delay parameter for adjusting the sensitivity of future prediction, and H b Represents the health index predicted for a future period of time.
4. The public health early warning system based on big data according to claim 3 is characterized in that: The steps for obtaining the health warning model are specifically as follows: The health trend index is used as input data, and the radial basis function kernel is applied to reconstruct the core architecture of the model, through the formula: K c (x,x′)=e -γ(||x-x′||)2 +b Generate kernel transformation data, where x, x′ represent differentiated data points, γ represents the shape of the control kernel function, β is the bias term added to the kernel function, and K c (x, x′) represents the adjusted kernel space data; Based on the kernel conversion data, the adjustment method of the kernel function parameters and the penalty parameters is refined, through the formula: C opt =argmin C (Σ(y i -the pred,i ) 2 +λC 2 ) Generate optimized parameter settings, where C is the penalty parameter and y i ,y pred,i They refer to the real-time observation value and the model prediction value respectively, λ is used to adjust the importance of the penalty term, C opt represents the optimized penalty parameter; Applying the optimized parameter settings, the formula is: Construct a health warning model, where M(x) is the health warning model, α i is the coefficient corresponding to the data point, y i is the target value of the data point, K c (x,x i ) is the kernel data after radial basis function processing, and θ is the threshold for adjusting the output result.
5. The public health early warning system based on big data according to claim 4 is characterized in that: The steps for obtaining the abnormal detection result are specifically as follows: The health warning model is used to analyze each data point in the current health data stream, and the difference between the model's output of the data point and the predetermined threshold is calculated to create an innovative formula: R(x)=M(x)-T+σ||x|| 2 Generate model response data points, where M(x) is the health warning model, T is the decision threshold, σ is the sensitivity adjustment coefficient added to the difference, and R(x) represents the adjusted model response; The model response data points are evaluated and the adaptive threshold judgment logic is used to determine whether the data points are abnormal, using the formula: Generate the abnormal state of the data point, where R(x) represents the adjusted model response, δ is the dynamic threshold adjusted based on the statistical characteristics of the data, and D(x) represents the abnormal state of the data point; Based on the abnormal status of the data points, the data points marked as abnormal are collected, and the formula is: A d ={x∣D(x)=1} Construct anomaly detection results, where A d is the set of abnormal data points, and D(x) is the abnormal state of the data point.
6. The public health early warning system based on big data according to claim 5 is characterized in that: The steps for obtaining the adjusted threshold are specifically as follows: Data points are collected from the anomaly detection results, and the sudden increase in hospital visit rate and the abnormal data of drug sales are quantitatively analyzed, and the standard deviation of the abnormal data points is calculated, through the formula: Generate standard deviation data points where x i is the value of the outlier data point, is the mean of the data points, and Nf represents the standard deviation calculated from the abnormal data; Using the standard deviation data points, combined with time series analysis, seasonal and cyclical adjustments are made to the abnormal data using the formula: Generate time series adjusted data, where Nf represents the standard deviation calculated from the abnormal data, α t ,ω,φ,β t are the amplitude, frequency, phase and amplitude of the cosine term of the periodic adjustment, respectively, and S(x) represents the data after seasonal and periodic adjustment; According to the adjusted data of the time series, the threshold is readjusted and the authenticity of the abnormal data is reflected by the formula: Vf=max(S(x))+κ Generate adjusted thresholds, where S(x) represents the data after seasonal and periodic adjustment, κ is a small amount used to increase the sensitivity of the threshold, and Vf is the newly set threshold.
7. The public health early warning system based on big data according to claim 6 is characterized in that: The steps for obtaining the abnormal verification result are specifically as follows: Using the adjusted threshold, the data points in the current data stream are evaluated, the difference between each data point and the threshold is calculated, and the outliers are preliminarily marked, using the formula: R i (x)=x-V f Generates the difference result of preliminary anomaly labeling, where x is the value of the data point and V f is the adjusted threshold, R i (x) is the difference result of the preliminary abnormal marking; A time series analysis is performed on the difference results of the preliminary anomaly markings to analyze whether the anomaly data points show deviations from the regular seasonal or cyclical patterns, using the formula: Generate time series analysis results, where R i (x) is the difference result of the preliminary abnormal marking, c j ,ω,φ j Adjusting the complexity and sensitivity of time series analysis, T a (x) represents the weighted time series anomaly index. According to the time series analysis results, the threshold is readjusted and abnormal data is captured and determined, through the formula: V n =V f +ζ·max(T a (x)) Generate anomaly verification results, where x is the value of the data point, V f is the adjusted threshold, T a (x) represents the weighted time series anomaly index, V n Indicates an abnormal validation result.
8. The public health early warning system based on big data according to claim 7 is characterized in that: The steps for obtaining the real-time health warning are specifically as follows: Summarize the data from the abnormal verification results, compare with the historical health data, and evaluate the current abnormal situation, calculate the deviation of the historical data, using the formula: Generate historical deviation analysis results, where x k Indicates the value of the current data point, h k represents the average value of the same period in history, α k represents the weight factor, and Pm represents the calculated historical deviation result; Based on the historical deviation analysis results, combined with public health research, the warning level is determined using the formula: Generate warning level results, where Pm represents the calculated historical deviation result, θ is the threshold, and Lg is the quantitative warning level; Using the warning level results, issue a warning through the operation interface using the formula: Pg=ifLg>1then'Publish'else'Monitor' Generate real-time health warnings, where Lg is the quantitative warning level and Pg represents the operational decision based on Lg.
Citation Information
Patent Citations
Identification of early diagnosis markers of lung adenocarcinoma based on co-expression similarity, and constructing method of risk prediction model
CN109841281A
Health emergency prevention and control and early warning system
CN112735586A
Medical information intelligent management system based on cloud platform
CN118213057A
Intelligent community resident health management method based on machine learning
CN119274802A
Method for Prediction for Nonlinear Seasonal Time Series
US20110160927A1