Disease risk prediction method and device

By collecting and processing disease case data, using multi-layer time domain convolution network and feature attention mechanism for dynamic risk level division, the problem of inefficiency of traditional methods is solved and efficient and accurate disease risk prediction is achieved.

CN120280137APending Publication Date: 2025-07-08ANHUI MEDICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510313945.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing infectious disease risk assessment methods rely on manual analysis, are inefficient and error-prone, cannot meet the timeliness and accuracy requirements of modern disease prevention and control, and fail to effectively capture the impact of time series dependence and multi-dimensional characteristics in case data.

Method used

Disease case data in the preset area is collected, abnormal detection and processing is performed, and dynamic risk level division is divided through multi-layer time domain convolution network and feature attention mechanism to generate low, medium and high risk levels, and a risk prediction model is trained to make predictions.

Benefits of technology

It realizes efficient processing of complex timing data and precise risk level classification, improving the accuracy and efficiency of disease risk prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120280137A_ABST
    Figure CN120280137A_ABST
Patent Text Reader

Abstract

The invention discloses a disease risk prediction method and device, and the method comprises the steps: collecting original data of a preset disease case in a preset region, and carrying out the anomaly detection and processing of the original data; carrying out dynamic risk grading based on data driving, and carrying out risk grading on the processed original data; dividing the original data subjected to risk grade division into a training set and a verification set, training by adopting the training set to obtain a risk prediction model, and inputting the verification set into the trained risk prediction model to obtain a risk prediction result; according to the method, complex time series data can be efficiently processed, and accurate risk grade division and prediction are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data analysis and processing, and particularly relates to a disease risk prediction method and device. Background Art

[0002] In the field of public health, the early identification and risk assessment of infectious diseases are crucial for the formulation of prevention and control strategies. In the context of globalization, with the frequent movement of the population, the transmission speed and geographical diffusion ability of infectious diseases have increased significantly. However, traditional risk assessment methods often rely on manual analysis, which is inefficient and error-prone when dealing with a large amount of historical case time-series data, and can no longer meet the timeliness and accuracy requirements of modern disease prevention and control.

[0003] In recent years, the rapid development of deep learning and big data technologies has brought new opportunities to solve this problem. However, most current methods still remain at the level of simple linear regression or single neural network structures, failing to effectively capture the time-series dependence relationships in case data and not fully considering the comprehensive impact of multiple-dimensional features. In addition, existing risk classification methods mostly rely on expert experience or simple threshold setting, unable to fully utilize the potential information in the data, resulting in inaccurate classification results. Summary of the Invention

[0004] To solve the deficiencies in the prior art, this application proposes a disease risk prediction method and device.

[0005] In a first aspect, this application proposes a disease risk prediction method, including:

[0006] Collect the original data of preset diseases in a preset area, and perform anomaly detection and processing on the original data;

[0007] Based on data-driven dynamic risk level classification, classify the processed original data into risk levels;

[0008] Divide the original data after risk level classification into a training set and a validation set, train a risk prediction model using the training set, and input the validation set into the trained risk prediction model to obtain a risk prediction result.

[0009] Optionally, the performing anomaly detection and processing on the original data includes:

[0010] Identify and label isolated anomaly data points as outliers by constructing multiple random trees;

[0011] Smooth the outliers using interpolation;

[0012] Repair the boundaries of the data after processing the outliers by filling with the previous and next values.

[0013] Optionally, for the data-driven dynamic risk level division, the processed original data is divided into risk levels, including:

[0014] Smoothing the processed original data;

[0015] Dividing the smoothed data into multiple clusters and automatically identifying different risk levels;

[0016] Dividing the processed original data into risk levels according to the different risk levels, and generating three risk levels of low, medium, and high.

[0017] Optionally, the dividing the smoothed data into multiple clusters is based on the following calculation formula:

[0018]

[0019] where J is the objective function, C i is the i-th cluster, μ i is the center of the i-th cluster, x j is the j-th data point, the data point belongs to the data set X, and the data set X = {x1, x2,.. x j .. x k}, and K is the total number of data points in the data set X.

[0020] Optionally, for the risk prediction model trained by using a training set, the validation set is input into the trained risk prediction model, and the risk prediction result is obtained, including:

[0021] Stacking multiple layers of time-domain convolutional network modules and extracting time series features from the multiple layers of time-domain convolutional network modules;

[0022] Through the attention mechanism, calculating the weights of the environmental features of the multiple layers of time-domain convolutional network modules, and weighting the environmental feature weights to the multiple layers of time-domain network modules to generate weighted features;

[0023] After splicing the time series features and the weighted features, inputting them into the fully connected layer of the multiple layers of time-domain convolutional network for fusion, outputting the probability distribution of the risk level, and selecting the category with the highest probability as the prediction result.

[0024] Optionally, the extracting time series features from the multiple layers of time-domain convolutional network modules includes performing causal convolution and dilated convolution on the data in the multiple layers of time-domain convolutional network modules.

[0025] Optionally, the calculation formula for calculating the weights of the environmental features of the multiple layers of time-domain convolutional network modules is as follows:

[0026]

[0027] where \(e\) i is the energy of the \(i\)-th environmental feature, and \(\alpha\) i is the attention weight of the \(i\)-th environmental feature, and \(n\) is the number of environmental features.

[0028] Optionally, the calculation formula for the probability distribution of the output risk level is:

[0029]

[0030] where \(M\) is the total number of data points in the training set or validation set, \(C\) is the number of risk categories, and \(y\) i,c is the \(c\)-th risk category label of the \(i\)-th data point, and \(\hat{y}_{ic}\) is the probability of the \(c\)-th risk category label of the \(i\)-th data point predicted by the model.

[0031] In a second aspect, a disease risk prediction device is proposed, including:

[0032] A raw data acquisition module, configured to acquire the raw data of preset disease cases in a preset area, and perform anomaly detection and processing on the raw data;

[0033] A risk level classification module, configured to perform risk level classification on the processed raw data based on data-driven dynamic risk level classification;

[0034] A risk prediction and evaluation module, configured to divide the raw data after risk level classification into a training set and a validation set, train a risk prediction model using the training set, and input the validation set into the trained risk prediction model to obtain a risk prediction result.

[0035] The beneficial effects brought by the technical solutions provided in some embodiments of the present application at least include:

[0036] The present application provides a disease risk prediction method and device. By acquiring the raw data of preset disease cases in a preset area, and performing anomaly detection and processing on the raw data; performing risk level classification on the processed raw data based on data-driven dynamic risk level classification; dividing the raw data after risk level classification into a training set and a validation set, training a risk prediction model using the training set, and inputting the validation set into the trained risk prediction model to obtain a risk prediction result; the present application can efficiently process complex time series data and achieve accurate risk level classification and prediction.

[0037] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application are realized and obtained by the structures specifically pointed out in the specification, claims, and drawings..

[0038] To make the above objects, features, and advantages of the present application more obvious and understandable, the following provides preferred embodiments in conjunction with the accompanying drawings and makes a detailed description as follows.

[0039] Advantages of additional aspects of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the description of the embodiments of the present invention or the prior art. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0041] Figure 1 Flowchart of the disease risk prediction method shown in the embodiments of the present application;

[0042] Figure 2 Principle block diagram of the disease risk prediction device shown in the specific examples of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the drawings.

[0044] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of methods and devices consistent with some aspects of the present application as detailed in the appended claims.

[0045] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise stated, "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0046] Embodiment 1

[0047] The following will be combined with the attached Figure 1 , and a disease risk prediction method provided by the embodiments of the present application will be introduced in detail. AsFigure 1 As shown in the figure, the method according to the embodiment of the present application may include the following steps:

[0048] Step S1: Collect the original data of the preset diseases in the preset area, and perform anomaly detection and processing on the original data;

[0049] Specifically in this embodiment, the preset area and the preset disease are the prefecture-level city and hand, foot and mouth disease respectively. Collect the real original data such as the number of hand, foot and mouth disease cases and environmental factors in the prefecture-level city, and perform anomaly detection and smooth data processing on the original data;

[0050] Specifically, first, collect the case data of hand, foot and mouth disease from various hospitals and public health monitoring systems, including environmental variables such as Date, Cases, AvgTemp, AvgPress, AvgRH, Prec, and SunHours. Secondly, perform preliminary cleaning on the collected original data, including removing invalid data, correcting obvious errors, and unifying the data format. Finally, perform anomaly detection and smooth data processing on the data based on methods such as the Isolation Forest algorithm and cubic spline interpolation. This process specifically includes the following steps:

[0051] Step S1.1: Identify and label the isolated abnormal data points as outliers by constructing multiple random trees;

[0052] Specifically, use the Isolation Forest algorithm to perform unsupervised anomaly detection on the case data of hand, foot and mouth disease, identify and remove the outliers in the data. Identify and label the outliers by constructing multiple random trees to "isolate" the abnormal data points. Assume the sample set is D = {d1, d2,...d i .., d N}, the Isolation Forest algorithm recursively partitions the data space, calculates the isolation degree of each data point, and uses it to evaluate whether it is an outlier. If a data point has a high isolation degree, it is an outlier. The formula is expressed as:

[0053]

[0054] Among them, Isolation Degree(d i ) is the isolation degree of point d i , and N is the total number of data points in the sample set D.

[0055] Set the proportion of outliers to 3%. This parameter is used to define the proportion of expected outliers in the data. The selection of 3% is based on the assumption of abnormal fluctuations in the time series data of hand, foot and mouth disease to ensure that most outliers are detected without misjudgment.

[0056] Step S1.2: Smooth the outliers using the interpolation method;

[0057] Specifically, use the cubic spline interpolation method to smooth the outliers to ensure the continuity of the data trend while removing the outliers. Cubic spline interpolation is a technique for smoothing data through piecewise cubic polynomials, which is particularly suitable for continuous and smooth data. Since the polynomial is fitted according to the local trend of the data points, it can better retain the original characteristics of the data.

[0058] For data points d i and d i+1 , the interpolation process is represented by the following formula:

[0059] S(x) = a0 + a1(x - x0) + a2(x - x0) 2 + a3(x - x0) 3

[0060] where S(x) is the interpolation function, a0, a1, a2, a3 are interpolation coefficients, and x0 is the position of the known data point. Step S1.3: Repair the boundaries of the data after processing the outliers using the method of filling with the previous and next values;

[0061] Since interpolation processing may introduce discontinuities or errors in the boundary part of the data, the method of filling with the previous and next values is used to repair the boundary to ensure the integrity and continuity of the data sequence. For boundary data points d1 and d N , the method of filling with the previous and next data is used:

[0062] d1 = d2, d N = d N-1 .

[0063] Step S2: Based on the data-driven dynamic risk level division, divide the processed original data into risk levels;

[0064] Specifically in this embodiment, the dynamic risk level of the number of hand, foot and mouth disease cases is divided by an unsupervised algorithm, and different risk levels of low, medium and high are automatically generated. The specific process is as follows:

[0065] Step S2.1: Smooth the processed original data;

[0066] To reduce the impact of short-term fluctuations on risk assessment, use the 15-day moving average method to smooth the case data. Let C t be the time series data of the number of cases. The formula for smoothing the data using the 15-day moving average is:

[0067]

[0068] Among them, is the number of cases after smoothing, C i ' is the number of the i-th original case, and t is the time point.

[0069] The method of this embodiment can effectively remove short-term fluctuations and highlight the long-term epidemic trend..

[0070] Step S2.2: Divide the smoothed data into multiple clusters to automatically identify different risk levels;

[0071] Use the K-means clustering algorithm to divide the dynamic distribution of case data into multiple clusters to automatically identify different risk levels. Let the data set after smoothing be X = {x1, x2,..x j ..x K}, and K is the total number of data points in the data set X;

[0072] The K-means algorithm performs clustering by minimizing the following objective function:

[0073]

[0074] Among them, J represents the objective function, C i represents the i-th cluster, μ i represents the center of the i-th cluster, and x j represents the j-th data point.

[0075] Through iterative optimization, minimize the distance between all sample points and their respective cluster centers..

[0076] Step S2.3: Divide the processed original data according to the different risk levels to generate three risk levels: low, medium, and high.

[0077] Divide the risk levels according to the clustering results to automatically generate three risk levels: low, medium, and high. The final risk level R can be automatically determined by the following formula:

[0078]

[0079] Among them, R is the risk level, represents the result of classifying the number of cases after smoothing through the clustering algorithm for classification.

[0080] Step S3: Divide the original data after risk level division into a training set and a validation set, use the training set to train to obtain a risk prediction model, and input the validation set into the trained risk prediction model to obtain a risk prediction result.

[0081] Specifically, the risk of hand, foot and mouth disease epidemic is predicted by combining TCN (Temporal Convolutional Network) and feature attention mechanism. Multidimensional features (such as the number of cases, temperature, humidity, etc.) that have been cleaned and standardized are input into the trained TCN-Attention model. TCN extracts the long-term dependencies of time series through causal convolution and dilated convolution, while the feature attention mechanism dynamically calculates the weights of each environmental feature to generate a weighted feature representation. The time series features and the weighted features are concatenated and then input into a fully connected layer for fusion, and finally the probability distributions of low, medium, and high risk levels are output. The model selects the category with the highest probability as the prediction result.

[0082] The specific steps are as follows:

[0083] Step S3.1: Stack multiple temporal convolutional network modules and extract time series features from the multiple temporal convolutional network modules;

[0084] Specifically, perform causal convolution and dilated convolution on the time series feature data.

[0085] Causal convolution ensures that the model only depends on past information and strictly follows the time causality. Its calculation formula is:

[0086]

[0087] where y t is the output at the t-th time step, x t-k is the input at the t - k-th time step, w k is the weight of the k-th convolutional kernel, and K is the convolutional kernel size.

[0088] Dilated convolution captures the long-term dependencies in the time series by increasing the dilation factor. Its calculation formula is:

[0089]

[0090] where d is the dilation factor, and t - d·k represents the input position of the dilated convolution.

[0091] Step S3.2: Through the attention mechanism, calculate the weights of the environmental features of the multiple temporal convolutional network modules, and weight the multiple temporal network modules with the weights of the environmental features to generate weighted features;

[0092] Specifically, based on TCN, a feature attention mechanism is introduced to dynamically weight the importance of multi-dimensional features (such as temperature, humidity, air pressure, etc.), and a TCN-Attention prediction model is constructed. By multiplying the attention weights with the time series features, a weighted feature vector can be obtained. Through the calculation of attention weights, the model can automatically focus on the features that have the greatest impact on the prediction target (the number of cases).

[0093] The calculation formula for the attention weights is as follows:

[0094]

[0095] Where, e i is the energy of the time series feature i, usually calculated through a fully connected layer, and α i is the attention weight of the i-th time series feature, and n is the number of time series features.

[0096] Step S3.3: Concatenate the time series features with the weighted features and input them into the fully connected layer of the multi-layer time domain convolutional network for fusion, output the probability distribution of the risk level, and select the category with the highest probability as the prediction result.

[0097] Use the divided dataset to train the TCN-Attention model, and optimize the model parameters through the backpropagation algorithm.

[0098] During the prediction process of the TCN-Attention model, the cross-entropy loss function (Cross-EntropyLoss) is used as the objective function:

[0099]

[0100] Where: M is the total number of data points in the training set or validation set, C is the number of risk categories, y i,c is the c-th risk category label of the i-th data point, is the probability of the c-th risk category label of the i-th data point predicted by the model.

[0101] It can be understood that the above steps S3.1 and S3.2 are both for model training on the training set. Finally, in S3.3, on the one hand, the training set is used to calculate the minimum value of the objective function to achieve the optimal effect of the model, and on the other hand, the validation set is used to calculate the objective function to calculate the prediction error.

[0102] The performance of the model is evaluated using metrics such as Accuracy, Precision, Recall, and F1-Score. The experimental results confirm that a disease risk prediction method provided by this application achieves efficient processing of complex time-series data and accurate risk level classification and prediction.

[0103] Example Two

[0104] This application proposes a device for establishing a disease risk prediction model, as Figure 2 shown, including:

[0105] A raw data acquisition module, which is used to acquire the raw data of preset diseases in a preset area and perform anomaly detection and processing on the raw data;

[0106] A risk level classification module, which is used to classify the processed raw data based on data-driven dynamic risk level classification;

[0107] A risk prediction and evaluation module, which is used to divide the raw data after risk level classification into a training set and a validation set, train a risk prediction model using the training set, and input the validation set into the trained risk prediction model to obtain a risk prediction result.

[0108] Each module is connected in sequence.

[0109] The module functions in the device of this embodiment have the same technical features as those in the method of Embodiment One, so the same technical effects are also achieved.

[0110] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0111] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the module, segment of a program, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0112] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

[0113] The applicant of this application has made a detailed description and illustration of the embodiments of this application in combination with the accompanying drawings of the specification. However, those skilled in the art should understand that the above embodiments are only the preferred implementation schemes of this application, and the detailed description is only to help readers better understand the spirit of this application, rather than a limitation on the protection scope of this application. On the contrary, any improvement or modification based on the inventive spirit of this application should fall within the protection scope of this application.

Claims

1. A disease risk prediction method, characterized in that, Including: Collecting the original data of preset disease cases in a preset area, and performing anomaly detection and processing on the original data; Based on data-driven dynamic risk level division, dividing the processed original data into risk levels; Dividing the original data after risk level division into a training set and a validation set, training a risk prediction model using the training set, and inputting the validation set into the trained risk prediction model to obtain a risk prediction result.

2. The disease risk prediction method according to claim 1, wherein The performing anomaly detection and processing on the original data includes: Constructing multiple random trees to isolate abnormal data points, identifying and marking the abnormal data points as outliers; Smoothing the outliers using interpolation; Repairing the boundaries of the data after processing the outliers by filling with previous and next values.

3. The disease risk prediction method according to claim 1, wherein The based on data-driven dynamic risk level division, dividing the processed original data into risk levels includes: Smoothing the processed original data; Dividing the smoothed data into multiple clusters and automatically identifying different risk levels; Dividing the processed original data into risk levels according to the different risk levels, generating three risk levels: low, medium, and high.

4. The disease risk prediction method according to claim 3, wherein The dividing the smoothed data into multiple clusters is based on the following calculation formula: Among them, J is the objective function, C i is the i-th cluster, μ i is the center of the i-th cluster, x j is the j-th data point, and the data point belongs to the data set X, where the data set X = {x1, x2,..x j ..x K}, and K is the total number of data points in the data set X.

5. The disease risk prediction method according to claim 1, wherein The training a risk prediction model using the training set and inputting the validation set into the trained risk prediction model to obtain a risk prediction result includes: Stacking multiple layers of time-domain convolutional network modules and extracting time series features from the multiple layers of time-domain convolutional network modules; Calculating the weights of the environmental features of the multiple layers of time-domain convolutional network modules through an attention mechanism, and weighting the environmental feature weights to the multiple layers of time-domain network modules to generate weighted features; Concatenating the time series features and the weighted features and inputting them into the fully connected layer of the multiple layers of time-domain convolutional network for fusion, outputting the probability distribution of the risk level, and selecting the category with the highest probability as the prediction result.

6. The disease risk prediction method according to claim 5, wherein The extracting time series features from the multiple layers of time-domain convolutional network modules includes: performing causal convolution and dilated convolution on the data in the multiple layers of time-domain convolutional network modules.

7. The disease risk prediction method according to claim 5, wherein The calculation formula for calculating the weights of the environmental features of the multiple layers of time-domain convolutional network modules is as follows: Among them, e i is the energy of the i-th environmental feature, and α i is the attention weight of the i-th environmental feature, and n is the number of environmental features.

8. The disease prediction method according to claim 5, characterized in that, The calculation formula for outputting the probability distribution of the risk level is: where M is the total number of data points in the training set or validation set, C is the number of risk categories, y i,c is the c-th risk category label of the i-th data point, is the probability of the c-th risk category label of the i-th data point.

9. A disease risk prediction device, characterized in that, Including: An original data acquisition module for collecting the original data of preset disease cases in a preset area and performing anomaly detection and processing on the original data; A risk level division module for dividing the processed original data into risk levels based on data-driven dynamic risk level division; A risk prediction and evaluation module for dividing the original data after risk level division into a training set and a validation set, training a risk prediction model using the training set, and inputting the validation set into the trained risk prediction model to obtain a risk prediction result.