Early warning method, device, medium and electronic equipment for infectious diseases
By quantitatively processing the data of close contacts in infectious disease and training the risk assessment model, the probability of close contacts being transferred to confirmed cases is predicted, and the problem of excessive granularity of close contacts in infectious disease prevention and control is solved, and refined risk management and early warning of the tested population is achieved.
Patent Information
- Application Number
- CN202210481944.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-05
AI Technical Summary
In the prevention and control of infectious diseases, the management particle size of close contacts is too coarse, resulting in management chaos. High-risk close contacts and low-risk close contacts may cause contact, increasing the risk of secondary infection.
By obtaining the sample data set, including the close contact information and diagnosis information of the target population, quantification processing and cross-verification, training the risk assessment model, and then risk assessment of the close contact information of the tested population, predict the probability of close contact being transferred to diagnosis, and early warning.
The refined and accurate risk score and management of the people being tested is achieved, the labor and time cost of investing in the risk of illness is reduced, and the life interference to the people being tested is reduced.
Smart Images

Figure CN114743690B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to an early warning method for infectious diseases, an early warning device for infectious diseases, a computer-readable storage medium, and an electronic device. Background Art
[0002] For important diseases such as infectious diseases, it is very important to be able to develop reasonable response strategies for target populations, such as close contacts (close contacts). For example, in the current process of infectious disease prevention and control, the management of close contacts is an important foundation and key link in preventing the spread of the virus, and its success directly affects the effectiveness of infectious disease prevention and control.
[0003] In traditional infectious disease prevention and control methods, the management of close contacts is mainly divided into two levels: close contacts and secondary close contacts, and unified management is carried out according to these two risk registrations. Obviously, this method has too coarse management granularity. When the scale of infectious diseases is large and the number of close contacts is large, management chaos is prone to occur, causing high-risk close contacts to come into contact with low-risk close contacts, and then secondary infection occurs. When designing response strategies, facing hundreds or thousands of close contacts, if the response strategies are too strict, people will complain; if the response strategies are too loose, the risk of recurrence of infectious diseases will increase.
[0004] In view of this, there is an urgent need in the art to develop a new early warning method and device for infectious diseases.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The purpose of the present disclosure is to provide an early warning method for infectious diseases, an early warning device for infectious diseases, a computer-readable storage medium and an electronic device, thereby at least to a certain extent overcoming the technical problem of imprecise and inaccurate target population division caused by the limitations of related technologies.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, there is provided an early warning method for infectious diseases, the method comprising: obtaining a sample data set, the sample data set comprising close contact information of a target population and confirmed information of the target population;
[0009] Quantifying the close contact information to obtain characteristic data;
[0010] Dividing the sample data set according to a cross-validation method, and training a model according to feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model;
[0011] Verifying the characteristic data in the initialization model to obtain a risk assessment model;
[0012] The close contact information of the population to be tested is obtained, and risk assessment is performed on the close contact information of the population to be tested according to the risk assessment model to obtain the probability that the close contacts of the population to be tested will be confirmed, so as to issue an early warning for the infectious disease based on the probability that the close contacts will be confirmed.
[0013] In an exemplary embodiment of the present disclosure, after acquiring the sample data set, the method further includes:
[0014] Constructing an information database corresponding to the sample data set;
[0015] Close contact information is collected in the information database, and the close contact information includes the target population's own information and the association information between confirmed patients related to the target population.
[0016] In an exemplary embodiment of the present disclosure, dividing the sample data set according to the cross-validation method, and training the model according to the feature data and confirmed information corresponding to the divided sample data set to obtain the initialization model includes:
[0017] Using a cross-validation algorithm to divide the sample data set into a training set and a verification set, and using the feature data and the confirmed information corresponding to the training set to solve and obtain initial parameters;
[0018] The model is trained using the initial parameters to obtain an initialized model.
[0019] In an exemplary embodiment of the present disclosure, verifying the feature data in the initialization model to obtain a risk assessment model includes:
[0020] Using the characteristic data and the confirmed information corresponding to the training set and the verification set to adjust the initial parameters in the initialization model to obtain target parameters;
[0021] A risk assessment model is obtained according to the target parameter, wherein the target parameter corresponds to the feature data included in the risk assessment model.
[0022] In an exemplary embodiment of the present disclosure, after verifying the feature data in the initialization model to obtain a risk assessment model, the method further includes:
[0023] Acquire a parameter threshold value corresponding to the target parameter, and compare the target parameter with the parameter threshold value to obtain a first comparison result;
[0024] The effect of the data feature corresponding to the target parameter on the probability of close contact of the tested population being confirmed is determined based on the first comparison result.
[0025] In an exemplary embodiment of the present disclosure, the risk assessment of the close contact information of the tested population according to the risk assessment model to obtain the probability of the close contact of the tested population being confirmed includes:
[0026] Quantitatively processing the close contact information of the population to be tested to obtain characteristic data to be evaluated;
[0027] The characteristic data to be evaluated is input into the risk assessment model so that the risk assessment model outputs the probability of close contacts of the tested population being confirmed.
[0028] In an exemplary embodiment of the present disclosure, the early warning of the infectious disease according to the probability of the close contact being diagnosed includes:
[0029] Obtaining a probability threshold corresponding to the probability that the close contacts of the tested population are confirmed, and comparing the probability that the close contacts of the tested population are confirmed with the probability threshold to obtain a second comparison result;
[0030] An early warning is issued for the infectious disease according to the second comparison result.
[0031] According to one aspect of the present disclosure, there is provided an early warning device for infectious diseases, the device comprising:
[0032] A sample acquisition module is configured to acquire a sample data set, wherein the sample data set includes close contact information of the target population and confirmed information of the target population;
[0033] A quantization processing module is configured to perform quantization processing on the close contact information to obtain feature data;
[0034] A model training module is configured to divide the sample data set according to a cross-validation method, and train a model according to feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model;
[0035] A model verification module is configured to verify the characteristic data in the initialization model to obtain a risk assessment model;
[0036] The probabilistic warning module is configured to obtain the close contact information of the population to be tested, and perform risk assessment on the close contact information of the population to be tested according to the risk assessment model to obtain the probability that the close contacts of the population to be tested will be confirmed, so as to issue an early warning for the infectious disease based on the probability that the close contacts will be confirmed.
[0037] According to one aspect of the present disclosure, there is provided an electronic device, comprising: a processor and a memory; wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the early warning method for infectious diseases of any of the above exemplary embodiments is implemented.
[0038] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the early warning method for infectious diseases in any of the above exemplary embodiments is implemented.
[0039] It can be seen from the above technical solutions that the early warning method for infectious diseases, the early warning device for infectious diseases, the computer storage medium and the electronic device in the exemplary embodiments of the present disclosure have at least the following advantages and positive effects:
[0040] In the method and apparatus provided in the exemplary embodiments of the present disclosure, the sample data set is divided according to the cross-validation method, so as to train the risk assessment model using the characteristic data and confirmed information corresponding to the sample data set, thereby obtaining the characteristics that affect the risk assessment results, realizing the streamlined processing of the characteristic data, and providing data guarantee and theoretical support for obtaining a risk assessment model with good interpretability. Furthermore, the risk assessment model is used to conduct a risk assessment on the close contact information of the tested population to obtain the probability of the close contact of the tested population being confirmed, providing an automated and intelligent disease risk assessment method, which can quickly, accurately and effectively predict the probability of infection after contact with a confirmed patient in the tested population, and realize the refined and accurate risk scoring and precise management of the tested population, greatly reducing the human and time costs invested in the disease risk of the tested population, while reducing the interference with the life of the tested population.
[0041] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0043] Figure 1A schematic diagram schematically illustrates a flow chart of an early warning method for infectious diseases in an exemplary embodiment of the present disclosure;
[0044] Figure 2 A schematic diagram schematically illustrates a flow chart of a method for collecting close contact information in a sample data set in an exemplary embodiment of the present disclosure;
[0045] Figure 3 A schematic diagram schematically illustrates a flow chart of a method for training and obtaining an initialization model in an exemplary embodiment of the present disclosure;
[0046] Figure 4 A schematic diagram schematically illustrates a flow chart of a method for obtaining a risk assessment model in an exemplary embodiment of the present disclosure;
[0047] Figure 5 A flowchart schematically illustrating a method for analyzing a risk assessment model in an exemplary embodiment of the present disclosure;
[0048] Figure 6 A schematic flow chart of a method for performing risk assessment on close contact information of a group of people to be tested in an exemplary embodiment of the present disclosure is schematically shown;
[0049] Figure 7 A schematic flowchart of a method for early warning of infectious diseases according to the probability of close contacts being diagnosed in an exemplary embodiment of the present disclosure is schematically shown;
[0050] Figure 8 A schematic diagram schematically shows the structure of an early warning device for infectious diseases in an exemplary embodiment of the present disclosure;
[0051] Fig. 9 An electronic device for implementing an early warning method for infectious diseases in an exemplary embodiment of the present disclosure is schematically shown;
[0052] Fig.10 A computer-readable storage medium for implementing an early warning method for infectious diseases in an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0053] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0054] The terms "a", "an", "the" and "said" are used in this specification to indicate the presence of one or more elements / components / etc.; the terms "including" and "having" are used to express an open-ended inclusion and mean that additional elements / components / etc. may exist in addition to the listed elements / components / etc.; the terms "first" and "second" etc. are used only as labels and are not intended to limit the quantity of their objects.
[0055] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0056] For important diseases such as infectious diseases, it is very important to be able to develop reasonable response strategies for target populations, such as close contacts. For example, in the current process of infectious disease prevention and control, the management of close contacts is an important foundation and key link in preventing the spread of the virus, and its success directly affects the effectiveness of infectious disease prevention and control.
[0057] In traditional infectious disease prevention and control methods, the management of close contacts is mainly divided into two levels: close contacts and secondary close contacts. Among them, close contacts are people who have direct contact with infected people. And secondary close contacts are people who have had contact with close contacts. Then, unified management is carried out based on the two risk registrations of close contacts and secondary close contacts.
[0058] Obviously, this method has too coarse management granularity. When the scale of infectious diseases is large and the number of close contacts is large, management chaos is prone to occur, which may lead to contact between high-risk close contacts and low-risk close contacts, and then secondary infection. When designing response strategies, facing hundreds or thousands of close contacts, if the strategy is too strict, it will have a negative impact; if the response strategy is too loose, it will increase the risk of repeated transmission of infectious diseases. Therefore, how to effectively predict whether you will be infected by infectious patients is an urgent problem to be solved.
[0059] In view of the problems existing in the related technologies, the present disclosure proposes an early warning method for infectious diseases. Figure 1 A flow chart showing an early warning method for infectious diseases is shown, Figure 1 As shown, the early warning method for infectious diseases includes at least the following steps:
[0060] Step S110: Obtain a sample data set, where the sample data set includes the close contact information of the target population and the confirmed information of the target population.
[0061] Step S120: quantize the close contact information to obtain feature data.
[0062] Step S130: Divide the sample data set according to the cross-validation method, and train the model according to the feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model.
[0063] Step S140: Verify the characteristic data in the initialization model to obtain a risk assessment model.
[0064] Step S150. Obtain the close contact information of the population to be tested, and perform risk assessment on the close contact information of the population to be tested according to the risk assessment model to obtain the probability of the close contacts of the population to be tested being confirmed, so as to issue an early warning for the infectious disease based on the probability of the close contacts being confirmed.
[0065] In an exemplary embodiment of the present disclosure, a sample data set is divided according to a cross-validation method, and the feature data and confirmed information corresponding to the sample data set are used to train a risk assessment model, thereby obtaining features that affect the risk assessment results, achieving streamlined processing of feature data, and providing data assurance and theoretical support for obtaining a risk assessment model with good interpretability. Furthermore, a risk assessment model is used to conduct a risk assessment on the close contact information of the tested population to obtain the probability of the close contact of the tested population being confirmed, providing an automated and intelligent disease risk assessment method, which quickly, accurately and effectively predicts the probability of infection after contact with a confirmed patient in the tested population, and achieves refined and accurate risk scoring and precise management of the tested population, greatly reducing the human and time costs invested in the risk of disease in the tested population, while reducing interference with the lives of the tested population.
[0066] The following is a detailed explanation of each step of the early warning method for infectious diseases.
[0067] In step S110, a sample data set is obtained, where the sample data set includes close contact information of the target population and confirmed information of the target population.
[0068] In an exemplary embodiment of the present disclosure, the risk assessment model may be established based on historical population data, that is, a sample data set. Specifically, multiple historical data sets are obtained to determine the sample data set.
[0069] When the target population is a group of people who have close contact with an infectious disease, the multiple historical data sets may be data sets corresponding to three historical stages.
[0070] The first can be the close contact data collected in the current infectious disease, the second can be the close contact data accumulated in the last infectious disease in this city, and the third can be the close contact data accumulated in similar infectious diseases in other cities.
[0071] In order to filter out a sample data set based on multiple historical data sets, the number of close contacts of this infectious disease and the number of close contacts of this infectious disease who have turned positive can be counted based on the close contact data collected in the first type of infectious disease.
[0072] Get the threshold number of confirmed cases corresponding to the number of close contacts of this infectious disease who turned positive.
[0073] Among them, the threshold value of the number of confirmed cases can be set to 10, or other threshold values can be set according to actual conditions and needs. This exemplary embodiment does not make any special limitations on this.
[0074] If the number of close contacts of this infectious disease who turn positive is greater than or equal to the threshold number of confirmed cases, the close contact data collected in this infectious disease are determined as the sample data set from multiple historical data sets.
[0075] When the number of close contacts of this infectious disease who turned positive is 15 and the confirmation number threshold is 10, the number of close contacts of this infectious disease who turned positive is greater than the corresponding confirmation number threshold. Therefore, the close contact data collected in this infectious disease in multiple historical data sets can be determined as the sample data set.
[0076] Therefore, when this infectious disease has produced a large number of close contacts, some of them may turn positive. The close contact data collected in this infectious disease can be used as the sample data set.
[0077] If the number of close contacts of this infectious disease who turned positive is less than the confirmed number threshold, and the total number of close contacts of this infectious disease is greater than or equal to the corresponding number threshold, the close contact data accumulated in the last infectious disease in this city are determined as the sample data set from multiple historical data sets.
[0078] When the number of close contacts of this infectious disease who turned positive is 5, the confirmation number threshold is 10, and the total number of close contacts of this infectious disease is 15, and the corresponding number threshold is 10, it is determined that the number of close contacts of this infectious disease who turned positive is less than the confirmation number threshold, and the total number of close contacts of this infectious disease is greater than or equal to the corresponding number threshold. Therefore, the close contact data accumulated in the last infectious disease in this city in multiple historical data sets can be determined as the sample data set.
[0079] Among them, the close contact data accumulated during the last infectious disease in the city may include the close contact data accumulated during the last infectious disease in the city, or the close contact data accumulated during a round of infectious diseases in the previous rounds in the city.
[0080] If the total number of close contacts of this infectious disease is less than the corresponding quantity threshold, the close contact data accumulated in similar infectious diseases in other cities are determined as a sample data set from multiple historical data sets.
[0081] When the total number of close contacts of this infectious disease is 5 and the corresponding quantity threshold is 10, it is determined that the total number of close contacts of this infectious disease is less than the corresponding quantity threshold. Therefore, the close contact data accumulated in similar infectious diseases in other cities in multiple historical data sets can be determined as the sample data set.
[0082] Among them, the sample data set may include close contact data accumulated in similar infectious diseases in other cities.
[0083] By determining the sample data set by statistically analyzing multiple historical data sets and comparing the corresponding thresholds, we can reasonably select the data source based on the current development of the disease, provide sufficient and accurate sample data sets for subsequent risk assessment of the population to be tested, and ensure the accuracy of disease risk assessment.
[0084] It is worth noting that the sample data set refers to the characteristic data corresponding to the people who had close contact with confirmed patients during the historical period. The sample data set includes the characteristic data corresponding to the patients who were diagnosed after contact with confirmed patients, and the characteristic data corresponding to the patients who were not diagnosed after contact with confirmed patients.
[0085] After the sample data set is obtained, close contact information may be collected from the sample data set.
[0086] In an alternative embodiment, Figure 2A flow chart of a method for collecting close contact information in a sample data set is shown, Figure 2 As shown, the method includes at least the following steps: In step S210, an information database corresponding to the sample data set is constructed.
[0087] Generally speaking, in order to manage close contacts, the disease control department will establish a corresponding information database for the sample data set of close contacts, and build the data set required for the risk assessment model based on the data in the information database. In other words, the information database is also the close contact information database.
[0088] In step S220, close contact information is collected in the information database, where the close contact information includes the target population's own information and the association information between confirmed patients related to the target population.
[0089] Close contact information can be collected from the established information database. Close contact information is mainly divided into two categories: one is the target population (close contact population)’s own information, such as age, gender, etc.; the other is the association information between the target population and the confirmed patients they have contacted, such as contact type, contact frequency, etc.
[0090] Specifically, close contact information may include gender, age, close contact / secondary close contact, contact type, relationship type with case / close contact, number of days since the last contact date, number of close contacts associated with associated cases, contact duration / minutes, contact frequency, contact protection / whether to wear a mask, whether the test turned positive, etc.
[0091] Among them, gender can include male, female, and unknown.
[0092] Age can be divided into 0-7, 7-22, 22-50, and 50+.
[0093] Contact types may include sharing meals, sharing a room / bed, living together, recreational activities (traveling together / participating in recreational activities together), working or studying in the same room, sharing the same plane / car / vehicle, diagnosis and treatment, and others (e.g., sharing the same time and space).
[0094] Relationship types with cases / close contacts can include family members, relatives, friends, colleagues, neighbors, and others.
[0095] The number of days from the last contact date to today is set between 0 and 30; the number of close contacts associated with associated cases is divided into 0-10, 10-20, 20-50, 50-100, 100-200, 200-500, 500-1000, 1000-2000, 2000-5000 and 5000-10000.
[0096] The contact time / minutes are divided into within 1 minute, within 10 minutes, within 1 hour, within 10 hours and more than 10 hours.
[0097] The contact frequency was divided into occasional, general and frequent.
[0098] Contact protection / whether to wear a mask is divided into yes and no.
[0099] Whether the test result has turned positive can also be divided into yes or no. Whether the test result has turned positive is the confirmed information in this application.
[0100] In addition, we can also adapt to local conditions and add more close contact information based on whether the target population has been to a high-risk place, is a member of a risk unit, or is engaged in a high-risk job, so as to facilitate the accurate collection of close contact information in different situations.
[0101] In this exemplary embodiment, by constructing an information database to collect close contact information, the management of sample data sets is facilitated, and data resources can also be provided for collecting close contact information or other data processing tasks, providing a convenient method for various data processing tasks and enriching the application scenarios of the information database.
[0102] In step S120, the close contact information is quantized to obtain feature data.
[0103] In an exemplary embodiment of the present disclosure, after the close contact information is collected, the close contact information can be quantified to obtain feature data.
[0104] Specifically, the close contact information may be quantized in the manner shown in Table 1.
[0105] Table 1
[0106]
[0107]
[0108]
[0109] For example, in the confirmed information corresponding to the collected close contact information, the positive conversion mark of the target population is 1, and the non-positive conversion mark is 0. The processing methods of other close contact information are shown in the value range and mapping value of Table 1.
[0110] In step S130, the sample data set is divided according to the cross-validation method, and the model is trained according to the feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model.
[0111] In an exemplary embodiment of the present disclosure, after the feature data is processed, a cross-validation process may be performed on the feature data to obtain an initialization model.
[0112] In an alternative embodiment, Figure 3A flow chart of the method for training and obtaining an initialization model is shown, such as Figure 3 As shown, the method includes at least the following steps: in step S310, a cross-validation algorithm is used to divide the sample data set into a training set and a verification set, and the characteristic data and confirmed information corresponding to the training set are used to solve and obtain initial parameters.
[0113] In order to make the constructed risk assessment model accurate, the characteristic data and the confirmed information can be verified and evaluated. Furthermore, in order to ensure the fairness of the verification and evaluation, a multi-fold cross-validation method can be used.
[0114] Generally, K-fold cross validation is used for model tuning to find the hyperparameter values that optimize the model's generalization performance. Once found, the model is retrained on the entire training set, and the model performance is finally evaluated using an independent test set.
[0115] K-fold cross validation uses the advantage of non-repeated sampling technology, that is, each sample point has only one chance to be included in the training set or test set in each iteration.
[0116] If the training data set is relatively small, increase the K value. Increasing the K value will result in more data being used for model training in each iteration, which can achieve the minimum deviation. At the same time, the algorithm time will be extended. In addition, the training blocks are highly similar, resulting in a higher variance in the evaluation results.
[0117] If the training set is relatively large, reduce the value of K. Reducing the value of K reduces the computational cost of evaluating the performance of the model by repeatedly fitting it on different data blocks, and obtains an accurate evaluation of the model based on the average performance.
[0118] Specifically, the specific steps of K-fold cross validation can be: the first step is to divide the original data set into K equal parts ("folds"); the second step is to use the first part as the test set and the rest as the training set; the third step is to train the model and calculate the accuracy of the model on the test set; the fourth step is to use a different part as the test set each time, and repeat steps 2 and 3K times; the fifth step is to use the average accuracy as the final model accuracy.
[0119] For example, the multi-fold cross-check may be a 5-fold cross-check.
[0120] The obtained sample data set is divided into 5 parts according to the 5-fold cross-validation method, in which the characteristic data and confirmed information corresponding to one sample data set are used as the verification set, and the characteristic data and confirmed information corresponding to the remaining 4 sample data sets are used as the training set.
[0121] It is worth mentioning that when performing the 5-fold cross-validation data division, the positive samples that have turned positive should be evenly divided into each verification set and training set.
[0122] In machine learning, you can almost always see an extra term added to the loss function. There are two common extra terms, generally called l 1 -norm and l 2 -norm, called L1 regularization and L2 regularization, or L1 norm and L2 norm in Chinese.
[0123] L1 regularization and L2 regularization can be regarded as penalty terms in the loss function. The so-called "penalty" refers to some restrictions on certain parameters in the loss function.
[0124] For the linear regression model, the model using L1 regularization is called Lasso regression, and the model using L2 regularization is called Ridge regression.
[0125] When the regularization parameter selects the L1 regularized model, the loss function of Lasso regression is shown in formula (1):
[0126]
[0127] In formula (1), the term after the plus sign is α‖w‖ 1 This is the L1 regularization term.
[0128] In general regression analysis, w represents the coefficient of the feature. From the above formula, we can see that the regularization term processes (restricts) the coefficient. L1 regularization refers to the sum of the absolute values of each element in the weight vector w, usually expressed as ‖w‖ 1 .
[0129] Generally, a coefficient is added before the regularization term, which can be represented by α or λ.
[0130] When the regularization function selects L1 regularization, L1 regularization can produce a sparse weight matrix, that is, a sparse model that can be used for feature selection.
[0131] Among them, a sparse matrix refers to a matrix in which many elements are 0 and only a few elements are non-zero values, that is, most of the coefficients of the obtained linear regression model are 0.
[0132] Usually, there are many features in machine learning. For example, when processing text, if a phrase is used as a feature, the number of features will reach tens of thousands. When predicting or classifying, too many features are obviously difficult to select. However, if the model obtained by substituting these features is a sparse model, it means that only a few features contribute to the model, and most of the features do not contribute, or contribute very little (because the coefficients in front of them are 0 or very small values, and even if they are removed, there is no effect on the model). At this time, we can only focus on features with non-zero coefficients. This is the relationship between sparse models and feature selection.
[0133] Therefore, L1 regularization can be used to simplify data features and remove the interference of redundant fields to obtain a risk assessment model with good interpretability.
[0134] In the process of solving the regularization function of L1 regularization using the training set, the solution of L1 regularization can be stopped when the conditions of this solution are met.
[0135] Specifically, the loss function with L1 regularization is shown in formula (2):
[0136] J=J 0 +α∑ w |w| (2)
[0137] Among them, J 0 is the original loss function, the item after the plus sign is the L1 regularization term, and α is the regularization coefficient.
[0138] Note that L1 regularization is the absolute value sum of the weights, and J is a function with an absolute value sign, so J is not completely differentiable.
[0139] The task of machine learning is to use some methods, such as gradient descent, to find the minimum value of the loss function.
[0140] When the original loss function J 0 When the L1 regularization term is added later, it is equivalent to the original loss function J 0 A constraint is made. Let L = α∑ w |w|, then J=J 0 +L, at this time, the original loss function J can be obtained under the constraint of L 0 Take the solution with the minimum value.
[0141] When the condition for ending the solution of the regularization function, that is, formula (2), is not met, the regularization coefficient α is adjusted, and the data features in the training set are again substituted into the L1 regularization function to obtain the loss value corresponding to the adjusted regularization coefficient α, and it is judged again whether the condition for ending the solution of the regularization function is met. This is repeated until the condition for ending the solution of the regularization function is met, and the solution is ended to obtain the initial parameters. Among them, the initial parameters are the regularization coefficient α for finally ending the solution of the regularization function.
[0142] In the process of 5-fold cross-validation, the sample data set is divided 5 times according to the ratio of 4:1 between the training set and the test set to obtain 5 training sets. Therefore, 5 sets of initial parameters can be obtained according to the characteristic data and confirmed information corresponding to the 5 training sets.
[0143] In step S320, the model is trained using the initial parameters to obtain an initialized model.
[0144] After solving 5 sets of initial parameters, the model can be trained using the initial parameters to obtain an initialized model.
[0145] Specifically, five groups of initial parameters may be substituted into the established model.
[0146] The model may be a logistic regression model, which is shown in formula (3):
[0147]
[0148] Among them, y is the label of whether it turns positive, x is the data feature, w is the initial parameter, and b is the bias term.
[0149] Then, the initial parameters are substituted into the model to obtain the initialized model. Therefore, 5 initialized models corresponding to 5 sets of initial parameters can be obtained respectively.
[0150] In step S140, the characteristic data in the initialization model is verified to obtain a risk assessment model.
[0151] In an exemplary embodiment of the present disclosure, after the initialization model is obtained through training, the feature data in the initialization model may be further verified to obtain a final risk assessment model.
[0152] In an alternative embodiment, Figure 4 A flow chart of the method for obtaining a risk assessment model is shown, such as Figure 4 As shown, the method may at least include the following steps: In step S410, the initial parameters in the initialization model are adjusted using the characteristic data and confirmed information corresponding to the training set and the verification set to obtain the target parameters.
[0153] Furthermore, the five initialization models are required to output the probability results of close contacts of the target population in the corresponding training set and verification set being confirmed, and the initial parameters in the probability results of close contacts of the target population being confirmed that are closest to the actual situation of the historical close contacts are determined as the target parameters.
[0154] For example, when the five groups of initial parameters are 0.1, 0.3, 1, 3 and 10, 0.1, 0.3, 1, 3 and 10 can be substituted into formula (3) respectively to obtain five groups of initialization models with different initial parameters.
[0155] In addition, the characteristic data corresponding to the close contact information in the sample data set are respectively input into 5 groups of initialization models with different initial parameters, so that the 5 groups of initialization models with different initial parameters can respectively output the probability results of close contacts of the 5 target populations being confirmed.
[0156] Since the confirmed information of close contacts in the sample data set is known, the probability results of the close contacts of the five target populations being confirmed can be compared with the confirmed information of the corresponding close contacts to determine the initial parameters corresponding to the probability results of the close contacts of the five target populations being confirmed that best fit the patient's condition as the target parameters. At this time, the target parameter is the optimal L1 regularization parameter.
[0157] In this exemplary embodiment, the optimal target parameters can be obtained by adjusting the initial parameters in the initial model, which provides a parameter basis for obtaining a risk assessment model with good interpretability.
[0158] In step S420, a risk assessment model is obtained according to target parameters, where the target parameters correspond to feature data included in the risk assessment model.
[0159] After determining the target parameters, the risk assessment model can be constructed using the target parameters.
[0160] Specifically, an initialization model corresponding to a target parameter is determined, and a risk assessment model is constructed through the target parameter, where the target parameter corresponds to feature data included in the risk assessment model.
[0161] The initialization model corresponding to the target parameter can be the established logistic regression model formula (3) without determining the target parameter. Therefore, after determining the target parameter, the target parameter can be substituted into formula (3) to construct a risk assessment model. At this time, w in the risk assessment model is the target parameter.
[0162] It is worth noting that since the risk assessment model can be a well-established logistic regression model with determined target parameters, w represents the regression coefficient. Moreover, in the risk assessment model, different data features correspond to a regression coefficient. Therefore, the target parameter is a set of regression coefficients.
[0163] After the risk assessment model is constructed, the risk assessment model can be saved and the regression coefficients of the risk assessment model can be analyzed to provide a deeper understanding of the spread of infectious diseases.
[0164] In an alternative embodiment, Figure 5 A flow chart of a method for analyzing a risk assessment model is shown, such as Figure 5 As shown, the method includes at least the following steps: in step S510, a parameter threshold corresponding to a target parameter is obtained, and the target parameter is compared with the parameter threshold to obtain a first comparison result.
[0165] Generally, the parameter threshold may be set to 0, or other values may be set according to actual conditions and requirements, and this exemplary embodiment does not specifically limit this.
[0166] After the parameter threshold is determined, the target parameter may be compared with the parameter threshold to obtain a first comparison result.
[0167] When the target parameter includes a set of regression coefficients, each regression coefficient may be compared to the parameter threshold.
[0168] In step S520, the effect of the data feature corresponding to the target parameter on the probability of close contact with the tested population being confirmed is determined based on the first comparison result.
[0169] If the first comparison result is that the target parameter is greater than the parameter threshold, it is determined that the data feature corresponding to the target parameter has a positive effect on the probability of close contact of the tested population being confirmed.
[0170] When the first comparison result of a regression coefficient in the target parameter and the parameter threshold is that the regression coefficient is greater than the parameter threshold, it indicates that the data characteristics corresponding to the regression coefficient have a positive effect on the probability of close contacts of the tested population being confirmed.
[0171] For example, when the regression coefficient corresponding to the data feature of whether or not to wear a mask is greater than the parameter threshold, it indicates that whether or not to wear a mask will have a positive effect on the probability of close contacts of the tested population being diagnosed.
[0172] If the first comparison result is that the target parameter is equal to the parameter threshold, it is determined that the data feature corresponding to the target parameter has no effect on the probability of close contact of the tested population being confirmed.
[0173] When the first comparison result of a regression coefficient in the target parameter and the parameter threshold is that the regression coefficient is equal to the parameter threshold, it indicates that the data characteristics corresponding to the regression coefficient have no effect on the probability of close contact of the tested population being confirmed.
[0174] When the regression coefficient corresponding to the data feature of gender is equal to the parameter threshold, it indicates that gender will not have any effect on the probability of close contacts of the tested population being confirmed.
[0175] If the first comparison result is that the target parameter is less than the parameter threshold, it is determined that the data feature corresponding to the target parameter has an adverse effect on the probability of close contact with the tested population being confirmed.
[0176] When the first comparison result of a regression coefficient in the target parameter and the parameter threshold is that the regression coefficient is less than the parameter threshold, it indicates that the data characteristics corresponding to the regression coefficient will have a reverse effect on the probability of close contact of the tested population being confirmed.
[0177] In this exemplary embodiment, by comparing the regression coefficient in the target parameter with the corresponding parameter threshold, the effect of data characteristics on disease risk outcomes can be determined, thereby improving the interpretability of the risk assessment model and helping personnel to form a deeper understanding of the spread of infectious diseases.
[0178] In addition, if the first comparison result is that the target parameter is less than the parameter threshold, it may be determined that the data feature corresponding to the target parameter has no effect on the probability of close contact with the tested population being confirmed.
[0179] When the first comparison result of a regression coefficient in the target parameter and the parameter threshold is that the regression coefficient is less than the parameter threshold, it indicates that the data characteristics corresponding to the regression coefficient will not have any effect on the probability of close contact of the tested population being confirmed.
[0180] Obviously, after analyzing the risk assessment model, managers can have a deeper understanding of what kind of contact methods are more likely to cause virus transmission, what kind of patients or scenarios are more likely to cause virus transmission, etc.
[0181] Moreover, when managers find that there are unreasonable aspects in the data features of the risk assessment model, they can also adjust or delete the data features to improve the generalization ability of the risk assessment model, making the risk assessment model more reasonable and accurate.
[0182] It is worth noting that when the risk assessment model is other logistic regression models or other models, the effect of data characteristics on disease risk results can also be interpreted according to the corresponding target parameters. This exemplary embodiment does not specifically limit this.
[0183] Furthermore, the risk assessment model can be used to predict the probability of close contacts becoming confirmed cases in the tested population.
[0184] In step S150, the close contact information of the population to be tested is obtained, and the risk assessment is performed on the close contact information of the population to be tested according to the risk assessment model to obtain the probability of the close contacts of the population to be tested being confirmed, so as to issue an early warning for the infectious disease based on the probability of the close contacts being confirmed.
[0185] In an exemplary embodiment of the present disclosure, after the risk assessment model is constructed, the risk assessment model can be used to perform risk assessment on the population to be tested.
[0186] In an alternative embodiment, Figure 6 A flow chart of a method for risk assessment of close contact information of a population to be tested is shown, Figure 6 As shown, the method at least includes the following steps: In step S610, the close contact information of the population to be tested is quantified to obtain characteristic data to be evaluated.
[0187] When the close contact information of the population to be tested is collected, the close contact information of the population to be tested can be quantified in the manner shown in Table 1 to obtain the characteristic data to be evaluated, which will not be repeated here.
[0188] In step S620, the characteristic data to be evaluated is input into the risk assessment model so that the risk assessment model outputs the probability of close contacts of the tested population being confirmed.
[0189] Furthermore, the characteristic data to be evaluated is input into the risk assessment model, and the risk assessment model can output the assessed risk scores of all close contacts of the infectious disease (the population to be tested) as the probability that the close contacts of the population to be tested will be confirmed.
[0190] In this exemplary embodiment, the disease risk results of the population to be tested can be obtained by quantitatively processing and risk evaluating the data to be evaluated, which improves the efficiency and accuracy of risk evaluation and provides data support for refined risk management.
[0191] After determining the probability of close contacts of the tested population being diagnosed, an early warning for the infectious disease can be determined based on the probability of close contacts being diagnosed.
[0192] In an alternative embodiment, Figure 7 A flow chart of a method for early warning of infectious diseases based on the probability of close contacts being confirmed is shown, as Figure 7 As shown, the method includes at least the following steps: in step S710, a probability threshold corresponding to the probability of close contacts of the population to be tested being confirmed is obtained, and the probability of close contacts of the population to be tested being confirmed is compared with the probability threshold to obtain a second comparison result.
[0193] Generally, the probability threshold may be 0.8, and probability thresholds of other values may also be set according to actual conditions and requirements, which is not particularly limited in this exemplary embodiment.
[0194] After obtaining the probability threshold, the probability of close contacts of the tested population being confirmed can be compared with the probability threshold to obtain a second comparison result.
[0195] In step S720, an early warning of the infectious disease is issued according to the second comparison result.
[0196] If the second comparison result is that the probability of close contacts of the tested population being confirmed is greater than or equal to the probability threshold, an early warning can be issued for the infectious disease. For example, based on the second comparison result, if the number of people whose probability of close contacts of the tested population being confirmed is greater than or equal to the probability threshold exceeds the number threshold, it is determined that the first measure will be adopted to issue an early warning for the infectious disease.
[0197] When the disease risk result is 0.9 and the risk threshold is 0.8, the second comparison result at this time is that the disease risk result is greater than the risk threshold, and it can be determined that the population to be tested is a high-risk population.
[0198] Specifically, the first measure may include completing as complete an epidemiological investigation as possible on the high-risk population as quickly as possible, and may also include expanding the scope of determination of secondary close contacts of the high-risk population, and may also include separately isolating and transferring the high-risk population, and may also include increasing the isolation time for the high-risk population and other measures.
[0199] Among them, expanding the scope of determining secondary close contacts of the high-risk group can be achieved by classifying the people in the entire building where the high-risk group lives as secondary close contacts.
[0200] In addition, separate isolation and transfer of this high-risk group can reduce the risk of transmission during isolation and transfer, and increasing the isolation time for this high-risk group can effectively respond to the emergence of patients in the incubation period.
[0201] If the second comparison result shows that the probability of close contacts of the tested population being confirmed is lower than the risk threshold, the second measure will be adopted.
[0202] When the disease risk result is 0.4 and the risk threshold is 0.8, the second comparison result is that the disease risk result is less than the risk threshold, and the population to be tested can be determined to be a low-risk population. Therefore, the second measure is taken for the low-risk population.
[0203] Specifically, the second measure may include reducing centralized isolation and adopting home isolation to reduce interference with the normal lives of low-risk groups.
[0204] Therefore, a simplified close contact risk score table can be formed with the risk threshold as the interval. For example, the close contact risk score table includes the level of 0-0.8 and the level of 0.8-1, which can help epidemic prevention personnel make risk judgments in the first time, improve the efficiency of taking measures against the tested population, and save the time cost invested in managing the tested population.
[0205] In this exemplary embodiment, through the second comparison result of the probability of close contact of the tested population being confirmed and the probability threshold, personalized treatment measures can be taken for different tested populations, thereby focusing on many links such as transfer, isolation, epidemiological investigation, nucleic acid testing and lifting of isolation for the tested population, and rationally planning epidemic prevention resources in response to a large number of tested populations, thereby ensuring the effectiveness of epidemic prevention and reducing interference with people's lives.
[0206] In the early warning method for infectious diseases in the exemplary embodiment of the present disclosure, the sample data set is divided according to the cross-validation method, so as to use the characteristic data and confirmed information corresponding to the sample data set to train the risk assessment model, thereby obtaining the characteristics that affect the risk assessment results, realizing the streamlined processing of the characteristic data, and providing data guarantee and theoretical support for obtaining a risk assessment model with good interpretability.
[0207] Furthermore, the risk assessment model is used to conduct risk assessment on the close contacts of the tested population to obtain the probability of the close contacts of the tested population being confirmed, providing an automated and intelligent disease risk assessment method, which can quickly, accurately and effectively predict the high-risk and low-risk populations in the tested population, and focus on the high-risk and low-risk populations in multiple links of transfer, isolation, epidemiological investigation, testing and release of isolation. When dealing with a large number of close contacts, it is necessary to rationally plan epidemic prevention resources, conduct refined risk scoring and precise management, reduce the contact between high-risk close contacts and low-risk close contacts, and thus reduce the occurrence of secondary infection. Therefore, the human cost and time cost of the risk of disease in the tested population are greatly reduced, while reducing the interference with the lives of the tested population.
[0208] In addition, in an exemplary embodiment of the present disclosure, an early warning device for infectious diseases is also provided. Figure 8 The schematic diagram of the structure of the early warning device for infectious diseases is shown in FIG. Figure 8 As shown, the early warning device 800 for infectious diseases may include: a sample acquisition module 810, a quantification processing module 820, a model training module 830, a model verification module 840 and a probability early warning module 850. Among them:
[0209] The sample acquisition module 810 is configured to acquire a sample data set, wherein the sample data set includes close contact information of the target population and confirmed information of the target population;
[0210] The quantization processing module 820 is configured to perform quantization processing on the close contact information to obtain feature data;
[0211] The model training module 830 is configured to divide the sample data set according to a cross-validation method, and train the model according to the feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model;
[0212] A model verification module 840 is configured to verify the feature data in the initialization model to obtain a risk assessment model;
[0213] The probability warning module 850 is configured to obtain the close contact information of the population to be tested, and perform risk assessment on the close contact information of the population to be tested according to the risk assessment model to obtain the probability that the close contacts of the population to be tested will be confirmed, so as to issue a warning for the infectious disease based on the probability that the close contacts will be confirmed.
[0214] In an exemplary embodiment of the present disclosure, after acquiring the sample data set, the method further includes:
[0215] Constructing an information database corresponding to the sample data set;
[0216] Close contact information is collected in the information database, and the close contact information includes the target population's own information and the association information between confirmed patients related to the target population.
[0217] In an exemplary embodiment of the present disclosure, dividing the sample data set according to the cross-validation method, and training the model according to the feature data and confirmed information corresponding to the divided sample data set to obtain the initialization model includes:
[0218] Using a cross-validation algorithm to divide the sample data set into a training set and a verification set, and using the feature data and the confirmed information corresponding to the training set to solve and obtain initial parameters;
[0219] The model is trained using the initial parameters to obtain an initialized model.
[0220] In an exemplary embodiment of the present disclosure, verifying the feature data in the initialization model to obtain a risk assessment model includes:
[0221] Using the characteristic data and the confirmed information corresponding to the training set and the verification set to adjust the initial parameters in the initialization model to obtain target parameters;
[0222] A risk assessment model is obtained according to the target parameter, wherein the target parameter corresponds to the feature data included in the risk assessment model.
[0223] In an exemplary embodiment of the present disclosure, after verifying the feature data in the initialization model to obtain a risk assessment model, the method further includes:
[0224] Acquire a parameter threshold value corresponding to the target parameter, and compare the target parameter with the parameter threshold value to obtain a first comparison result;
[0225] The effect of the data feature corresponding to the target parameter on the probability of close contact of the tested population being confirmed is determined based on the first comparison result.
[0226] In an exemplary embodiment of the present disclosure, the risk assessment of the close contact information of the tested population according to the risk assessment model to obtain the probability of the close contact of the tested population being confirmed includes:
[0227] Quantitatively processing the close contact information of the population to be tested to obtain characteristic data to be evaluated;
[0228] The characteristic data to be evaluated is input into the risk assessment model so that the risk assessment model outputs the probability of close contacts of the tested population being confirmed.
[0229] In an exemplary embodiment of the present disclosure, the early warning of the infectious disease according to the probability of the close contact being diagnosed includes:
[0230] Obtaining a probability threshold corresponding to the probability that the close contacts of the tested population are confirmed, and comparing the probability that the close contacts of the tested population are confirmed with the probability threshold to obtain a second comparison result;
[0231] An early warning is issued for the infectious disease according to the second comparison result.
[0232] The specific details of the above-mentioned early warning device 800 for infectious diseases have been described in detail in the corresponding early warning method for infectious diseases, so they will not be repeated here.
[0233] It should be noted that, although several modules or units of the early warning device 800 for infectious diseases are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0234] In addition, in an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.
[0235] Refer to the following Fig. 9 An electronic device 900 according to such an embodiment of the present invention will be described. Fig. 9 The electronic device 900 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0236] like Fig. 9As shown, the electronic device 900 is in the form of a general computing device. The components of the electronic device 900 may include, but are not limited to: the at least one processing unit 910, the at least one storage unit 920, a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910), and a display unit 940.
[0237] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification.
[0238] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922 , and may further include a read-only memory unit (ROM) 923 .
[0239] The storage unit 920 may also include a program / utility 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0240] Bus 930 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0241] The electronic device 900 may also communicate with one or more external devices 1100 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 950. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 960. As shown, the network adapter 940 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0242] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the embodiment of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiment of the present disclosure.
[0243] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible embodiments, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to perform the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of the present specification.
[0244] refer to Fig.10 As shown, a program product 1000 for implementing the above method according to an embodiment of the present invention is described, which can adopt a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, a readable storage medium can be any tangible medium containing or storing a program, which can be used by or in combination with an instruction execution system, an apparatus or a device.
[0245] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0246] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0247] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0248] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0249] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A method for early warning of infectious diseases, characterized in that: The method comprises: Acquire a historical data set to determine a sample data set, wherein the sample data set includes close contact information of a target population and confirmed information of the target population, wherein the target population includes close contacts of an infectious disease, and the confirmed information of the target population is used to indicate whether the target population is confirmed, and the historical data set includes one or more of close contact data collected for the current infectious disease in the target area, close contact data accumulated in the previous infectious disease in the target area, and close contact data accumulated in infectious diseases in other areas different from the target area; Quantifying the close contact information to obtain characteristic data; Dividing the sample data set according to a cross-validation method, and training a model according to feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model; Verifying the characteristic data in the initialization model to obtain a risk assessment model; Acquire the close contact information of the population to be tested, and perform risk assessment on the close contact information of the population to be tested according to the risk assessment model to obtain the probability that the close contact of the population to be tested will be confirmed, so as to issue an early warning for the infectious disease according to the probability that the close contact will be confirmed; The early warning of the infectious disease according to the probability of the close contact being confirmed includes: obtaining a probability threshold corresponding to the probability of the close contact of the tested population being confirmed, and comparing the probability of the close contact of the tested population being confirmed with the probability threshold to obtain a second comparison result; when the second comparison result is that the number of people whose probability of the close contact of the tested population being confirmed is greater than or equal to the probability threshold exceeds the number threshold, taking a first prevention and control measure for the tested population to issue an early warning of the infectious disease; When the second comparison result is that the probability of close contacts of the tested population being confirmed is less than the probability threshold, the second prevention and control measure is taken for the tested population.
2. The early warning method for infectious diseases according to claim 1, characterized in that: After obtaining the sample data set, the method further includes: Constructing an information database corresponding to the sample data set; Close contact information is collected in the information database, and the close contact information includes the target population's own information and the association information between confirmed patients related to the target population.
3. The early warning method for infectious diseases according to claim 1, characterized in that: The method of dividing the sample data set according to the cross-validation method and training the model according to the feature data and confirmed information corresponding to the divided sample data set to obtain the initialization model includes: Using a cross-validation algorithm to divide the sample data set into a training set and a verification set, and using the feature data and the confirmed information corresponding to the training set to solve and obtain initial parameters; The model is trained using the initial parameters to obtain an initialized model.
4. The early warning method for infectious diseases according to claim 3, characterized in that: The verifying the characteristic data in the initialization model to obtain a risk assessment model includes: Using the characteristic data and the confirmed information corresponding to the training set and the verification set to adjust the initial parameters in the initialization model to obtain target parameters; A risk assessment model is obtained according to the target parameter, wherein the target parameter corresponds to the feature data included in the risk assessment model.
5. The early warning method for infectious diseases according to claim 4, characterized in that: After verifying the characteristic data in the initialization model to obtain a risk assessment model, the method further includes: Acquire a parameter threshold value corresponding to the target parameter, and compare the target parameter with the parameter threshold value to obtain a first comparison result; The manner in which the data feature corresponding to the target parameter affects the probability of close contacts of the tested population being confirmed is determined based on the first comparison result.
6. The early warning method for infectious diseases according to claim 1, characterized in that: The risk assessment of the close contact information of the tested population according to the risk assessment model to obtain the probability of the close contact of the tested population being confirmed includes: Quantitatively processing the close contact information of the population to be tested to obtain characteristic data to be evaluated; The characteristic data to be evaluated is input into the risk assessment model so that the risk assessment model outputs the probability of close contacts of the tested population being confirmed.
7. An early warning device for infectious diseases, characterized in that: include: A sample acquisition module is configured to acquire a historical data set to determine a sample data set, wherein the sample data set includes close contact information of a target population and confirmed information of the target population, wherein the target population includes close contact population of an infectious disease, and the confirmed information of the target population is used to characterize whether the target population is confirmed, and the historical data set includes one or more of close contact data collected for the current infectious disease in the target area, close contact data accumulated in the last infectious disease in the target area, and close contact data accumulated in infectious diseases in other areas different from the target area; A quantization processing module is configured to perform quantization processing on the close contact information to obtain feature data; A model training module is configured to divide the sample data set according to a cross-validation method, and train a model according to feature data and confirmed information corresponding to the divided sample data set to obtain an initialized model; A model verification module is configured to verify the characteristic data in the initialization model to obtain a risk assessment model; A probability warning module is configured to obtain the close contact information of the tested population, and perform risk assessment on the close contact information of the tested population according to the risk assessment model to obtain the probability that the close contact of the tested population will be confirmed, so as to issue an early warning for the infectious disease according to the probability that the close contact will be confirmed; The early warning of the infectious disease according to the probability of the close contact being confirmed includes: obtaining a probability threshold corresponding to the probability of the close contact of the tested population being confirmed, and comparing the probability of the close contact of the tested population being confirmed with the probability threshold to obtain a second comparison result; when the second comparison result is that the number of people whose probability of the close contact of the tested population being confirmed is greater than or equal to the probability threshold exceeds the number threshold, taking a first prevention and control measure for the tested population to issue an early warning of the infectious disease; When the second comparison result is that the probability of close contacts of the tested population being confirmed is less than the probability threshold, the second prevention and control measure is taken for the tested population.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a transmitter, the early warning method for infectious diseases described in any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: include: Transmitter; A memory, used to store executable instructions of the transmitter; Wherein, the transmitter is configured to execute the early warning method for infectious diseases described in any one of claims 1-6 by executing the executable instructions.
Citation Information
Patent Citations
Major infectious disease propagation risk early warning and prevention and control analysis system for COVID-19
CN111863271A
Infectious disease propagation scale simulation method and device and electronic equipment
CN112365998A