Disease information management method and system based on multi-source data

Through multi-source data management methods and comprehensive platforms, infectious disease and chronic disease data are integrated for risk assessment and trend prediction, solving the problems of untimely detection and imperfect management in the existing technology, achieving more timely and accurate disease detection and early warning, and improving the efficiency of infectious disease prevention and control and the effectiveness of chronic disease management.

CN120412973APending Publication Date: 2025-08-01GUANGDONG ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY LAB (GUANGZHOU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510481559.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art has problems such as missing information, inaccurate data and lag in infectious disease detection and chronic disease information management, resulting in untimely detection and imperfect management, and the inability to effectively warn and provide reasonable suggestions.

Method used

Through multi-source data management methods, infectious disease and chronic disease data are integrated, logistic regression models and deep learning algorithms are used to conduct risk assessment and trend prediction, visual analysis is carried out in combination with geographic information systems, and a comprehensive platform is built for disease information management.

Benefits of technology

It has achieved more timely and accurate disease detection and early warning, improved the efficiency of prevention and control of major infectious diseases, optimized chronic disease management, and improved the efficiency of medical resource utilization and the accuracy of medical advice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412973A_ABST
    Figure CN120412973A_ABST
Patent Text Reader

Abstract

The invention relates to a disease information management method and system based on multi-source data, and relates to the technical field of medical artificial intelligence, and the method comprises the steps: forming multi-source heterogeneous data through employing the obtained infectious disease data and chronic disease data, and carrying out the preprocessing with the corresponding data features as a reference, on one hand, risk prediction and clustering analysis are performed on infectious disease related data through a logistic regression model to obtain risk assessment information, and a propagation trend is predicted through the model; on the other hand, chronic disease data are predicted and evaluated through the model, a medical advice result is obtained, and finally disease control information management is conducted in combination with all the data. Therefore, a set of perfect disease prevention and control information management is realized, on one hand, more timely and accurate infectious disease detection and early warning are realized, and the efficiency and effect of major infectious disease prevention and control work are improved, and on the other hand, chronic disease management is perfected, and the problems of diagnosis and treatment errors and lagging of patient treatment consciousness caused by incomplete data are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical artificial intelligence, and particularly to a method and system for disease information management based on multi-source data. Background Art

[0002] At present, disease information management is of great significance. For example, in the application of major infectious diseases, it can effectively prevent and control the spread of infectious diseases, and in the management of chronic disease information, it can effectively help patients control chronic diseases.

[0003] In the related art, on the one hand, major infectious diseases have characteristics such as fast transmission speed, wide influence range, and high fatality rate. The detection of major infectious diseases plays a major role in blocking the spread of infectious diseases, can issue early warnings to areas where infectious diseases may break out in advance, mobilize human and material resources for disease screening, effectively cut off the transmission chain of infectious diseases, and prevent the further expansion of the scope of infectious diseases.

[0004] On the other hand, chronic diseases such as hypertension and coronary heart disease have gradually become a concern for patients. Due to the problems of long disease courses and the concealment of some diseases in chronic diseases, the treatment of chronic diseases is relatively cumbersome. In the diagnosis, treatment, and management of chronic diseases, doctors often need to give comprehensive treatment and health management suggestions based on the results of each physical examination.

[0005] There are certain defects in the existing technologies for both infectious disease detection and chronic disease information management. Specifically, the existing infectious disease detection technologies mainly rely on hospital visit data, so usually patients can only be detected when they have severe symptoms, and most of the patient's condition information depends on the patient's self-reported information, which is prone to problems such as information missing, inaccurate / incomplete data; on the other hand, chronic disease information management can only directly feedback the current situation data collected, the experience is not perfect, and reasonable suggestions cannot be provided based on the information, showing lag. Summary of the Invention

[0006] This application provides a method and system for disease information management based on multi-source data to simultaneously achieve disease information management for both infectious disease and chronic disease data. On the one hand, using data related to infectious diseases to predict the transmission trend of infectious diseases, it can detect the signs of infectious disease transmission before patients have severe symptoms, timely lock in susceptible populations and risk areas, realize early warning of infectious diseases, and ultimately achieve more timely and accurate disease detection, risk assessment and warning, improving the efficiency and effect of major infectious disease prevention and control work; on the other hand, using data related to infectious diseases to generate medical advice, it improves problems such as diagnostic errors caused by poor information management and untimely follow-up due to the lag in patients' awareness of seeking medical treatment.

[0007] In a first aspect, this application provides a method for disease information management based on multi-source data, including:

[0008] Obtain the infectious disease prediction data of the target area from the preset first data source, and combine it with the chronic disease-related data obtained from the second data source to form multi-source heterogeneous data;

[0009] Based on each data feature in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to the preset data integration and preprocessing rules for data preprocessing, and obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data;

[0010] Perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information;

[0011] Select the data containing time series from the first data to be evaluated as input, and perform transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result;

[0012] Respectively perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset advice assistance model, and obtain a feature recognition result output by the chronic disease prediction decision model and a medical advice result output by the advice assistance model;

[0013] Based on the first data to be evaluated and the second data to be evaluated, combine the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result for disease control information management.

[0014] Optionally, based on each data feature in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to the preset data integration and preprocessing rules for data preprocessing, and obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data, including:

[0015] Analyze the infectious disease prediction data features in the multi-source heterogeneous data to obtain data features, and the data features include time features, association features, and data type features;

[0016] Based on the data features, perform data preprocessing on the infectious disease prediction data according to the preset data integration and preprocessing rules to obtain the first data to be evaluated;

[0017] Remove outliers and errors from the chronic disease-related data, screen out missing items with data missing from the chronic disease-related data, and select data to be fused from the chronic disease-related data;

[0018] Adopt a data filling algorithm, using the same type of data at adjacent time points of the missing item as a benchmark, and according to calculate the filled data of the missing item;

[0019] Adopt a data fusion algorithm, and according to perform weighted average fusion on different data to obtain fused data;

[0020] Based on the chronic disease-related data, the filled data, and the fused data, determine the second data to be evaluated;

[0021] wherein, t1 and t2 are time points adjacent to the missing item, G1 is the quantity corresponding to the time point t1, G2 is the quantity corresponding to the time point t2, G is the filled data calculated corresponding to the time point t, and t1 < t < t2, E i represents the i-th data to be fused, and w i is the weight corresponding to the data to be fused.

[0022] Optionally, based on the data characteristics, preprocess the infectious disease prediction data according to preset data integration preprocessing rules to obtain the first data to be evaluated, including:

[0023] Extract the data corresponding to the time characteristics from the infectious disease prediction data for time format alignment to obtain the first feature-processed data with unified data time dimension;

[0024] According to the preset data association rules, select the association keys representing the association characteristics from the first feature-processed data, and based on the association keys, merge the data corresponding to the association characteristics and related to the same subject in the first feature-processed data to obtain the second feature-processed data;

[0025] Adopt the Z-score standardization method, and according to standardize the numerical data corresponding to the data type characteristics in the second feature-processed data, and extract the text data belonging to the data type characteristics from the second feature-processed data, and perform classification transformation on the text data to obtain the third feature-processed data;

[0026] According to the preset data threshold and outlier detection algorithm, perform error removal and duplicate removal processing on the third feature-processed data to obtain the first data to be evaluated;

[0027] wherein, X is the original data value, μ is the data mean, and σ is the data standard deviation.

[0028] Optionally, perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select the disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information, including:

[0029] Select characteristic variables related to the spread of infectious diseases from the first data to be evaluated, where the characteristic variables at least include the growth rate of medical consultations, the growth rate of medicines, and the absenteeism rate caused by infectious diseases;

[0030] Using the characteristic variables, according to the formula Construct a logistic regression model;

[0031] Perform risk prediction on the first data to be evaluated through the logistic regression model to obtain the risk area evaluation information of the target area;

[0032] Adopt the K-means clustering algorithm to extract personnel signaling data, entry-exit data, and patient-related data from the first data to be evaluated as samples to be evaluated, and according to Perform clustering analysis to obtain the objective function;

[0033] Analyze according to the objective function to obtain the risk classification level;

[0034] Use the risk area evaluation information and the risk classification level as risk evaluation information;

[0035] Among them, K is the preset number of clusters, the goal is to minimize the sum of the squares of the distances from each sample to the center of its affiliated cluster, J is the objective function, C i is the i-th cluster, μ i is the center of the i-th cluster, x j is the sample data, that is, the sample to be evaluated, P(Y = 1) represents the probability that the area is a risk area, Y is the dependent variable, Y = 1 represents a risk area, Y = 0 represents a non-risk area, X i is the independent variable, that is, the selected characteristic variables, β i is the model parameter, which is solved by the maximum likelihood estimation method.

[0036] Optionally, perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset advice assistance model respectively, and obtain the characteristic recognition result output by the chronic disease prediction decision model and the medical advice result output by the advice assistance model, including:

[0037] Adopt a deep learning algorithm to construct a chronic disease prediction decision model, and select patient detection data with time series from the second data to be evaluated as the model input;

[0038] The chronic disease prediction decision model extracts features and patterns from the patient detection data, predicts the development trend of chronic diseases, and obtains the characteristic recognition result;

[0039] Taking the second data to be evaluated as the input, with the goal of minimizing the patient's medical treatment cost and maximizing the utilization efficiency of medical resources, a multi-objective optimization algorithm is adopted, and through the constructed recommended auxiliary model, the medical treatment recommendation result is generated.

[0040] Optionally, taking the second data to be evaluated as the input, with the goal of minimizing the patient's medical treatment cost and maximizing the utilization efficiency of medical resources, a multi-objective optimization algorithm is adopted, and through the constructed recommended auxiliary model, the medical treatment recommendation result is generated, including:

[0041] Extract the patient's medical treatment related data from the second data to be evaluated, and according to Calculate the patient's medical treatment cost;

[0042] Extract the medical resource related data from the second data to be evaluated, and according to Calculate the utilization efficiency of medical resources;

[0043] According to the patient's medical treatment cost and the utilization efficiency of medical resources, adopt a multi-objective optimization algorithm to generate the medical treatment recommendation result;

[0044] Among them, c j is the j-th medical treatment cost, x j is the decision variable for whether to select medical treatment services, x j is 0 or 1, e k is the utilization efficiency weight of the k-th medical resource, z k is the decision variable for whether to use medical resources, z k is 0 or 1, and continuously iterate and optimize the decision variables through the genetic algorithm to obtain the optimal medical treatment and physical examination recommendations as the medical treatment recommendation result.

[0045] Optionally, based on the first data to be evaluated and the second data to be evaluated, combining the risk assessment information, the transmission trend prediction result, the medical treatment recommendation result, and the feature recognition result for disease control information management, including:

[0046] Using geographic information system technology, analyze the data to be analyzed containing geographic location information in the first data to be evaluated, and perform factor correlation analysis based on the first data to be evaluated to obtain the correlation dimension information of each data;

[0047] Based on the correlation dimension information, use the first data to be evaluated for visual processing of the correlation to obtain a correlation visualization chart;

[0048] Based on the established infectious disease prediction and early warning model, an auxiliary decision-making platform is constructed, and the correlation visualization chart and the infectious disease transmission prediction result are rendered on the visualization operation interface of the auxiliary decision-making platform, and the prevention and control feedback management of the infectious disease transmission prediction result is carried out;

[0049] Using data fusion technology, according to F = [F1, F2, F3] or F = w1F1 + w2F2 + w3F3, the features identified from the second data to be evaluated, the medical advice result, and the feature recognition result are fused to obtain the target medical advice;

[0050] The target medical advice is displayed on the visualization page for disease information management of chronic diseases;

[0051] Among them, the prevention and control feedback management is used to provide decision support including the best prevention and control measures. F1 is the feature extracted from the medical advice result, F2 is the feature corresponding to the feature recognition result, F3 is the feature extracted from the second data to be evaluated, and w1, w2, and w3 are all preset weights, and w1 + w2 + w3 = 1.

[0052] Optionally, based on the association dimension information, the first data to be evaluated is used for correlation visualization processing to obtain a correlation visualization chart, including:

[0053] In the preset map data, according to the correlation between the data in the first data to be evaluated, visualization in the spatial dimension and time dimension is performed to obtain a spatial distribution chart and a time trend change chart;

[0054] Based on the association dimension information, variable data with correlation is extracted from the first data to be evaluated, and the potential data relationship between each variable data is analyzed;

[0055] Based on the potential data relationship, in the form of a scatter plot matrix, each variable data is rendered in the visualization chart to obtain a scatter plot;

[0056] According to the spatial distribution chart, the time trend change chart, and the scatter plot, the correlation visualization chart is determined.

[0057] Optionally, it further includes:

[0058] Obtain the first sample data and historical data related to infectious diseases, and obtain the second prediction data related to chronic diseases;

[0059] The first sample data and the second sample data are respectively preprocessed to obtain the first training sample to be trained, the historical data corresponding to the first training sample to be trained, and the second training sample to be trained including time series;

[0060] Perform knowledge graph analysis based on the historical data and the first training sample to obtain knowledge graph data;

[0061] Utilize recurrent neural networks and memory networks in deep learning to construct an infectious disease prediction and early warning model and a chronic disease prediction and decision-making model, and preset a weight matrix and parameters in the models;

[0062] Perform model training on the infectious disease prediction and early warning model based on the first training sample, the historical data, and the knowledge graph data, and perform feature extraction and model training on the chronic disease prediction and decision-making model according to the second training sample;

[0063] According to the evaluation indicators preset for each model, use the prediction results output by the infectious disease prediction and early warning model for model evaluation, and use the prediction results output by the chronic disease prediction and decision-making model for model evaluation;

[0064] Optimize the corresponding model parameters according to the model evaluation results until the model training of each model is completed;

[0065] Among them, during the construction of the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model, a forgetting gate, an input gate, an output gate, memory unit update, and hidden layer state are respectively preset.

[0066] In a second aspect, the present application provides a disease information management system based on multi-source data, including:

[0067] A multi-source heterogeneous data acquisition module for acquiring infectious disease prediction data of a target area from a preset first data source and combining it with chronic disease-related data obtained from a second data source to form multi-source heterogeneous data;

[0068] A preprocessing module for taking the data features in the multi-source heterogeneous data as a benchmark and extracting corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing to obtain first data to be evaluated corresponding to the infectious disease prediction data and second data to be evaluated corresponding to the chronic disease-related data;

[0069] A risk assessment and prediction module for performing risk prediction on the first data to be evaluated through a preset logistic regression model and performing clustering analysis on the disease condition data selected from the first data to be evaluated to obtain risk assessment information;

[0070] A transmission prediction module for selecting data including time series from the first data to be evaluated as input and performing transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result;

[0071] A recommendation prediction module, configured to respectively perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset recommendation assistance model, so as to obtain a feature recognition result output by the chronic disease prediction decision model and a medical advice result output by the recommendation assistance model;

[0072] A disease control information management module, configured to perform disease control information management based on the first data to be evaluated and the second data to be evaluated, in combination with the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result.

[0073] In summary, in the embodiments of the present application, infectious disease data and chronic disease data are obtained from different data sources to form multi-source heterogeneous data, and preprocessing is performed based on their corresponding data characteristics. On the one hand, a logistic regression model is used to perform risk prediction and clustering analysis on infectious disease-related data to obtain risk assessment information, and the transmission trend is predicted through the model to obtain the transmission trend prediction result of the infectious disease; on the other hand, the second data to be evaluated is respectively predicted and evaluated through a chronic disease prediction decision model and a recommendation assistance model to obtain a feature recognition result and a medical advice result. Finally, disease control information management is performed in combination with the preprocessed data to be evaluated, the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result. It can be seen that the present application realizes a complete set of disease prevention and control information management. On the one hand, it realizes more timely and accurate detection and early warning of infectious diseases, improving the efficiency and effectiveness of major infectious disease prevention and control work. On the other hand, it improves chronic disease management, avoiding problems such as diagnostic errors caused by imperfect data and the lag of patients' awareness of seeking medical treatment. Description of the Drawings

[0074] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0075] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0076] Figure 1 It is a schematic flowchart of a disease information management method based on multi-source data provided by an embodiment of the present application;

[0077] Figure 2 It is a schematic step flowchart of a disease information management method based on multi-source data provided by an optional embodiment of the present application;

[0078] Figure 3It is a flowchart of infectious disease prediction management in disease information management based on multi-source data provided by an optional example of this application;

[0080] Figure 4 It is a flowchart of chronic disease management in disease information management based on multi-source data provided by an optional example of this application;

[0081] Figure 5 It is a block diagram of the structure of a disease information management system based on multi-source data provided by an embodiment of this application;

[0082] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners

[0083] To make the objectives, technical solutions and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application.

[0084] For the convenience of understanding the embodiments of this application, further explanatory descriptions will be given below with reference to the accompanying drawings and specific embodiments. The embodiments do not constitute a limitation to the embodiments of this application.

[0085] In specific implementation, the detection and evaluation solution of this embodiment is mainly applied to major infectious diseases, including but not limited to: viral diseases, bacterial diseases, parasitic and fungal diseases, and newly discovered infectious diseases in infectious diseases. Specifically, it can be applied to the detection of major infectious diseases such as dengue fever, pneumonia, and AIDS.

[0086] Figure 1 It is a flowchart of a method for disease information management based on multi-source data provided by an embodiment of this application. As Figure 1 shown, the method for disease information management based on multi-source data provided by the embodiment of this application may specifically include the following steps:

[0087] Step 110, obtain the infectious disease prediction data of the target area from a preset first data source, and combine the chronic disease-related data obtained from the second data source to form multi-source heterogeneous data.

[0088] In specific implementation, both the first data source and the second data source can be composed of multiple data sources that provide data. Among them, the infectious disease prediction data is mainly composed of collecting medical system data (including but not limited to disease transmission methods, disease incubation periods, the number of patients seeking medical treatment, and the conditions of the patients seeking medical treatment), and multi-party social data (including but not limited to the number of people absent from work due to illness, pharmacy drug purchase situations, personnel signaling data, customs entry and exit situations, and global disease outbreak situations).

[0089] The chronic disease-related data is mainly composed of collecting the following data:

[0090] Widely collect basic knowledge and research literature on various chronic diseases, covering information such as pathogenesis, treatment methods and their side effects, commonly used drugs and their side effects, and disease development trends. At the same time, collect past objective test reports (such as blood glucose content G, heart rate H, urine composition U, etc.) and subjective performance data (such as headache, chest tightness, etc.) of chronic disease patient cases; collect medical resource data of local hospitals and community health service centers, including the types and quantities E of medical testing equipment, the specialties and numbers M of medical staff, the types and quantities D of drugs, and the maximum carrying capacity C of the hospital, etc.

[0091] Among them, this embodiment can obtain the corresponding above-mentioned data for the target area, or divide the obtained data according to regions for subsequent processing.

[0092] In the related art, on the one hand, in the detection of major diseases, the actions of patients are mainly based on the self-reported information of patients, which has defects such as missing information and inaccurate data; on the other hand, the chronic disease information management lacks key information on medical resources and cannot reasonably formulate scientific treatment and follow-up plans based on patients and local medical levels, resulting in insufficient utilization efficiency of medical resources. It can be seen that there are obvious deficiencies in the existing technology that relies on a single data source.

[0093] To solve this problem, this embodiment collects multi-touch data from different data sources. By broadening the data sources, the information is made more comprehensive and timely. On the one hand, for the prediction of infectious disease transmission, it can achieve more timely and accurate disease detection, risk assessment and early warning in the process of risk assessment prediction, etc., making up for the deficiencies of the existing technology that relies on a single data source. On the other hand, widely collect chronic disease-related information, including information such as medical resources and medical treatment costs when patients are treated, improve the utilization efficiency of medical resources, reduce medical treatment costs, enable patients and doctors to quickly obtain data, shorten the follow-up and treatment processes, and improve the quality of medical services, thus breaking through the limitations of traditional reliance on hospital visit data, broadening the data sources, and making the information more comprehensive and timely.

[0094] Step 120, based on the data features in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to preset data integration preprocessing rules to perform data preprocessing, and obtain first data to be evaluated corresponding to the infectious disease prediction data and second data to be evaluated corresponding to the chronic disease related data.

[0095] Specifically, the data integration preprocessing rules are mainly composed of multiple data preprocessing processes and their corresponding data processing methods. In this embodiment, different data preprocessing processes may be different. For example, the preprocessing process of infectious disease prediction data may include: data alignment, data fusion, standardization, structuring and data cleaning, etc.; the preprocessing process of chronic disease-related data may include: removing outliers and erroneous data, filling in data gaps, data alignment and deduplication, etc.

[0096] In practice, multi-source heterogeneous data includes data in different forms / formats from different data sources, and each piece of data typically possesses different characteristics, known as data signatures. Therefore, the solution of this embodiment primarily uses data signatures as a benchmark to preprocess the multi-source heterogeneous data from multiple dimensions.

[0097] To improve data processing efficiency, illustratively, the preprocessing order of infectious disease prediction data can be: data alignment - data fusion - standardization and structuring - data cleaning.

[0098] In this embodiment, data requiring preprocessing can be extracted for preprocessing. For example, data alignment processes data according to certain format standards. If some data already conforms to a certain format, preprocessing is unnecessary. Therefore, to improve data processing efficiency, this embodiment primarily analyzes the data characteristics of each data source in multi-source heterogeneous data to determine the data requiring corresponding preprocessing. Data that does not require preprocessing is then extracted for preprocessing. Data that does not require preprocessing can be left unprocessed, reducing resource consumption.

[0099] Therefore, by preprocessing each item of multi-source heterogeneous data in sequence, this embodiment obtains data to be evaluated with a unified data format, standard, and structure.

[0100] Step 130 : performing risk prediction on the first data to be evaluated using a preset logistic regression model, and selecting disease condition data from the first data to be evaluated for cluster analysis to obtain risk assessment information.

[0101] Among them, the risk area assessment information includes risk area assessment information and risk classification levels. The risk classification levels mainly include three risk levels: high, medium, and low. The first data to be evaluated involves data related to the population. In this embodiment, the population is divided according to the risk level, namely high-risk population, medium-risk population, and low-risk population. The high-risk population refers to those who have had close contact with infectious disease patients and show related symptoms. The medium-risk population refers to those whose activity trajectories intersect with high-risk areas.

[0102] In specific implementation, this embodiment mainly constructs a risk area prediction algorithm based on logistic regression to form a logistic regression model. Then, using the first data to be evaluated as input parameters, characteristic variables related to the spread of infectious diseases are selected from them, and risk prediction is carried out through the logistic regression model. The logistic regression model performs risk prediction on the input characteristic variables and outputs the risk area probability of the area, that is, the probability that the area is a risk area, as the risk area assessment information.

[0103] In actual implementation, data related to the condition is selected from the first data to be evaluated, and clustering analysis is carried out using clustering algorithms (such as K-means clustering algorithm, mean shift clustering algorithm, etc.). Clustering analysis mainly takes the selected data as samples and inputs them into the algorithm for clustering the samples. In this embodiment, the data related to the condition is usually related to patients. Therefore, through clustering processing, patients or populations of the same type can be clustered, that is, the population is divided into different risk levels.

[0104] Thus, this embodiment conducts risk assessment for regions and populations respectively, provides early warnings for areas where infectious diseases may break out and risk populations, and clarifies the relationships between infectious diseases, regions, and populations, as well as the infection situations of regions and populations within the regions. The risk assessment for regions and populations proposed in this embodiment can timely identify susceptible populations and detect major diseases in risk areas, starting from the most critical links in the spread of infectious diseases to achieve timely detection.

[0105] Step 140: Select the data containing time series from the first data to be evaluated as input, and perform a spread trend prediction through a preset infectious disease prediction and warning model to obtain a spread trend prediction result.

[0106] In this embodiment, the infectious disease prediction result is mainly used to predict the spread trend of infectious diseases, so as to achieve more timely and accurate disease detection, risk assessment, and warning.

[0107] In related technologies, timely identifying susceptible populations and risk areas is the most critical link in major disease detection. However, the currently detected data relies on hospital visit data and needs to wait until patients show severe symptoms to be detected, which has the problem of being insufficiently timely.

[0108] To solve this technical problem, this embodiment mainly uses neural network technology to construct a prediction and early warning model. Using the trained prediction and early warning model, with time series data as input data, the prediction and early warning model is used for trend prediction, including analyzing the signs of infectious disease transmission and accurately grasping the transmission path of infectious diseases.

[0109] This embodiment uses the analyzed signs of infectious disease transmission to lock in susceptible populations and risk areas, and uses the transmission path of infectious diseases mastered by the complete first data to be evaluated to avoid data defects caused by patients' self-reporting or incomplete data.

[0110] Step 150: Respectively use a preset chronic disease prediction and decision-making model and a preset advice assistance model to perform prediction and evaluation on the second data to be evaluated, and obtain the feature recognition result output by the chronic disease prediction and decision-making model and the medical advice result output by the advice assistance model.

[0111] In specific implementation, deep learning algorithms such as recurrent neural network (RNN) and its variant long short-term memory network (LSTM) are used to construct an online intelligent chronic disease prediction and decision-making model, and a medical treatment and physical examination advice assistance model is constructed. The second data to be evaluated is respectively input into the two models for prediction and evaluation. On the one hand, the optimal medical treatment and physical examination advice are predicted and output through the advice assistance model to obtain the medical advice result; on the other hand, the feature recognition result is predicted and output through the prediction and decision-making model for predicting the development trend of chronic diseases.

[0112] Step 160: Based on the first data to be evaluated and the second data to be evaluated, combine the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result for disease control information management.

[0113] In specific implementation, this embodiment can combine multiple models together, use visualization technology to construct a comprehensive platform, and display various predicted data results in the comprehensive platform to perform disease control information management based on multi-source heterogeneous data, combining risk assessment information, transmission trend prediction results, medical advice results, and feature recognition results.

[0114] It can be seen that in view of the problems of untimely detection and data defects in the prior art, on the one hand, the embodiments of the present application collect a large amount of multi-source data from different channels to form multi-source heterogeneous data, perform preprocessing according to data characteristics, use multi-channel data to complement and verify each other, and improve the data to improve the accuracy of the data, so as to be more timely and accurate in disease detection and avoid data defects; on the other hand, on the basis of obtaining data, the embodiments use models and clustering algorithms for risk prediction and assessment, and perform trend prediction through a prediction and early warning model, combining multiple models and algorithms together as a detection model to improve disease detection and realize risk assessment and early warning of infectious diseases. The ultimate technical effects achieved include: realizing more timely and accurate disease detection, risk assessment and early warning, improving the efficiency and effect of major infectious disease prevention and control work, and effectively ensuring public health and social stability.

[0115] Furthermore, limited by the data form and the differences in prediction management for different diseases, there is currently no unified solution for combining infectious disease prediction and chronic disease management for reasonable information management, that is, there is currently no technical solution for both infectious disease transmission prediction and chronic disease management, and only single prediction / management can be achieved.

[0116] In response to this, the embodiments perform unified preprocessing on the collected multi-source data, standardize each item of data, solve the problem of different data forms, and use model prediction to evaluate the data, and combine the prediction results of each item to jointly perform information management.

[0117] Refer to Figure 2 , which shows a schematic flow chart of the steps of a disease information management method based on multi-source data provided by an optional embodiment of the present application. The method may specifically include the following steps:

[0118] Step 210, obtain infectious disease prediction data of a target area from a preset first data source, and combine the chronic disease-related data obtained from a second data source to form multi-source heterogeneous data.

[0119] In a specific implementation, refer to Figure 3 shown, the embodiments of the present application collect relevant data from the medical system and relevant situation data collected from multiple parties in society to form infectious disease prediction data.

[0120] Refer to Figure 4 shown, the embodiments of the present application widely collect chronic disease-related knowledge and diagnosis and treatment results from data sources, and combine the distribution of medical resources to form chronic disease-related data.

[0121] Step 220: Taking each data feature in the multi-source heterogeneous data as a reference, extracting corresponding data from the multi-source heterogeneous data according to the preset data integration and preprocessing rules for data preprocessing, to obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data.

[0122] In an optional embodiment, taking each data feature in the multi-source heterogeneous data as a reference, extracting corresponding data from the multi-source heterogeneous data according to the preset data integration and preprocessing rules for data preprocessing, to obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data, which may specifically include: analyzing the infectious disease prediction data features in the multi-source heterogeneous data to obtain data features, where the data features include time features, association features, and data type features; taking the data features as a reference, performing data preprocessing on the infectious disease prediction data according to the preset data integration and preprocessing rules to obtain the first data to be evaluated; removing outliers and errors from the chronic disease-related data, screening out missing items with data missing from the chronic disease-related data, and selecting data to be fused from the chronic disease-related data; using a data filling algorithm, taking the same-type data at adjacent time points of the missing item as a reference, and according to calculating the filled data of the missing item; using a data fusion algorithm, according to fusing different data by the weighted average method to obtain fused data; determining the second data to be evaluated based on the chronic disease-related data, the filled data, and the fused data; where t1 and t2 are time points adjacent to the missing item, G1 is the quantity corresponding to the time point t1, G2 is the quantity corresponding to the time point t2, G is the filled data calculated corresponding to the time point t, and t1 < t < t2, E i represents the i-th data to be fused, and w i is the weight corresponding to the data to be fused.

[0123] Referring to Figure 4 as shown, in a specific implementation, screening the chronic disease data in the collected chronic disease-related data, and removing outliers and incorrect data. Using a data filling algorithm to fill in the missing data. For example, for blood glucose content data, linear interpolation can be used for filling, calculating the missing blood glucose content according to the above formula, and then the data can be classified according to different chronic diseases, and different pathological phenomena of the same chronic disease can be further subdivided. Finally, the data is aligned and structured to establish an association bridge between the data.

[0124] For the medical resource data collected from various regions, the collected data is screened and aligned to remove duplicate and invalid data. A data fusion algorithm is used to integrate data from different sources. For example, for the medical equipment data of different hospitals, it can be fused by the weighted average method. Assume that the number of devices in hospital i is E i , and the weight is w i , then the fused number of devices is E. Then, the data is standardized and structured, and visualized based on GIS technology combined with remote sensing maps for subsequent analysis and application.

[0125] In the actual implementation, for the infectious disease prediction data, preprocessing is mainly carried out based on the data characteristics of each item of data. Specifically, different infectious disease prediction data reflects different data characteristics. For example, data related to time usually has time characteristics; data with certain correlations usually has correlation characteristics; different types of data (such as numerical and text) usually have data type characteristics. In this embodiment, data analysis and feature analysis are performed on each item of data in the multi-source heterogeneous data to identify and classify the data and determine the data characteristics possessed by the data. Then, based on this, preprocessing is carried out on each item of data to obtain the data to be evaluated.

[0126] Optionally, in this embodiment, based on the data characteristics, the infectious disease prediction data is preprocessed according to the preset data integration and preprocessing rules to obtain the first data to be evaluated, including: extracting the data corresponding to the time characteristics from the infectious disease prediction data for time format alignment to obtain the first feature-processed data with a unified data time dimension; according to the preset data association rules, selecting the association keys representing the association characteristics from the first feature-processed data, and based on the association keys, merging the data corresponding to the association characteristics and related to the same subject in the first feature-processed data to obtain the second feature-processed data; using the Z-score standardization method, according to standardize the numerical data corresponding to the data type characteristics in the second feature-processed data, and extract the text data belonging to the data type characteristics from the second feature-processed data, and perform classification transformation on the text data to obtain the third feature-processed data; according to the preset data threshold and outlier detection algorithm, perform error and duplicate removal processing on the third feature-processed data to obtain the first data to be evaluated; where X is the original data value, μ is the data mean, and σ is the data standard deviation.

[0127] Exemplarily, the infectious disease prediction data can be preprocessed in the preprocessing order of data alignment - data fusion - data standardization and structuring - data cleaning. The following describes each preprocessing process in combination.

[0128] For data alignment: The infectious disease prediction data consists of multiple data obtained from different data sources. Due to certain differences in each data source, the obtained data may also have certain differences. For example, the time of each data source is inconsistent, resulting in non-uniform time dimensions for each data obtained. The non-uniform data formats are likely to affect the fusion of data. Therefore, to solve the data fusion problem, in this embodiment, the formats of each data are unified through data alignment.

[0129] Exemplarily, when unifying data from the time dimension, it is possible to analyze the data with time characteristics in multi-source heterogeneous data, analyze its current time format, determine the preprocessed data format according to the number of data belonging to the same time format, or preset a standard time format (such as the ISO6601 format) as the data format. Subsequently, the time data of each data source is uniformly converted to obtain data with a unified time dimension, which is convenient for subsequent fusion.

[0130] For data fusion: To fully explore the relevance between data, in this embodiment, different data with relevance are fused. Exemplarily, for each data with correlation characteristics, the correlation relationship between the data is analyzed, and then according to the data correlation rules, using the area code or personnel identifier, etc. as the correlation key, the data related to the same subject in different data sources are merged. For example, the data such as the number of medical consultations, the number of absences due to illness, and the pharmacy drug purchase situation in this area are fused, so as to obtain the second feature-processed data.

[0131] For data standardization and structuring: The data with data type characteristics are mainly of two types: numerical and text. For the data with numerical characteristics (such as the number of medical consultations, the number of drug purchases, etc.), mainly use the standardization method to preprocess it; for the data with text characteristics (such as symptom descriptions, disease names, etc.), use the method of text vectorization in natural language processing technology to perform data conversion in the form of classification coding and convert it into a numerical form, thereby obtaining the third feature-processed data after standardization and structuring, which is convenient for subsequent analysis and processing.

[0132] Furthermore, to achieve the standardized conversion of numerical data, in this embodiment, the Z-score standardization method is mainly used to complete the data conversion in combination with a preset formula.

[0133] It should be noted that during standardization, different numerical data can be calculated separately. For example, the number of medical consultations is the number of medical consultations in different hospitals and different institutions for standardization, and the same is true for the number of drug purchases. The standardized numerical data is used to unify different magnitudes of data into the same numerical range, that is, after the above different numerical data are standardized separately, the mean is unified to 0 and the standard deviation is 1.

[0134] For data cleaning: There may still be some incorrect or duplicate data in the infectious disease prediction data. To improve the efficiency of subsequent data analysis, reasonable data thresholds and outlier detection algorithms can be set to remove incorrect and duplicate data. For example, when the number of clinic visits exceeds a certain multiple of the historical maximum without reasonable reasons, it is determined as an outlier and deleted; through data fingerprint technology (such as calculating data fingerprints using hash functions), corresponding identifiers are assigned to the data as fingerprints. If two data fingerprints are the same, they are considered the same data, and one of the data records can be deleted, thus realizing the identification and deletion of duplicate records.

[0135] In actual implementation, referring to Figure 3 , after the data preprocessing of this embodiment is completed, a database can be constructed to facilitate data retrieval and analysis during subsequent risk assessment and prediction and early warning.

[0136] Exemplarily, the Hadoop HDFS distributed file storage system is used to store the original and preprocessed data to meet the large-scale data storage requirements; the MongoDB database is used at the application layer to store structured key data, which is stored in the form of documents, facilitating the storage and query of complex data structures. For example, storing infectious disease-related data documents by region, including fields of each data source, meeting the large-scale data storage requirements, facilitating the management and query of complex data structures, that is, achieving the purpose of convenient retrieval and analysis.

[0137] Step 230, perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information.

[0138] In an alternative embodiment, the above-mentioned performing risk prediction on the first data to be evaluated through a preset logistic regression model, and selecting disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information includes: selecting characteristic variables related to the spread of infectious diseases from the first data to be evaluated, where the characteristic variables at least include the visit growth rate, drug growth rate, and absenteeism rate caused by infectious diseases; using the characteristic variables, according to the formula construct a logistic regression model; perform risk prediction on the first data to be evaluated through the logistic regression model to obtain risk area assessment information of the target area; adopt the K-means clustering algorithm to extract subscriber signaling data, entry and exit data, and patient-related data from the first data to be evaluated as samples to be evaluated, and according to perform clustering analysis to obtain an objective function; analyze according to the objective function to obtain a risk classification level; use the risk area assessment information and the risk classification level as risk assessment information; where K is a preset number of clusters, the goal is to minimize the sum of the squares of the distances from each sample to the center of the cluster it belongs to, J is the objective function, Ci is the i-th cluster, and μ i is the center of the i-th cluster, and x j is the sample data, that is, the sample to be evaluated. P(Y = 1) represents the probability that the area is a risk area. Y is the dependent variable, Y = 1 represents the risk area, Y = 0 represents the non-risk area, and X i is the independent variable, that is, the selected characteristic variable, and β i is the model parameter, which is solved by the maximum likelihood estimation method

[0139] In the specific implementation, data related to the spread of infectious diseases is selected from the first sample to be evaluated as the characteristic variables, including but not limited to: the growth rate of the number of medical consultations, the growth rate of the sales of specific drugs in pharmacies, the absenteeism rate due to illness, etc. Then, a logistic regression model is constructed to predict the characteristic variables. Among them, to implement the construction of the logistic regression model based on the risk area prediction algorithm, this embodiment uses the formula to implement the construction of the logistic regression model and use this formula to implement the prediction of the probability that the area belongs to the risk area

[0140] In the specific implementation, this embodiment mainly uses the risk population assessment algorithm based on cluster analysis. Using the personnel signaling data, entry-exit information, and the condition data of medical patients as samples, and adopting the K-means algorithm, each sample is clustered to determine the corresponding objective function, so as to divide the population into corresponding risk levels through clustering

[0141] In the actual implementation, this embodiment uses the clustering algorithm and combines the clustering formula

[0142] to perform cluster analysis on the input samples to achieve accurate division of risk levels

[0143] Therefore, aiming at the deficiencies of the existing infectious disease detection technologies, in the existing technologies, for the behavior of patients, only the self-reported information of patients, hospital medical data, etc. can be relied on, resulting in some defects such as missing information and inaccurate data

[0144] In this embodiment, timely judgment of the susceptible population and the risk area is the most critical link in the detection of major diseases. By collecting a large amount of multi-channel and multi-link data, including non-traditional medical data such as the number of absentees due to illness, entry-exit, personnel signaling data, and drug purchase situations, multi-source data is formed, breaking through the limitation of traditional reliance on hospital medical data, making the risk warning more timely, accurate, and comprehensive, being able to detect the signs of the spread of infectious diseases before the patients show severe symptoms, timely locking in the susceptible population and the risk area, and realizing the early warning of infectious diseases

[0145] Meanwhile, by using multi-channel data to complement and verify each other, data defects caused by patients' self-reporting are avoided, and the transmission routes of infectious diseases are accurately grasped. The ultimate technical effect achieved is to realize more timely and accurate disease detection, risk assessment and early warning, improve the efficiency and effectiveness of major infectious disease prevention and control work, and effectively guarantee public health and social stability.

[0146] In an optional embodiment, the K-means clustering algorithm is adopted in this embodiment to extract personnel signaling data, entry-exit data, and patient-related data from the first data to be evaluated for clustering, and an objective function for cluster analysis is obtained, which specifically includes: extracting personnel signaling data, entry-exit data, and patient-related data from the first data to be evaluated as samples to be evaluated; using the samples to be evaluated as a benchmark, according to perform cluster analysis to obtain the objective function; where K is the preset number of clusters, the goal is to minimize the sum of the squared distances from each sample to the cluster center it belongs to, J is the objective function, C i is the i-th cluster, μ i is the i-th cluster center, and x j is the sample data, that is, the sample to be evaluated.

[0147] In this embodiment, the population is divided into different risk levels of high, medium, and low through clustering, such as high-risk populations (in close contact with infectious disease patients and showing related symptoms), medium-risk populations (whose activity trajectories intersect with high-risk areas), and low-risk populations.

[0148] It should be noted that in fact, multiple different methods can be adopted when dealing with multiple scenarios. Which specific data to use mainly depends on the scenario, and suggestions are mainly provided by industry experts and incorporated by algorithm personnel, and finally the corresponding algorithm development is completed. Including the data collection stage, the data categories collected are only used as examples, and more categories of data can be collected according to a broader scenario. This application emphasizes the solution framework proposed for the application scenario, and the focus is on the overall framework proposed.

[0149] Step 240: Select the data containing time series from the first data to be evaluated as the input, and perform a transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result.

[0150] Exemplarily, the input of the infectious disease prediction and early warning model may include the data content mentioned in the previous preprocessing process, such as the number of clinic visits and the quantity of purchased drugs in the previous days, and may also include the output results of the aforementioned risk model, such as the risk level of this area in the previous days and other information. After processing these time series signals, the prediction results for the current day or multiple future days are finally obtained. For example, based on the data of the previous days, it can be predicted whether there will be a large-scale increase in the number of clinic visits tomorrow, providing a data reference basis for the medical resource allocation or disease prevention and control department to formulate relevant policies.

[0151] In specific implementation, before inputting the data into the infectious disease prediction and early warning model for predicting the spread trend, or extracting features and patterns from the data through the chronic disease prediction and decision-making model to predict the development trend of chronic diseases, corresponding models can be constructed and model training can be carried out.

[0152] Optionally, this embodiment further includes: obtaining first sample data and historical data related to infectious diseases, and obtaining second prediction data related to chronic diseases; respectively preprocessing the first sample data and the second sample data to obtain a first sample to be trained, the historical data corresponding to the first sample to be trained, and a second sample to be trained including time series; performing knowledge graph analysis based on the historical data and the first sample to be trained to obtain knowledge graph data; using the recurrent neural network and memory network in deep learning to construct an infectious disease prediction and early warning model and a chronic disease prediction and decision-making model, and presetting a weight matrix and parameters in the models; training the infectious disease prediction and early warning model based on the first sample to be trained, the historical data, and the knowledge graph data, and performing feature extraction and model training on the chronic disease prediction and decision-making model according to the second sample to be trained; evaluating the models using the prediction results output by the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model according to the evaluation indicators preset for each model; optimizing the corresponding model parameters according to the model evaluation results until each model completes model training; wherein, in the construction process of the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model, a forgetting gate, an input gate, an output gate, memory unit update, and hidden layer state are respectively preset.

[0153] In specific implementation, based on the knowledge graph and artificial intelligence technology, combined with the infectious disease transmission mechanism and historical data, an infectious disease prediction and early warning model is constructed. For example, the recurrent neural network (RNN) in deep learning and its variants (such as long short-term memory network LSTM) are used to model time series data to predict the infectious disease spread trend. Taking the LSTM model as an example, its core formula (for the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model) is as follows:

[0154] Forgetting gate: ft = σ(W if · x t + b if + W hf · h t-1 + b hf ); Input gate: i t = σ(W ii · x t + b ii + W hi · h t-1 + b hf ); Memory cell update: C t = f t ⊙ C t-1 + i t ⊙ tanh(W ic · x t + b ic + W hc · h t-1 + b hc ); Hidden layer state: h t = o t ⊙ tanh(C t )); Prediction result: y t = W iy · h t + b iy ; where y t is the prediction result at the current time, x t is the input data at the current time, h t is the hidden state at the current time, C t is the memory cell at the current time, σ is the sigmoid function, ⊙ is element-wise multiplication, and W and b are both model parameters.

[0155] Exemplarily, the collected data samples are divided into a training set, a validation set, and a test set, for example, divided according to the ratio of 70%, 15%, 15%.

[0156] During the model training and validation process, the training set is used to train the model. By adjusting the model parameters, such as the weights and biases of the neural network, the model achieves better performance on the training set. The training process is monitored on the validation set to prevent overfitting of the model. The cross-validation method, such as K-fold cross-validation with K = 5 or K = 0, is adopted to further improve the accuracy of model evaluation.

[0157] During model testing, use sample data that was not involved in training (such as reserved historical data or newly collected data) to evaluate the accuracy and generalizability of the model. Among them, the evaluation metrics include accuracy, recall, F1-score, etc. Among them, accuracy: Accuracy = number of correctly predicted samples / total number of samples; recall: Recall = number of correctly predicted samples / number of actual positive samples; F1-score: Accuracy = 2×Precision×Recall / Precision + Recall; Precision = number of correctly predicted samples / number of predicted positive samples. By analyzing the test results, optimize and improve the model to ensure that the model performance meets the actual application requirements.

[0158] Step 250, adopt a deep learning algorithm to construct a chronic disease prediction and decision-making model, and select patient detection data with a time series from the second data to be evaluated as the model input.

[0159] Step 260, the chronic disease prediction and decision-making model extracts features and patterns from the patient detection data, predicts the development trend of chronic diseases, and obtains a feature recognition result.

[0160] A unified description of steps 250 - 260:

[0161] In a specific implementation, adopt a deep learning algorithm, combine the above RNN and its variant LSTM to construct an online intelligent chronic disease prediction and decision-making model. Using patient detection data with a time series as the input, the model learns the features and patterns in the data and predicts the development trend of chronic diseases.

[0162] For example, assume the input data is X = [x1, x2, …, x T , where x T represents the patient data at the t-th moment, the hidden layer state of the model is h t and the output prediction result is y t .

[0163] Step 270, using the second data to be evaluated as the input, with the goal of minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources, adopt a multi-objective optimization algorithm, and through the constructed recommendation assistance model, generate a medical advice result.

[0164] In a specific implementation, this embodiment fully considers the patient's economic situation and the situation of nearby medical resources. Using the constructed medical treatment and physical examination recommendation assistance model, adopt a multi-objective optimization algorithm, such as a genetic algorithm, with the goal of minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources, to generate reasonable medical treatment and physical examination recommendations.

[0165] Thus, based on the preprocessed chronic disease-related knowledge literature and hospital diagnosis and treatment data, this embodiment uses machine learning algorithms to deeply mine key information such as the incidence law and disease development trend of chronic diseases, and predicts and evaluates the chronic disease status of patients.

[0166] This embodiment uses a medical visit and physical examination advice assistance model to process the preprocessed medical resource data from various regions, and generates reasonable medical visit and physical examination advice considering the patient's personal economic situation and the accessibility of nearby medical resources.

[0167] In an alternative embodiment, taking the second data to be evaluated as the input, with the goal of minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources, a multi-objective optimization algorithm is adopted, and through the constructed advice assistance model, a medical visit advice result is generated. Specifically, it may include: extracting patient medical visit-related data from the second data to be evaluated, and calculating the patient's medical cost according to... calculating the patient's medical cost; extracting medical resource-related data from the second data to be evaluated, and calculating the utilization efficiency of medical resources according to... calculating the utilization efficiency of medical resources; according to the patient's medical cost and the utilization efficiency of medical resources, adopting a multi-objective optimization algorithm to generate a medical visit advice result; where c j is the j-th medical cost, x j is a decision variable for whether to choose a medical service, x j is 0 or 1, e k is the utilization efficiency weight of the k-th medical resource, z k is a decision variable for whether to use a medical resource, z k is 0 or 1, and the decision variables are continuously iteratively optimized through a genetic algorithm to obtain the optimal medical visit and physical examination advice as the medical visit advice result.

[0168] Specifically, this embodiment uses formulas to achieve minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources. Among them, taking... as the patient's medical cost function, the medical cost includes but is not limited to: registration fees, examination fees, etc.; taking... as the medical resource utilization efficiency function.

[0169] Step 280, based on the first data to be evaluated and the second data to be evaluated, combined with the risk assessment information, the spread trend prediction result, the medical visit advice result, and the feature recognition result, perform disease control information management.

[0170] Optionally, the above-mentioned performing disease control information management based on the first data to be evaluated and the second data to be evaluated, combined with the risk assessment information, the spread trend prediction result, the medical visit advice result, and the feature recognition result, may include the following sub-steps:

[0171] Sub-step 2801: Using geographic information system technology, analyze the data to be analyzed containing geographic location information in the first data to be evaluated, and perform factor correlation analysis based on the first data to be evaluated to obtain the correlation dimension information of each data.

[0172] Sub-step 2802: Using the first data to be evaluated for visualization processing of relevance with the correlation dimension information as the benchmark, to obtain a visualization chart of relevance.

[0173] Sub-step 2803: Based on the established infectious disease prediction and early warning model, construct an auxiliary decision-making platform, and render the visualization chart of relevance and the infectious disease transmission prediction result on the visualization operation interface of the auxiliary decision-making platform, and conduct prevention and control feedback management on the infectious disease transmission prediction result.

[0174] A unified description of sub-steps 2801 - 2803:

[0175] In actual implementation, this embodiment mainly uses geographic information system (GIS) technology to extract data containing geographic location information from the first data to be evaluated for visualization, including but not limited to: personnel signaling data, customs entry and exit situations, etc. Then, analyze the correlation between data, determine the correlation dimension, and based on this, perform visualization processing on the data.

[0176] During visualization, this embodiment adds a multi-factor correlation analysis dimension. In addition to separately displaying the data of each data source, the medical system data and multi-party social data are correlated and analyzed and then visually presented.

[0177] Optionally, with the correlation dimension information as the benchmark, this embodiment uses the first data to be evaluated for visualization processing of relevance to obtain a visualization chart of relevance, including: in the preset map data, perform visualization in the spatial dimension and time dimension according to the relevance between the data in the first data to be evaluated to obtain a spatial distribution chart and a time trend change chart; based on the correlation dimension information, extract the variable data of relevance from the first data to be evaluated, and analyze the potential data relationship between each variable data; based on the potential data relationship, use a scatter plot matrix method to render each variable data in the visualization chart to obtain a scatter plot; determine the visualization chart of relevance according to the spatial distribution chart, the time trend change chart, and the scatter plot.

[0178] Specifically, in this embodiment, hotspots of personnel activities, ports of entry and exit, areas with a high incidence of infectious diseases, etc. are marked on the map to visually present the spatial distribution characteristics; for time series data (such as the change in the number of medical consultations over time), line charts, bar charts and other charts are used to display the change trends. Then, during the visualization process, correlation analysis is performed on data that is correlated, such as the number of medical consultations, the situation of purchasing drugs at pharmacies, and the number of absences due to illness. Through the scatter plot matrix, the pairwise relationships between these three variables are shown. In the scatter plot, it can be observed that when the number of medical consultations increases, the change trends of the purchase volume of pharmacy-related drugs and the number of absences due to illness, so as to deeply explore the potential connections between different factors and provide a more comprehensive perspective for the analysis of the spread of infectious diseases.

[0179] Referring to Figure 3 , in this embodiment, visual data is used to construct a visual operation interface, and on this basis, an auxiliary decision-making platform is built based on the established database and prediction and early warning model. The platform provides a visual operation interface to facilitate disease control personnel to view real-time infectious disease data, prediction results and risk assessment information; through data mining and analysis functions, it provides decision-making support for disease control personnel, such as recommending the best prevention and control measures (such as the demarcation of isolation areas, vaccination strategies, etc.).

[0180] The auxiliary decision-making evaluation can be used to realize the application of the comprehensive disease control business management platform. Specifically, the auxiliary decision-making platform is connected to the comprehensive disease control business management platform to achieve real-time data sharing and business collaboration. When the prediction and early warning model issues a warning, the comprehensive disease control business management platform automatically triggers corresponding prevention and control processes, such as notifying relevant departments to conduct personnel investigations, material allocation, etc., and real-time feedback on the progress and effects of the prevention and control work, so as to adjust the prevention and control strategies in a timely manner. Combining real-time multi-source heterogeneous data, continuously optimize the prediction and early warning model and the auxiliary decision-making plan to realize the dynamic optimization and continuous improvement of the infectious disease detection system.

[0181] Sub-step 2804, using data fusion technology, according to F = [F1, F2, F3] or F = w1F1 +

[0182] w2F2 + w3F3, fuse the features identified from the second data to be evaluated, the medical advice results, and the feature recognition results to obtain the target medical advice.

[0183] Sub-step 2805, display the target medical advice on the visualization page to manage the disease information of chronic diseases.

[0184] Among them, the prevention and control feedback management is used to provide decision support including the best prevention and control measures. F1 is a feature extracted from the medical advice result, F2 is the corresponding feature of the feature recognition result, F3 is a feature extracted from the second data to be evaluated, and w1, w2, and w3 are all preset corresponding weights, and w1 + w2 + w3 = 1.

[0185] A unified description is given for sub-step 2804 - sub-step 2805:

[0186] In the related art, there is currently a preliminary established digital health management platform, which collects hospital inspection and medical information, physical examination information of community health service centers, etc., assists the chronic disease follow-up process, and provides help for patients and doctors during the chronic disease treatment process. However, the management platform in the existing technology can only feedback part of the current situation. Patients still need to spend time going to the hospital for medical treatment each time, and there may be problems such as diagnostic errors caused by incomplete physical examination data and untimely follow-up caused by the lag of patients' awareness of seeking medical treatment. The existing technology also does not consider the patient's own economic situation and the nearby medical resource situation, and the provided offline medical advice ignores the feasibility of the patient, thus restricting the next step of the chronic disease treatment process of the patient.

[0187] In view of the above defects of the existing technology in chronic disease management, in this embodiment, multi-modal data is collected and sorted out, combined with an intelligent chronic disease prediction decision model and a medical treatment and physical examination advice assistance model to generate a chronic disease intelligent agent. By fusing the output results of the online intelligent chronic disease prediction decision model and the medical treatment and physical examination advice assistance model, and combining the patient's personal report data at the same time, multi-modal data is collected and sorted out. The data fusion technology, such as feature splicing and weighted fusion, is used to integrate the output features of different models. On this basis, a chronic disease intelligent agent is constructed to display the target medical advice on the visualization page and manage the disease information of chronic diseases.

[0188] Thus, according to the decision result of the chronic disease intelligent agent, this embodiment provides detailed and feasible medical advice for patients, including the time, place, examination items, and treatment plan for medical treatment, etc.

[0189] In actual implementation, this embodiment uses Figure 3 the shown comprehensive disease control business management platform, combined with Figure 4 the shown chronic disease intelligent agent, to jointly construct a comprehensive management platform for realizing the prediction management of infectious disease transmission and the medical treatment management of chronic diseases, and solving the technical obstacles existing when the existing technology combines the prediction of infectious diseases and the management of chronic diseases.

[0190] In summary, the embodiment of the present application proposes a disease prevention and control information management method comprehensively applied to infectious disease prediction and chronic disease management, which collects a large amount of multi-source data from different channels, forms multi-source heterogeneous data, and performs corresponding preprocessing, and can build a distributed database with the obtained data to meet the query and analysis requirements.

[0191] On the one hand, aiming at the deficiencies of the existing technology in the aspect of infectious disease prediction / detection, data visualization and risk assessment are carried out on the preprocessed infectious disease prediction data, an early warning model is built, multi-source data is visualized by using GIS technology, and then the auxiliary decision-making analysis platform is connected to the comprehensive disease control business management platform, and the prediction and early warning feedback is realized by combining real-time multi-source heterogeneous data, so as to achieve more timely and accurate disease detection, risk assessment and early warning, assist in the prevention and control work of major infectious diseases, realize real-time data sharing and business collaboration, automatically trigger the prevention and control process according to the early warning and feedback the prevention and control effect, and continuously optimize the detection system. The final technical effects achieved include: realizing more timely and accurate disease detection, risk assessment and early warning, improving the efficiency and effect of the prevention and control work of major infectious diseases, and effectively ensuring public health and social stability.

[0192] On the other hand, aiming at the deficiencies of the existing technology in the aspect of chronic disease management, an innovative chronic disease intelligent agent is constructed. By integrating chronic disease-related data, algorithms such as machine learning and deep learning are used to deeply explore the incidence law and disease development trend of chronic diseases, and accurately predict and evaluate the chronic disease status of patients. Considering the economic conditions of patients and the accessibility of medical resources comprehensively, reasonable medical treatment and physical examination suggestions are generated. The multi-modal data is fused and analyzed to form a comprehensive decision-making basis, and detailed and feasible medical treatment suggestions are provided for patients, covering the time, place, examination items and treatment plans of medical treatment. It has a variety of beneficial effects: ① Patients can obtain follow-up feedback in a timely manner, ensure the timeliness of each step of diagnosis and treatment, and effectively improve the treatment effect; ② Patients can comprehensively understand the situation of nearby hospitals, reasonably plan the time for medical treatment, and reduce the medical treatment cost and waiting time; ③ The system can provide scientific and reasonable treatment and follow-up plans according to the medical level of patients and their vicinity, and improve the utilization efficiency of medical resources; ④ Various data are centralized on the platform, and patients and relevant doctors can quickly obtain detection data, make a rough judgment on the condition and diagnosis and treatment decisions in a timely manner, significantly shorten the follow-up and treatment process time, and improve the quality of medical services.

[0193] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously.

[0194] Such as Figure 5As shown in the figure, an embodiment of the present application further provides a disease information management system 500 based on multi-source data, including:

[0195] A multi-source heterogeneous data acquisition module 510, configured to obtain infectious disease prediction data of a target area from a preset first data source, and combine the chronic disease-related data obtained from a second data source to form multi-source heterogeneous data;

[0196] A preprocessing module 520, configured to, based on the data features in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing, and obtain first data to be evaluated corresponding to the infectious disease prediction data and second data to be evaluated corresponding to the chronic disease-related data;

[0197] A risk assessment and prediction module 530, configured to perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information;

[0198] A transmission prediction module 540, configured to select the data including time series from the first data to be evaluated as input, and perform transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result;

[0199] A recommendation prediction module 550, configured to perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset recommendation assistance model respectively, and obtain a feature recognition result output by the chronic disease prediction decision model and a medical treatment recommendation result output by the recommendation assistance model;

[0200] A disease control information management module 560, configured to perform disease control information management based on the first data to be evaluated and the second data to be evaluated, in combination with the risk assessment information, the transmission trend prediction result, the medical treatment recommendation result, and the feature recognition result.

[0201] Optionally, the preprocessing module includes:

[0202] A data feature analysis sub-module, configured to analyze the infectious disease prediction data features in the multi-source heterogeneous data to obtain data features, where the data features include time features, association features, and data type features;

[0203] An infectious disease data prediction and processing sub-module, configured to, based on the data features, perform data preprocessing on the infectious disease prediction data according to preset data integration and preprocessing rules to obtain first data to be evaluated;

[0204] A chronic disease data preprocessing sub-module for removing outliers and errors from the chronic disease-related data, screening for missing items with missing data from the chronic disease-related data, and selecting data to be fused from the chronic disease-related data; using a data filling algorithm, with the same-type data at adjacent time points of the missing item as a benchmark, and according to calculate the filled data for the missing item; using a data fusion algorithm, according to fuse different data by the weighted average method to obtain fused data; based on the chronic disease-related data, the filled data, and the fused data, determine the second data to be evaluated; where t1 and t2 are time points adjacent to the missing item, G1 is the quantity corresponding to the time point t1, G2 is the quantity corresponding to the time point t2, G is the filled data calculated for the time point t, and t1 < t < t2, E i represents the i-th data to be fused, and w i is the weight corresponding to the data to be fused.

[0205] Optionally, the infectious disease data prediction processing sub-module includes:

[0206] A first preprocessing unit for extracting each data corresponding to the time feature from the infectious disease prediction data for time format alignment to obtain first feature-processed data with a unified data time dimension;

[0207] A second preprocessing unit for selecting association keys representing association features from the first feature-processed data according to preset data association rules, and based on the association keys, merging the data corresponding to the association features and related to the same subject in the first feature-processed data to obtain second feature-processed data;

[0208] A third preprocessing unit for using the Z-score standardization method, according to perform data standardization on the numerical data corresponding to the data type features in the second feature-processed data, and extract the text data belonging to the data type features from the second feature-processed data, and perform classification transformation on the text data to obtain third feature-processed data;

[0209] A fourth preprocessing unit for performing error and duplicate removal processing on the third feature-processed data according to preset data thresholds and outlier detection algorithms to obtain the first data to be evaluated; where X is the original data value, μ is the data mean, and σ is the data standard deviation.

[0210] Optionally, the risk assessment prediction module includes:

[0211] a characteristic variable selection submodule, configured to select characteristic variables related to the spread of infectious diseases from the first data to be evaluated, wherein the characteristic variables include at least a growth rate of medical consultations, a growth rate of drugs, and an absenteeism rate caused by infectious diseases;

[0212] Logistic regression model construction submodule is used to use the characteristic variables according to the formula Build a logistic regression model;

[0213] a risk prediction submodule, configured to perform risk prediction on the first data to be evaluated using the logistic regression model to obtain risk area assessment information of the target area;

[0214] The clustering submodule uses the K-means clustering algorithm to extract personnel signaling data, entry and exit data, and patient-related data from the first data to be evaluated as samples to be evaluated, and Perform cluster analysis to obtain the objective function; perform analysis based on the objective function to obtain the risk classification level; use the risk area assessment information and the risk classification level as risk assessment information; where K is the pre-set number of clusters, the goal is to minimize the sum of the squares of the distances from each sample to the cluster center, J is the objective function, and C i is the i-th cluster, μ i is the i-th cluster center, x j is the sample data, i.e. the sample to be evaluated, P(Y=1) represents the probability that the region is a risk region, Y is the dependent variable, Y=1 represents a risk region, Y=0 represents a non-risk region, X i is the independent variable, that is, the selected characteristic variable, β i are model parameters, which are solved by maximum likelihood estimation.

[0215] Optionally, the suggestion prediction module includes:

[0216] A development trend prediction submodule is configured to use a deep learning algorithm to construct a chronic disease prediction decision model, and select patient test data with a time series from the second data to be evaluated as model input; the chronic disease prediction decision model extracts features and patterns from the patient test data, predicts the development trend of the chronic disease, and obtains a feature recognition result;

[0217] The medical advice result determination submodule is used to take the second data to be evaluated as input, minimize the patient's medical costs and maximize the efficiency of medical resource utilization as the goals, adopt a multi-objective optimization algorithm, and generate medical advice results through the constructed advice auxiliary model.

[0218] Optionally, the medical advice result determination sub-module is specifically configured to: extract patient medical-related data from the second data to be evaluated, and calculate the patient's medical cost according to Extract medical resource-related data from the second data to be evaluated, and calculate the medical resource utilization efficiency according to Generate a medical advice result by using a multi-objective optimization algorithm based on the patient's medical cost and the medical resource utilization efficiency; where c j Is the jth medical cost, x j Is the decision variable for whether to select medical services, x j Is 0 or 1, e k Is the utilization efficiency weight of the kth medical resource, z k Is the decision variable for whether to use medical resources, z k Is 0 or 1, continuously iterate and optimize the decision variables through a genetic algorithm to obtain the optimal medical treatment and physical examination advice as the medical advice result.

[0219] Optionally, the disease control information management module includes:

[0220] A dimension analysis sub-module for using geographic information system technology to analyze the data to be analyzed containing geographical location information in the first data to be evaluated, and performing factor correlation analysis based on the first data to be evaluated to obtain the correlation dimension information of each data;

[0221] A visualization sub-module for performing visualization processing of relevance based on the correlation dimension information by using the first data to be evaluated to obtain a relevance visualization chart;

[0222] An infectious disease prevention and control management sub-module for constructing an auxiliary decision-making platform based on the established infectious disease prediction and early warning model, rendering the relevance visualization chart and the infectious disease transmission prediction result on the visualization operation interface of the auxiliary decision-making platform, and performing prevention and control feedback management on the infectious disease transmission prediction result;

[0223] A data fusion sub-module for using data fusion technology to perform feature fusion on the features identified from the second data to be evaluated, the medical advice result, and the feature recognition result according to F = [F1, F2, F3] or F =

[0224] w1F1 + w2F2 + w3F3 to obtain the target medical advice;

[0225] The chronic disease information management sub-module is used to display the target medical advice on the visualization page and manage the disease information of chronic diseases; among them, the prevention and control feedback management is used to provide decision-making support including the best prevention and control measures. F1 is a feature extracted from the medical advice result, F2 is the feature corresponding to the feature recognition result, F3 is a feature extracted from the second data to be evaluated, and w1, w2, and w3 are all corresponding preset weights, and w1 + w2 + w3 = 1.

[0226] Optionally, the visualization sub-module includes:

[0227] The first visualization unit is used to perform visualization in the spatial dimension and the time dimension according to the relevance between the data in the first data to be evaluated in the preset map data, and obtain a spatial distribution chart and a time trend change chart;

[0228] The variable analysis unit is used to extract the relevant variable data from the first data to be evaluated based on the correlation dimension information and analyze the potential data relationship between the variable data;

[0229] The second visualization unit is used to render each variable data in the visualization chart in the form of a scatter plot matrix based on the potential data relationship to obtain a scatter plot;

[0230] The relevance visualization chart determination unit is used to determine the relevance visualization chart according to the spatial distribution chart, the time trend change chart, and the scatter plot.

[0231] Optionally, the disease information management system based on multi-source data further includes:

[0232] The sample set collection and preprocessing module is used to obtain the first sample data and historical data related to infectious diseases, and obtain the second prediction data related to chronic diseases; preprocess the first sample data and the second sample data respectively to obtain the first training sample to be trained, the historical data corresponding to the first training sample to be trained, and the second training sample to be trained including time series;

[0233] The knowledge graph analysis module is used to perform knowledge graph analysis according to the historical data and the first training sample to be trained to obtain knowledge graph data;

[0234] The model construction module is used to construct an infectious disease prediction and early warning model and a chronic disease prediction and decision-making model by using the recurrent neural network and the memory network in deep learning, and preset a weight matrix and parameters in the model;

[0235] A model training module, configured to train the infectious disease prediction and early warning model based on the first sample to be trained, the historical data, and the knowledge graph data, and extract features and train the chronic disease prediction and decision-making model according to the second sample to be trained;

[0236] A model evaluation and optimization module, configured to evaluate the models by using the prediction results output by the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model according to the preset evaluation indicators corresponding to each model; wherein, during the construction of the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model, a forgetting gate, an input gate, an output gate, a memory cell update, and a hidden layer state are respectively preset.

[0237] It should be noted that the multi-source data-based disease information management system provided in the embodiments of the present application can execute the multi-source data-based disease information management method provided in any embodiment of the present application, and has the corresponding functions and beneficial effects of executing the method.

[0238] In specific implementation, the above multi-source data-based disease information management system can be integrated into a device, so that the device can obtain multi-source heterogeneous data with multiple touch points. On the one hand, it can conduct prediction and evaluation of infectious disease transmission, and on the other hand, it can provide prediction suggestions for chronic diseases. As an electronic device, it can achieve perfect disease prevention and control information management. The electronic device can be composed of two or more physical entities, or can be composed of one physical entity. For example, the electronic device can be a personal computer (PC), a computer, a server, etc. The embodiments of the present application do not make specific limitations in this regard.

[0239] Such as Figure 6As shown in the figure, an embodiment of the present application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114; the memory 113 is used to store computer programs; when the processor 111 executes the programs stored on the memory 113, it implements the steps of the disease information management method based on multi-source data provided by any one of the foregoing method embodiments. Exemplarily, the steps of the disease information management method based on multi-source data may include the following steps: obtaining infectious disease prediction data of a target area from a preset first data source, and combining the chronic disease-related data obtained from the second data source to form multi-source heterogeneous data; taking each data feature in the multi-source heterogeneous data as a benchmark, extracting corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing, to obtain first data to be evaluated corresponding to the infectious disease prediction data and second data to be evaluated corresponding to the chronic disease-related data; performing risk prediction on the first data to be evaluated through a preset logistic regression model, and selecting disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information; selecting the data containing time series from the first data to be evaluated as input, and performing transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result; respectively performing prediction evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset advice assistance model to obtain a feature recognition result output by the chronic disease prediction decision model and a medical advice result output by the advice assistance model; based on the first data to be evaluated and the second data to be evaluated, combining the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result for disease control information management.

[0240] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the disease information management method based on multi-source data provided by any one of the foregoing method embodiments.

[0241] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0242] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A disease information management method based on multi-source data, characterized in that Including: Obtain the infectious disease prediction data of the target area from a preset first data source, and combine the chronic disease-related data obtained from the second data source to form multi-source heterogeneous data; Based on each data feature in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing, and obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data; Perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information; Select the data containing time series from the first data to be evaluated as input, and perform transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result; Perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset advice assistance model respectively, and obtain a feature recognition result output by the chronic disease prediction decision model and a medical advice result output by the advice assistance model; Based on the first data to be evaluated and the second data to be evaluated, combine the risk assessment information, the transmission trend prediction result, the medical advice result, and the feature recognition result for disease control information management.

2. The method according to claim 1, characterized in that, Based on each data feature in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing, and obtain the first data to be evaluated corresponding to the infectious disease prediction data and the second data to be evaluated corresponding to the chronic disease-related data, including: Analyze the infectious disease prediction data features in the multi-source heterogeneous data to obtain data features, and the data features include time features, association features, and data type features; Based on the data features, perform data preprocessing on the infectious disease prediction data according to preset data integration and preprocessing rules to obtain the first data to be evaluated; Remove outliers and errors from the chronic disease-related data, screen out the missing items with data missing from the chronic disease-related data, and select the data to be fused from the chronic disease-related data; Adopt a data filling algorithm, taking the same type of data at adjacent time points of the missing item as the benchmark, and calculate the filled data of the missing item according to the calculation; Using a data fusion algorithm, according to Fuse different data by weighted average method to obtain fused data; Based on the chronic disease-related data, the supplementary data, and the fused data, determine the second data to be evaluated; Among them, t1 and t2 are time points adjacent to the missing item, G1 is the quantity corresponding to the time point t1, G2 is the quantity corresponding to the time point t2, G is the supplementary data calculated corresponding to the time point t, and t1 < t < t2, E i represents the i-th data to be fused, w i is the weight corresponding to the data to be fused.

3. The method according to claim 2, wherein Based on the data features, perform data preprocessing on the infectious disease prediction data according to preset data integration and preprocessing rules to obtain the first data to be evaluated, including: Extract each data corresponding to the time feature from the infectious disease prediction data for time format alignment to obtain the first feature processing data with unified data time dimension; According to the preset data association rules, select the association keys representing the association features from the first feature processing data, and based on the association keys, merge the data corresponding to the association features and related to the same subject in the first feature processing data to obtain the second feature processing data; Using the Z-score normalization method, according to perform data standardization on the numerical data corresponding to the data type features in the second feature processing data, extract the text data belonging to the data type features from the second feature processing data, and perform classification transformation on the text data to obtain the third feature processing data; According to the preset data threshold and outlier detection algorithm, perform error and duplicate removal processing on the third feature processing data to obtain the first data to be evaluated; Wherein, X is the original data value, μ is the data mean, and σ is the data standard deviation.

4. The method according to claim 1, wherein Perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information, including: Select feature variables related to the spread of infectious diseases from the first data to be evaluated, and the feature variables at least include the growth rate of medical consultations, the growth rate of drugs, and the absenteeism rate caused by infectious diseases; Using the characteristic variables, according to the formula Construct a logistic regression model; Perform risk prediction on the first data to be evaluated through the logistic regression model to obtain risk area assessment information of the target area; Using the K-means clustering algorithm, extract the personnel signaling data, entry-exit data, and patient-related data from the first data to be evaluated as the samples to be evaluated, and according to perform clustering analysis to obtain the objective function; Analyze according to the objective function to obtain the risk classification level; Use the risk area assessment information and the risk classification level as risk assessment information; Among them, K is the preset number of clusters, and the goal is to minimize the sum of the squares of the distances from each sample to the center of its belonging cluster. J is the objective function, C i is the i-th cluster, and μ i is the center of the i-th cluster, x j is the sample data, that is, the sample to be evaluated. P(Y = 1) represents the probability that the area is a risk area. Y is the dependent variable, Y = 1 represents the risk area, and Y = 0 represents the non-risk area. X i is the independent variable, that is, the selected feature variable, and β i is the model parameter, which is solved by the maximum likelihood estimation method.

5. The method according to claim 1, wherein Perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset advice assistance model respectively, and obtain the feature recognition result output by the chronic disease prediction decision model and the medical advice result output by the advice assistance model, including: Adopt a deep learning algorithm to construct a chronic disease prediction decision model, and select patient detection data with time series from the second data to be evaluated as the model input; The chronic disease prediction decision model extracts features and patterns from the patient detection data to predict the development trend of chronic diseases and obtain a feature recognition result; Taking the second data to be evaluated as the input, with the goal of minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources, adopt a multi-objective optimization algorithm, and generate a medical advice result through the constructed advice assistance model.

6. The method according to claim 5, wherein Taking the second data to be evaluated as the input, with the goal of minimizing the patient's medical cost and maximizing the utilization efficiency of medical resources, adopt a multi-objective optimization algorithm, and generate a medical advice result through the constructed advice assistance model, including: Extract the patient's medical treatment-related data from the second data to be evaluated, and calculate the patient's medical treatment cost according to ​ Extract medical resource-related data from the second data to be evaluated, and calculate the medical resource utilization efficiency according to ​ Generate a medical advice result according to the patient's medical cost and the utilization efficiency of medical resources by adopting a multi-objective optimization algorithm; Among them, c j is the j-th medical cost, x j is the decision variable for whether to choose medical services, x j is 0 or 1, e k is the utilization efficiency weight of the k-th medical resource, z k is the decision variable for whether to use medical resources, z k is 0 or 1. By continuously iterating and optimizing the decision variables through the genetic algorithm, the optimal medical treatment and physical examination suggestions are obtained as the results of medical treatment suggestions.

7. The method according to claim 1, characterized in that, Based on the first data to be evaluated and the second data to be evaluated, combine the risk assessment information, the spread trend prediction result, the medical advice result, and the feature recognition result for disease control information management, including: Use geographic information system technology to analyze the data to be analyzed containing geographic location information in the first data to be evaluated, and perform factor correlation analysis based on the first data to be evaluated to obtain the correlation dimension information of each data; Taking the correlation dimension information as the benchmark, perform visualization processing of the correlation using the first data to be evaluated to obtain a correlation visualization chart; Based on the established infectious disease prediction and early warning model, construct an auxiliary decision-making platform, and render the correlation visualization chart and the infectious disease spread prediction result on the visualization operation interface of the auxiliary decision-making platform, and perform prevention and control feedback management on the infectious disease spread prediction result; Adopt data fusion technology to perform feature fusion on the features identified from the second data to be evaluated, the medical advice result, and the feature recognition result according to F = [F1, F2, F3] or F = w1F1 + w2F2 + w3F3 to obtain the target medical advice; Display the target medical advice on the visualization page for disease information management of chronic diseases; Among them, the prevention and control feedback management is used to provide decision-making support including the best prevention and control measures. F1 is a feature extracted from the medical advice result, F2 is the corresponding feature of the feature recognition result, F3 is a feature extracted from the second data to be evaluated, and w1, w2, and w3 are all corresponding preset weights, and w1 + w2 + w3 = 1.

8. The method according to claim 7, wherein Based on the associated dimension information, perform visualization processing of the association using the first data to be evaluated to obtain an association visualization chart, including: In the preset map data, perform visualization in the spatial dimension and time dimension according to the association between the data in the first data to be evaluated to obtain a spatial distribution chart and a time trend change chart; Based on the associated dimension information, extract the variable data of the association from the first data to be evaluated and analyze the potential data relationship between the variable data; Based on the potential data relationship, use the scatter plot matrix method to render the variable data in the visualization chart to obtain a scatter plot; Determine the association visualization chart according to the spatial distribution chart, the time trend change chart, and the scatter plot.

9. The method according to claim 1, characterized in that, It also includes: Obtain the first sample data and historical data related to infectious diseases, and obtain the second prediction data related to chronic diseases; Preprocess the first sample data and the second sample data respectively to obtain the first training sample to be trained, the historical data corresponding to the first training sample to be trained, and the second training sample to be trained including time series; Perform knowledge graph analysis according to the historical data and the first training sample to be trained to obtain knowledge graph data; Use the recurrent neural network and memory network in deep learning to construct an infectious disease prediction and early warning model and a chronic disease prediction and decision-making model, and preset a weight matrix and parameters in the model; Based on the first training sample to be trained, the historical data, and the knowledge graph data, train the infectious disease prediction and early warning model, and extract features and train the chronic disease prediction and decision-making model according to the second training sample to be trained; According to the evaluation indicators preset for each model, use the prediction results output by the infectious disease prediction and early warning model to evaluate the model, and use the prediction results output by the chronic disease prediction and decision-making model to evaluate the model; Optimize the corresponding model parameters according to the model evaluation results until each model completes model training; Among them, during the construction of the infectious disease prediction and early warning model and the chronic disease prediction and decision-making model, a forgetting gate, an input gate, an output gate, memory unit update, and hidden layer state are preset respectively.

10. A disease information management system based on multi-source data, characterized in that, It includes: A multi-source heterogeneous data acquisition module, which is used to acquire the infectious disease prediction data of the target area from the preset first data source and combine it with the chronic disease-related data acquired from the second data source to form multi-source heterogeneous data; A preprocessing module, configured to benchmark with each data feature in the multi-source heterogeneous data, extract corresponding data from the multi-source heterogeneous data according to preset data integration and preprocessing rules for data preprocessing, and obtain first data to be evaluated corresponding to infectious disease prediction data and second data to be evaluated corresponding to chronic disease-related data; A risk assessment and prediction module, configured to perform risk prediction on the first data to be evaluated through a preset logistic regression model, and select disease condition data from the first data to be evaluated for clustering analysis to obtain risk assessment information; A transmission prediction module, configured to select data including time series from the first data to be evaluated as input, and perform transmission trend prediction through a preset infectious disease prediction and early warning model to obtain a transmission trend prediction result; A recommendation prediction module, configured to perform prediction and evaluation on the second data to be evaluated through a preset chronic disease prediction decision model and a preset recommendation assistance model respectively, and obtain a feature recognition result output by the chronic disease prediction decision model and a medical treatment recommendation result output by the recommendation assistance model; A disease control information management module, configured to perform disease control information management based on the first data to be evaluated and the second data to be evaluated, in combination with the risk assessment information, the transmission trend prediction result, the medical treatment recommendation result, and the feature recognition result.

Citation Information

Cited By

  • Disease control big data analysis method and system

    CN121506533A