Big Data-Based Analysis Method for the Risk of Cross-Border Transmission of Animal Diseases
Through multimodal data fusion and infectious disease transmission analysis, the data heterogeneity and timeliness in cross-border animal epidemic transmission risk analysis are solved, efficient risk assessment and prevention and control measures are achieved, and the scientificity and intelligence level of epidemic prevention and control are improved.
Patent Information
- Application Number
- CN202510324280.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-19
AI Technical Summary
In the existing cross-border animal disease transmission risk analysis methods, the processing complexity of multi-source heterogeneous data, data quality assurance and risk prediction accuracy are insufficient, which has affected the effectiveness of epidemic prevention and control measures.
A multimodal data fusion algorithm based on deep learning is used to integrate epidemiological monitoring data, social media information, meteorological data and shipping logistics information, build a cross-border transmission path model, combine infectious disease transmission model and GIS technology, and use machine learning algorithms to evaluate cross-border transmission risks, and propose prevention and control intervention measures.
It improves the accuracy and real-time analysis of cross-border animal epidemic transmission risk, can detect potential risks in the early stage, take precise prevention and control measures, reduce the risk of epidemic spread, and improve public health security.
Smart Images

Figure CN119851971B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to a method for analyzing the risk of cross-border transmission of animal diseases based on big data. Background Art
[0002] With the continuous advancement of globalization, the cross-border transmission of animal diseases has become an important issue in the field of global public health. Factors such as large-scale animal trade, personnel flow, and climate change have led to the rapid spread of animal diseases across national borders, bringing serious impacts on global agricultural production and ecosystems. Traditional animal disease monitoring systems mostly focus on single monitoring points and local transmission chains, and are unable to comprehensively and dynamically predict and evaluate the risks of cross-border transmission effectively.
[0003] In recent years, with the rapid development of big data technology, especially the wide application of multi-dimensional data such as sensor technology, satellite remote sensing, social media information, and meteorological data, related fields have begun to attempt to use big data analysis methods for the monitoring and prediction of animal diseases. By integrating multi-source heterogeneous data, the epidemic trends, transmission paths, and potential transmission risks of animal diseases can be obtained in real time. However, existing cross-border transmission risk analysis methods have many deficiencies, mainly reflected in aspects such as data processing complexity, data quality assurance, and risk prediction accuracy.
[0004] The existing technologies have the following deficiencies:
[0005] The data types involved in the cross-border transmission of animal diseases are very extensive, including both real-time epidemiological data from sensors and monitoring devices, unstructured data from social media, news reports, etc., and external environment data from meteorological satellites, shipping logistics, etc. These data have significant differences in formats, granularities, timeliness, etc. How to effectively integrate and process these multi-source heterogeneous data is a highly challenging technical problem. Especially in cross-border transmission risk analysis, the timeliness and accuracy of data directly affect the effectiveness of risk prediction. If these data cannot be processed in real time, or different sources of data cannot be effectively cleaned and docked, the analysis results may have large errors, affecting the implementation of epidemic prevention and control measures. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for analyzing the risk of cross-border transmission of animal diseases based on big data to solve the deficiencies in the background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions: A method for analyzing the risk of cross-border transmission of animal diseases based on big data, including:
[0008] Real-time obtain heterogeneous data including epidemiological surveillance data, social media information, meteorological data, and shipping and logistics information, and use a multi-modal data fusion algorithm based on deep learning to integrate heterogeneous data from different sources to form a unified data set;
[0009] Build a cross-border transmission path model based on the infectious disease transmission model and epidemiological theory, and combine the spatio-temporal information and transmission characteristics in big data to simulate the potential transmission paths of animal diseases;
[0010] During the simulation process, use machine learning algorithms to perform fusion processing and analysis on the integrated heterogeneous data, and evaluate the degree of fusion of cross-border transmission heterogeneous data between different regions;
[0011] Based on the fusion evaluation results of multi-source data, real-time predict the high-risk areas and transmission speeds of the cross-border transmission of animal diseases, judge the potential cross-border transmission risks according to the prediction results, and propose corresponding prevention and control intervention measures.
[0012] Preferably, the multi-modal data fusion algorithm uses a deep learning model, including but not limited to convolutional neural network, recurrent neural network, multi-layer perceptron or Transformer model, and fuses multi-source heterogeneous data through feature extraction and attention mechanism.
[0013] Preferably, the constructed cross-border transmission path model is based on an improved SIR infectious disease transmission model and combines with a GIS geographic information system to dynamically simulate the disease transmission paths under different time and space conditions.
[0014] Preferably, after analyzing the migration paths and migration times of different animal species, a species migration anomaly index is generated. The acquisition method of the species migration anomaly index is as follows:
[0015] There is a directed graph G=(V,E), where: V represents the set of all habitats, are all habitat nodes, and E represents the set of migration paths between habitats, is the migration path from habitat to habitat ; represents the weight of edge , that is, the migration intensity from to ;
[0016] Set the PageRank values of all nodes to be equal initially, that is , where n is the number of habitats; The iterative formula for the PageRank value is: ; where: is habitat The PageRank value; α is the damping factor, representing the probability that a viewer jumps to a random node. Indicates pointing to the habitat All habitat sets; Indicates from the habitat All migration path sets emitted; Is from the habitat To The migration intensity; Initialize the PageRank value of each habitat, update the PageRank value of each habitat according to the iterative formula of the PageRank value, repeat the update until the PageRank value converges. Once the PageRank value of each habitat is calculated, calculate the mean μ and standard deviation σ of the PageRank values of all habitats, and calculate the species migration anomaly index. The expression is: ; In the formula, AE is the species migration anomaly index.
[0017] Preferably, after analyzing the differences in logistics flow in different regions, a logistics flow difference index is generated. The acquisition method of the logistics flow difference index is:
[0018] Obtain the inbound and outbound volume Lij, the number of transportation times Li, and the total transportation volume Q of a certain region within different time ranges; Calculate the mean of the logistics flow of all regions, that is, the average level of the logistics flow of all regions. The expression is: ; Among them: Is the mean of the logistics flow of all regions, m is the total number of regions. The standard deviation of the logistics flow measures the fluctuation of the logistics flow. The expression is: ; Among them: Is the standard deviation of the logistics flow, Is the total logistics flow of region i, Is the mean of the logistics flow of all regions; Obtained by calculating the difference normalization of the logistics flow of each region with respect to the overall flow mean. The expression is: ; Among them: DI is the logistics flow difference index.
[0019] Preferably, convert the species migration anomaly index and the logistics flow difference index into a comprehensive feature vector, use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes predicting the fusion degree value label of cross-border transmission heterogeneous data between different regions as the prediction target, and takes minimizing the sum of the prediction errors of the fusion degree value labels of cross-border transmission heterogeneous data between all different regions as the training target. Train the machine learning model until the sum of the prediction errors reaches convergence and then stop the model training. Determine the fusion degree value of cross-border transmission heterogeneous data between different regions according to the model output result. Among them, the machine learning model is a polynomial regression model.
[0020] Preferably, compare the obtained fusion degree value of cross-border transmission heterogeneous data between different regions with a preset threshold. If the fusion degree value of cross-border transmission heterogeneous data is greater than or equal to the preset threshold, it indicates that the fusion degree of cross-border transmission heterogeneous data is high, and at this time, a data fusion normal signal is generated; if the fusion degree value of cross-border transmission heterogeneous data is less than the preset threshold, it indicates that the fusion degree of cross-border transmission heterogeneous data is low, and at this time, a data fusion abnormal signal is generated.
[0021] Preferably, define the speed at which an animal disease spreads from one region to another as the transmission speed S, mark the high-risk region as Ahigh, which represents the region with a high transmission risk, and define the infection probability of the high-risk region as R;
[0022] Through the fusion of multi-source data, the fusion degree value LR of cross-border transmission heterogeneous data between different regions is obtained, and a mathematical formula based on the infectious disease transmission model will be used to predict the transmission speed and risk;
[0023] The calculation formula for the transmission speed is: ; where: represents the infectious ability of the source region of transmission, represents the susceptibility of the receiving region;
[0024] The calculation expression for the cross-border transmission risk R is: ; where: represents the transmission probability of the source region, represents the input probability of the receiving region;
[0025] The determination of the high-risk region is based on the predicted value of the risk R. If the R value of a certain region exceeds the preset threshold Trisk, the region is regarded as a high-risk region: Once the high-risk region Ahigh and the transmission speed S are obtained, further judge the potential cross-border transmission risk, and the potential cross-border transmission risk is evaluated through the formula, and the formula is: ; where: represents the potential cross-border transmission risk, which is the accumulation of the transmission risks of all high-risk regions According to the potential risk , different prevention and control intervention measures are proposed.
[0026] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0027] 1. Through multi-source data fusion, deep learning modeling, and infectious disease transmission analysis, the present invention effectively solves the problems of data heterogeneity, timeliness, and accurate prediction in the risk analysis of cross-border transmission of animal diseases. By using a multi-modal data fusion algorithm, heterogeneous data such as epidemiological surveillance data, social media information, meteorological data, and shipping logistics information are integrated to form a high-quality data set, ensuring the comprehensiveness and timeliness of information. Through an improved SIR infectious disease transmission model combined with GIS technology, the cross-border transmission path of the disease is dynamically simulated to improve the prediction accuracy of the epidemic transmission trend. Further, based on a machine learning algorithm, the species migration anomaly index and the logistics flow difference index are calculated to quantify the degree of data fusion between different regions. Finally, a polynomial regression model is used to realize the intelligent assessment of cross-border transmission risks, and corresponding prevention and control measures are proposed according to the risk level.
[0028] 2. By calculating the disease transmission speed and cross-border transmission risk index in real time, this method can actively generate normal or abnormal signals for data fusion, quickly discover potential risks in the early stage of the epidemic, and take precise prevention and control measures such as regional lockdown, strengthened quarantine, vaccination, etc., to reduce the risk of epidemic spread. Overall, the present invention significantly improves the intelligent, automated, and scientific level of cross-border animal disease transmission risk assessment, providing a strong guarantee for public health security. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0032] Embodiment, please refer to Figure 1 As shown, the method for analyzing the risk of cross-border transmission of animal diseases based on big data in this embodiment includes:
[0033] Real-time obtain heterogeneous data including epidemiological surveillance data, social media information, meteorological data, and shipping and logistics information, and adopt a multi-modal data fusion algorithm based on deep learning to integrate heterogeneous data from different sources to form a unified data set;
[0034] Construct a cross-border transmission path model based on the infectious disease transmission model and epidemiological theory, and combine the spatio-temporal information and transmission characteristics in big data to simulate the potential transmission paths of animal diseases;
[0035] During the simulation process, use machine learning algorithms to perform fusion processing and analysis on the integrated heterogeneous data, and evaluate the degree of fusion of cross-border transmission heterogeneous data between different regions;
[0036] Based on the fusion evaluation results of multi-source data, real-time predict the high-risk regions and transmission speed of the cross-border transmission of animal diseases. According to the prediction results, judge the potential cross-border transmission risks and propose corresponding prevention and control intervention measures.
[0037] The specific steps for real-time obtaining and fusing heterogeneous data include:
[0038] Real-time obtain epidemiological data such as animal disease case data, epidemic area information, transmission speed, and infection status through various sensors, monitoring devices, and public reports of health organizations.
[0039] Use natural language processing (NLP) technology to crawl public discussions, posts, news reports, etc. related to animal diseases from social media platforms (such as Twitter, Facebook, Weibo, etc.), and extract information related to the epidemic (such as the source of infection, epidemic area, transmission route, etc.).
[0040] Obtain meteorological data related to the transmission of animal diseases, including meteorological parameters such as temperature, humidity, precipitation, and wind speed, which may affect the spread of diseases or the occurrence of emergencies.
[0041] Obtain information such as cross-border animal trade flows, transportation routes, types of goods, and transportation timeliness from shipping companies, logistics platforms, customs, border monitoring, etc. These data are crucial for predicting the transmission paths of animal diseases.
[0042] Structured data (such as sensor data, meteorological data, logistics information, etc.) usually has a relatively fixed format. Ensure the accuracy of the data through methods such as data cleaning, denoising, outlier processing, and data standardization. Unstructured data (such as social media information, news reports, etc.) needs to be processed using text mining techniques, sentiment analysis, keyword extraction, etc. to convert the information into structured data and extract useful information related to animal diseases.
[0043] According to the timeliness requirements of the data, timestamp the real-time data to ensure that all data sources are synchronized within a unified time window, and avoid analysis errors caused by timeliness differences.
[0044] Select appropriate deep learning algorithms: Use deep learning-based multimodal data fusion algorithms (such as multi-layer perceptron (MLP), convolutional neural network (CNN), or recurrent neural network (RNN), etc.) to design a fusion model. Convolutional neural network (CNN): Suitable for processing data with spatial relationships, such as meteorological data or epidemic data related to geographical locations. Recurrent neural network (RNN): Suitable for processing time-series data, such as real-time epidemic spread data, shipping logistics data, meteorological data, etc. Transformer model: Suitable for parallel processing of multimodal data, and can handle multiple time-series data and complex multi-source heterogeneous data sets.
[0045] By encoding the data from different data sources, convert different types of heterogeneous data such as epidemiological data, social media information, meteorological data, and logistics information into a unified embedding space, ensuring that various types of data can be effectively compared and fused within the same feature space.
[0046] Use the feature extraction layer in the deep learning model to extract key features (such as epidemic outbreak areas, transmission speeds, weather changes, etc.) from each data source, and perform alignment processing to ensure that the data from different sources can be effectively corresponding.
[0047] Through deep learning methods such as weighting, concatenation, and attention mechanisms, fuse the features extracted from different data sources. Especially when facing heterogeneous data sources such as meteorology, social media, and logistics, adopt the multi-head attention mechanism or weighted fusion algorithm to automatically adjust the importance of different data sources.
[0048] Integrate the fused multimodal features into a unified data set (usually in tensor or vector format), and each data item contains comprehensive information from different sources, providing input for subsequent model training and analysis.
[0049] Conduct quality assessment on the fused data set, including inspections of data accuracy, integrity, consistency, etc. Ensure the high quality of the data set by calculating indicators such as the correlation between data sources and information redundancy.
[0050] Verify the effect of the fusion algorithm through cross-validation and test set, evaluate the improvement of data fusion on model performance, and ensure that the fused data can accurately reflect the transmission law of animal diseases.
[0051] To ensure the timeliness and accuracy of the dataset, a dynamic update mechanism is adopted to obtain the latest epidemiological surveillance data, social media information, meteorological data, and shipping and logistics information in real time, and the deep learning model is adjusted and optimized in real time through online learning methods (such as incremental learning or online training). Through the feedback and monitoring of the prediction results, the deep learning model and data fusion algorithm are continuously improved to make them more adaptable to the changes in the epidemic situation and new data sources.
[0052] Construct a cross-border transmission path model based on the infectious disease transmission model and epidemiological theory, specifically:
[0053] SIR model: The classic SIR (Susceptible-Infected-Recovered) model is suitable for describing the spread of infectious diseases in a population, where: S (Susceptible): refers to individuals who have not been infected and may be infected by contact with infected individuals. I (Infected): infected individuals who can transmit the disease. R (Recovered): refers to individuals who have recovered and have immunity or have died.
[0054] In the cross-border transmission path model, an improved SIR model or its variants (such as SEIR, SIS models, etc.) is adopted to consider more complex transmission processes, such as incubation period, loss of immunity, and other factors.
[0055] Considering the geographical transmission characteristics of the epidemic, diffusion models (such as Laplacian model, random walk model, spread model, etc.) are used to analyze the transmission paths between different regions.
[0056] Based on the transmission dynamics theory in epidemiology, analyze how diseases cross the border through different transmission routes (such as animal trade, traffic flow, human activities, etc.).
[0057] Basic reproduction number R0: Introduce the concept of the basic reproduction number (R0) to calculate the disease transmission potential of each region. An R0 value greater than 1 indicates that the epidemic is likely to spread in that region, and vice versa indicates that the epidemic is likely to terminate. Based on the change of the R0 value of the cross-border transmission path, high-risk regions can be determined.
[0058] Consider the geographical, climatic, socio-economic, and cultural characteristics of different regions, which affect the characteristics of the spread of animal diseases. For example, climate change may affect the seasonality of the spread of diseases, and trade flows and population gatherings may affect the spread speed and scope of the epidemic.
[0059] Use spatio-temporal information in big data, such as meteorological data, logistics flow, animal movement, epidemic history, etc., to enhance the spatio-temporal dimension of the model. Use big data to model the epidemic trends in different regions, times, and spaces.
[0060] Analyze the geographical transmission paths of animal diseases by combining GIS technology, integrating disease information with geographical coordinates, boundary data, trade routes, etc., and conduct visualization and dynamic simulation of cross-border transmission paths.
[0061] Construct a network model of disease transmission. By analyzing the transmission relationships between various nodes (such as countries, regions, animal populations), predict the cross-border transmission paths of diseases. Through a weighted graph structure, analyze the transmission risks and key transmission nodes in each region.
[0062] Combine meteorological data (such as temperature, humidity, precipitation, etc.) and the external environment (such as information on key logistics hubs like seaports and airports) to construct a cross-border transmission model. External environment data helps identify climate conditions prone to transmission, animal transportation routes, and key transmission nodes.
[0063] Based on the established cross-border transmission path model, combine spatio-temporal data and transmission characteristics to simulate the transmission paths of animal diseases in different time periods and different geographical regions. Use multi-objective optimization methods (such as particle swarm algorithm, genetic algorithm) to optimize the transmission paths, and combine actual factors such as transportation flow, trade flow, and population flow to predict the possible routes of disease transmission.
[0064] Use the infectious disease transmission model and big data to simulate the speed and direction of the epidemic spread. Based on the interconnections between various nodes, predict the spread speed of the epidemic, and continuously adjust the prediction results of the transmission paths according to real-time data.
[0065] Animal migration behaviors (such as migration paths, cycles, habitat changes, etc.) have an important impact on the transmission of animal diseases. Especially in cross-border transmission, migrating animals may become new sources of transmission. Use the global animal migration database (such as animal GPS tracking data, ecological monitoring data, etc.) to obtain the migration paths and migration times of different animal species. Integrate animal activity data and disease infection rate information to establish an association model between animal migration and disease transmission. Analyze the migration patterns of different regions and species, and evaluate their transmission risks under different seasons and climate changes.
[0066] Use the global animal migration database (such as animal GPS tracking data, ecological monitoring data, etc.) to obtain the migration paths and migration times of different animal species. Integrate animal activity data and disease infection rate information to establish an association model between animal migration and disease transmission. Analyze the migration patterns of different regions and species, and evaluate their transmission risks under different seasons and climate changes.
[0067] After analyzing the migration paths and migration times of different animal species, generate a species migration anomaly index. The method for obtaining the species migration anomaly index is as follows:
[0068] Define the graph structure. Nodes: habitats, migration routes, or ecological blocks. Edges: the migration paths of species from one node (habitat) to another node (habitat). Edge weight: The edge weight can be defined according to factors such as migration flow, animal quantity, and migration duration. The greater the weight, the more frequent the migration activity, and vice versa. Let there be a directed graph G=(V,E), where: V represents the set of all habitats, are all habitat nodes, and E represents the set of paths for migration between habitats, is the migration path from habitat to habitat . represents the weight of the edge , that is, the migration intensity from to .
[0069] The calculation of the PageRank value is based on the random walk model. Assume that a random "browser" jumps from one habitat to another with a certain probability. The core idea of the algorithm is that the importance of each habitat comes from the "recommendation" of its neighbor nodes (migration paths). Initially, set the PageRank values of all nodes to be equal, that is , where n is the number of habitats; the iterative formula for the PageRank value is: ; where: is the PageRank value of habitat ; α is the damping factor (usually set to 0.15), representing the probability that the browser jumps to a random node. represents the set of all habitats pointing to habitat (that is, the habitats migrating to ); represents the set of all migration paths starting from habitat (that is, the habitats starting from ); is the migration intensity from habitat to .
[0070] Initialize the PageRank value of each habitat.
[0071] Update the PageRank value of each habitat according to the iterative formula of the PageRank value.
[0072] Repeat the update until the PageRank value converges (that is, the change amplitude is less than the set threshold).
[0073] Once the PageRank values of each habitat are calculated, these values can be used to evaluate migration anomalies. Calculate the mean μ and standard deviation σ of the PageRank values of all habitats, and calculate the species migration anomaly index, with the expression: ; where AE is the species migration anomaly index, and the larger the value, the more abnormal the migration behavior of the habitat and the more deviated from the normal migration pattern.
[0074] The flow patterns and frequencies of cross-border logistics flow data (including animal trade, veterinary drugs, feed transportation, etc.) are not only closely related to the spread speed and path of the epidemic, but also can reflect the epidemic transmission risks in the trade and transportation processes. There are significant differences in the logistics flow among different regions, especially the heterogeneity of factors such as different trade agreements, customs control measures, and cross-border transportation methods, which also need to be carefully captured and analyzed.
[0075] By collecting trade and logistics flow data of different countries (such as customs data, international trade data, transportation route information, etc.), analyze the epidemic transmission paths in the cross-border transportation process. Obtain different types of cross-border logistics data (such as animal trade, veterinary product transportation, etc.), conduct spatio-temporal pattern analysis on them, and evaluate which logistics routes become high-risk channels for epidemic transmission. Consider the impact of transportation timeliness, transportation density, and the heterogeneity of transportation routes on the epidemic spread.
[0076] After analyzing the differences in the logistics flow among different regions, generate a logistics flow difference index. The method for obtaining the logistics flow difference index is:
[0077] Obtain the inbound and outbound volume Lij, transportation times Li, and total transportation volume Q of a certain region within different time ranges; calculate the mean of the logistics flow of all regions, that is, the average level of the logistics flow of all regions, with the expression: ; where: is the mean of the logistics flow of all regions, and m is the total number of regions. The standard deviation of the logistics flow measures the fluctuation of the logistics flow. Through the standard deviation calculation, the dispersion degree of the logistics flow of each region relative to the average level can be understood, with the expression: ; where: is the standard deviation of the logistics flow, is the total logistics flow of region i. is the mean of the logistics flow of all regions.
[0078] The logistics flow difference index is an indicator used to represent the difference degree of the logistics flow among different regions. It is obtained by calculating the difference of the logistics flow of each region relative to the overall flow mean and standardizing it, with the expression: ; where: DI is the logistics flow difference index.
[0079] High-difference index region: If DI > 2 (or other thresholds), it can be considered that there are significant abnormal fluctuations in the logistics flow in the region, which may be caused by logistics bottlenecks, surging demand, traffic disruptions, etc. Low-difference index region: If DI is close to zero, it indicates that the logistics flow in this region is similar to that in other regions and the flow is relatively stable.
[0080] Convert the species migration anomaly index and the logistics flow difference index into a comprehensive feature vector, and use the comprehensive feature vector as the input of the machine learning model. The machine learning model takes the prediction of the fusion degree value label of the cross-border transmission heterogeneous data between different regions for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of the prediction errors of the fusion degree value labels of the cross-border transmission heterogeneous data between all different regions as the training target to train the machine learning model until the sum of the prediction errors reaches convergence and then stops the model training. Determine the fusion degree value of the cross-border transmission heterogeneous data between different regions according to the model output result, where the machine learning model is a polynomial regression model.
[0081] The method for obtaining the fusion degree value of the cross-border transmission heterogeneous data between different regions is: obtain the corresponding function expression from the comprehensive feature vector training data of the trained machine learning model: ; where is the output function of the model, AE is the species migration anomaly index, DI is the logistics flow difference index, is the fusion degree value of the cross-border transmission heterogeneous data between different regions.
[0082] Compare the obtained fusion degree value of the cross-border transmission heterogeneous data between different regions with the preset threshold. If the fusion degree value of the cross-border transmission heterogeneous data is greater than or equal to the preset threshold, it indicates that the fusion degree of the cross-border transmission heterogeneous data is high, and at this time, a data fusion normal signal is generated; if the fusion degree value of the cross-border transmission heterogeneous data is less than the preset threshold, it indicates that the fusion degree of the cross-border transmission heterogeneous data is low, and at this time, a data fusion abnormal signal is generated.
[0083] Based on the fusion evaluation results of multi-source data, real-time predict the high-risk regions and the transmission speed of the cross-border transmission of animal diseases. According to the prediction results, judge the potential cross-border transmission risks and propose corresponding prevention and control intervention measures, specifically:
[0084] Define the speed at which an animal disease spreads from one region to another as the transmission speed S,
[0085] Mark the high-risk region as Ahigh, indicating the region with high transmission risk, which usually needs to be prioritized for prevention and control, and define the infection probability of the high-risk region as R.
[0086] The integration degree value LR of cross-border transmission heterogeneous data between different regions is obtained through the integration of multi-source data (including epidemiological data, social media information, meteorological data, logistics flow, etc.). Mathematical formulas based on the infectious disease transmission model will be used to predict the transmission speed and risk.
[0087] The transmission speed S is mainly determined by two factors: the transmission ability of the source of infection and the susceptibility of the receiving area. The calculation formula for the transmission speed is: ; where: represents the infectious ability of the source area of transmission (such as the epidemic outbreak point), which usually depends on the infection rate, disease characteristics, and animal density in the source area. represents the susceptibility of the receiving area, which usually depends on the sanitary conditions, immune level, and animal epidemic prevention measures in this area.
[0088] The calculation expression for the cross-border transmission risk R is: ; where: represents the transmission probability of the source area, which usually depends on the severity of the epidemic in the source area (such as the number of infected people, transmission rate, etc.). represents the input probability of the receiving area, which usually depends on the probability of disease exposure in the receiving area and reflects the mobility between this area and the source area (such as transportation flow, personnel flow, etc.).
[0089] The judgment of high-risk areas is based on the predicted value of the risk R. If the R value of a certain area exceeds the preset threshold Trisk, the area is regarded as a high-risk area:
[0090] Once the high-risk area Ahigh and the transmission speed S are obtained, further judge the potential cross-border transmission risk and propose corresponding prevention and control intervention measures.
[0091] The potential cross-border transmission risk is evaluated through a formula, and the formula is: ; where: represents the potential cross-border transmission risk, which is the accumulation of the transmission risks of all high-risk areas of.
[0092] According to the potential risk , different prevention and control intervention measures can be proposed, such as:
[0093] Blockade and isolation of high-risk areas: For high-risk areas with a relatively high transmission speed, prevention and control measures such as blockade and isolation can be taken to restrict the movement of animals and reduce the spread of the virus.
[0094] Strengthen quarantine and monitoring: Strengthen quarantine measures and monitoring in high-risk areas and potential transmission paths to ensure the timely discovery and handling of potential epidemics.
[0095] Vaccination and Immunization: Conduct large-scale animal vaccination in high-risk areas to improve the immunization level and reduce the risk of animal infection.
[0096] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data and performing software simulation to get a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0097] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0098] It should be understood that the term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0099] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A method for analyzing the risk of cross-border transmission of animal diseases based on big data, characterized in that: Including: Real-time obtain heterogeneous data including epidemiological surveillance data, social media information, meteorological data, and shipping logistics information, and adopt a multi-modal data fusion algorithm based on deep learning to integrate heterogeneous data from different sources to form a unified data set; Construct a cross-border transmission path model based on the infectious disease transmission model and epidemiological theory, and combine the spatio-temporal information and transmission characteristics in big data to simulate the potential transmission paths of animal diseases; During the simulation process, use machine learning algorithms to perform fusion processing and analysis on the integrated heterogeneous data to evaluate the fusion degree of cross-border transmission heterogeneous data between different regions; Based on the fusion evaluation results of multi-source data, real-time predict the high-risk regions and transmission speed of the cross-border transmission of animal diseases. According to the prediction results, judge the potential cross-border transmission risks and propose corresponding prevention and control intervention measures.
2. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 1, wherein: The multi-modal data fusion algorithm uses a deep learning model, including a convolutional neural network, a recurrent neural network, a multi-layer perceptron, or a Transformer model, to fuse multi-source heterogeneous data through feature extraction and attention mechanism.
3. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 2, wherein: The constructed cross-border transmission path model is based on an improved SIR infectious disease transmission model and combines a GIS geographic information system to dynamically simulate the disease transmission paths under different time and space conditions.
4. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 3, wherein: After analyzing the migration paths and migration times of different animal species, generate a species migration anomaly index. The method for obtaining the species migration anomaly index is: There is a directed graph \(G=(V, E)\), where: \(V\) represents the set of all habitats, which are all habitat nodes, and \(E\) represents the set of migration paths between habitats, which is the migration path from habitat to habitat . The weight of the edge is denoted as , that is, the migration intensity from to . Set the PageRank values of all nodes to be equal initially, i.e., , where n is the number of habitats; the iterative formula for the PageRank value is: ; where: is the PageRank value of habitat ; α is the damping factor, representing the probability that a viewer jumps to a random node, represents the set of all habitats pointing to habitat ; represents the set of all migration paths starting from habitat ; is the migration intensity from habitat to ; Initialize the PageRank values of each habitat, update the PageRank values of each habitat according to the iterative formula of the PageRank value, and repeat the update until the PageRank value converges. Once the PageRank values of each habitat are calculated, calculate the mean μ and standard deviation σ of the PageRank values of all habitats, and calculate the species migration anomaly index. The expression is: ; In the formula, AE is the species migration anomaly index.
5. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 4, wherein: After analyzing the differences in logistics flows in different regions, generate a logistics flow difference index. The method for obtaining the logistics flow difference index is: Obtain the inbound and outbound volume of goods Lpq, total logistics flow and total transportation volume Q within a certain area in different time ranges; calculate the mean value of the logistics flow of all areas, that is, the average level of the logistics flow of all areas, and the expression is: where: is the mean value of the logistics flow of all areas, m is the total number of areas, and the standard deviation of the logistics flow measures the fluctuation of the logistics flow, and the expression is: where: is the standard deviation of the logistics flow, is the total logistics flow of area p, is the mean value of the logistics flow of all areas; it is obtained by normalizing the difference between the logistics flow of each area and the overall flow mean value, and the expression is: where: DI is the logistics flow difference index.
6. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 5, wherein: Convert the species migration anomaly index and the logistics flow difference index into a comprehensive feature vector, and use the comprehensive feature vector as the input of a machine learning model. The machine learning model takes predicting the fusion degree value label of cross-border transmission heterogeneous data between different regions for each group of comprehensive feature vectors as the prediction target, and takes minimizing the sum of prediction errors for the fusion degree value labels of cross-border transmission heterogeneous data between all different regions as the training target. Train the machine learning model until the sum of prediction errors reaches convergence and then stop the model training. Determine the fusion degree value of cross-border transmission heterogeneous data between different regions according to the model output results. Among them, the machine learning model is a polynomial regression model.
7. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 6, wherein: Compare the obtained fusion degree value of cross-border transmission heterogeneous data between different regions with a preset threshold. If the fusion degree value of cross-border transmission heterogeneous data is greater than or equal to the preset threshold, it indicates that the fusion degree of cross-border transmission heterogeneous data is high. At this time, generate a data fusion normal signal; If the fusion degree value of cross-border transmission heterogeneous data is less than the preset threshold, it indicates that the fusion degree of cross-border transmission heterogeneous data is low. At this time, generate a data fusion abnormal signal.
8. The method for analyzing the risk of cross-border transmission of animal diseases based on big data according to claim 7, wherein: Define the speed at which an animal disease spreads from one region to another as the transmission speed S, mark the high-risk region as Ahigh, which represents the region with high transmission risk, and define the infection probability of the high-risk region as R; The fusion degree value LR of cross-border transmission heterogeneous data between different regions is obtained through the fusion of multi-source data, and a mathematical formula based on the infectious disease transmission model will be used to predict the transmission speed and risk; The calculation formula for the propagation speed is as follows: ; where: represents the infectivity of the source area of propagation, represents the susceptibility of the receiving area; The calculation expression for the cross-border transmission risk R is as follows: ; where: represents the transmission probability of the source region, represents the input probability of the receiving region; The determination of high-risk areas is based on the predicted value of risk R. If the R value of a certain area exceeds the preset threshold Trisk, the area is regarded as a high-risk area: Once the high-risk area Ahigh and the transmission speed S are obtained, the potential cross-border transmission risk is further judged. The potential cross-border transmission risk is evaluated by the formula: ; where: represents the potential cross-border transmission risk, which is the accumulation of the transmission risks of all high-risk areas According to the potential risk , different prevention and control intervention measures are proposed.
Citation Information
Patent Citations
Data quality control method and system for traditional Chinese and western medicine medical big data
CN110827935A
Multi-source heterogeneous medical test and examination data processing method and device, equipment and medium
CN113488182A