Big data-based animal epidemic detection data intelligent collection and supervision system and method
Patent Information
- Application Number
- CN202611081939.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]为解决上述现有系统普遍采用同质化的安全防护策略,防护标准难以统一,存在生物安全数据泄露的公共风险,也易造成养殖主体商业信息泄露;缺乏覆盖数据全生命周期的分级防护机制,在数据传输、存储、共享各环节采用无差别的加密与管控方式,普遍存在高敏感数据防护强度不足、低敏感数据防护冗余的问题;跨主体数据共享与协同分析往往需要直接调取原始数据,无法在不暴露原始敏感信息的前提下完成联合运算,制约了多源大数据的融合分析效能,打击了养殖主体的数据共享意愿,难以形成高效协同的监管闭环的问题,实现以上通过三级分布式采集网络实现全域疫病数据的规范采集与统一管理,双维度分级标注破解了生物安全与商业信息双重敏感数据的防护标准技术难题;全生命周期分级防护机制兼顾数据安全与算力效率,联邦学习结合动态脱敏实现跨主体数据安全协同;时空卷积预警模型配套平滑校验机制保障预警精准平稳,构建全流程闭环监管体系,显著提升动物疫病防控的数字化、智能化水平的目的
[0046]本发明提供了基于大数据的动物疫病检测数据智能采集监管系统及方法。具备以下有益效果:
Smart Images

Figure CN122822401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of disease prevention and control management, specifically to an intelligent data collection and monitoring system and method for animal disease detection based on big data. Background Technology
[0002] With the rapid development of livestock farming towards large-scale and intensive operations, the digital and intelligent transformation of animal disease prevention and control has become an industry consensus. Traditional animal disease supervision relies heavily on manual on-site inspections and offline laboratory sampling, which not only suffers from significant time lags but also makes it difficult to achieve real-time dynamic monitoring across the entire area. To improve the efficiency of disease prevention and control, multi-source data from the farming, testing, and regulatory ends are typically integrated, and big data analytics is used to achieve disease risk monitoring and trend prediction, thereby improving the response efficiency of disease prevention and control to a certain extent. Currently, significant technical barriers remain when facing real-time monitoring scenarios across the entire domain. Firstly, animal disease detection data possesses dual sensitive attributes of biosafety and commercial information. Existing systems generally adopt homogeneous security protection strategies, making it difficult to unify protection standards. This poses a public risk of biosafety data leakage and also easily leads to the leakage of commercial information of farming entities. Secondly, there is a lack of a hierarchical protection mechanism covering the entire data lifecycle. Indiscriminate encryption and control methods are used in all stages of data transmission, storage, and sharing, resulting in insufficient protection strength for highly sensitive data and redundant protection for low-sensitivity data. Thirdly, cross-entity data sharing and collaborative analysis often require direct retrieval of raw data, making it impossible to complete joint calculations without exposing original sensitive information. This restricts the efficiency of multi-source big data fusion analysis, discourages farming entities from sharing data, and makes it difficult to form an efficient and collaborative regulatory closed loop.
[0003] Chinese invention patent application CN121481243A, published on February 6, 2026, discloses an animal disease prevention and control monitoring and management system. This system comprises a multi-dimensional data acquisition module, a data management module, and a risk warning module. The multi-dimensional data acquisition module collects animal living data and animal living environment data in real time and uploads this data to the data management module. The data management module preprocesses the collected data, stores and manages the preprocessed data, and sends it to the risk warning module. Simultaneously, it dynamically adjusts the data acquisition frequency of the multi-dimensional data acquisition module based on fluctuations in the animal living data and animal living environment data. The risk warning module predicts animal disease risks based on the data and issues warnings based on the prediction results. However, this technical solution cannot provide differentiated and confidential supervision of animal disease detection data, thus reducing the security of animal disease prevention and control monitoring. Summary of the Invention
[0004] (a) Technical problems to be solved
[0005] To address the issues raised by existing systems, which generally employ homogeneous security protection strategies, lack standardized protection standards, pose a public risk of biosafety data leakage, and are prone to leaking commercial information of livestock farmers; lack a tiered protection mechanism covering the entire data lifecycle, and employ indiscriminate encryption and control methods at each stage of data transmission, storage, and sharing, resulting in insufficient protection for highly sensitive data and redundant protection for low-sensitive data; and the fact that cross-entity data sharing and collaborative analysis often require direct retrieval of raw data, making it impossible to complete joint calculations without exposing original sensitive information, thus hindering the efficiency of multi-source big data fusion analysis, discouraging livestock farmers' willingness to share data, and making it difficult to form an efficient and collaborative regulatory loop, the following solutions are proposed: A three-level distributed acquisition network enables standardized collection and unified management of disease data across the entire domain; dual-dimensional tiered labeling solves the technical challenge of protection standards for both biosafety and commercially sensitive data; a full lifecycle tiered protection mechanism balances data security and computing efficiency; federated learning combined with dynamic desensitization achieves secure cross-entity data collaboration; and a spatiotemporal convolutional early warning model with a smoothing verification mechanism ensures accurate and stable early warnings. This comprehensive closed-loop regulatory system significantly improves the digitalization and intelligence of animal disease prevention and control.
[0006] (II) Technical Solution
[0007] This invention is achieved through the following technical solution: a method for intelligent collection and supervision of animal disease detection data based on big data, the method comprising the following steps:
[0008] S1. Construct a multi-level distributed data acquisition terminal network, connect to multi-source data nodes of breeding entities, disease detection institutions, and animal health supervision departments, automatically collect multi-dimensional raw data of disease detection, and perform unified timestamp alignment to construct a multi-source disease detection dataset;
[0009] S2. Based on the multi-source disease detection dataset, perform dual-dimensional attribute hierarchical labeling on the original disease detection data, identify and label the biosafety risk level and commercial information sensitivity level of the data respectively, and generate corresponding data hierarchical labels.
[0010] S3. Based on the data classification label, establish a data lifecycle classification protection mechanism, and configure differentiated encryption rules and access control strategies for data transmission, data storage, and data sharing.
[0011] S4. Based on the federated learning framework and dynamic data anonymization, a privacy and security interaction middleware is built to realize collaborative regulatory computation of cross-subject data that is available but not visible;
[0012] S5. Construct a big data epidemic prevention and control analysis model, input the compliant dataset that has passed the graded protection verification into the model, identify and output regional epidemic risk warning and regulatory decision results, and form a closed-loop system of data collection-protection-analysis-regulation.
[0013] Preferably, the multi-level distributed data acquisition terminal network constructed in S1 includes a three-level architecture of edge acquisition nodes, regional aggregation nodes, and central monitoring nodes;
[0014] The edge acquisition nodes are deployed in large-scale farms, township veterinary stations and grassroots testing points. They have built-in standardized data interfaces and IoT acquisition adapters for real-time acquisition and preliminary format verification of on-site disease detection data.
[0015] The regional aggregation nodes are deployed at the municipal-level disease prevention and control institutions and are used for the aggregation, cleaning and standardization preprocessing of multi-source heterogeneous disease detection data within their jurisdictions.
[0016] The central monitoring node is deployed in provincial and above-level animal health regulatory departments for the overall management, scheduling and comprehensive analysis of disease detection data across the region.
[0017] Preferably, step S2 includes the following steps:
[0018] S21. Based on the animal disease classification and management standards, and combined with the disease type and data content attributes, classify the biosafety risk level to obtain the risk level label corresponding to each disease detection data.
[0019] S22. Based on the data, the depth and impact of the business information of the breeding entity can be inferred, the sensitivity level of commercial information can be divided, and the sensitivity level label corresponding to each piece of disease detection data can be obtained.
[0020] S23. Combine and map the biosafety risk level identifier with the commercial information sensitivity level identifier to generate a unique data classification label and irreversibly bind it to the corresponding disease detection data unit. The label flows with the entire life cycle of the data.
[0021] Preferably, step S3 includes the following steps:
[0022] S31. For the data transmission link, according to the graded labels of the epidemic detection data, the corresponding strength of the encryption algorithm and transmission channel are matched to achieve graded transmission protection;
[0023] S32. For the data storage stage, the storage media and access permissions corresponding to the isolation level are matched according to the hierarchical labels of the epidemic detection data to achieve hierarchical storage protection.
[0024] S33. Regarding the data sharing process, set sharing permission thresholds and circulation approval procedures based on the hierarchical labels of disease detection data to achieve hierarchical sharing and protection.
[0025] Preferably, step S4 includes the following steps:
[0026] S41. Build a horizontal federated learning framework, where each disease data subject acts as a local node to complete data feature training locally, and only uploads the encrypted model parameter gradient to the aggregation node without transmitting the original disease detection data, so as to realize cross-subject joint modeling operation.
[0027] S42. Configure a dynamic desensitization engine to automatically match desensitization rules based on the graded labels of disease detection data. Perform identifier replacement, numerical generalization and masking processing on low- and medium-level disease detection data that are transferred across entities. After desensitization, the data cannot be reversed to identify a specific breeding entity.
[0028] Preferably, step S5 includes the following steps:
[0029] S51. Construct an epidemic risk early warning model based on spatiotemporal convolutional neural networks, and define the input feature dimension and output result dimension;
[0030] S52. Construct a training sample set, use historical epidemic data and related environmental data to train the model, and iteratively optimize the network parameters through backpropagation until the loss function converges.
[0031] S53. Input the real-time compliance dataset that has undergone graded protection verification into the trained model, and output the regional epidemic risk level, potential transmission routes and regulatory handling recommendations.
[0032] Preferably, the specific process of S52 includes:
[0033] S521. Collect historical records of animal disease outbreaks, laboratory test data of diseases during the corresponding time periods, spatial distribution data of breeding entities, and regional environmental meteorological data to construct a model input feature set;
[0034] S522. Based on the official animal disease emergency response level and the treatment level of the epidemic site and epidemic area, the risk level of the samples is labeled as a real label for training the disease risk early warning model.
[0035] S523. Normalize the sample data and augment the time series data, divide it into training set, validation set and test set according to the preset ratio, and iteratively optimize the model parameters through the joint loss function until they are below the convergence threshold.
[0036] Preferably, before outputting the risk warning result in step S53, a warning smoothing verification step is also included, the specific process of which is as follows:
[0037] S531. Read the animal disease risk level results output from the previous monitoring cycle;
[0038] S532. Calculate the difference in disease risk level between the current model output and the risk level of the previous period.
[0039] S533. If the difference in the level jump does not exceed the preset maximum level jump threshold for a single period, the disease risk level result calculated by the model is directly output; if the difference in the level jump exceeds the maximum level jump threshold for a single period, the risk level output this time is forcibly limited to the maximum jump threshold range, and a second verification of the multi-source disease detection data is triggered.
[0040] A big data-based intelligent data collection and monitoring system for animal disease detection is provided to implement the aforementioned big data-based intelligent data collection and monitoring method for animal disease detection; the system includes:
[0041] The multi-level data acquisition and networking module is responsible for connecting to multi-source disease data nodes, collecting raw disease detection data in real time and building a standardized dataset.
[0042] The hierarchical labeling and protection module is used to perform hierarchical labeling of two-dimensional disease data, as well as the configuration and execution of hierarchical protection strategies throughout the entire life cycle.
[0043] The privacy and security interaction module has a built-in federated learning engine and dynamic de-identification engine to build a cross-entity privacy and security interaction middleware.
[0044] The intelligent regulatory analysis module is used to run big data epidemic regulatory analysis models and output epidemic risk warnings and regulatory decision-making results.
[0045] (III) Beneficial Effects
[0046] This invention provides an intelligent data collection and monitoring system and method for animal disease detection based on big data. It has the following beneficial effects:
[0047] I. Significantly improved capabilities in the collection and management of disease data across the entire region. Relying on a three-tiered distributed collection network of "edge collection nodes - regional aggregation nodes - central monitoring nodes," the network vertically connects the four-tiered epidemic prevention system at the provincial, municipal, county, and township levels, and horizontally covers three types of data nodes: breeding entities, testing institutions, and regulatory departments. Coupled with a unified timestamp alignment and standardized preprocessing mechanism, it has achieved standardized collection, efficient aggregation, and unified management of multi-source heterogeneous disease detection data. This has solved the pain points of traditional disease data collection, such as scattered, heterogeneous formats, and poor timeliness, providing a high-quality and comprehensive data foundation for subsequent security protection and intelligent analysis, and greatly improving the overall coverage and data quality of disease detection data collection and management.
[0048] Second, the dual-sensitivity data protection standards are clear and unified. A technical solution is adopted to independently rate and combine mapping and labeling the two dimensions of biosafety risk and commercial information sensitivity. This accurately defines the dual sensitivity attributes of each piece of disease detection data and generates a unique graded label that flows with the entire data life cycle. This solves the core problem of inconsistent standards between biosafety control and commercial information protection. It not only ensures the high level of biosafety control for major disease data, but also fully protects the commercial privacy of aquaculture entities, achieving a balance between data security and circulation efficiency.
[0049] Third, the accuracy and resource efficiency of protection throughout the entire lifecycle are both improved. A graded protection mechanism covering the entire process of transmission, storage and sharing is established based on data classification and labeling. Different encryption rules, physical isolation methods and access control strategies are matched for different levels of data. High-sensitivity data is configured with high-strength protection and low-sensitivity data is equipped with lightweight protection. This fundamentally eliminates the problem of insufficient protection for high-risk data and redundant protection for low-risk data. While ensuring the security of the entire chain of epidemic detection data, it effectively reduces the system's computing power consumption and significantly improves the utilization efficiency of protection resources.
[0050] Fourth, cross-entity collaborative supervision balances data security and business efficiency. Based on the federated learning framework and dynamic desensitization technology, a privacy-secure interactive middleware is built. Through the horizontal federated learning model of "local training and parameter aggregation", cross-entity joint modeling of original disease detection data can be achieved without leaving the domain. With the help of the dynamic desensitization engine, the safe transfer of low- and medium-level data is ensured. Under the premise of completely eliminating the risk of data leakage and privacy, the barriers to cross-entity data sharing are broken down, which effectively enhances the willingness of breeding entities to share data, consolidates the data foundation for full-domain big data supervision, and significantly improves the comprehensive efficiency of multi-department collaborative prevention and control.
[0051] V. Intelligent disease supervision is precise, stable, and controllable. A disease risk early warning model based on spatiotemporal convolutional neural networks is constructed. A pre-level protection and verification mechanism ensures the compliance and security of input data. It can accurately output the regional disease risk level, potential transmission path, and regulatory and disposal recommendations. The supporting early warning smoothing verification mechanism effectively avoids industry fluctuations and waste of regulatory resources caused by drastic changes in early warning levels. It realizes early detection, early warning, and early disposal of disease prevention and control, and comprehensively improves the accuracy, stability, and timeliness of animal disease supervision. Attached Figure Description
[0052] Figure 1 A flowchart of the intelligent collection and supervision method for animal disease detection data based on big data provided by the present invention;
[0053] Figure 2 A schematic diagram of the modules of the intelligent collection and supervision system for animal disease detection data based on big data provided by the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] An example of the intelligent data collection and monitoring system and method for animal disease detection based on big data is as follows:
[0056] Example 1:
[0057] Please see Figures 1-2 A method for intelligent collection and supervision of animal disease detection data based on big data, comprising the following steps:
[0058] S1. Construct a multi-level distributed data acquisition terminal network, connect to multi-source data nodes of breeding entities, disease detection institutions, and animal health supervision departments, automatically collect multi-dimensional raw data of disease detection, and perform unified timestamp alignment to construct a multi-source disease detection dataset;
[0059] The multi-level distributed data acquisition terminal network built in S1 includes a three-level architecture of edge acquisition nodes, regional aggregation nodes, and central monitoring nodes.
[0060] Edge acquisition nodes are deployed in large-scale farms, township veterinary stations and grassroots testing points. They are equipped with standardized data interfaces and IoT acquisition adapters for real-time acquisition and preliminary format verification of on-site disease detection data.
[0061] Regional aggregation nodes are deployed at the municipal-level disease prevention and control institutions to aggregate, clean, and standardize the preprocessing of multi-source heterogeneous disease detection data within their jurisdictions.
[0062] The central monitoring nodes are deployed in provincial and above-level animal health regulatory departments for the overall management, scheduling and comprehensive analysis of disease detection data across the region.
[0063] The multi-level distributed data acquisition terminal network built in S1 includes a three-level architecture of edge acquisition nodes, regional aggregation nodes and central monitoring nodes. All nodes are networked through the national government extranet and the animal disease prevention and control network, and rely on the Network Time Protocol (NTP) to achieve millisecond-level clock synchronization across the entire domain.
[0064] Specifically, in Example S1, edge acquisition nodes are deployed in three scenarios within the county: large-scale farms, township veterinary stations, and grassroots animal disease testing points. For large-scale farms, the nodes connect to the farm's livestock production management system and IoT monitoring equipment within the pens via the MQTT protocol. Data collected includes livestock inventory, age structure, immunization time and vaccine type, number of clinical cases, and the number and time of harmless disposal of dead livestock. For grassroots testing points, the nodes connect to rapid animal disease testing equipment via RS232 / USB hardware interfaces. Data collected includes sample number, sampling subject, testing item, testing result, and testing time. The default acquisition frequency of the edge acquisition nodes is set to once per hour, but can be switched to a high-frequency acquisition mode of once every 15 minutes during an emergency outbreak. All data is locally validated for field validity, removed null values, and standardized in format before being transmitted upwards.
[0065] Regional aggregation nodes are deployed at the municipal-level animal disease prevention and control institutions. On the one hand, they receive basic data uploaded by all edge nodes within their jurisdiction via HTTPS interfaces. On the other hand, they connect to the Laboratory Information Management System (LIS) of municipal-level animal disease testing laboratories through the dedicated epidemic prevention network to collect accurate data from serological testing and pathogen nucleic acid testing. Fields include unique sample identifier, submitting unit, disease type, testing method, Ct value / antibody titer, test conclusion, and testing personnel. The regional aggregation nodes perform deduplication, missing value completion, field mapping, and format standardization on the aggregated multi-source heterogeneous data to form a standardized municipal-level dataset, which is then uploaded to the provincial-level central monitoring node at a fixed interval of once every 2 hours.
[0066] The central monitoring node is deployed in the provincial animal health regulatory department. On the one hand, it connects to the provincial government regulatory platform through the government data sharing and exchange interface to collect regulatory and enforcement data such as the issuance of animal quarantine certificates, emergency response to diseases, epidemic prevention and control supervision and inspection, and registration of introduction and transportation of animals. Fields include quarantine certificate number, owner information, transportation quantity, quarantine results, disposal measures, and execution time. On the other hand, it receives standardized datasets uploaded by all prefecture-level city and regional nodes to complete the aggregation of data across the entire region.
[0067] The system has an internal clock synchronization service that uses the system clock of the central monitoring node as a reference to bind millisecond-level timestamps and spatial geographic codes (corresponding to the latitude and longitude and administrative division codes of the breeding entity / testing point) to each piece of collected data. According to the rule of "timestamp alignment and spatial code matching", data from different sources and with different collection frequencies are associated and aligned to form a multi-source disease detection dataset that includes three major categories: "breeding production data, laboratory testing data, and regulatory enforcement data" and covers both time and space dimensions. This dataset serves as a unified data base for subsequent hierarchical labeling and analysis.
[0068] The ability to collect and manage disease data across the entire region has been significantly improved. Relying on a three-tiered distributed collection network of "edge collection nodes - regional aggregation nodes - central supervision nodes," it vertically connects the four-tiered epidemic prevention system at the provincial, municipal, county, and township levels, and horizontally covers three types of data nodes: breeding entities, testing institutions, and regulatory departments. With the help of a unified timestamp alignment and standardized preprocessing mechanism, it has achieved standardized collection, efficient aggregation, and unified management of multi-source heterogeneous disease detection data. It has solved the pain points of traditional disease data collection, such as scattered, heterogeneous formats, and poor timeliness, and provided a high-quality, comprehensive data foundation for subsequent security protection and intelligent analysis, greatly improving the overall coverage and data quality of disease detection data collection and management.
[0069] S2. Based on the multi-source disease detection dataset, perform dual-dimensional attribute hierarchical labeling on the original disease detection data, identify and label the biosafety risk level and commercial information sensitivity level of the data respectively, and generate corresponding data hierarchical labels.
[0070] S2 includes the following steps:
[0071] S21. Based on the animal disease classification and management standards, and combined with the disease type and data content attributes, classify the biosafety risk level to obtain the risk level label corresponding to each disease detection data.
[0072] Specifically, in Example S21, the system incorporates an "Animal Disease Classification and Risk Level Mapping Table," which strictly adheres to the classification standards for Class I, II, and III diseases in the "Animal Epidemic Prevention Law of the People's Republic of China" to formulate mapping rules. After data input, the system automatically identifies the corresponding disease classification through keyword matching and disease field identification: pathogen detection data, outbreak data, and data related to epidemic sites and areas for Class I animal diseases (such as African swine fever and highly pathogenic avian influenza) are marked as high-risk; detection and monitoring data for Class II and III animal diseases are marked as medium-risk; routine non-disease data such as routine health checks and immune antibody monitoring are marked as low-risk. Finally, the system outputs the biosafety risk level identifier for each data point.
[0073] S22. Based on the data, the depth and impact of the business information of the breeding entity can be inferred, the sensitivity level of commercial information can be divided, and the sensitivity level label corresponding to each piece of disease detection data can be obtained.
[0074] Specifically, in Example S22, the system conducts a quantitative assessment from two dimensions: “subject identifiability” and “operational impact”. Precise inventory data and individual positive test data that can directly estimate the breeding scale, core production capacity, and disease infection rate are marked as core sensitive. Internal operating information such as routine production records and daily monitoring data are marked as internal sensitive. Public information that has been filed with regulatory authorities and industry aggregated statistical data are marked as public. Finally, the system outputs the commercial information sensitivity level identifier for each piece of data.
[0075] S23. Combine and map the biosafety risk level identifier with the commercial information sensitivity level identifier to generate a unique data classification label and irreversibly bind it to the corresponding disease detection data unit. The label flows with the entire life cycle of the data.
[0076] Specifically, in embodiment S23, the system combines and maps the level identifiers of the two dimensions to generate a unique grade label in the format of "Biosafety Level - Commercial Sensitivity Level", such as "High - Core Sensitivity", "Medium - Internal Sensitivity" and "Low - Public Level". The grade label is irreversibly bound to the corresponding data unit ID through SHA-256 hash operation. The label accompanies the data throughout its entire lifecycle from collection, transmission, storage to sharing and destruction, and serves as the sole basis for determining all security protection strategies.
[0077] The dual-sensitivity data protection standard is clear and unified. It adopts a technical solution of independent rating and combined mapping labeling of biosafety risk and commercial information sensitivity in two dimensions to accurately define the dual sensitivity attributes of each piece of disease detection data and generate a unique graded label that flows with the entire data life cycle. It solves the core problem of inconsistent standards for biosafety control and commercial information protection. It not only ensures the high level of biosafety control for major disease data, but also fully protects the commercial privacy of breeding entities, and achieves a balance between data security and circulation efficiency.
[0078] S3. Establish a data lifecycle-wide hierarchical protection mechanism based on data hierarchical labels, and configure differentiated encryption rules and access control strategies for data transmission, data storage, and data sharing.
[0079] S3 includes the following steps:
[0080] S31. For the data transmission link, according to the graded labels of the epidemic detection data, the corresponding strength of the encryption algorithm and transmission channel are matched to achieve graded transmission protection;
[0081] Specifically, in the data transmission stage of embodiment S31:
[0082] For critical disease monitoring data marked as "High-Core Sensitive," the highest level of transmission protection is implemented, employing the national standard SM4 symmetric encryption algorithm for end-to-end full-link encryption. The transmission channel uses a dedicated IPsec VPN, and two-way digital certificate authentication is performed on both the sending and receiving ends to ensure a unique and trustworthy transmission link, preventing data interception and decryption during transit. For routine disease monitoring data marked as "Medium-Internal Sensitive," the standard TLS 1.3 transport layer encryption protocol is used, coupled with the OAuth 2.0 dynamic token interface authentication mechanism, balancing transmission efficiency and security strength. For immunization monitoring statistics data marked as "Low-Public," HTTPS is used for basic transmission, supplemented with MD5 hash verification and digital signature mechanisms, prioritizing the integrity and tamper-proof nature of data transmission while reducing transmission computing power consumption.
[0083] S32. For the data storage stage, the storage media and access permissions corresponding to the isolation level are matched according to the hierarchical labels of the epidemic detection data to achieve hierarchical storage protection.
[0084] Specifically, in the data storage stage of embodiment S32:
[0085] High-level sensitive epidemic disease data is stored on physically isolated dedicated encrypted storage nodes certified by the national cryptography management department. Before data is written to disk, all fields are transparently encrypted, and a full operation audit log is configured. All access operations require two levels of manual approval from the data administrator and the business manager, and access is only supported by limited IP addresses within the whitelist. Medium-level epidemic disease monitoring data is stored in a logically isolated dedicated partition of the database, using a role-based access control (RBAC) mechanism. Personnel in different positions can only access data fields within their authorized scope. A full data backup and integrity verification are performed every 7 days. Low-level public data is stored in a shared business database, configured with a three-replica disaster recovery backup mechanism, and supports batch queries and statistical calls within the authorized scope.
[0086] S33. Regarding the data sharing process, set sharing permission thresholds and circulation approval procedures based on the hierarchical labels of disease detection data to achieve hierarchical sharing and protection.
[0087] Specifically, in the data sharing phase of embodiment S33:
[0088] High-sensitivity major disease raw data are strictly prohibited from being shared across entities or departments. Aggregated regional statistical results can only be output within the authorized scope, and differential privacy noise-adding processing must be performed before output, adding Laplace noise to avoid inferring original individual data through multiple statistical analyses. Before sharing medium-sensitivity disease monitoring data across entities, the desensitization engine is automatically triggered to perform desensitization processing, and it must be approved by both the data owner and the regulatory department to confirm that the sharing scope and purpose are compliant before it can be transferred. Low-sensitivity public data is authorized to be called through standardized open interfaces. All sharing operations automatically retain tamper-proof operation logs, including information such as the calling entity, calling time, and the scope of data called, supporting full-chain traceability of data flow.
[0089] The system achieves a dual improvement in the accuracy and resource efficiency of protection throughout the entire lifecycle. Based on data classification and labeling, it establishes a graded protection mechanism covering the entire process of transmission, storage, and sharing. Differentiated encryption rules, physical isolation methods, and access control strategies are matched for different levels of data. High-sensitivity data is configured with high-strength protection, while low-sensitivity data adopts lightweight protection. This fundamentally eliminates the problem of insufficient protection for high-risk data and redundant protection for low-risk data. While ensuring the security of disease detection data throughout the entire chain, it effectively reduces the system's computing power consumption and significantly improves the utilization efficiency of protection resources.
[0090] S4. Based on the federated learning framework and dynamic data anonymization, a privacy and security interaction middleware is built to realize collaborative regulatory computation of cross-subject data that is available but not visible;
[0091] S4 includes the following steps:
[0092] S41. Build a horizontal federated learning framework, where each disease data subject acts as a local node to complete data feature training locally, and only uploads the encrypted model parameter gradient to the aggregation node without transmitting the original disease detection data, so as to realize cross-subject joint modeling operation.
[0093] Specifically, in embodiment S41, the privacy-secure interaction middleware is deployed at the provincial-level central monitoring node, and the horizontal federated learning framework covers all prefecture-level regional nodes and key aquaculture enterprise nodes. When it is necessary to train a global disease risk early warning model, each local node completes local model training using its own original disease detection data. After each round of local training, only the homomorphically encrypted model parameter gradient is uploaded to the provincial aggregation node. The aggregation node securely aggregates the parameters uploaded by multiple parties, updates the global model parameters, and then distributes them to each local node for the next round of training. After 15-20 iterations, a converged global model is obtained. Throughout the entire process, no original disease detection data is transmitted across subjects or regions; only the encrypted model parameters are transmitted, truly realizing collaborative computation where "data remains stationary while value moves." This not only protects the data sovereignty and privacy security of each subject but also achieves the analytical efficiency of global data joint modeling.
[0094] S42. Configure a dynamic desensitization engine to automatically match desensitization rules based on the graded labels of disease detection data. Perform identifier replacement, numerical generalization and masking processing on low- and medium-level disease detection data that are transferred across entities. After desensitization, the data cannot be reversed to identify a specific breeding entity.
[0095] Specifically, in embodiment S42, the dynamic desensitization engine automatically matches the corresponding desensitization rules based on the data classification labels for cross-entity flow scenarios of low- and medium-level disease detection data: for identifier fields that can identify specific entities, such as farm names, unified social credit codes, and detailed addresses, one-way hash replacement or masking is performed; for numerical sensitive fields such as stock quantity, infection rate, and slaughter quantity, interval generalization or Laplace noise addition is performed; the desensitized data retains statistical analysis value, but cannot be reversed to locate specific farming entities, thus balancing data usability and privacy security.
[0096] Cross-entity collaborative supervision balances data security and business efficiency. Based on a federated learning framework and dynamic desensitization technology, a privacy-secure interactive middleware is built. Through a horizontal federated learning model of "local training and parameter aggregation," cross-entity joint modeling of original disease detection data can be achieved without leaving the domain. In conjunction with a dynamic desensitization engine, the secure transfer of low- and medium-level data is ensured. While completely eliminating the risk of data leakage and privacy, the barriers to cross-entity data sharing are broken down, effectively enhancing the willingness of aquaculture entities to share data, consolidating the data foundation for comprehensive big data supervision, and significantly improving the overall efficiency of multi-departmental collaborative prevention and control.
[0097] S5. Construct a big data epidemic prevention and control analysis model, input the compliant dataset that has passed the graded protection verification into the model, identify and output regional epidemic risk warning and regulatory decision results, and form a closed-loop system of data collection-protection-analysis-regulation.
[0098] S5 includes the following steps:
[0099] S51. Construct an epidemic risk early warning model based on spatiotemporal convolutional neural networks, and define the input feature dimension and output result dimension;
[0100] S52. Construct a training sample set, use historical epidemic data and related environmental data to train the model, and iteratively optimize the network parameters through backpropagation until the loss function converges.
[0101] S53. Input the real-time compliance dataset that has undergone graded protection verification into the trained model, and output the regional epidemic risk level, potential transmission routes and regulatory handling recommendations.
[0102] Specifically, in Example S5, a disease risk early warning model based on spatiotemporal convolutional neural network (ST-CNN) is constructed. The model input dimensions include three major categories and a total of 18 features: time dimension (number of positive tests in the past 7 days, immunization coverage rate, disease incidence rate), spatial dimension (regional breeding density, distance from the epidemic site, and transportation activity), and environmental dimension (average daily temperature, relative humidity, and cumulative rainfall). The model output dimensions include the disease risk level of each administrative region, the spatial distribution of high-risk areas, the potential transmission path of the disease, the list of key monitored objects, and prevention and control recommendations.
[0103] The specific process of S52 includes:
[0104] S521. Collect historical records of animal disease outbreaks, laboratory test data of diseases during the corresponding time periods, spatial distribution data of breeding entities, and regional environmental meteorological data to construct a model input feature set;
[0105] S522. Based on the official animal disease emergency response level and the treatment level of the epidemic site and epidemic area, the risk level of the samples is labeled as a real label for training the disease risk early warning model.
[0106] S523. Normalize the sample data and augment the time series data, divide it into training set, validation set and test set according to the preset ratio, and iteratively optimize the model parameters through the joint loss function until they are below the convergence threshold.
[0107] Specifically, in Example S52, during the model training phase, a sample set is constructed by collecting historical disease outbreak records from the past 10 years, detection data for the corresponding time periods, aquaculture data, and environmental meteorological data. The officially released disease level and emergency response level are used as the true labels, and the samples are labeled with four levels of risk (low risk / moderate risk / relatively high risk / high risk). The samples are divided into training set, validation set, and test set in a ratio of 8:1:1. The error between the prediction result and the true label is calculated using the cross-entropy loss function. The network weights and bias parameters are iteratively updated using the Adam optimizer and backpropagation algorithm until the loss function converges to below the preset threshold of 0.02, thus completing the model training and parameter solidification.
[0108] S53 includes a warning smoothing verification step before outputting the risk warning result. The specific process is as follows:
[0109] S531. Read the animal disease risk level results output from the previous monitoring cycle;
[0110] S532. Calculate the difference in disease risk level between the current model output and the risk level of the previous period.
[0111] S533. If the difference in level jump does not exceed the preset maximum level jump threshold for a single period, the disease risk level result calculated by the model is directly output; if the difference in level jump exceeds the maximum level jump threshold for a single period, the risk level output this time is forcibly limited to the maximum jump threshold range, and a second verification of the multi-source disease detection data is triggered.
[0112] Specifically, in embodiment S53, the four risk levels are first assigned numerical values: low risk = 1, moderate risk = 2, relatively high risk = 3, and high risk = 4; the maximum level jump threshold in a single cycle is set to level 1. This threshold is set according to the work specifications for adjusting the level of animal disease prevention and control, so as to avoid frequent changes in prevention and control policies from affecting the order of breeding production.
[0113] If the difference between the model output and the level jump of the previous period is ≤1 (not exceeding the threshold), the corresponding risk level result will be output directly, and the corresponding regulatory handling suggestions will be pushed at the same time.
[0114] If the difference between the warning levels is greater than 1 (exceeding the threshold), the system will forcibly limit the warning level issued to the public to a level of 1. At the same time, it will automatically activate the multi-source data secondary verification mechanism, calling up multi-source information such as laboratory verification data, on-site investigation data, and monitoring data of surrounding areas for secondary verification, eliminating misjudgments caused by factors such as local sample anomalies and false positives, and ensuring that the warning results are stable and reliable.
[0115] The intelligent disease monitoring system is precise, stable, and controllable. It constructs a disease risk early warning model based on spatiotemporal convolutional neural networks, and a pre-level protection and verification mechanism ensures the compliance and security of input data. It can accurately output the regional disease risk level, potential transmission routes, and regulatory and disposal suggestions. The supporting early warning smoothing verification mechanism effectively avoids industry fluctuations and waste of regulatory resources caused by drastic changes in early warning levels. It realizes early detection, early warning, and early disposal of disease prevention and control, and comprehensively improves the accuracy, stability, and timeliness of animal disease monitoring.
[0116] For example, consider a specific scenario of monitoring and supervising African swine fever outbreaks at a pig farm in a certain province:
[0117] The provincial regulatory authorities have set the goal of maintaining a low-risk level for routine prevention and control across the entire region, and the key monitoring areas are to maintain a general risk warning threshold. The preset higher risk triggering benchmark is ≥5 positive samples per week.
[0118] Once the platform is launched, nodes at all levels collect data at a predetermined frequency; in the detection data uploaded by a node in a certain city during a certain week, three cases of positive African swine fever test results from small-scale farms were found.
[0119] Following the S2 grading labeling, this batch of positive test data was marked as "high-core sensitive".
[0120] Implementing S3-level protection, this batch of data is transmitted through a dedicated IPsec encrypted channel and stored on a physically isolated dedicated storage node. Access is only authorized to municipal-level or higher epidemic prevention personnel, and cross-departmental sharing of raw data is prohibited.
[0121] When executing S4, the provincial central node updates the overall disease risk early warning model by combining data from various cities through a horizontal federated learning framework, without retrieving the original test data from each farm.
[0122] After executing S5 and inputting the compliant dataset into the model, it was initially determined that the city's risk level had risen from low risk to higher risk, with a level jump difference of 2, exceeding the single-cycle maximum jump threshold of level 1; the early warning smoothing verification mechanism was triggered, and the system first adjusted the publicly released early warning level to general risk, while simultaneously initiating a secondary data review.
[0123] After laboratory verification and on-site investigation, the three positive samples were confirmed to be false positives. The model was corrected to a low risk level, which ultimately did not trigger an upgrade of the prevention and control level and avoided unnecessary industry panic and investment of prevention and control resources. If the verification confirms that the risk is real, the risk level will be officially upgraded in the next monitoring cycle, and precise control measures will be pushed out at the same time to achieve a smooth and orderly closed-loop supervision.
[0124] Example 2:
[0125] Please see Figures 1-2 A big data-based intelligent data collection and monitoring system for animal disease detection, used to implement a big data-based intelligent data collection and monitoring method for animal disease detection; the system includes:
[0126] The multi-level data acquisition and networking module is responsible for connecting to multi-source disease data nodes, collecting raw disease detection data in real time and building a standardized dataset.
[0127] The hierarchical labeling and protection module is used to perform hierarchical labeling of two-dimensional disease data, as well as the configuration and execution of hierarchical protection strategies throughout the entire life cycle.
[0128] The privacy and security interaction module has a built-in federated learning engine and dynamic de-identification engine to build a cross-entity privacy and security interaction middleware.
[0129] The intelligent regulatory analysis module is used to run big data epidemic regulatory analysis models and output epidemic risk warnings and regulatory decision-making results.
[0130] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent collection and monitoring of animal disease detection data based on big data, characterized in that, The method includes the following steps: S1. Construct a multi-level distributed data acquisition terminal network, connect to multi-source data nodes of breeding entities, disease detection institutions, and animal health supervision departments, automatically collect multi-dimensional raw data of disease detection, and perform unified timestamp alignment to construct a multi-source disease detection dataset; S2. Based on the multi-source disease detection dataset, perform dual-dimensional attribute hierarchical labeling on the original disease detection data, identify and label the biosafety risk level and commercial information sensitivity level of the data respectively, and generate corresponding data hierarchical labels. S3. Based on the data classification label, establish a data lifecycle classification protection mechanism, and configure differentiated encryption rules and access control strategies for data transmission, data storage, and data sharing. S4. Based on the federated learning framework and dynamic data anonymization, a privacy and security interaction middleware is built to realize collaborative regulatory computation of cross-subject data that is available but not visible; S5. Construct a big data epidemic prevention and control analysis model, input the compliant dataset that has passed the graded protection verification into the model, identify and output regional epidemic risk warning and regulatory decision results, and form a closed-loop system of data collection-protection-analysis-regulation.
2. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 1, characterized in that, The multi-level distributed data acquisition terminal network constructed in S1 includes a three-level architecture of edge acquisition nodes, regional aggregation nodes, and central monitoring nodes. The edge acquisition nodes are deployed in large-scale farms, township veterinary stations and grassroots testing points. They have built-in standardized data interfaces and IoT acquisition adapters for real-time acquisition and preliminary format verification of on-site disease detection data. The regional aggregation nodes are deployed at the municipal-level disease prevention and control institutions and are used for the aggregation, cleaning and standardization preprocessing of multi-source heterogeneous disease detection data within their jurisdictions. The central monitoring node is deployed in provincial and above-level animal health regulatory departments for the overall management, scheduling and comprehensive analysis of disease detection data across the region.
3. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 1, characterized in that, S2 includes the following steps: S21. Based on the animal disease classification and management standards, and combined with the disease type and data content attributes, classify the biosafety risk level to obtain the risk level label corresponding to each disease detection data. S22. Based on the data, the depth and impact of the business information of the breeding entity can be inferred, the sensitivity level of commercial information can be divided, and the sensitivity level label corresponding to each piece of disease detection data can be obtained. S23. Combine and map the biosafety risk level identifier with the commercial information sensitivity level identifier to generate a unique data classification label and irreversibly bind it to the corresponding disease detection data unit. The label flows with the entire life cycle of the data.
4. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 1, characterized in that, S3 includes the following steps: S31. For the data transmission link, according to the graded labels of the epidemic detection data, the corresponding strength of the encryption algorithm and transmission channel are matched to achieve graded transmission protection; S32. For the data storage stage, the storage media and access permissions corresponding to the isolation level are matched according to the hierarchical labels of the epidemic detection data to achieve hierarchical storage protection. S33. Regarding the data sharing process, set sharing permission thresholds and circulation approval procedures based on the hierarchical labels of disease detection data to achieve hierarchical sharing and protection.
5. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 1, characterized in that, S4 includes the following steps: S41. Build a horizontal federated learning framework, where each disease data subject acts as a local node to complete data feature training locally, and only uploads the encrypted model parameter gradient to the aggregation node without transmitting the original disease detection data, so as to realize cross-subject joint modeling operation. S42. Configure a dynamic desensitization engine to automatically match desensitization rules based on the graded labels of disease detection data. Perform identifier replacement, numerical generalization and masking processing on low- and medium-level disease detection data that are transferred across entities. After desensitization, the data cannot be reversed to identify a specific breeding entity.
6. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 1, characterized in that, S5 includes the following steps: S51. Construct an epidemic risk early warning model based on spatiotemporal convolutional neural networks, and define the input feature dimension and output result dimension; S52. Construct a training sample set, use historical epidemic data and related environmental data to train the model, and iteratively optimize the network parameters through backpropagation until the loss function converges. S53. Input the real-time compliance dataset that has undergone graded protection verification into the trained model, and output the regional epidemic risk level, potential transmission routes and regulatory handling recommendations.
7. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 6, characterized in that, The specific process of S52 includes: S521. Collect historical records of animal disease outbreaks, laboratory test data of diseases during the corresponding time periods, spatial distribution data of breeding entities, and regional environmental meteorological data to construct a model input feature set; S522. Combining the official animal disease emergency response levels and the treatment levels of epidemic sites and areas, the samples are labeled with risk levels to serve as real labels for training the disease risk early warning model. S523. Normalize the sample data and augment the time series data, divide it into training set, validation set and test set according to the preset ratio, and iteratively optimize the model parameters through the joint loss function until they are below the convergence threshold.
8. The intelligent collection and monitoring method for animal disease detection data based on big data as described in claim 6, characterized in that, Before outputting the risk warning result in S53, a warning smoothing verification step is also included. The specific process is as follows: S531. Read the animal disease risk level results output from the previous monitoring cycle; S532. Calculate the difference in disease risk level between the current model output and the risk level of the previous period. S533. If the difference in the level jump does not exceed the preset maximum level jump threshold for a single cycle, the disease risk level result calculated by the model is directly output. If the difference in risk level exceeds the maximum risk level threshold for a single cycle, the risk level output for this time will be forcibly limited to the maximum risk level threshold range, and a second verification of the multi-source disease detection data will be triggered.
9. A big data-based intelligent data collection and monitoring system for animal disease detection, used to implement the big data-based intelligent data collection and monitoring method for animal disease detection as described in any one of claims 1-8; characterized in that, The system includes: The multi-level data acquisition and networking module is responsible for connecting to multi-source disease data nodes, collecting raw disease detection data in real time and building a standardized dataset. The hierarchical labeling and protection module is used to perform hierarchical labeling of two-dimensional disease data, as well as the configuration and execution of hierarchical protection strategies throughout the entire life cycle. The privacy and security interaction module has a built-in federated learning engine and dynamic de-identification engine to build a cross-entity privacy and security interaction middleware. The intelligent regulatory analysis module is used to run big data epidemic regulatory analysis models and output epidemic risk warnings and regulatory decision-making results.
Citation Information
Patent Citations
Animal epidemic disease prevention and control monitoring management system
CN121481243A