A mosquito-borne disease risk prediction federated learning method and system based on waterlogging distribution

CN122842980APending Publication Date: 2026-09-29广州市疾病预防控制中心(广州市卫生监督所)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611191371.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种基于积水分布的蚊媒疾病风险预测联邦学习方法及系统,以解决现有技术中环境感知与疾病预警数据割裂、时序错位,以及异构终端异步协同训练中因语义不对齐与特征分布漂移导致下游疾病模型性能退化的技术问题

Benefits of technology

[0033](1)通过动态计算蚊虫潜伏期偏移量,将积水环境数据与病例数据按真实生物学滞后关系对齐,避免了传统方法因固定时间窗口导致的信号错位,显著提升风险预测的因果准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122842980A_ABST
    Figure CN122842980A_ABST
Patent Text Reader

Abstract

The application discloses a mosquito-borne disease risk prediction federated learning method and system based on water accumulation distribution, solves the problems of time sequence dislocation of environment and case data, and asynchronous training feature drift. The method constructs a cloud aggregator, an environment perception end and a pathological data end cascaded federated learning architecture: the environment perception end extracts a water accumulation state representation vector and confidence and uploads; the cloud aggregator combines meteorological data to calculate a latent period offset and encapsulates and issues; the pathological data end retrieves historical environmental data with the offset, constructs an environmental risk prior distribution by confidence weighting, and fuses with pathological features to obtain a disease risk prediction value; the cloud aggregator independently and asynchronously aggregates water accumulation recognition model gradients and disease model gradients respectively, updates mutually independent global water accumulation recognition model parameters and global disease model parameters, and suppresses feature drift through confidence gating and backtracking window retrieval, giving consideration to privacy security and module decoupling iteration, and improving prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and public health, and more specifically, to a federated learning method and system for predicting the risk of mosquito-borne diseases based on water accumulation distribution. Background Technology

[0002] The spread of mosquito-borne infectious diseases (such as dengue fever and Zika virus disease) is closely related to the distribution of stagnant water in the environment, which is a key breeding ground for mosquitoes. Traditional methods for predicting the risk of mosquito-borne diseases often rely on manual Breteau index surveys or modeling based on single meteorological data, which suffers from problems such as monitoring lag, insufficient spatial coverage, and separation of environmental and pathological data. In recent years, with the development of IoT and AI technologies, automatic monitoring of stagnant water based on image recognition has become possible. However, existing solutions typically deploy environmental perception and disease early warning as independent systems, lacking a time-series alignment mechanism for the dynamic characteristics of mosquitoes' external incubation period. This results in a temporal misalignment between environmental risk signals and case occurrence, making it difficult to achieve accurate causal relationship modeling.

[0003] Furthermore, environmental sensing devices and medical institutions are managed by different entities, and raw images and case data cannot be centrally aggregated due to privacy and compliance requirements, making traditional centralized training models unsuitable. While existing federated learning methods can address the data silo problem, they generally employ synchronous aggregation strategies, requiring all participants to complete computations in the same round. However, the sampling frequencies, communication conditions, and computational loads of the environmental sensing end and the pathological data end differ significantly, making forced synchronization prone to resource waste and model update delays. More critically, if asynchronous independent training is simply adopted, the semantic misalignment between the visual feature space and the disease prediction space, coupled with feature distribution drift caused by high-frequency updates of the environmental model, can lead to instability in the input of downstream disease models, performance degradation, and even catastrophic amnesia. Therefore, there is an urgent need for a federated learning method that can integrate dynamic distribution of water accumulation and epidemiological data, support asynchronous collaborative training of heterogeneous terminals, and possess resistance to feature drift, in order to achieve real-time, accurate, and privacy-secure prediction of mosquito-borne disease risks. Summary of the Invention

[0004] The purpose of this invention is to provide a federated learning method and system for predicting mosquito-borne disease risk based on water accumulation distribution, in order to solve the technical problems in the prior art of data fragmentation and temporal misalignment between environmental perception and disease early warning, as well as the performance degradation of downstream disease models caused by semantic misalignment and feature distribution drift in asynchronous collaborative training of heterogeneous terminals.

[0005] To achieve the above objectives, the present invention provides the following two aspects:

[0006] In a first aspect, the present invention provides a federated learning method for predicting the risk of mosquito-borne diseases based on water accumulation distribution, characterized by comprising the following steps:

[0007] S1. Construct a federated learning system, which includes a cloud aggregator, multiple environmental sensing terminals, and multiple pathological data terminals; pre-configure target monitoring area mapping relationships so that each pathological data terminal is associated with at least one target monitoring area, and each target monitoring area contains at least one environmental sensing terminal.

[0008] S2. During the monitoring and operation phase, each environmental sensing terminal acquires current environmental image data at a preset acquisition frequency, extracts the water accumulation state representation vector and the corresponding water accumulation confidence level through the locally deployed water accumulation recognition model, and uploads the water accumulation state representation vector, the water accumulation confidence level and its acquisition timestamp to the cloud aggregator; the environmental sensing terminal also calculates the gradient of the water accumulation recognition model based on the local water accumulation annotation data, and uploads the water accumulation recognition model gradient to the cloud aggregator according to a preset environmental training cycle;

[0009] S3. The cloud aggregator acquires meteorological data of the target monitoring area, dynamically calculates the incubation period offset Δ of the current round based on the preset incubation period model, and encapsulates the water accumulation status representation vector, water accumulation confidence, collection timestamp and the incubation period offset Δ uploaded by each environmental sensing terminal belonging to the same target monitoring area into an environmental semantic data packet, and sends it to the pathological data terminal associated with the target monitoring area.

[0010] S4. Each pathological data terminal acquires local epidemiological data and the corresponding case collection timestamp t2, and extracts pathological features via a pathological feature encoder; using t2-Δ as the baseline timestamp, it retrieves from the local time-series buffer the water accumulation state representation vector sequence and the corresponding water accumulation confidence sequence within a preset backtracking window length centered on the baseline timestamp; based on the water accumulation confidence sequence, it constructs an environmental risk prior distribution, and inputs the pathological features and the environmental risk prior distribution into a multimodal fusion module to obtain a disease risk prediction value; based on the disease risk prediction value and disease label, it calculates the disease model gradient and uploads it to the cloud aggregator, whereby the disease model includes the pathological feature encoder and the multimodal fusion module;

[0011] S5. When disease model gradients uploaded by each pathological data terminal are received within a preset disease training period, the cloud aggregator performs independent aggregation on the disease model gradients to update the global disease model parameters, and distributes the updated global disease model parameters to each pathological data terminal according to the preset disease training period; when water accumulation recognition model gradients uploaded by each environmental perception terminal are received within a preset environment training period, the water accumulation recognition model gradients perform independent aggregation on the water accumulation recognition model gradients to update the global water accumulation recognition model parameters, and distributes the updated global water accumulation recognition model parameters to each environmental perception terminal according to the preset environment training period; wherein, the global water accumulation recognition model parameters are updated only based on the water accumulation recognition model gradients uploaded by the environmental perception terminal, and the global disease model parameters are updated only based on the disease model gradients uploaded by the pathological data terminal, and the aggregation and updates of the two are independent of each other and do not require synchronous waiting.

[0012] Preferably, in step S2, the water accumulation state representation vector is a fixed-dimensional semantic embedding vector obtained by the water accumulation recognition model through a fully connected bottleneck layer or a projection head, and the semantic embedding vector supports temporal aggregation operations; the water accumulation confidence is the area ratio or classification probability value of the water accumulation region output by the water accumulation recognition model; the local water accumulation annotation data is obtained in the following way: before the deployment of the environmental perception terminal, historical environmental images are manually annotated as the local basic dataset for the first round of training of federated learning; during the operation of the environmental perception terminal, image samples whose prediction confidence of the water accumulation recognition model is lower than a preset threshold are manually reviewed and annotated, and continuously supplemented into the local water accumulation annotation data.

[0013] Preferably, in step S3, before the cloud aggregator sends out the environmental semantic data packet, it further includes: performing regional weighted aggregation on the water accumulation status representation vectors uploaded by multiple environmental sensing terminals belonging to the same target monitoring area to obtain a regional water accumulation semantic vector; wherein, the weight of the weighted aggregation is determined by the water accumulation confidence score uploaded by each environmental sensing terminal; performing an arithmetic average on the water accumulation confidence scores uploaded by the multiple environmental sensing terminals to obtain a regional water accumulation confidence score; and encapsulating the regional water accumulation semantic vector and the regional water accumulation confidence score into the environmental semantic data packet instead of the data from a single environmental sensing terminal and sending it out.

[0014] Preferably, in step S4, the length of the preset backtracking window is determined by the difference between the maximum and minimum values ​​of the latency offset Δ within the target monitoring area during a preset historical period; the construction of the environmental risk prior distribution based on the water accumulation confidence sequence specifically includes at least one of the following methods:

[0015] Using the water accumulation confidence sequence as weights, weighted temporal pooling is performed on the water accumulation state characterization vector sequence to obtain the distributed characterization vector;

[0016] Based on the water accumulation confidence sequence and the collection timestamp, kernel density estimation is performed to generate an embedded environmental risk curve.

[0017] The sequence of water accumulation state representation vectors and the sequence of water accumulation confidence are input into a temporal attention network to output a context-aware distribution representation.

[0018] Preferably, in step S3, the dynamic calculation method for the latency offset Δ is as follows:

[0019] Acquire time-series data of daily average temperature and relative humidity for the target monitoring area within a preset backtracking time window;

[0020] The time series data is input into the preset incubation period model, which is a mapping function constructed based on the pathogen external incubation period accumulated temperature rule;

[0021] The estimated incubation period of the pathogen under the current meteorological conditions is output as the incubation period offset Δ.

[0022] Preferably, in step S1, the method for establishing the target monitoring area mapping relationship includes at least one of the following:

[0023] A pre-configured attribution table based on administrative division codes and medical institution practice addresses;

[0024] The results of dynamic clustering analysis based on historical case address distribution data and environmental sensing terminal geographic location data;

[0025] A single pathological data endpoint is associated with multiple target monitoring areas, and a single environmental sensing endpoint simultaneously belongs to multiple target monitoring areas.

[0026] Preferably, in step S4, the local time-series buffer adopts a circular queue structure, and its storage unit is a triple containing the acquisition timestamp, the water accumulation state representation vector, and the water accumulation confidence level; the capacity of the circular queue is determined by the sum of the maximum expected value of the latency offset Δ and the half-width of the preset backtracking window; when the pathological data terminal receives a new environmental semantic data packet, it writes it into the circular queue in the order of acquisition timestamps and automatically overwrites historical data that exceeds the capacity.

[0027] Preferably, in step S5, the independent aggregations all adopt a weighted average strategy, and the independent aggregations of the gradient of the water accumulation identification model and the independent aggregations of the gradient of the disease model respectively adopt independent weight calculation rules; the global disease model parameters include the parameters of the disease model, and the global water accumulation identification model parameters include all the parameters of the water accumulation identification model; in step S4, the water accumulation state representation vector and the water accumulation confidence are used as fixed environmental context conditions input during training, and the disease model gradient is only used to update the parameters of the disease model.

[0028] Preferably, the method further includes an exception handling step:

[0029] When the cloud aggregator does not receive a water accumulation status representation vector uploaded by a certain environmental sensing terminal within a preset timeout period, it sends a missing marker to the pathological data terminal associated with that environmental sensing terminal.

[0030] If the pathological data terminal matches the timestamp corresponding to the missing marker during the retrieval in step S4, the confidence level of the water accumulation at the corresponding position will be set to zero when constructing the environmental risk prior distribution, or a preset default vector and zero confidence level will be used to participate in the distribution construction.

[0031] Secondly, this invention provides a federated learning system for predicting mosquito-borne disease risk based on water accumulation distribution. This system is configured to execute a federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution as described in the first aspect, comprising: a cloud aggregator for acquiring meteorological data and dynamically calculating the incubation period offset Δ, encapsulating the incubation period offset Δ and environmental data together into an environmental semantic data package and distributing it, and independently aggregating the gradient of the water accumulation identification model and the gradient of the disease model to update the corresponding global model parameters; multiple environmental sensing terminals for acquiring current environmental images at a preset acquisition frequency during the monitoring operation phase, extracting the water accumulation state representation vector and water accumulation confidence and uploading them, and also for calculating the gradient of the water accumulation identification model based on local water accumulation annotation data and uploading it according to a preset environmental training cycle; the locally deployed water accumulation identification model is a model updated via a federated learning mechanism; and multiple pathological data terminals for maintaining a local temporal buffer, performing window retrieval and constructing the environmental risk prior distribution based on the distributed incubation period offset Δ, and calculating and uploading the disease model gradient according to a preset disease training cycle.

[0032] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0033] (1) By dynamically calculating the mosquito incubation period offset, the data of water accumulation environment and case data are aligned according to the real biological lag relationship, avoiding the signal misalignment caused by the fixed time window in the traditional method, and significantly improving the causal accuracy of risk prediction.

[0034] (2) The gradient of the water accumulation identification model and the gradient of the disease model are independently and asynchronously aggregated to adapt to the differences in sampling frequency, communication and computing power between the environmental perception end and the pathological data end. There is no need to wait for the slow node to synchronize, thus avoiding resource waste and model update delay.

[0035] (3) By using confidence gating and backtracking window retrieval mechanisms, the feature distribution drift caused by semantic misalignment and high-frequency water accumulation identification model iteration during asynchronous update is effectively suppressed, ensuring the stability of downstream disease model input and improving the robustness and long-term reliability of cross-domain collaborative modeling. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart of a federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution, provided by an embodiment of the present invention.

[0038] Figure 2 This is a schematic diagram of a module of a federated learning system for predicting the risk of mosquito-borne diseases based on water accumulation, provided by an embodiment of the present invention. Detailed Implementation

[0039] To make the technical solution of the present invention clearer, the present invention will be described completely and clearly below in conjunction with specific embodiments. The following embodiments are only some examples of the present invention. Other embodiments obtained by those skilled in the art based on these embodiments without creative effort are all within the protection scope of the present invention.

[0040] Example:

[0041] The core design idea of ​​the federated learning method and system for mosquito-borne disease risk prediction based on water accumulation distribution provided in this embodiment is to decouple the water accumulation identification task and the disease risk prediction task into two relatively independent optimization objectives, and perform cascaded modeling, rather than using end-to-end joint training or direct feature concatenation. Specifically, the water accumulation identification model deployed at the monitoring site is updated only based on its own visual recognition task, without using the occurrence of disease as a training objective; the disease model deployed at the medical end is updated only based on the disease prediction task, and does not backpropagate gradients to the water accumulation identification model. The two types of models communicate unidirectionally only through a set of structured and semantically relatively stable intermediate representations output by the former (i.e., the water accumulation state representation vector and water accumulation confidence, as described below). This design approach is primarily based on the following considerations: If the two types of features are jointly trained end-to-end using a common approach, on the one hand, water accumulation identification and disease prediction often belong to different deployment and management entities (the former may be operated by municipal sanitation or third-party monitoring agencies, while the latter belongs to medical institutions), making it difficult to centralize the original data due to privacy compliance requirements; on the other hand, even if centralized, the water accumulation identification model, for its own optimization needs, typically updates much more frequently than the disease model. If the disease model directly relies on these frequently changing visual features as input, the continuous drift of upstream feature distribution will lead to unstable performance of the downstream model, and may even cause previously learned disease prediction patterns to be overwritten by new features, resulting in catastrophic forgetting. The confidence weighting, temporal backtracking window, and independent aggregation period in subsequent steps of this embodiment all share the common goal of suppressing this feature drift without gradient backpropagation or requiring synchronous waiting between the two types of terminals. The specific mechanism will be explained in conjunction with the following steps.

[0042] Please see Figure 1 Firstly, this embodiment provides a federated learning method for predicting the risk of mosquito-borne diseases based on water accumulation distribution:

[0043] S1. Construction of Federated Learning System and Establishment of Target Monitoring Area Mapping Relationship

[0044] The federated learning system constructed in this embodiment includes three types of nodes: a cloud aggregator, multiple environmental sensing terminals, and multiple pathological data terminals. The environmental sensing terminals are deployed at key monitoring locations, such as sewer entrances, areas around park water bodies, and construction sites—places prone to water accumulation and mosquito breeding. The pathological data terminals are deployed in medical institution information systems, disease control reporting platforms, and other entities capable of acquiring local epidemiological data. The cloud aggregator, acting as a central node, does not collect raw environmental or case data itself; it is only responsible for coordinating information transmission between the environmental sensing terminals and the pathological data terminals, as well as the parameter aggregation performed independently by the water accumulation identification model and the disease model, as described later.

[0045] The construction of this federated learning system requires pre-configuration of target monitoring area mapping relationships, ensuring that each pathological data endpoint is associated with at least one target monitoring area, and each target monitoring area contains at least one environmental sensing endpoint. The establishment of these target monitoring area mapping relationships can be achieved through at least one of the following methods, which can be used individually or in combination: First, based on a pre-configured attribution table of administrative division codes and medical institution practice addresses, a pathological data endpoint (such as a district CDC) is statically bound directly to an environmental sensing endpoint within its administrative jurisdiction, suitable for scenarios with clear administrative boundaries and medical resources allocated according to district boundaries; Second, based on dynamic clustering analysis results of historical case address distribution data and environmental sensing endpoint geographic location data, such as using the DBSCAN equal-density clustering algorithm, monitoring areas are automatically delineated with the geographic clustering center of historical cases as the core, suitable for scenarios with frequent population movement and where administrative boundaries and actual epidemic distribution do not completely overlap; Third, allowing one pathological data endpoint to be associated with multiple target monitoring areas (e.g., a municipal CDC simultaneously covering multiple streets), and allowing one environmental sensing endpoint to belong to multiple target monitoring areas simultaneously (e.g., sensors deployed at the administrative district boundaries), to enhance the system's robustness in spatial coverage. The above mapping relationship can be stored in the regional configuration database of the cloud aggregator and updated regularly as administrative divisions are adjusted or the epidemic situation changes.

[0046] Establishing this mapping relationship is a prerequisite for the cloud aggregator to determine which environmental sensing data should be packaged and sent to which pathological data terminal in subsequent steps. Without this step, the water accumulation information generated by the environmental sensing terminals cannot be routed to the relevant disease prediction tasks, and the environmental data and case data association modeling described in steps S3 to S4 will lose its data foundation. At the same time, the combination of static attribution tables and dynamic clustering, and the allowance for many-to-many overlapping designs, enables the mapping relationship to cover both the relatively fixed routine scenarios of administrative management and the reality of inconsistent case distribution with administrative boundaries. This avoids the omission of environmental risk information that should be associated due to an overly rigid mapping relationship. This is an advantage of step S1 compared to simply dividing by administrative region.

[0047] S2, Water accumulation identification and data upload at the environmental sensing end

[0048] During the monitoring and operation phase, each environmental sensing terminal acquires current environmental image data at a preset collection frequency (e.g., once per hour). The water accumulation status representation vector and the corresponding water accumulation confidence score are extracted by a locally deployed water accumulation recognition model. This water accumulation recognition model is a lightweight semantic segmentation or object detection model that is continuously updated through a federated learning mechanism. For example, it can use MobileNet as the backbone network and combine it with a DeepLabV3+ segmentation head to adapt to the actual conditions of limited computing power of the deployed equipment on site.

[0049] Specifically, the water accumulation state representation vector is taken from the fixed-dimensional semantic embedding vector (e.g., 128-dimensional) output by the fully connected bottleneck layer or projection head of the model, and is L2 normalized to support temporal aggregation operations in subsequent steps. The water accumulation confidence is the area ratio or classification probability value of the water accumulation region output by the model, ranging from 0 to 1. The physical meaning of the water accumulation confidence is the reliability of the water accumulation identification model's judgment that water accumulation exists in the current observation, rather than the level of risk of mosquito-borne diseases caused by the water accumulation. For example, visually large areas of clear water with clear boundaries may obtain a high confidence (the water accumulation identification model has a high degree of certainty in judging that water accumulation exists in the area), but its actual risk of mosquito larvae breeding may be lower than that of small areas of turbid organic water with blurred boundaries (the latter may not have a high confidence, but the risk may be higher). This deliberate semantic decoupling between confidence level and risk intensity is the foundation for all anti-feature drift designs in subsequent steps of this embodiment: when constructing the environmental risk prior distribution in step S4, the confidence level plays a gating role in determining the weight that the visual observation should receive based on its credibility, rather than directly serving as the risk measure itself. This avoids the direct injection of immature or biased risk judgments from the water accumulation identification model into the disease prediction process.

[0050] The environmental sensing terminal packages and uploads the water accumulation state representation vector, water accumulation confidence score, and precise acquisition timestamp to the cloud aggregator. The upload process uses TLS encryption and does not include any original image data to ensure environmental privacy at the monitoring site. In addition to the data upload for disease prediction tasks, the environmental sensing terminal also calculates the gradient of the water accumulation recognition model based on local water accumulation annotation data and uploads this gradient to the cloud aggregator according to a preset environmental training cycle. This gradient is used for the federated aggregation update of the water accumulation recognition model in subsequent step S5, which is a necessary prerequisite for the continuous iterative update of the water accumulation recognition model. The local water accumulation annotation data is obtained in the following ways: Before deployment of the environmental sensing terminal, professional personnel manually annotate historical environmental images, including pixel-level segmentation masks of water accumulation areas or classification labels indicating whether water accumulation is present or absent, serving as the local basic dataset for the first round of federated learning training. During the operation of the environmental sensing terminal, image samples with prediction confidence scores below a preset threshold are manually reviewed and annotated, and continuously supplemented into the local water accumulation annotation data to maintain the iterative effect of the water accumulation recognition model. Using confidence-triggered verification instead of random sampling verification allows for priority annotation of samples where the water accumulation identification model is currently performing poorly, thus concentrating annotation resources on the data that is most helpful for the iteration of the water accumulation identification model.

[0051] S3, cloud aggregator's processing of environmental data and encapsulation of environmental semantic data packets

[0052] After receiving the water accumulation status representation vector, water accumulation confidence level and collection timestamp uploaded by the environmental sensing terminal, the cloud aggregator does not directly forward it to the pathological data terminal as is. Instead, it needs to complete two processes: calculating the latency offset and data encapsulation.

[0053] First, the cloud aggregator acquires meteorological data for the associated target monitoring area within a preset backtracking time window, including time-series data such as daily average temperature and relative humidity, and inputs this time-series data into a preset incubation period model. This incubation period model is a mapping function constructed based on the accumulated temperature rule for pathogen incubation periods. The accumulated temperature rule is a mature modeling method in epidemiology and entomology, used to describe the time required for mosquito-borne pathogens to complete their external development under specific temperature and humidity conditions. It maps meteorological conditions into a duration estimate in days or hours. The incubation period estimate output by the incubation period model based on the current meteorological conditions serves as the incubation period offset Δ for this cycle.

[0054] The introduction of the Δ calculation step addresses the issue of temporal misalignment: water accumulation observed at a specific moment in the environmental sensing system typically does not immediately lead to disease cases at the same time. Instead, it requires a series of processes, including the pathogen completing its incubation period within mosquitoes, the mosquito transmitting the virus through bites, and the onset and recording of symptoms, before it becomes reflected in case data. Without addressing this time lag, directly correlating water accumulation observation data with case data at the time of case collection will introduce incorrect causal relationships, causing the disease model to learn noise rather than the true environment-disease association. The existence of Δ allows subsequent steps to find the historical water accumulation observation window that is biologically relevant to the case by working backward from the case collection time to Δ.

[0055] While calculating Δ, if multiple environmental sensing terminals are deployed within a target monitoring area, the cloud aggregator can first perform regional-level weighted aggregation of the water accumulation status representation vectors uploaded by each environmental sensing terminal within that area to obtain the regional water accumulation semantic vector V. region The weighted aggregation weights are determined by the confidence scores of water accumulation uploaded by each environmental sensing terminal:

[0056] V region =Σ(C i ·V i ) / ΣC i ;

[0057] Where V i Let C be the water accumulation state representation vector uploaded by the i-th environmental sensing terminal. iThe corresponding water accumulation confidence score is summed, covering all participating environmental sensing terminals within the target monitoring area. Correspondingly, the arithmetic mean of the water accumulation confidence scores uploaded by each environmental sensing terminal within the target monitoring area is performed to obtain the regional water accumulation confidence score. The advantage of using regional-level weighted aggregation is that it reduces the interference of observation noise from a single sensor due to factors such as viewpoint obstruction and localized contamination on the overall regional judgment, making the distributed environmental semantic data packets more reflective of the overall condition of the monitoring area.

[0058] After completing the above calculations, the cloud aggregator encapsulates the water accumulation state representation vector (or regional water accumulation semantic vector when regional aggregation exists), water accumulation confidence (corresponding to regional water accumulation confidence), collection timestamp, and latency offset Δ into an environmental semantic data packet. This packet can be generated using structured serialization methods such as Protobuf, and includes a version number field to identify the version of the water accumulation identification model on which the data packet was generated. This version number field works in conjunction with the feature drift monitoring mechanism in the optional embodiments described later. After encapsulation, the data packet is sent to the pathological data terminal associated with the target monitoring area.

[0059] S4. Local buffer maintenance for pathology data, disease risk prediction, and gradient calculation.

[0060] Each pathology data endpoint continuously maintains a local time-series buffer to cache successively received environmental semantic data packets. This time-series buffer is implemented using a circular queue structure, with each storage unit consisting of a triple containing the collection timestamp, a water accumulation status representation vector, and a water accumulation confidence level. Each time a new environmental semantic data packet is received, the pathology data endpoint writes it into the circular queue in the order of its collection timestamp. When the queue is full, it automatically overwrites the oldest historical data. The capacity of the circular queue is determined by the sum of the maximum expected value of the latency offset Δ and the half-width of the preset backtracking window. For example, if the maximum expected value of Δ is 21 days, the half-width of the backtracking window is 7 days, and the environmental sensing endpoint collects data hourly, the queue capacity should be no less than (21+7) days × 24 records / day, or 672 records, to ensure that the time range required for any retrieval request falls within the cached data range. Using a circular queue instead of an infinitely growing storage structure allows for long-term operation with a defined storage overhead, and its automatic overwriting of the oldest data naturally aligns with the practical requirement of this invention to retain only historical data within the latency timescale.

[0061] Based on this, the pathology data terminal acquires local epidemiological data (such as the number of outpatient fever cases, the positive rate of mosquito vector monitoring, etc.) and the corresponding case collection timestamp t2, and extracts pathological features through a pathological feature encoder (such as an LSTM or Transformer encoder). Simultaneously, the pathology data terminal uses t2-Δ as the baseline timestamp, that is, subtracting the incubation period offset Δ calculated in step S3 from the case collection time point, and retrospectively calculates back to the biologically corresponding exposure time point. It then retrieves the water accumulation state characterization vector sequence and the corresponding water accumulation confidence sequence within a preset backtracking window centered on this baseline timestamp from the local circular queue. This method of retrospectively calculating the baseline timestamp and then symmetrically selecting the window is a concrete implementation of this embodiment to solve the problem of temporal misalignment, ensuring that the water accumulation observation data used for judgment is no longer data close to the calendar of the case collection time, but rather data within a historical time window that is truly biologically relevant to the case. The length of the backtracking window is determined by the difference between the maximum and minimum values ​​of the latency offset Δ within the target monitoring area over a preset historical period. If the historically calculated Δ value for a certain area fluctuates significantly (the difference between the maximum and minimum values ​​is large), it indicates that there is considerable uncertainty in the current estimate of Δ, and the window should be widened accordingly to cover a wider range of candidate observation data. Conversely, if the historical Δ value for the area is relatively stable, the window can be narrowed to focus on a more precise temporal neighborhood. This method does not require the introduction of additional statistical assumptions; it only requires continuously recording the Δ value calculated for the area by the cloud aggregator in each round, making it simple to implement.

[0062] Based on the retrieved water accumulation confidence sequence, the pathological data side further constructs a priori distribution of environmental risk. Specifically, at least one of the following methods can be used, either individually or in combination: First, using the water accumulation confidence sequence as weights, perform weighted temporal pooling (weighted average or weighted max pooling) on ​​the water accumulation state representation vector sequence to obtain a fixed-dimensional distribution representation vector. This is simple to implement, has low overhead, and is suitable for scenarios with high real-time requirements. Second, perform kernel density estimation based on the water accumulation confidence sequence and collection timestamps to fit discrete temporal observations into a continuous environmental risk curve and encode it into a fixed-dimensional embedding vector, which can more meticulously depict the temporal distribution of risk. Third, input the water accumulation state representation vector sequence and the water accumulation confidence sequence into a temporal attention network (masking is performed on missing confidence positions). The network adaptively learns the contribution weights of observation data at different time points and outputs a context-aware distribution representation, which has the strongest expressive power but relatively higher overhead. Regardless of the method used, the confidence sequence plays a gating / weighting role: the lower the confidence of an observation, the weaker its contribution to the final distribution representation. This is the specific embodiment of the design idea of ​​decoupling confidence and risk semantics in step S2 in this step. Even if the confidence of a certain observation is not high, the disease model will not treat it as low risk directly, but will reduce its weight in the construction of the distribution representation to avoid unreliable single observations causing too much disturbance to the final prediction.

[0063] Subsequently, the pathological data end inputs pathological features and prior environmental risk distribution representations into a multimodal fusion module (such as a cross-attention layer or a gated fusion network) to obtain a disease risk prediction value. Based on this prediction value and disease labels, a loss is calculated and backpropagated to obtain the disease model gradient. The disease model includes a pathological feature encoder and a multimodal fusion module. During this backpropagation process, the water accumulation state representation vector and water accumulation confidence, which are used as inputs for forward computation, are always treated as fixed environmental context conditions and do not participate in backpropagation, nor do they generate gradients pointing to the water accumulation recognition model (the parameter update boundary corresponds to step S5). The obtained disease model gradient is only used to update the parameters of the disease model. This is because the update rhythm of the water accumulation recognition model is usually faster than that of the disease model. If the gradient's scope is not limited, each local training on the pathological data end may be disturbed by subtle changes in the upstream model version. By strictly limiting the environmental output to non-trainable conditional inputs, the disease model actually learns how to make predictions based on given environmental risk cues—a relatively stable task. The iteration of the water accumulation recognition model is isolated and only exerts its influence indirectly and mildly through the confidence value. The gradient is uploaded to the cloud aggregator after differential privacy processing.

[0064] S5 and cloud aggregator for independent aggregation and parameter distribution of two types of gradients.

[0065] The cloud aggregator maintains two completely independent sets of global model parameters for the two data links: the global disease model parameters and the global water accumulation identification model parameters. These two parameters are not shared and do not affect each other.

[0066] When the cloud aggregator receives disease model gradients uploaded from various pathological data endpoints within a preset disease training cycle (e.g., weekly), it uses a weighted averaging strategy to independently aggregate these gradients, update the global disease model parameters (the global disease model parameters are the same as the disease model parameters, which are the parameters of the pathological feature encoder and the multimodal fusion module), and distributes them to each pathological data endpoint according to the cycle. Similarly, when the cloud aggregator receives water accumulation recognition model gradients uploaded from various environmental sensing endpoints within a preset environment training cycle (e.g., daily), it also uses a weighted averaging strategy (such as FedAvg or its variants) to independently aggregate the water accumulation recognition model gradients, update the global water accumulation recognition model parameters, and distribute them to each environmental sensing endpoint according to the cycle. The global water accumulation recognition model parameters specifically include all parameters of the water accumulation recognition model. The weight calculation rules for these two types of aggregations are independent of each other. For example, the environmental side can be weighted by the sample size or data quality indicators of each endpoint, while the disease side can use similar but independently set rules; the two do not need to be the same.

[0067] The two types of aggregation processes described above are independent of each other and do not require synchronous waiting. Correspondingly, the global water accumulation identification model parameters are updated only based on the gradient of the water accumulation identification model uploaded by the environmental perception end, and the global disease model parameters are updated only based on the gradient of the disease model uploaded by the pathology data end, with no cross-updates. This design directly addresses two issues: First, the environmental perception end and the pathology data end have significant differences in sampling frequency, communication conditions, and computational load. Forcing aggregation in the same round and mutual waiting would lead to resource waste and update delays. This design allows the water accumulation identification model and the disease model to iterate independently according to their own data characteristics without blocking each other. Second, by strictly limiting the update source of the two types of global parameters to their respective gradients and disallowing cross-updates, the frequent iterations on the environmental side directly impact the trained parameters on the disease side.

[0068] Exception handling mechanism

[0069] The aforementioned steps S1 to S5 describe the data flow process between the environmental sensing terminal, the cloud aggregator, and the pathological data terminal under normal operating conditions. However, in actual deployment environments, environmental sensing terminals are often edge devices deployed at locations such as sewer openings or around park water bodies. Unexpected situations such as equipment failure, network communication interruptions, and power supply anomalies are inevitable, causing a particular environmental sensing terminal to fail to upload data in a timely manner according to the preset collection frequency. If such situations are not addressed, the pathological data terminal may encounter a missing data point at a certain timestamp in its local circular queue during the retrieval in step S4, thus affecting the stability of disease risk prediction.

[0070] Therefore, this embodiment also includes an anomaly handling step: when the cloud aggregator does not receive a water accumulation status representation vector uploaded by a certain environmental sensing terminal within a preset timeout period, it sends a missing marker to the pathological data terminal associated with that environmental sensing terminal to inform that the pathological data terminal does not have valid environmental observation data at the corresponding timestamp position. When the pathological data terminal performs window retrieval in step S4, if the retrieved timestamp matches the missing marker, it sets the water accumulation confidence corresponding to that position to zero when constructing the environmental risk prior distribution, or uses a preset default vector and zero confidence to participate in the distribution construction.

[0071] This approach works without requiring an additional compensation logic because it reuses the confidence gating mechanism established in step S4: regardless of the construction method of the environmental risk prior distribution, the confidence sequence determines the weight assigned to each observation. Once the confidence at a certain position is set to zero, the contribution of that position to the final distribution representation will also approach zero, essentially being ignored numerically, eliminating the need for a separate missing value imputation algorithm. The advantage of this approach is that when individual environmental sensing devices experience brief offline periods or malfunctions, the disease risk prediction for the affected area will not fluctuate drastically, and the cloud aggregator can continue distributing data at its original pace without waiting for the device to recover. This aligns with the design principle in step S5 where the environmental and pathological sides do not need to wait synchronously.

[0072] Please see Figure 2 Secondly, this embodiment provides a federated learning system for predicting the risk of mosquito-borne diseases based on water accumulation distribution:

[0073] The aforementioned steps S1 to S5 and the exception handling mechanism are implementation methods of the present invention at the method level. From the system level, the federated learning system that implements the above method also constitutes an implementation method of the present invention. The three types of nodes it contains and the functions undertaken by each node correspond one-to-one with the method implementation methods described above.

[0074] The system is configured to execute a federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution, as described above. The system includes a cloud aggregator for acquiring meteorological data and dynamically calculating the incubation period offset Δ (corresponding to the calculation process of Δ in step S3), encapsulating Δ and environmental data into an environmental semantic data package and distributing it (corresponding to the encapsulation and distribution process in step S3), and independently aggregating the gradients of the water accumulation identification model and the disease model to update the corresponding global model parameters (corresponding to the independent aggregation process in step S5). It also includes multiple environmental sensing terminals for collecting data at a preset frequency. The system collects environmental images, extracts water accumulation state representation vectors and water accumulation confidence scores, and uploads them (corresponding to the data collection and upload process in step S2). It is also used to calculate the gradient of the water accumulation identification model based on local water accumulation annotation data and upload it according to a preset environmental training cycle (corresponding to step S2). The water accumulation identification model deployed locally is a model that is continuously updated through a federated learning mechanism. It also includes multiple pathological data terminals, which are used to maintain local temporal buffers, perform retrieval based on the issued Δ execution window and construct environmental risk prior distribution (corresponding to step S4), and calculate and upload disease model gradients according to a preset disease training cycle (corresponding to step S4).

[0075] Furthermore, it also includes optional embodiments: feature drift monitoring and recalibration prompts.

[0076] The confidence gating, backtracking window, and independent aggregation cycle designs in steps S4 and S5 above constitute the mechanism for passively absorbing the impact of feature drift in this embodiment. That is, even when features have already drifted to a certain extent, the weight of unreliable observations is reduced to smooth cross-version transition data, preventing drastic fluctuations in the downstream disease model. Furthermore, this embodiment can also introduce an optional implementation method of active monitoring and on-demand recalibration as a supplement to the aforementioned passive absorption mechanism.

[0077] Specifically, after each independent aggregation based on the gradient of the water accumulation recognition model and updating the parameters of the global water accumulation recognition model, the cloud aggregator can calculate the cosine distance between the representation vectors output by the two versions of the global water accumulation recognition model before and after the update on a preset standard validation set. This cosine distance intuitively reflects the degree of change in the semantic embedding space after this round of updates: the smaller the distance, the more gradual the fine-tuning for downstream applications; the larger the distance, the more drastic the semantic space change. If the cosine distance exceeds a preset threshold (e.g., 0.3), a significant feature drift is determined, and the cloud aggregator sends a recalibration prompt to each associated pathological data endpoint.

[0078] Upon receiving the prompt, the pathology data provider can temporarily increase the width of the preset backtracking window in the next training round, or, in the case of using a temporal attention network to construct a prior distribution of environmental risks, temporarily increase the temperature coefficient of the attention network (increasing the temperature coefficient makes the distribution of attention weights smoother and more uniform, weakening the dominant role of observations at individual time points). The common effect of these two response methods is to proactively reduce the dependence on a single observation point, allowing the pathology data provider to adapt more smoothly to the new feature distribution, rather than having old and new features directly impact the trained parameters without any buffer.

[0079] Compared to the passive drift absorption mechanism in steps S4 and S5, this alternative implementation method has the advantage of changing from passive to active: it does not have to wait until the feature drift has actually affected the prediction accuracy before gradually correcting it, but can actively intervene as soon as a significant drift is detected, further reducing the impact on the stability of disease prediction in scenarios with large version jumps in the water accumulation identification model.

[0080] In summary, the federated learning method and system for predicting mosquito-borne disease risk based on water accumulation distribution in this embodiment has the following advantages:

[0081] Firstly, by dynamically calculating Δ in step S3 and using window retrieval based on t2-Δ in step S4, the environmental data and case data are aligned in a biological sense, thus solving the temporal misalignment problem in the background technology and improving prediction accuracy.

[0082] Secondly, by independently and asynchronously aggregating the gradients of the water accumulation identification model and the disease model in step S5, the two types of terminals can iterate at their own adaptive rhythms without needing to wait for synchronization, thus avoiding the waste of resources and update delays caused by forced synchronization.

[0083] Third, by using confidence gating and backtracking window retrieval mechanisms, combined with an anomaly handling mechanism and optional enhanced implementation of feature drift active monitoring, the feature drift caused by asynchronous updates and high-frequency iterations of the water accumulation identification model is effectively suppressed, ensuring the stability of downstream disease model input, avoiding catastrophic forgetting, and improving the robustness and long-term reliability of the system in scenarios of node failure or drastic model version changes.

[0084] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution, characterized in that, Includes the following steps: S1. Construct a federated learning system, which includes a cloud aggregator, multiple environmental sensing terminals, and multiple pathological data terminals; pre-configure target monitoring area mapping relationships so that each pathological data terminal is associated with at least one target monitoring area, and each target monitoring area contains at least one environmental sensing terminal. S2. During the monitoring and operation phase, each environmental sensing terminal acquires current environmental image data at a preset acquisition frequency, extracts the water accumulation state representation vector and the corresponding water accumulation confidence level through the locally deployed water accumulation recognition model, and uploads the water accumulation state representation vector, the water accumulation confidence level and its acquisition timestamp to the cloud aggregator; the environmental sensing terminal also calculates the gradient of the water accumulation recognition model based on the local water accumulation annotation data, and uploads the water accumulation recognition model gradient to the cloud aggregator according to a preset environmental training cycle; S3. The cloud aggregator acquires meteorological data of the target monitoring area and dynamically calculates the incubation period offset Δ of the current round based on the preset incubation period model. The water accumulation status representation vector, water accumulation confidence, collection timestamp, and latency offset Δ uploaded by each environmental sensing terminal belonging to the same target monitoring area are encapsulated into an environmental semantic data packet and sent to the pathological data terminal associated with the target monitoring area. S4. Each pathological data terminal obtains local epidemiological data and corresponding case collection timestamp t2, and extracts pathological features through the pathological feature encoder. Using t2-Δ as the baseline timestamp, a sequence of water accumulation state representation vectors and corresponding water accumulation confidence sequences centered on the baseline timestamp and within a preset backtracking window length are retrieved from the local time-series buffer. An environmental risk prior distribution is constructed based on the water accumulation confidence sequences. The pathological features and the environmental risk prior distribution are input into a multimodal fusion module to obtain a disease risk prediction value. The disease model gradient is calculated based on the disease risk prediction value and disease labels and uploaded to the cloud aggregator. The disease model includes the pathological feature encoder and the multimodal fusion module. S5. When disease model gradients uploaded by each pathological data terminal are received within a preset disease training period, the cloud aggregator performs independent aggregation on the disease model gradients to update the global disease model parameters, and distributes the updated global disease model parameters to each pathological data terminal according to the preset disease training period; when water accumulation recognition model gradients uploaded by each environmental perception terminal are received within a preset environment training period, the water accumulation recognition model gradients perform independent aggregation on the water accumulation recognition model gradients to update the global water accumulation recognition model parameters, and distributes the updated global water accumulation recognition model parameters to each environmental perception terminal according to the preset environment training period; wherein, the global water accumulation recognition model parameters are updated only based on the water accumulation recognition model gradients uploaded by the environmental perception terminal, and the global disease model parameters are updated only based on the disease model gradients uploaded by the pathological data terminal, and the aggregation and updates of the two are independent of each other and do not require synchronous waiting.

2. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S2, the water accumulation state representation vector is a fixed-dimensional semantic embedding vector obtained by the water accumulation recognition model through a fully connected bottleneck layer or a projection head, and the semantic embedding vector supports temporal aggregation operations; the water accumulation confidence is the area ratio or classification probability value of the water accumulation region output by the water accumulation recognition model; the local water accumulation annotation data is obtained in the following way: before the deployment of the environmental perception terminal, historical environmental images are manually annotated as the local basic dataset for the first round of training of federated learning; during the operation of the environmental perception terminal, image samples whose prediction confidence of the water accumulation recognition model is lower than a preset threshold are manually reviewed and annotated, and continuously supplemented into the local water accumulation annotation data.

3. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S3, before the cloud aggregator sends out the environmental semantic data packet, it further includes: performing regional weighted aggregation on the water accumulation status representation vectors uploaded by multiple environmental sensing terminals belonging to the same target monitoring area to obtain a regional water accumulation semantic vector; wherein, the weight of the weighted aggregation is determined by the water accumulation confidence score uploaded by each environmental sensing terminal; performing an arithmetic average on the water accumulation confidence scores uploaded by the multiple environmental sensing terminals to obtain a regional water accumulation confidence score; and encapsulating the regional water accumulation semantic vector and the regional water accumulation confidence score into the environmental semantic data packet instead of the data from a single environmental sensing terminal and sending it out.

4. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S4, the length of the preset backtracking window is determined by the difference between the maximum and minimum values ​​of the latency offset Δ within the target monitoring area during a preset historical period; the construction of the environmental risk prior distribution based on the water accumulation confidence sequence specifically includes at least one of the following methods: Using the water accumulation confidence sequence as weights, weighted temporal pooling is performed on the water accumulation state characterization vector sequence to obtain the distributed characterization vector; Based on the water accumulation confidence sequence and the collection timestamp, kernel density estimation is performed to generate an embedded environmental risk curve. The sequence of water accumulation state representation vectors and the sequence of water accumulation confidence are input into a temporal attention network to output a context-aware distribution representation.

5. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S3, the dynamic calculation method for the latency offset Δ is as follows: Acquire time-series data of daily average temperature and relative humidity for the target monitoring area within a preset backtracking time window; The time series data is input into the preset incubation period model, which is a mapping function constructed based on the pathogen external incubation period accumulated temperature rule; The estimated incubation period of the pathogen under the current meteorological conditions is output as the incubation period offset Δ.

6. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S1, the method for establishing the target monitoring area mapping relationship includes at least one of the following: A pre-configured attribution table based on administrative division codes and medical institution practice addresses; The results of dynamic clustering analysis based on historical case address distribution data and environmental sensing terminal geographic location data; A single pathological data endpoint is associated with multiple target monitoring areas, and a single environmental sensing endpoint simultaneously belongs to multiple target monitoring areas.

7. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S4, the local time-series buffer adopts a circular queue structure, and its storage unit is a triple containing the collection timestamp, the water accumulation state representation vector, and the water accumulation confidence. The capacity of the circular queue is determined by the sum of the maximum expected value of the latency offset Δ and the half-width of the preset backtracking window. When the pathological data terminal receives a new environmental semantic data packet, it writes it into the circular queue in the order of the collection timestamp and automatically overwrites the historical data that exceeds the capacity.

8. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, In step S5, the independent aggregations all adopt a weighted average strategy, and the independent aggregations of the gradients of the water accumulation identification model and the independent aggregations of the gradients of the disease model adopt independent weight calculation rules respectively; the global disease model parameters include the parameters of the disease model, and the global water accumulation identification model parameters include all the parameters of the water accumulation identification model; in step S4, the water accumulation state representation vector and the water accumulation confidence are used as fixed environmental context conditions input during training, and the disease model gradient is only used to update the parameters of the disease model.

9. The federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution according to claim 1, characterized in that, The method also includes an exception handling step: When the cloud aggregator does not receive a water accumulation status representation vector uploaded by a certain environmental sensing terminal within a preset timeout period, it sends a missing marker to the pathological data terminal associated with that environmental sensing terminal. If the pathological data terminal matches the timestamp corresponding to the missing marker during the retrieval in step S4, the confidence level of the water accumulation at the corresponding position will be set to zero when constructing the environmental risk prior distribution, or a preset default vector and zero confidence level will be used to participate in the distribution construction.

10. A federated learning system for predicting mosquito-borne disease risk based on water accumulation distribution, characterized in that, The system is configured to execute a federated learning method for predicting mosquito-borne disease risk based on water accumulation distribution as described in any one of claims 1 to 9, comprising: a cloud aggregator for acquiring meteorological data and dynamically calculating the incubation period offset Δ, encapsulating the incubation period offset Δ and environmental data together into an environmental semantic data package and distributing it, and independently aggregating the gradient of the water accumulation identification model and the gradient of the disease model to update the corresponding global model parameters; multiple environmental sensing terminals for acquiring current environmental images at a preset acquisition frequency during the monitoring operation phase, extracting the water accumulation state representation vector and water accumulation confidence and uploading them, and also for calculating the gradient of the water accumulation identification model based on local water accumulation annotation data and uploading it according to a preset environmental training cycle; the locally deployed water accumulation identification model is a model updated via a federated learning mechanism; and multiple pathological data terminals for maintaining a local temporal buffer, performing window retrieval and constructing environmental risk prior distribution based on the distributed incubation period offset Δ, and calculating and uploading the disease model gradient according to a preset disease training cycle.