Visualized early warning platform based on management of vector ecology
By constructing a multi-scale encoder-decoder model and a cascaded prediction method, the problems of delayed timeliness and insufficient visualization of vector-borne disease monitoring data were solved, enabling real-time prediction and visualization decision support for vector-borne disease risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING CENT FOR DISEASE CONTROL & PREVENTION (CHONGQING EMERGENCY TREATMENT CENT FOR DISASTER RELIEF & DISEASE PREVENTION)
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing vector-borne disease monitoring models suffer from time lag, insufficient data analysis capabilities, and an inability to achieve real-time risk perception and early warning. Furthermore, the low level of data visualization fails to provide managers with intuitive decision-making support.
A method for predicting the risk of vector-borne disease transmission based on a multi-scale encoder-decoder model is constructed. This method combines multi-scale feature construction with weakly supervised risk label generation, employs cascaded prediction and hybrid sequence correction, identifies peak windows through dual thresholds, generates graded early warnings, and displays the risk levels in real time on a visualization platform.
It enables real-time prediction and visualization of vector-borne disease monitoring data, improves the timeliness and accuracy of risk warnings, provides intuitive decision support, and meets the needs of forward-looking management.
Smart Images

Figure CN122117470A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network data communication for natural focal diseases, and specifically discloses a visualized early warning platform based on vector-borne biological ecology management. Background Technology
[0002] Currently, the monitoring of disease-carrying organisms such as mosquitoes, flies, rats, and cockroaches mainly adopts an independent, fixed-point monitoring model based on administrative divisions (such as districts and counties). Specifically, each regional disease control department sets up scattered fixed monitoring points (such as residential areas and parks) within its jurisdiction and conducts regular manual monitoring according to national standard methods (such as mosquito-attracting lamps and cage traps). Technicians need to set up and retrieve equipment on-site, manually count and classify the captured organisms, and record environmental parameters. Finally, the data is summarized and reported in paper or electronic form. Some regions have introduced intelligent monitoring terminals that can automatically upload data and have established regional independent databases, realizing the initial digitization of single-point or single-area data collection.
[0003] However, the existing monitoring models mentioned above still have significant shortcomings in meeting the needs of modern, efficient, and precise public health management. The core problems are mainly reflected in the lag in timeliness and the lack of analytical and judgment capabilities. First, the traditional monitoring methods, which are mainly based on manual labor, have long cycles and low frequencies. Data collection, reporting, processing and analysis reports are often delayed by several days or even weeks, making it impossible to achieve real-time risk perception and early warning. Second, monitoring data, laboratory data and other multi-source information are usually scattered and lack effective integration and in-depth analysis. At the same time, the existing systems focus on data recording and post-event statistics, lacking prediction and intuitive visualization of spatiotemporal dynamic risks, making it difficult to support forward-looking decision-making. Summary of the Invention
[0004] In view of this, one of the objectives of the present invention is to provide a method for predicting the risk of vector-borne disease transmission, comprising the following steps:
[0005] S1: Collect monitoring point identification, administrative division code and environmental meteorological data, simultaneously collect vector biological monitoring indicators and population epidemiological data, and preprocess the collected data;
[0006] S2: Multi-scale feature construction and weakly supervised risk label generation: First, cumulative effect features based on time decay operators are constructed for average temperature, relative humidity and rainfall data to characterize the lagged impact of environmental factors on vector breeding. At the same time, based on expert rules, vector density component, biological transmission component and population response component are integrated to synthesize a continuous infectious risk index and normalize it as a weakly supervised ground truth label for the second stage of model training.
[0007] S3: Construct a multi-scale encoder-decoder model with stage identifiers: Construct a multi-scale encoder-decoder model that includes a multi-scale embedding layer, a temporal encoder, and a channel encoder. Use multi-resolution convolutional slices and gating networks to capture temporal changes at different granularities. At the same time, introduce task-aware stage identifier vectors and spatial identifier vectors as conditional inputs. Adjust the feature scaling and translation of each layer through conditional normalization layers so that a single model can adaptively adapt to different prediction tasks and different spatial regions under shared parameters.
[0008] S4: Perform two-stage cascaded prediction and mixed sequence correction, using a cascaded inference strategy. In the first stage, the density stage identifier drives the model to output a future vector density prediction sequence. In the second stage, the prediction sequence is concatenated with the original input, and the risk stage identifier drives the model to output an infection risk index.
[0009] S5: Peak and window identification based on dual thresholds: Set dual thresholds for density warning and infection risk, and transform the continuous prediction sequence into a binary state sequence to identify the starting point of the future peak and the peak window that continuously exceeds the threshold. Further, through spatiotemporal coupling analysis, detect whether the density peak and the risk peak overlap or are closely connected, thereby determining the high-confidence transmission outbreak period and distinguishing between simple biological growth and actual transmission risk.
[0010] S6: Graded Early Warning Generation and Visual Decision Support: Based on the identified peak time, duration window, and peak intensity, early warning signals are generated in combination with preset graded rules. In the front-end display module, the real-time risk level of each area is presented as map color blocks, and the predicted curves of density and risk are overlaid on the trend analysis chart.
[0011] The second objective of this invention is to provide a visual early warning platform based on vector-borne organism ecological management, in order to solve the technical problem that existing vector monitoring data has a low degree of visualization and cannot provide managers with intuitive and timely decision-making basis.
[0012] To achieve the above objectives, the present invention provides the following technical solution:
[0013] It includes a data acquisition module, a central server, and an interaction module;
[0014] The data acquisition module is used to collect raw data from monitoring areas at the municipal, district, and county levels.
[0015] The central server includes a vector-borne disease data processing module, a vector-borne disease transmission risk prediction module, and a map visualization module. The vector-borne disease data processing module is used to process the data collected by the collection module.
[0016] The vector-borne disease transmission risk prediction module, based on the vector-borne disease transmission risk prediction method described above, enables the prediction of vector density and the risk of vector-borne disease transmission.
[0017] The map visualization module is used to convert the data from the vector-borne organism data processing module into hierarchical color block illustrations on an electronic map.
[0018] The interactive module is used to provide a specialized data interaction interface for specific disease vectors.
[0019] This platform first collects raw data on disease vectors through various monitoring devices deployed on-site. This data is then coordinated by hierarchical uploading submodules based on administrative levels, and transmitted to the central server via an edge data aggregation server or directly. Subsequently, the disease vector data processing module is activated. Its data standardization and preprocessing submodule automatically reviews and standardizes the imported data. Data that passes the review is distributed to mosquito, rodent, and tick submodules. The mosquito submodule calculates the Breteau index and tent attraction index and determines the risk level. The rodent submodule calculates the capture rate and path index and performs a risk assessment. The tick submodule calculates the tick density index and generates risk alerts. These results drive the map visualization module. The system uses a geographic rendering submodule to render risk levels on an electronic map using standard color gradations. Users can perform drill-down analysis through hierarchical display submodules or replay historical data through timeline control submodules. Meanwhile, real-time data card submodules, risk ring chart submodules, and trend analysis submodules provide multi-dimensional auxiliary views. Finally, users can access specialized interfaces for refined data filtering and querying through mosquito, rodent, and tick interactive submodules in the interactive module. The entire process achieves full-link integration from data collection, processing, visualization to interactive analysis, solving the technical problem of low visualization of existing vector monitoring data, which cannot provide managers with intuitive and timely decision-making basis. Attached Figure Description
[0020] Figure 1 This is a flowchart of a visual early warning platform based on vector-borne organism ecological management.
[0021] Figure 2 This is a schematic diagram of the data collection module of a visual early warning platform based on vector-borne organism ecological management.
[0022] Figure 3 This is a schematic diagram of the vector-borne organism data processing module of a visualization-based early warning platform for vector-borne organism ecological management.
[0023] Figure 4 This is a schematic diagram of the map visualization module of a visualization early warning platform based on vector-borne organism ecological management.
[0024] Figure 5This is a schematic diagram of the interactive module of a visual early warning platform based on vector-borne organism ecological management.
[0025] Figure 6 This is a schematic diagram of the framework of the vector-borne disease transmission risk prediction method in Example 2. Detailed Implementation
[0026] Example 1
[0027] like Figures 1 to 5 As shown in the figure, a visual early warning platform based on vector-borne organism ecological management is disclosed. This embodiment of the visual early warning platform based on vector-borne organism ecological management is mainly configured and applied to vector-borne organism monitoring and early warning management within a municipal administrative region. It includes a data acquisition module, a central server, and an interaction module. The data acquisition module serves as the platform's data source, responsible for automatically or manually collecting raw vector-borne organism data through various monitoring devices installed on-site. The central server includes a vector-borne organism data processing module, a vector-borne infection risk prediction module, and a map visualization generation module. The vector-borne infection risk prediction module is based on the vector-borne infection risk prediction method described in Embodiment 2. The system employs various methods to predict vector density and vector-borne disease transmission risk. A vector-borne organism data processing module receives raw data from the acquisition module, performs review and standardization processing, while a map visualization module receives and processes geographic information and monitoring data from the vector-borne organism data processing module. This transforms digitized monitoring indices into intuitive, hierarchical color-block illustrations overlaid on an electronic map. An interactive module serves as the front-end interface for users to interact with various platform functions. Dedicated control buttons allow users to activate and switch to independent interactive interfaces for specific vectors, providing specialized filtering and query functions for monitoring data of that type of vector within each interface.
[0028] like Figure 2 As shown, the data collection module includes multiple vector-borne disease detection devices, an edge data aggregation server, and a hierarchical upload submodule. The vector-borne disease detection devices include mosquito-attracting lamps for mosquito monitoring, rat traps and cages for rodent monitoring, and cloth flags for tick collection. These devices constitute the platform's front-end data source, enabling initial data collection such as vector capture, counting, and identification through manual methods. The hierarchical upload submodule automatically executes different upload paths based on the administrative level of the data source to achieve hierarchical data aggregation and transmission. For city-level monitoring points, data is directly transmitted to the central server. For district / county-level and lower-level monitoring points, data is first uploaded to the edge data aggregation server configured for that area. The edge data aggregation server is deployed at district / county-level network nodes and is responsible for receiving, temporarily storing, and initially integrating raw data from all monitoring devices within its jurisdiction, before synchronizing the data to the central server according to a set cycle or trigger conditions.
[0029] In this solution, the initial acquisition of data is achieved through vector-borne disease detection devices. As physical terminals directly deployed at the monitoring site, these devices can continuously and objectively capture first-hand information on vector-borne disease activity, thus ensuring the objectivity of the data source. The edge data aggregation server effectively solves the problems of network congestion and excessive pressure on the central server caused by the direct concurrent upload of massive amounts of terminal data by aggregating data locally at various district and county-level nodes, thereby improving the stability and efficiency of data transmission. The central server, as the destination of all data streams, ensures unified access and centralized storage of data. The hierarchical upload submodule, through intelligent path selection, distributes and optimizes data streams according to administrative levels, realizing an efficient transmission mode of direct transmission at the city level and aggregation at the district and county levels. Compared with the traditional mode of independent monitoring and decentralized data management by different institutions, this solution, through its intelligent hierarchical upload architecture, ensures the timeliness of all monitoring data while effectively avoiding network congestion caused by the direct concurrent access of massive amounts of terminal data to the central server, thereby significantly optimizing the overall network resource utilization. The entire module constitutes an efficient, stable, and scalable unified data acquisition and aggregation system.
[0030] like Figure 3 As shown, the vector-borne disease data processing module includes a data standardization and preprocessing submodule, a regional statistics submodule, a mosquito submodule, a rodent submodule, and a tick submodule. The data standardization and preprocessing submodule serves as the platform's unified data processing backend, used to perform automatic and manual review of the received raw data. The regional statistics submodule dynamically counts and displays the total number of monitoring districts and counties with valid reported data on the current platform and their proportion of all districts and counties in the city. It also summarizes and displays the total number of all active monitoring points on the platform in real time, reflecting the scale of monitoring resource investment, and counts the number of monitoring points that have completed data reporting in the latest reporting period, quickly presenting the overall progress of the data collection task. The mosquito submodule directly calculates the BI index and the tent attraction index. The rodent submodule calculates the capture rate and the path index. The tick submodule displays the tick index and the density index.
[0031] The data standardization preprocessing submodule includes a data receiving unit, an audit and judgment unit, and a data maintenance unit. The data receiving unit provides a standardized data entry interface to receive data from the acquisition module, including but not limited to key fields such as monitoring point location, time, vector type, capture quantity, and environmental parameters. The audit and judgment unit performs an audit process on the raw data. Its audit logic is based on preset data integrity rules (such as whether required fields are complete), numerical rationality rules (such as whether the capture quantity is within the historical normal fluctuation range or exceeds the theoretical maximum value), and logical consistency rules (such as whether the monitoring time matches the task cycle). The data first undergoes automatic rule verification. If it passes all rules, it is marked as approved and enters the processing queue. If any rule verification fails, it is automatically marked as awaiting manual review and pushed to the administrator terminal for final judgment. The data maintenance unit maintains and corrects the data in the audit process. For data that fails the audit, the original submitter can modify it according to the feedback reasons and resubmit it.
[0032] In this solution, the data receiving unit automatically acquires and distributes data from the central server, replacing the inefficient traditional method of manual export and distribution. The review and judgment unit performs automatic review based on preset rules, with most data verification work completed automatically by the system. Compared with the traditional method of relying entirely on manual review, this greatly improves processing efficiency and reliability while ensuring consistency. The automatic-to-manual mechanism retains the necessary human intervention channel for complex situations, balancing efficiency and accuracy. The data maintenance unit provides a standardized channel for modifying and resubmitting data that fails the review, forming a closed-loop processing system and changing the drawbacks of traditional processes where data errors are difficult to trace and correct.
[0033] The mosquito submodule includes a mosquito index calculation unit and a mosquito risk assessment unit. The mosquito index calculation unit is used to receive and process standardized raw data that has been reviewed and approved by the data standardization preprocessing submodule, and automatically calculates the Breteau index and the tent attraction index according to the national standard formula.
[0034] The BI index is calculated using the formula BI index = (number of positive containers / number of households inspected) × 100, and is used to quantify the risk of Aedes mosquito larvae breeding (generally, an index > 20 indicates a risk of transmission). The tent trap index is calculated using the formula tent trap index = total number of mosquitoes captured / number of mosquito nets (traps) deployed, and is used to directly reflect the density level of adult mosquitoes and the risk of biting.
[0035] The mosquito risk assessment unit is used to compare the index results obtained by the mosquito index calculation unit with the preset risk threshold and automatically determine the risk level.
[0036] The rodent submodule includes a rodent index calculation unit and a rodent risk assessment unit. The rodent index calculation unit is used to receive and process standardized raw data that has been reviewed and approved by the data standardization preprocessing submodule, and automatically calculates the rodent density capture rate and path index based on the number of captured rodents and rodent track survey data.
[0037] The capture rate is calculated using the formula: capture rate (%) = (number of mice captured / number of effective traps or cages) × 100%, which directly reflects the absolute population density of mice per unit space. The path index is calculated using the formula: path index (locations / km) = number of mouse tracks found / length of the inspection path (km), which reflects the range and frequency of mouse activity in the environment.
[0038] The rodent risk assessment unit is used to assess the intensity of rodent activity and potential risks based on the indicators output by the rodent indicator calculation unit.
[0039] The tick submodule includes a tick density calculation unit and a tick risk assessment unit. The tick density calculation unit is used to receive and process standardized raw data that has been reviewed and approved by the data standardization and preprocessing submodule, process data from the flag method and the host surface tick detection method, and calculate the environmental free tick density index and the parasitic tick density index.
[0040] Tick density index is usually calculated using the formula Tick density index (ticks / standard unit) = total number of ticks captured / flag distance (km) or number of animals examined (ticks). The flag method (unit: ticks / km) is used to assess the density of free ticks in the environment, while the host surface tick detection method (unit: ticks / host) is used to assess the density of parasitic ticks.
[0041] The tick risk assessment unit is used to assess tick density levels and potential risks based on the density index output by the tick density calculation unit.
[0042] In this solution, the data standardization preprocessing submodule conducts unified review of multi-source raw data, overcoming the shortcomings of traditional manual processing, such as low efficiency and difficulty in standardization, thus laying the foundation for data quality. The regional statistics submodule realizes dynamic statistics and display of monitoring districts and counties, monitoring points, and reporting progress, changing the situation of traditional statistical lag and unclear overall situation. The dedicated mosquito, rodent, and tick submodules calculate and display their professional indices, such as BI index, capture rate, and tick density index, respectively, for core disease vectors. This breaks through the limitations of traditional comprehensive reports or simple data listings in terms of in-depth analysis and rapid judgment, and realizes precise classification and intuitive presentation of data. All submodules work together to transform raw data into standardized, statistical, and specialized decision-making information, significantly improving the systematicness, timeliness, and professionalism of data processing.
[0043] like Figure 4As shown, the map visualization module includes a geographic rendering submodule, a hierarchical display submodule, a timeline control submodule, a real-time data card submodule, a risk ring chart submodule, and a trend analysis submodule. The geographic rendering submodule is used to visualize the vector risk level of each district and county in the form of a block-based color-gradient map. The hierarchical display submodule is used to respond to user clicks to achieve a step-by-step drill-down display from districts and counties to villages and towns. The timeline control submodule is used to provide historical data that can be filtered by year, month, and day. The real-time data card module is used to display the latest monitoring summary of each district and county in real time in a fixed area of the sidebar of the map view. The risk ring chart submodule is used to generate a proportional ring chart based on the risk level distribution of each district and county within the current view range. The trend analysis submodule is used to work in conjunction with the risk ring chart submodule to collaboratively display the time-series evolution trend of the number of districts and counties at each risk level.
[0044] The geographic rendering submodule includes a base map loading unit and a risk overlay unit. The base map loading unit is used to call third-party map APIs or offline vector maps to load the basic geographic base map. The platform performs coordinate system unification and secure call encapsulation on third-party map services and ensures that all displayed map base maps are derived from nationally approved surveying and mapping results. The base map hierarchy includes at least five levels of administrative boundary data: provincial, municipal, district / county, township, and village. The risk overlay unit is used to convert the risk data of each district / county into map colors. The risk overlay unit reads the risk-assessed data from the vector-borne disease data processing module and divides it into four categories according to a unified standard. The risk level is determined by the BI index. A BI index of less than 5 is low risk and is rendered in green; 5 to 10 is transmission risk and is rendered in yellow; 10 to 20 is cluster outbreak risk and is rendered in orange; and above 20 is local outbreak risk and is rendered in red. Other vector-borne disease monitoring indicators (such as rodent capture rate and tick density index) are converted into a unified risk level (i.e., low risk, transmission risk, cluster outbreak risk, and local outbreak risk) based on their corresponding national standards or preset thresholds. Then, the corresponding colors (green, yellow, orange, and red) are used for rendering based on this unified level to ensure the consistency of the visual expression of risk across the entire map.
[0045] The hierarchical display submodule includes a click detection unit, a boundary switching unit, and a detail loading unit. The click detection unit reads the geocode of a colored area and triggers a drill-down command when a user clicks on it. The boundary switching unit dynamically switches the administrative boundary level displayed on the map after receiving a drill-down command. When a district or county-level click command is received, it switches from the district or county boundary data source to the subordinate township boundary data and re-renders the map. When a township-level click command is received, it switches to the administrative village or community boundary. Drill-down to the natural village or villager group level is supported. The detail loading unit synchronously loads the monitoring data of the lower-level area and updates the risk chromatogram during the drill-down process.
[0046] The timeline control submodule includes a time-series data storage unit, a calendar selection unit, and a data backtracking unit. The time-series data storage unit receives and stores standardized monitoring data and risk level results output by the vector-borne disease data processing module, forming a historical monitoring time-series database that can be quickly queried. The calendar selection unit supports single date selection, continuous date range selection, and quick selection of preset time windows. After the user selects a date, the timestamp is output to the data backtracking unit. The data backtracking unit receives the timestamp parameter, extracts the corresponding historical monitoring data from the time-series data storage unit, and calls the geographic rendering submodule to regenerate the risk thematic layer for that time point.
[0047] The real-time data card submodule includes a data carousel unit and a details viewing unit. The data carousel unit is used to display the real-time data of all monitored districts and counties in the form of a card list. Each card displays the district / county name, the latest BI index value, and the corresponding risk level label. The cards are arranged in descending order of risk level, with high-risk cards displayed at the top. The list scrolls automatically and pauses when the mouse hovers over it. The details viewing unit is used to respond to user clicks on cards. After clicking, the map is automatically located to the corresponding district / county and the hierarchical display submodule is triggered to display the details of the townships under the jurisdiction of that district.
[0048] The risk ring chart submodule includes a data integration unit and a graphics rendering unit. The data integration unit is used to statistically analyze the distribution of the number of each risk level in all visible districts and counties within the current map viewport, and calculate the percentage of the four levels: low risk, transmission risk, cluster epidemic risk, and local outbreak risk. The graphics rendering unit is used to call the chart library to render the ring chart, with the four arcs filled with green, yellow, orange, and red respectively, and each arc labeled with a percentage value.
[0049] The trend analysis submodule includes a trend calculation unit and a line graph rendering unit. The trend calculation unit is used to extract the daily risk level data of each district and county within the current statistical period from the time series data storage unit, count the number of districts and counties at each level on a daily basis, and generate a time series. The line graph rendering unit is used to render a single Y-axis line graph, with time on the horizontal axis. The line is dynamically segmented and colored according to the daily risk level, strictly following the color mapping of the ring graph. The color change of a single line reflects the dynamic trend of the proportion of risk level components.
[0050] In this solution, the geographic rendering submodule transforms abstract risk level data into an intuitive, block-based color-coded map. Compared to the traditional table or text report format, this provides a clear overview of the spatial distribution of risks, greatly improving the efficiency of overall situation perception. The hierarchical display submodule supports drill-down display from districts and counties to villages and towns, changing the rigid mode of traditional systems that can only view data at a single administrative level. It provides flexible and in-depth data exploration capabilities, meeting different needs from macro-level decision-making to micro-level investigation. The timeline control submodule provides time-based filtering and historical data playback functions, overcoming the shortcomings of traditional methods that can only view static snapshots of the current or specific time points. This enables dynamic tracing and comparative analysis of risk situations over time.
[0051] like Figure 5 As shown, the interaction module includes mosquito-related interaction sub-modules, rodent-related interaction sub-modules, and tick-related interaction sub-modules. The mosquito-related interaction sub-module includes a first button unit and a first filtering control unit. The first button unit serves as the entry point for mosquito data interaction, responding to user clicks, controlling the switching of the platform's main interface, and launching the mosquito-specific interaction interface. The first filtering control unit is integrated into the mosquito-specific interaction interface, providing a multi-level filtering panel for mosquito monitoring data. It allows users to sequentially select administrative regions and monitoring indicators (such as the BI index or the tent attraction index), and generate query commands based on the filtering conditions to retrieve and display the corresponding detailed mosquito monitoring data list and statistical charts.
[0052] In this scheme, the base map loading unit strictly loads a standard geographic base map, ensuring the authority and accuracy of the spatial benchmark. This fundamentally avoids the geographic distortion caused by the use of simplified or self-made sketches in traditional methods. The risk overlay unit classifies risk data according to a unified standard and renders it with corresponding colors, achieving standardization of risk expression and consistency in visualization. This completely changes the visual confusion and interpretation difficulties caused by relying on manual coloring or arbitrary classification in traditional methods. The two units work together to overlay the standard spatial foundation with unified risk semantics, generating a risk situation map that is both accurate and reliable and easy to understand, providing a solid foundation for accurate spatial analysis and decision-making.
[0053] In this solution, the click detection unit allows users to directly interact with the map visualization blocks and generate commands, replacing the indirect method of relying on sidebar menus or drop-down lists for area selection in traditional systems. This makes the operation path shorter and more intuitive. The boundary switching unit can respond instantly to the user's drill-down command, seamlessly switching the map view from the current administrative boundary to the next level administrative boundary and rendering it. This solves the inefficiency problem of traditional systems that often require jumping to different map pages or manually zooming to view lower-level areas. The detail loading unit automatically loads and displays the corresponding lower-level area monitoring data and risk information after boundary switching, realizing real-time linkage updates between the map view and data content. This avoids the disconnect that requires additional operations or waiting after map switching in traditional methods to obtain data.
[0054] The rodent interaction submodule includes a second button unit and a second filtering control unit. The second button unit serves as the entry point for rodent data interaction, responding to user clicks, controlling the switching of the platform's main interface, and launching the rodent-specific interactive interface. The second filtering control unit is integrated into the rodent-specific interactive interface, providing a multi-level filtering panel for rodent monitoring data. Users can sequentially select administrative regions and monitoring indicators (such as capture rate or path index), generate query commands based on filtering conditions, and retrieve and display the corresponding detailed rodent monitoring data list and statistical charts.
[0055] The tick interaction submodule includes a third button unit and a third filtering control unit. The third button unit serves as the entry point for tick data interaction, responding to user clicks, controlling the switching of the platform's main interface, and launching the tick-specific interactive interface. The third filtering control unit is integrated into the tick-specific interactive interface, providing a multi-level filtering panel for tick monitoring data. Users can sequentially select administrative regions and monitoring indicators (such as tick density index), generate query commands based on filtering conditions, and retrieve and display the corresponding detailed tick monitoring data list and statistical charts.
[0056] First, monitoring data is collected and aggregated. At each monitoring point, on-site monitoring personnel use the vector detection devices included in the collection module, such as mosquito lamps for mosquito monitoring, rat traps and cages for rodent monitoring, and cloth flags for tick collection, to complete the initial data collection work, including the capture, counting, and identification of vector organisms. Then, the data is reported through the path defined by the hierarchical upload submodule. For monitoring points at the district and county level and below, the data is first uploaded to the edge data aggregation server deployed in the area for temporary storage and preliminary aggregation. For city-level monitoring points, the data is directly transmitted to the central server.
[0057] Next, the data standardization and indicator calculation stage begins. The vector-borne disease data processing module starts working, with its data standardization preprocessing submodule starting first. The data receiving unit within the data standardization preprocessing submodule receives raw data from the acquisition module. Subsequently, the review and judgment unit verifies whether the mandatory key information in the monitoring data records (such as monitoring point, time, vector type, and quantity) is complete, determines whether the raw values and calculation results are within the reliable threshold and normal range, and checks whether the logical relationships within and between data records (such as whether the time and task cycle match, whether the monitoring method and data items correspond, and whether the geographical location and administrative affiliation are consistent) are correct. Data that passes the review is marked as approved, while data that fails the review is marked as pending manual review and pushed to the administrator terminal. The data maintenance unit provides a channel for modifying and resubmitting data that fails the manual review. Data that passes the review is automatically distributed by the data standardization preprocessing submodule to the corresponding professional calculation submodule, namely the mosquito submodule, rodent submodule, or tick submodule, based on the vector type field in the record, for subsequent special indicator calculation and risk assessment.
[0058] The mosquito index calculation unit in the mosquito submodule calculates the Breteau index and the tent attraction index based on the data. The mosquito risk assessment unit automatically determines the risk level by comparing the calculation results with preset thresholds. The rodent index calculation unit in the rodent submodule calculates the capture rate and the path index. The rodent risk assessment unit assesses the risk based on these. The tick density calculation unit in the tick submodule calculates the density index of free ticks and parasitic ticks in the environment. The tick risk assessment unit generates risk warnings by combining epidemiological thresholds. At the same time, the regional statistics submodule dynamically counts and displays the number of effective monitoring counties, the total number of active monitoring points, and the progress of data reporting.
[0059] Next, users can view and analyze the map globally through the map visualization module. The base map loading unit in the geographic rendering submodule loads the standard map base map, while the risk overlay unit reads the risk assessment results generated by the data processing module and renders different risk levels (low risk, transmission risk, cluster epidemic risk, local outbreak risk) into corresponding colors (green, yellow, orange, red) according to a unified standard, thereby generating a city-wide vector risk block color-coded map on the electronic map.
[0060] Users can perform drill-down analysis through the hierarchical display submodule. Clicking the detection unit captures user interactions with map blocks. The boundary switching unit then switches the administrative boundary level. The detail loading unit simultaneously loads and updates the monitoring data and risk chromatograms of lower-level areas (such as drilling down from districts and counties to townships and villages). Users can also perform historical backtracking through the timeline control submodule. Its calendar selection unit allows users to select specific dates, and the data backtracking unit extracts the corresponding historical data from the time-series data storage unit, driving the geographic rendering submodule to regenerate historical risk thematic maps.
[0061] The real-time data card submodule's data carousel unit displays the latest monitoring summary of each district and county in the form of cards in the map sidebar. The details viewing unit responds to card click events, automatically locates the map, and displays the details. The risk pie chart submodule's data integration unit counts the proportion of districts and counties at each risk level within the current map viewport, and the graphics rendering unit renders it as a pie chart. The trend analysis submodule's trend calculation unit extracts historical data to generate a time series of the number of districts and counties at each risk level, and the line rendering unit renders it as a dynamically segmented and colored trend line chart, which works in conjunction with the pie chart.
[0062] Finally, users can perform specific data queries through the interactive modules. On the platform's main interface, users can switch to the mosquito-specific interactive interface by clicking the first button unit of the mosquito-related interactive sub-module. In this interface, the first filtering control unit provides a filtering panel, allowing users to select administrative regions and monitoring indicators, thereby retrieving and displaying the corresponding detailed data list and charts for mosquitoes. Similarly, users can initiate and perform specific filtering and queries for rodent data through the second button unit and the second filtering control unit of the rodent-related interactive sub-module, and initiate and perform specific filtering and queries for tick data through the third button unit and the third filtering control unit of the tick-related interactive sub-module.
[0063] Example 2
[0064] As attached Figure 6 As shown, this embodiment, based on Embodiment 1, further introduces a vector-borne disease transmission risk prediction method based on multivariate time-series data. This method uniformly cleans, aligns, and standardizes multi-source heterogeneous time-series data from environmental meteorology, vector surveillance, and population epidemiology, and constructs multi-scale features and weakly supervised continuous transmission risk labels, enabling stable training of the risk prediction model even in the absence of direct "true risk values." Furthermore, it introduces a conditional modulation mechanism for task stage identifiers and spatial identifiers, achieving adaptive modeling of the same model for the two stages of "density prediction / transmission risk prediction" and differences in different regions under shared parameters. Simultaneously, it employs two-stage cascaded prediction and mixed sequence correction to reduce distribution bias during inference. At the output end, it combines density and risk dual thresholds to identify the peak start point and duration window, and determines the transmission outbreak period through spatiotemporal coupling. Finally, it generates a four-level warning system (green / yellow / orange / red) and visualizes it using color blocks on an administrative division map and a semi-transparent highlighted window overlay on a trend chart. This allows for earlier and more robust identification of "when, where, and at what level" the future transmission risk will be, supporting disease control departments in proactively allocating resources and deploying targeted prevention and control measures.
[0065] A method for predicting vector density and infection risk based on multivariate time-series data includes the following steps:
[0066] S1: Collect monitoring point identifiers, administrative division codes, and environmental meteorological data. Simultaneously collect vector-borne disease monitoring indicators (Bretu index, tent-trapping index, capture rate, path index, environmental free tick density index, parasitic tick density index) and population epidemiological data (pathogen positivity rate, number of cases in the population, number of fever visits). During the collection phase, record spatial identifiers for each data point and obtain spatial condition vectors through coding mapping to support multi-regional differentiated predictions. Time-align the time series of each indicator to form a unified multivariate time series. Use linear interpolation to fill in missing values and construct a standardized historical input tensor.
[0067] S2: Multi-scale Feature Construction and Weakly Supervised Risk Label Generation: This step first constructs cumulative effect features based on time decay operators for average temperature, relative humidity, and rainfall data to characterize the lagged impact of environmental factors on vector breeding. Simultaneously, based on expert rules, a continuous infectious risk index is synthesized and normalized by integrating vector density, biological transmission, and population response components, serving as the weakly supervised ground truth labels for the second-stage model training.
[0068] S3: Constructing a multi-scale encoder-decoder model that introduces stage identifiers
[0069] A multi-scale encoder-decoder model is constructed, which includes a multi-scale embedding layer, a temporal encoder, and a channel encoder. Multi-resolution convolutional slices and gating networks are used to capture temporal changes at different granularities. At the same time, task-aware stage identifier vectors and spatial identifier vectors are introduced. These two vectors are used as conditional inputs. The feature scaling and translation of each layer are adjusted by normalization through conditionalization layers, so that a single model can adaptively adapt to different prediction tasks (density / risk) and different spatial regions under shared parameters.
[0070] S4: Perform a two-stage cascaded prediction and hybrid sequence correction, employing a cascaded inference strategy. In the first stage, a density stage identifier drives the model to output a predicted sequence of future vector density. In the second stage, this predicted sequence is concatenated with the original input, and a risk stage identifier drives the model to output an infectious risk index. To address distribution bias caused by reliance on predicted values during inference, a hybrid sequence strategy is introduced during the training phase. Real observations or model predictions are randomly sampled according to a Bernoulli distribution probability as input for the second stage, and the two stages are simultaneously optimized using a joint loss function.
[0071] S5: Peak and Window Identification Based on Dual Thresholds: Set dual thresholds for density warning and infection risk, and transform the continuous prediction sequence into a binary state sequence to identify the starting point of the future peak and the peak window that continuously exceeds the threshold. Further, through spatiotemporal coupling analysis, detect whether the density peak and the risk peak overlap or are closely connected, thereby determining the high-confidence transmission outbreak period and distinguishing between simple biological growth and actual transmission risk.
[0072] S6: Tiered Early Warning Generation and Visualized Decision Support
[0073] Based on the identified peak time, duration window, and peak intensity, and combined with preset grading rules (green / yellow / orange / red), early warning signals are generated. In the front-end display module, the real-time risk level of each region is presented using map color blocks, and the predicted curves of density and risk are overlaid on the trend analysis chart. Semi-transparent highlighted color blocks are used to mark the predicted peak window and key time nodes, intuitively displaying the future evolution trend of transmission risk and assisting disease control departments in making advance resource allocation and prevention and control deployments.
[0074] Furthermore, S1 specifically includes the following sub-steps:
[0075] S11: Data Acquisition. Utilizing the data from the existing acquisition module's indicator calculation unit and external interfaces on the platform, three types of basic time-series data are synchronously collected at a preset time granularity, with the monitoring area as the unit:
[0076] The first category is environmental meteorological data: average temperature is denoted as... Relative humidity is denoted as Rainfall is recorded as ;
[0077] The second category is vector-borne disease monitoring indicators: the Breteau index is denoted as... The account-trapping index is denoted as The capture rate is denoted as The path index is denoted as The environmental free tick density index is denoted as The parasitic tick density index is denoted as ;
[0078] The third category is population epidemiological data: the positive rate of vector-borne pathogens is recorded as... The number of confirmed cases in the population is recorded as The number of visits to fever clinics is recorded as .
[0079] in Indicates a discrete sampling time marker.
[0080] During the data collection phase, spatial identification fields, including monitoring point identification and administrative division codes, are recorded for each time series data to enable the differentiation, management, and spatial attribution of data from different regions and subsequent prediction results.
[0081] The monitoring point identifier is used to uniquely identify the monitoring point or monitoring area unit from which the data originates; the administrative division code is used to identify the administrative division to which the monitoring point / monitoring area belongs. By storing the above spatial identifiers and observation features together at each sampling time, subsequent predictions can associate the input data with the output results based on the monitoring point identifier / administrative division code, thereby supporting prediction and display for multiple administrative division units.
[0082] Furthermore, the monitoring point identifiers and administrative division codes are mapped separately, and a spatial condition vector is obtained through learnable embedding and fusion, denoted as... This is used for the dynamic scale gating and generation of conditional modulation parameters for each layer in the subsequent step S3, thereby achieving differentiated adaptation of predictions for different regions while keeping the shared model parameters unchanged.
[0083] S12: Missing value completion and time alignment
[0084] To address the issues of inconsistent sampling frequencies from different data sources or data gaps due to equipment malfunctions, a linear interpolation algorithm is employed for alignment and completion. A unified timeline is then constructed. For any target time If there are missing values in the original data sequence, then the completion value is calculated. The calculation formula is as follows:
[0085]
[0086] in It is the most recent valid observation before the missing time. It is the most recent valid observation after the missing time point; and These are the corresponding valid observation times; This represents the target time step that needs to be completed. After alignment, the original input matrix is formed. .
[0087] S13: Data Standardization
[0088] To eliminate the influence of different physical dimensions on model convergence, the Z-Score standardization method is used to process continuous variables. Standardization... Each element Perform this element-wise. Standardized eigenvalues. The calculation formula is as follows:
[0089]
[0090] in These are the original observations; This is the average value of the indicator over a historical window; This represents the standard deviation of the indicator within the historical window; To prevent the use of tiny positive numbers with a denominator of zero, the final result is the standardized history input tensor. Used for subsequent feature extraction and model input.
[0091] Step S2 specifically includes the following steps:
[0092] S21: Construction of Environmental Cumulative Effect Characteristics
[0093] Considering the lag and cumulative nature of vector-borne disease proliferation, cumulative effect characteristics are constructed for average temperature, relative humidity, and rainfall data. A time-decay cumulative operator is defined, with a length of... Within the historical window Weighted averages were applied to calculate the cumulative environmental index of temperature. Humidity cumulative environmental index With Rainfall Accumulation Environmental Index The formula is as follows:
[0094]
[0095] in, Indicates a discrete sampling time marker. Indicates from the current moment The number of time steps to backtrack. To accumulate window length, This is a time decay factor used to assign higher weight to recent observations; among which, Indicates the time elapsed since the current time Forward The decay weight of each time step observation, as... Increase the weight by decreasing exponentially, thus giving higher weight to more recent observations. , , Representing time respectively The average temperature, relative humidity, and rainfall were observed; based on this... , , Average temperature, relative humidity, and rainfall at the window, respectively. Internal Press The environmental accumulation index, obtained by attenuation-weighted accumulation, is used to characterize the lag and accumulation of environmental factors.
[0096] S22: Calculation of Vector Density and Transmission Risk Components
[0097] Since the level of infection risk in the real environment is difficult to quantify directly, this step uses a weakly supervised method based on expert rules to construct an infection risk index as the training ground value for the second-stage model.
[0098] First, define the three core components of infection risk:
[0099] Vector density risk component Based on historical quantile determination, if the current vector density index set If the maximum historical percentile level corresponding to each indicator is not lower than 75%, then... If the risk level is between 50% and 75%, it is recorded as a high-risk value; if it is between 50% and 75%, it is recorded as a medium-risk value; otherwise, it is recorded as a low-risk value.
[0100] Biological propagation components The main basis is the pathogen positivity rate. If the pathogen positivity rate is found to be greater than zero, the component is set as a high-risk value. Otherwise, it is normalized and mapped according to the historical average.
[0101] Crowd response components Based on abnormal fluctuations in the number of cases and visits, calculate the current time. number of cases The deviation from the historical average for the same period is calculated using the following formula:
[0102]
[0103] in This is the historical baseline number of cases for the same period; As a linear rectified function, negative values in the deviation are truncated to 0. A response is only generated when the number of cases / visits exceeds the historical baseline for the same period, retaining only positive abnormal increase signals and avoiding misinterpretation of decreases or normal fluctuations. This applies to the number of visits to fever clinics. Calculate the deviation in the same way and incorporate it. (Take the maximum value of the two).
[0104] S23: Synthesis and Normalization of Infection Risk Index
[0105] Based on the above three components, a comprehensive infection risk scoring function is constructed. The final weakly supervised label sequence is generated through normalization. :
[0106] The formula for calculating the overall score is as follows:
[0107]
[0108] in These are the weighting coefficients for each risk component, preset by disease control experts based on the characteristics of the main disease vectors in the area.
[0109] Final Risk Index Label The calculation formula is as follows:
[0110]
[0111] in and These are the maximum and minimum risk scores calculated within the historical period, respectively.
[0112] This step generates a continuous infectious risk index sequence with values between 0 and 1, which will serve as the target dependent variable in the subsequent second-stage prediction model. The supervised true value is used to address the lack of direct observations for the infection risk index.
[0113] Step S3 specifically includes the following steps:
[0114] S31: Multi-resolution slice projection and feature embedding
[0115] To capture the changing patterns of vector and environmental data at different time granularities, a multi-scale embedding layer was constructed. This is because, on an hourly / dayly scale, mosquito density may fluctuate intraday with environmental factors such as temperature and humidity, thus affecting the instantaneous risk of virus transmission; on a weekly / monthly scale, factors such as rainfall and breeding cycles drive the regular rise and fall of vector populations, corresponding to the phased changes in the number of infectious disease cases.
[0116] Let the standardized history input tensor obtained in step S1 be denoted as... ,set up At different slice scales, spatial identifier vectors are constructed based on the monitoring point identifiers and administrative division codes recorded during the data collection phase. It is used to characterize the spatial differences between different monitoring points / administrative division units, and serves as a conditional input to participate in subsequent dynamic scale gating and conditional modulation, thereby enabling a single model to achieve multi-region prediction under shared parameters.
[0117] use A group of one-dimensional convolutional filters with different kernel sizes and strides, shared across the channel dimension, are used to extract multi-granularity features from the normalized historical input tensor. Each scale generates a slice feature sequence. and its scale-adaptive embedding after dynamic gating fusion The calculation formula is as follows:
[0118]
[0119] in, Indicates the first Each scale of convolutional mapping operation has a convolution kernel size corresponding to a different time observation window; It is a linear projection matrix used to uniformly map slice features of different scales to the latent feature dimension; This is a gated network used to output the position of each scale. Rating ; The scale weights are normalized by Softmax, enabling soft selection that dynamically favors short-term or long-term scales depending on the sample / season / task stage. The current task stage identifier vector (density stage takes...) Risk phase ), used to make gating weights sensitive to task type. It is a spatial identifier vector used to make the gating weights sensitive to differences in different monitoring points / administrative division units.
[0120] S32: Task Awareness Phase Identifier Fusion
[0121] To enable a single model to adapt to two different task stages—vector density prediction and infection risk prediction—a learnable task-aware stage identifier vector is introduced. A set of stage identifiers is defined, including the first-stage identifier vector representing density prediction. and the second-stage identification vector representing risk prediction Compared to directly... Unlike the input embedding, this implementation uses conditional modulation to ensure that task information is integrated into each layer of feature transformation.
[0122] First, the slices obtained from multi-scale gated fusion are embedded... Adding the positional encoding to form the input embedding; then at each layer, it is processed by the task stage identifier. Spatial signage jointly generated modulation parameters The normalized output is scaled and translated using the following formula:
[0123]
[0124] in, The input embedding sequence after adding location information; It is a sinusoidal position encoding vector used to preserve the temporal position information of the sequence; These are the weighting coefficients for positional encoding; For the first The modulation parameter generation network (a two-layer MLP) outputs parameters consistent with the feature dimensions. and ; For LayerNorm operations; This represents multiplication by dimension. Indicates the first Layer conditionalization and layer normalization operations, Indicates entering the first The input features of the layer normalization module. Through the above conditional modulation, the model can adaptively adjust the feature distribution and focus under different task stages and different spatial units, without having to hard-add extra vectors at the input.
[0125] S33: Multi-scale temporal dependency capture
[0126] The fused features are input into a temporal encoder, and dependencies across time steps are captured through a multi-head self-attention mechanism. To enable task-aware attention computation, conditional modulation can be applied to the input features before computing the query matrix, key matrix, and value matrix. As input to the linear mapping, where, The input fused feature sequence is the time encoder; In the first Layer based on task phase identifier Spatial signage Generate modulation parameters, for The feature representation after conditional layer normalization modulation is used for subsequent linear mapping construction. This enables temporal attention to possess task / spatial awareness capabilities.
[0127] Constructing the query matrix Key matrix Sum matrix Calculate the temporal attention output The formula is as follows:
[0128]
[0129] in, The dimension of the attention head; Indicates matrix transpose; This is a normalized exponential function. This step outputs the hidden state containing global temporal features. .
[0130] S34: Channel Feature Dimensionality Reduction and Correlation Modeling
[0131] To address the issue of numerous and redundant input variables, a dimensionality reduction attention mechanism is introduced into the channel encoder to extract implicit relationships between variables. Before calculating the channel attention, spatial compression is performed on the key and value matrices, and a compression mapping function is defined. Calculate channel attention output The formula is as follows:
[0132] , ,
[0133] in, By using a one-dimensional convolutional layer with a stride greater than 1, the features are downsampled in the time dimension, thereby reducing computational complexity while preserving key information. This is the uncompressed channel query matrix; and These are the compressed key matrix and value matrix, respectively; For channel attention dimension. To enhance task adaptability, , , It can be generated from conditionally modulated features, thus enabling variable correlation modeling to have a differentiated attention pattern for different stages.
[0134] S35: Segmented Recursive Decoding Prediction
[0135] A multi-step decoder is used to generate prediction results in segments to mitigate error accumulation in long-sequence predictions. The future time window to be predicted is divided into segments. The sub-segment, for the first Predicted output of each segment The calculation formula is as follows:
[0136]
[0137] in, It is a fully connected output layer; This is the prediction result for the previous segment; the generation of the complete future sequence is achieved through recursive calls. This recursive decoding process can be corrected using a hybrid sequence strategy during the training phase, and generated using an autoregressive approach during the inference phase. This indicates that a concatenation operation is performed along the feature dimension.
[0138] Furthermore, step S4 specifically includes the following steps:
[0139] S41: Phase One Vector Density Prediction Inference Initiated
[0140] The first stage of the reasoning process involves constructing the standardized historical input tensor in step S1. The input is to a multi-scale codec model, and simultaneously, a spatial identifier vector generated from monitoring point identifiers and administrative division codes is used. As a conditional input; set the stage identifier vector as Both participate in the dynamic scale gating calculation and the generation of conditional modulation parameters at each layer in step S3, thereby instructing the model to perform vector density prediction tasks on the corresponding spatial units. The model outputs a vector density prediction sequence for a future period. The formula is expressed as follows:
[0141]
[0142] in, This represents the constructed prediction model; For models to share a set of parameters; It includes multi-step predicted values such as Breteau index, capture rate, and tick density index.
[0143] S42: Feature Concatenation and Input Reconstruction
[0144] The vector density prediction sequence generated in the first stage is regarded as a prediction prior feature. To facilitate a unified training and inference process, a density sequence is defined for the input of the second stage. : Take during the reasoning stage During the training phase (See the definition of the mixed sequence in equation (16),) With the original standardized history input tensor Alignment and concatenation along the time dimension are performed to construct the expanded second-stage input tensor. The splicing operation is defined as follows:
[0145]
[0146] in, This indicates a concatenation operation along the feature channel dimensions; This means shifting the density sequence backward by time step, so that it can be used as a known input for the next stage.
[0147] S43: Phase Two Infection Risk Prediction Initiation
[0148] The second stage of the inference process will expand the input tensor. When input into the same multi-scale codec model again, the spatial identifier vector... Keep it unchanged to represent the spatial unit to which the current prediction belongs; switch the stage identifier vector to Both serve as conditional inputs in step S3, participating in the generation of dynamic scale gating and conditional modulation parameters at each layer. This instructs the model to extrapolate future infection risks based on historical environment and predicted biological density, and the model outputs a predicted infection risk index sequence. The formula is as follows:
[0149]
[0150] The sequence This refers to the future evolution trend of the infection risk index defined in step S2.
[0151] S44: Mixed Sequence Training Strategy and Joint Loss Optimization
[0152] In the model training phase, to address the input distribution bias caused by relying on the first-stage predictions during the second-stage inference, a mixed sequence correction strategy is introduced, defining a mixed input sequence. The calculation formula is as follows:
[0153]
[0154] in This represents the actual historical observation value of vector density; These are the predicted values from the first stage of the model; It is a random variable that follows a Bernoulli distribution. express Obtain the parameter as Bernoulli distribution, The probability is a mixture, which gradually decreases with increasing training epochs, forcing the model to gradually adapt to using predicted values for inference. Here, δ can be sampled progressively over time or uniformly at the sequence level. During the training phase, it is taken according to equation (14). Build To drive
[0155] The second stage of learning; the reasoning stage is taken according to formula (14). Complete cascaded prediction.
[0156] Constructing a joint loss function Simultaneously optimize the tasks in both phases:
[0157]
[0158] in The weak supervision risk label generated in step S2; This is the task balancing coefficient, used to adjust the weight of the risk prediction task in the total loss. Through joint backpropagation, the shared parameters, dynamic scale gating network parameters, conditional modulation network parameters at each layer, and the two-stage identifier vectors are updated. , And update the spatial identifier vector generated from the monitoring point identifier and administrative division code. Embedding / mapping parameters.
[0159] Step S5 specifically includes the following sub-steps:
[0160] S51: Threshold Setting and State Binarization
[0161] To transform continuous predicted values into discrete risk states, a dual-judgment threshold is first established. The density warning threshold is defined. This threshold is set based on national standards for vector-borne disease density control; defining the transmission risk threshold. This threshold is set based on historical epidemiological survey data.
[0162] The future multi-step vector density prediction sequence output from step S4 and the predicted sequence of the infection risk index Calculate the binarized state sequence respectively and For any time within the prediction time window The state determination formula is as follows:
[0163]
[0164] in and Representing time respectively Density exceeding the standard status indicator and risk exceeding the standard status indicator; and These are the corresponding model prediction values.
[0165] S52: Peak Start Point and Continuity Window Extraction
[0166] Based on a binary state sequence, the time boundaries of risk events are identified. A "peak start point" is defined. The moment when the state changes from 0 to 1, i.e., satisfying The timing; defining the "peak window". This refers to the time interval during which the state continuously exceeds the threshold. For vector density and transmission risk, density peak window sets are extracted separately. With risk peak window set :
[0167]
[0168]
[0169] This step can accurately pinpoint specific time periods from continuously fluctuating time series to identify potential future abnormal clusters, providing a time coordinate for the precise allocation of prevention and control resources.
[0170] S53: Spatiotemporal Coupling Analysis and Outbreak Phase Determination
[0171] To improve the confidence level of the early warning, a coupled analysis is performed on the two types of peak windows identified. This involves detecting whether the density peak window and the risk peak window overlap or are closely connected on the future prediction timeline. A high-confidence propagation outbreak period is defined. The decision logic is as follows:
[0172]
[0173] in It represents the intersection of time intervals, that is, the period when both density and risk are high; This indicates the peak risk period that occurs within a preset lag time after the density peak ends. If it exists... If so, then that period will be designated as a key period for prevention and control; if only [the virus] exists... And without If so, it is determined to be a simple biological breeding period.
[0174] Step S6 specifically includes the following steps:
[0175] S61: Dynamic hierarchical early warning signal generation
[0176] Based on the peak characteristics identified in step S5, and combined with preset grading rules, four levels of early warning signals—green, yellow, orange, and red—are generated. For multi-region prediction scenarios, the administrative division codes recorded in step S11 are used. (or monitoring point identification) () serves as the spatial index key for each region unit. Generate corresponding graded early warning signals respectively. The graded early warning signal To ensure consistent risk level output with the unified risk level system in Example 1, this is used for subsequent map rendering and linked display. Hierarchical logic function. The definition is as follows:
[0177] (twenty three)
[0178] 1. : indicates in spatial unit On the future prediction timeline, there exists a high-confidence propagation outbreak window determined by the overlap or close connection between the "density peak window" and the "risk peak window".
[0179] 2. : Represents a spatial unit The predicted risk index peak exceeded the set extremely high risk circuit breaker threshold. The corresponding output is Red(I).
[0180] 3. : Represents a spatial unit There is a window of opportunity for a potential outbreak, but the peak risk has not yet exceeded [a certain threshold]. The corresponding output is Orange(II), which is used to distinguish it from Red(I).
[0181] 4. : This indicates that a transmission outbreak has not yet occurred, and only the peak window of vector density was detected, but no peak window of risk was detected. This corresponds to the "simple biological growth period", and the output is Yellow(III).
[0182] 5. otherwise: If none of the above conditions are met, output Green(IV).
[0183] in, The threshold for triggering a circuit breaker is set as follows: Red alert corresponds to extremely high risk, indicating the need for immediate emergency response and enhanced disinfection and isolation; Orange alert corresponds to high risk, indicating the need to focus on potential transmission and strengthen vector control; Yellow alert corresponds to medium risk, indicating the need to strengthen monitoring and implement targeted control measures; Green alert corresponds to low risk and normal status.
[0184] superscript This means that the same rule is calculated and output separately on different spatial units, thereby enabling the generation of multi-regional early warnings. It is an empty set.
[0185] S62: Visualized Layer Rendering and Decision Support
[0186] The generated early warning signals and forecast data are pushed to the map visualization module, which then calls the geographic rendering submodule to perform visualization rendering to aid decision-making. Map layer rendering is based on administrative divisions, using administrative division codes. Each regional unit Warning level It associates with the corresponding geographic boundary elements and calls the geographic rendering submodule to fill the map blocks of each region with the corresponding colors to intuitively present the differences in regional risk distribution; at the same time, it can generate "prediction and early warning thematic layers" to present the changes in early warning levels within the future time window in an overlay display without changing the existing map display, and supports switching between historical and future windows via the timeline.
[0187] Trend and window overlay display: In the trend analysis submodule, you can press... and Select a target region cell and plot the density prediction sequence for that region over multiple steps. Risk Index Prediction Series Simultaneously, translucent highlight blocks were used to mark the identified outbreak windows on the timeline background. and the starting point of the peak This visualization method not only shows the evolution trend of values, but also intuitively informs disease control personnel when, where, and at what level the risk of transmission occurs through highlighted windows. This supports them in formulating targeted resource allocation and prevention and control deployment plans in advance, and can also output early warning prompts and handling suggestions for corresponding areas through interactive modules.
[0188] The above descriptions are merely embodiments of the present invention, and common knowledge regarding specific structures and characteristics in the solutions is not described in detail here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or its practicality.
Claims
1. A method for predicting the risk of vector-borne disease transmission, characterized in that, Includes the following steps: S1: Collect monitoring point identification, administrative division code and environmental meteorological data, simultaneously collect vector biological monitoring indicators and population epidemiological data, and preprocess the collected data; S2: Multi-scale feature construction and weakly supervised risk label generation: First, cumulative effect features based on time decay operators are constructed for average temperature, relative humidity and rainfall data to characterize the lagged impact of environmental factors on vector breeding. At the same time, based on expert rules, vector density component, biological transmission component and population response component are integrated to synthesize a continuous infectious risk index and normalize it as a weakly supervised ground truth label for the second stage of model training. S3: Construct a multi-scale encoder-decoder model with stage identifiers: Construct a multi-scale encoder-decoder model that includes a multi-scale embedding layer, a temporal encoder, and a channel encoder. Use multi-resolution convolutional slices and gating networks to capture temporal changes at different granularities. At the same time, introduce task-aware stage identifier vectors and spatial identifier vectors as conditional inputs. Adjust the feature scaling and translation of each layer through conditional normalization layers so that a single model can adaptively adapt to different prediction tasks and different spatial regions under shared parameters. S4: Perform two-stage cascaded prediction and mixed sequence correction, using a cascaded inference strategy. In the first stage, the density stage identifier drives the model to output a future vector density prediction sequence. In the second stage, the prediction sequence is concatenated with the original input, and the risk stage identifier drives the model to output an infection risk index. S5: Peak and window identification based on dual thresholds: Set dual thresholds for density warning and infection risk, and transform the continuous prediction sequence into a binary state sequence to identify the starting point of the future peak and the peak window that continuously exceeds the threshold. Further, through spatiotemporal coupling analysis, detect whether the density peak and the risk peak overlap or are closely connected, thereby determining the high-confidence transmission outbreak period and distinguishing between simple biological growth and actual transmission risk. S6: Graded Early Warning Generation and Visual Decision Support: Based on the identified peak time, duration window, and peak intensity, early warning signals are generated in combination with preset graded rules. In the front-end display module, the real-time risk level of each area is presented as map color blocks, and the predicted curves of density and risk are overlaid on the trend analysis chart.
2. The method for predicting the risk of vector-borne disease transmission according to claim 1, characterized in that, Step S2 specifically includes the following steps: S21: Construction of environmental cumulative effect characteristics, defining a time-decay cumulative operator, in a length of... Within the historical window Weighted averages were applied to calculate the cumulative environmental index of temperature. Humidity cumulative environmental index With Rainfall Accumulation Environmental Index The formula is as follows: , in, Indicates a discrete sampling time marker. Indicates from the current moment The number of time steps to backtrack. To accumulate window length, This is a time decay factor used to assign higher weight to recent observations; among which, Indicates the time elapsed since the current time Forward The decay weight of each time step observation, as... Increases exponentially decreasing. , , Representing time respectively Average temperature, relative humidity, and rainfall observations; S22: Calculation of vector density and transmission risk components, defining three core components of transmission risk: Vector density risk component Biological propagation component ; Crowd response components : Calculate the current time based on abnormal fluctuations in the number of cases and visits. number of cases The deviation from the historical average for the same period is calculated using the following formula: , in This is the historical baseline number of cases for the same period; As a linear rectified function, negative values in the deviation are truncated to 0. A response is only generated when the number of cases / visits exceeds the historical baseline for the same period, retaining only positive abnormal increase signals and avoiding misinterpretation of decreases or normal fluctuations, for fever clinic visit volume. Calculate the deviation in the same way and incorporate it. ; S23: Synthesis and Normalization of Infection Risk Index: Constructing a Comprehensive Infection Risk Scoring Function Based on the Three Components Mentioned Above. The final weakly supervised label sequence is generated through normalization. : The formula for calculating the overall score is as follows: , in These are the weighting coefficients for each risk component; Final Risk Index Label The calculation formula is as follows: , in and These are the maximum and minimum risk scores calculated within the historical period, respectively.
3. The method for predicting the risk of vector-borne disease transmission according to claim 1, characterized in that, Step S3 specifically includes the following steps: S31: Multi-resolution slice projection and feature embedding. The normalized historical input tensor obtained in step S1 is denoted as... ,set up At different slice scales, spatial identifier vectors are constructed based on the monitoring point identifiers and administrative division codes recorded during the data collection phase. It is used to characterize the spatial differences between different monitoring points / administrative division units, and serves as a conditional input for subsequent dynamic scale gating and conditional modulation; use A group of one-dimensional convolutional filters with different kernel sizes and strides, shared across the channel dimension, are used to extract multi-granularity features from the normalized historical input tensor. Each scale generates a slice feature sequence. and its scale-adaptive embedding after dynamic gating fusion The calculation formula is as follows: , in, Indicates the first Convolutional mapping operations at various scales; This is a linear projection matrix used to uniformly map slice features of different scales to the latent feature dimension. ; This is a gated network used to output the position of each scale. Rating ; These are the scale weights after Softmax normalization; This is the current task stage identifier vector, used to make the gating weights sensitive to the task type; This is a spatial identifier vector used to make the gating weights sensitive to differences between different monitoring points / administrative division units; S32: Task-aware stage identifier fusion, defining a set of stage identifiers, including the first-stage identifier vector representing density prediction. and the second-stage identification vector representing risk prediction First, the slices obtained by multi-scale gated fusion are embedded Adding the positional encoding to form the input embedding; then at each layer, it is processed by the task stage identifier. Spatial signage The jointly generated modulation parameters are scaled and shifted on the normalized output, calculated using the following formula: , in, The input embedding sequence after adding location information; It is a sinusoidal position encoding vector used to preserve the temporal position information of the sequence; These are the weighting coefficients for positional encoding; For the first The modulation parameter generation network of the layer outputs the same as the feature dimension. and ; For LayerNorm operations; This represents dimension-wise multiplication; Indicates the first Layer conditionalization and layer normalization operations; Indicates entering the first Input features of the layer normalization module; S33: Based on multi-scale temporal dependency capture, the fused features are input into the temporal encoder, and the dependencies across time steps are captured through a multi-head self-attention mechanism; a query matrix is constructed. Key matrix Sum matrix Calculate the temporal attention output The formula is as follows: , in, The dimension of the attention head; Indicates matrix transpose; It is a normalized exponential function; S34: Channel Feature Dimensionality Reduction and Correlation Modeling. Addressing the issue of numerous and redundant input variables, a dimensionality reduction attention mechanism is introduced into the channel encoder to extract implicit correlations between variables. Before calculating channel attention, spatial compression is performed on the key and value matrices, defining a compression mapping function. Calculate channel attention output The formula is as follows: , , in, By using a one-dimensional convolutional layer with a stride greater than 1, the features are downsampled in the time dimension, thereby reducing computational complexity while preserving key information. This is the uncompressed channel query matrix; and These are the compressed key matrix and value matrix, respectively; For channel attention dimension; S35: Segmented recursive decoding prediction, using a multi-step decoder to generate prediction results in segments, dividing the future time window to be predicted into segments. The sub-segment, for the first Predicted output of each segment The calculation formula is as follows: , in, It is a fully connected output layer; Based on the prediction result of the previous segment, the generation of the complete future sequence is achieved through recursive calls. This indicates that a concatenation operation is performed along the feature dimension.
4. The method for predicting the risk of vector-borne disease transmission according to claim 1, characterized in that: S41: First-stage vector density prediction inference, the first-stage inference process, which uses the standardized historical input tensor constructed in step S1. The input is to a multi-scale codec model, and simultaneously, a spatial identifier vector generated from monitoring point identifiers and administrative division codes is used. As a conditional input; Set the stage identifier vector to Both participate in the dynamic scale gating calculation and the generation of conditional modulation parameters at each layer in step S3, thereby instructing the model to perform vector density prediction tasks on the corresponding spatial units. The model outputs a vector density prediction sequence for a future period of time. The formula is expressed as follows: , in, This represents the constructed prediction model; For models to share a set of parameters; Multi-step predicted values including Breteau index, capture rate, tick density index, etc. S42: Feature Concatenation and Input Reconstruction. The vector density prediction sequence generated in the first stage is regarded as a prediction prior feature. To facilitate the unification of the training and inference processes, a density sequence is defined for the input of the second stage. : Take during the reasoning stage During the training phase ,Will With the original standardized history input tensor Alignment and concatenation along the time dimension are performed to construct the expanded second-stage input tensor. The splicing operation is defined as follows: , in, This indicates a concatenation operation along the feature channel dimensions; This means shifting the density sequence backward by time step, so that it can be used as a known input to the second stage; S43: Second-stage infection risk prediction reasoning, the second-stage reasoning process, the expanded input tensor When input into the same multi-scale codec model again, the spatial identifier vector... Keep it unchanged to represent the spatial unit to which the current prediction belongs; switch the stage identifier vector to Both serve as conditional inputs in step S3, participating in the generation of dynamic scale gating and conditional modulation parameters at each layer. This instructs the model to extrapolate future infection risks based on historical environment and predicted biological density, and the model outputs a predicted infection risk index sequence. The formula is as follows: , The sequence This refers to the future evolution trend of the infection risk index defined in step S2; S44: Hybrid Sequence Training Strategy and Joint Loss Optimization. In the model training phase, to address the input distribution bias caused by reliance on first-stage predictions during second-stage inference, a hybrid sequence correction strategy is introduced, defining a hybrid input sequence. The calculation formula is as follows: , in This represents the actual historical observation value of vector density; These are the predicted values from the first stage of the model; It is a random variable that follows a Bernoulli distribution. express Obtain the parameter as Bernoulli distribution, The probability is mixed, gradually decreasing with each training epoch, forcing the model to gradually adapt to using predicted values for inference. The training phase takes... Build To drive the second stage of learning; the reasoning stage takes Complete cascaded prediction.
5. A visual early warning platform based on vector-borne organism ecological management, characterized in that: It includes a data acquisition module, a central server, and an interaction module; The data acquisition module is used to collect raw data from monitoring areas at the municipal, district, and county levels. The central server includes a vector-borne disease data processing module, a vector-borne disease transmission risk prediction module, and a map visualization module. The vector-borne disease data processing module is used to process the data collected by the collection module. The vector-borne disease transmission risk prediction module is based on the vector-borne disease transmission risk prediction method according to any one of claims 1-4, and realizes the prediction of vector density and vector-borne disease transmission risk. The map visualization module is used to convert the data from the vector-borne organism data processing module into hierarchical color block illustrations on an electronic map. The interactive module is used to provide a specialized data interaction interface for specific disease vectors.
6. The visualized early warning platform based on vector-borne organism ecological management according to claim 5, characterized in that: The acquisition module includes multiple vector-borne organism detection devices, an edge data aggregation server, and a hierarchical upload submodule; All of the aforementioned vector-borne disease detection devices are used to collect raw monitoring data of vector-borne diseases; The edge data aggregation server is deployed at each district and county-level node to aggregate data uploaded by each vector-borne disease detection device within its jurisdiction. The hierarchical upload submodule is used to execute different upload paths according to the administrative level of the data source. It determines that municipal data is directly uploaded to the central server, while district and county level and below data are first sent to the edge data aggregation server, aggregated by the server, and then synchronously uploaded to the central server.
7. The visualized early warning platform based on vector-borne organism ecological management according to claim 5, characterized in that: The map visualization module includes a geographic rendering submodule, a hierarchical display submodule, and a timeline control submodule. The geographic rendering submodule is used to visualize the risk levels of each district and county in the form of a block-based color-gradient map. The hierarchical display submodule is used to respond to user clicks and implement a step-by-step drill-down display from districts and counties to villages and towns; The time axis control submodule is used to filter and replay historical data by time.
8. The visualized early warning platform based on vector-borne organism ecological management according to claim 7, characterized in that: The geographic rendering submodule includes a base map loading unit and a risk overlay unit; The base map loading unit is used to load a standard geographic base map; The risk overlay unit is used to classify risk data according to a unified standard and render the corresponding colors.
9. The visualization early warning platform based on vector-borne organism ecological management according to claim 7, wherein the hierarchical display submodule includes a click detection unit, a boundary switching unit, and a detail loading unit; The click detection unit is used to capture user interaction with the map visualization area and generate drill-down instructions; The boundary switching unit is used to respond to the drill-down command and control the map to switch from the current administrative boundary to the next level of administrative boundary data for rendering. The detailed loading unit is used to synchronously load and display the monitoring data and risk information of the corresponding lower-level area after the boundary switch.
10. The visualized early warning platform based on vector-borne organism ecological management according to claim 5, characterized in that: The interaction module includes a mosquito interaction submodule, a rodent interaction submodule, and a tick interaction submodule; The mosquito-related interactive submodule is used to launch a mosquito-specific interface and provide data filtering and query functions. The mouse interaction submodule is used to launch a mouse-specific interface and provide data filtering and query functions; The tick-related interaction submodule is used to launch the tick-specific interface and provide data filtering and query functions.