Hygienic rodent damage risk prediction method based on big data

By integrating multi-source data and using spatiotemporal graph convolutional networks, the timeliness and spatial coverage issues of rodent infestation risk assessment were resolved, enabling dynamic rodent infestation risk prediction at the city level and improving the scientific rigor and timeliness of public health management.

CN120875591AActive Publication Date: 2025-10-31FUJIAN INT TRAVEL HEALTH CARE CENT +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511396608.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-10-31
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing rodent monitoring methods lack timeliness and spatial coverage, making it difficult to achieve dynamic prediction and early warning of large-scale rodent risks. Furthermore, the spatiotemporal characteristics of the garbage collection process have not been effectively integrated, resulting in inaccurate rodent risk assessments.

Method used

By acquiring and spatially gridding multi-source data, the system calculates the cleaning delay duration, load anomaly, and efficiency decay coefficient, constructs a spatiotemporal graph structure, and uses graph neural networks to predict rodent infestation risk, generating future spatiotemporal probability distributions.

Benefits of technology

It enables dynamic modeling and accurate prediction of rodent infestation risks, improves the efficiency and applicability of risk assessment, provides continuous and dynamic risk monitoring results, and facilitates the formulation of prevention and control measures and the optimization of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875591A_ABST
    Figure CN120875591A_ABST
Patent Text Reader

Abstract

The invention discloses a big data-based sanitary rodent damage risk prediction method, particularly relates to the field of rodent damage risk prediction, and is used for solving the problem of rodent damage risk prediction due to neglect of correlation analysis of garbage collection and rodent damage in the prior art. The method is realized by introducing multi-source heterogeneous data and carrying out spatial gridding modeling; in the rodent pest period date sequence, calculating a clearance delay time length and a load abnormity index, and generating a clearance efficiency attenuation coefficient in combination with the traffic flow data; constructing a space-time diagram structure, taking the features as dynamic node features, and establishing spatial adjacency and traffic flow topological connection; inputting the space-time diagram and the rodent damage risk identification into a graph neural network, and training a prediction model through a space-time convolution and circulation unit; and finally, under real-time data stream input, node features are dynamically updated, and rodent damage risk probability distribution of a future preset time step is generated, so that prospective prediction of urban rodent damage risks is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rodent pest risk prediction technology, and more specifically, to a method for predicting sanitary rodent pest risks based on big data. Background Technology

[0002] With the acceleration of urbanization and the continuous increase in population density, the amount of urban domestic waste generated is increasing year by year, and the complexity of waste collection and disposal is constantly increasing. The efficiency and quality of sanitation operations have a direct impact on the level of public environmental sanitation. At the same time, if there are delays or a decline in collection efficiency in the waste accumulation and collection process, it is easy to create environmental conditions suitable for rodent survival and reproduction, causing rodent problems to show a phased outbreak trend, posing a serious threat to urban public health and safety and residents' health.

[0003] Current rodent monitoring methods largely rely on manual reporting or the deployment of fixed-point rodent traps. While these methods can reflect rodent density in localized areas, their timeliness and spatial coverage are insufficient, making it difficult to dynamically predict and warn of large-scale rodent infestation risks. Furthermore, the garbage collection process exhibits distinct spatiotemporal characteristics. Collection plans, vehicle trajectories, and traffic flow patterns at different collection points are coupled in time and space, making it difficult to reveal the formation patterns of rodent infestation risks through simple single-dimensional statistical analysis. Therefore, there is an urgent need for a comprehensive analytical method that integrates garbage collection data with historical rodent infestation reports to construct a city-level dynamic risk assessment. This method should accurately reflect the impact of abnormal collection operations and traffic factors on rodent breeding conditions and enable early prediction of future rodent infestation risks through spatiotemporal forecasting, thereby providing scientific support for urban environmental governance and public health prevention and control. Summary of the Invention

[0004] In order to overcome the above-mentioned deficiencies of the prior art, embodiments of the present invention provide a method for predicting sanitary rodent pest risks based on big data to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for predicting sanitary rodent infestation risk based on big data includes the following steps: S1. Obtain regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data, and perform spatial grid mapping on the above multi-source data; S2. Based on the date sequence of the rodent infestation cycle, calculate the garbage collection delay time and load anomaly index, and integrate traffic flow data to generate the garbage collection efficiency attenuation coefficient. S3. Construct a spatiotemporal graph structure for the spatial grid, and use the cleaning delay time, load anomaly index, and efficiency decay coefficient as dynamic node features; S4. Train a rodent risk prediction model by inputting the spatiotemporal graph structure of the spatial grid and the rodent risk label into the graph neural network. S5. Based on the real-time inflow of multi-source data streams, dynamically update the feature vectors of each node in the spatiotemporal graph, use the rodent risk prediction model to infer rodent risk online, and generate the spatiotemporal probability distribution of rodent risk within a preset time step in the future.

[0006] In a preferred embodiment, step S1, acquiring regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data, and performing spatial gridding mapping on the above multi-source data, specifically includes: Acquire historical rodent infestation report data, map the rodent infestation report data to the corresponding spatial grid cells according to geographical location, and establish a historical rodent infestation event database for each spatial grid; Waste collection data for the accessed area, including the planned service time window and collection volume for each collection point; Collect historical trajectory data of waste collection vehicles and map the vehicle trajectories and waste collection point locations into spatial grid cells through map matching.

[0007] In a preferred embodiment, step S2, which involves calculating the collection delay duration and load anomaly indicators based on the collection vehicle trajectory data and garbage collection data corresponding to the date sequence of the rodent infestation cycle, and integrating traffic flow data to generate a collection efficiency attenuation coefficient, specifically includes: Extract the observation time window before the occurrence of rodent infestation events from the historical rodent infestation event database and mark it as a rodent infestation cycle, then convert the rodent infestation cycle into the corresponding date sequence; Extract the planned removal time window for each removal point within the rodent infestation cycle, and calculate the removal delay time for each removal point by combining the removal vehicle trajectory. Obtain historical waste collection data for each waste collection point and establish a waste collection benchmark model to calculate the deviation between the current waste collection volume and the historical benchmark and obtain load anomaly indicators. Obtain traffic flow data for the road segments corresponding to the trajectories of waste collection vehicles, establish a mapping relationship between traffic flow data and the degree of waste collection efficiency attenuation, and output the waste collection efficiency attenuation coefficient.

[0008] In a preferred embodiment, establishing the mapping relationship between traffic flow data and the degree of waste disposal efficiency attenuation, and outputting the waste disposal efficiency attenuation coefficient, specifically includes: Based on the historical cleaning delay duration and load anomaly index data within the date series corresponding to the historical rodent infestation cycle, a regression model is constructed by combining the traffic flow data of the road segments where the corresponding cleaning vehicle trajectories are located. The traffic flow data includes average driving speed of road segments, traffic flow density, and congestion index characteristics. The influence coefficients between traffic flow characteristics, waste removal delay time, and abnormal load indicators are calculated by fitting regression. After normalizing the influence coefficients, the delay influence coefficients and load influence coefficients are weighted and fused according to preset weights to obtain the waste removal efficiency attenuation coefficient.

[0009] In a preferred embodiment, step S3, constructing the spatiotemporal graph structure of the spatial grid, specifically including the collection delay duration, load anomaly indicators, and efficiency decay coefficient as dynamic node features, includes: A spatiotemporal graph structure with spatial grid cells as nodes is constructed. The node feature vectors include the cleaning delay duration, load anomaly index, and cleaning efficiency attenuation coefficient. Based on the spatial adjacency relationship of spatial grid cells, undirected connection edges are established between nodes to connect adjacent spatial grid cells that share common edges; Directed connections between nodes are established based on the road network topology between each waste collection point, with the direction of the edges consistent with the travel direction of the waste collection vehicles.

[0010] In a preferred embodiment, step S4, training the rodent risk prediction model by combining the spatiotemporal graph structure of the spatial grid with the rodent risk identifier input graph neural network, specifically includes: A spatiotemporal graph convolutional network is used as the core computing unit. The spatiotemporal graph convolutional network includes graph convolution operations in the spatial dimension and gated recurrent units in the temporal dimension. Spatial graph convolutional layers use Chebyshev polynomial approximation of spectral graph convolution to capture the spatial dependencies between dynamic nodes in the spatiotemporal graph; The time-gated cyclic unit layer processes the time-series changes of dynamic node features in the spatiotemporal graph and learns the temporal evolution law; The rodent risk identifiers are converted into binary label sequences, and the rodent risk prediction model is trained using a masked binary cross-entropy loss function.

[0011] In a preferred embodiment, the rodent risk label is marked as follows: Extract the ratio of the total number of rats caught to the number of rat traps deployed in each spatial grid cell from the rat infestation report data, and calculate the rat density in each spatial grid cell. Spatial grid cells with rodent density exceeding a set rodent risk threshold are defined as rodent risk cells and marked with a rodent risk label.

[0012] In a preferred embodiment, step S5, which involves dynamically updating the feature vectors of each node in the spatiotemporal graph based on the real-time inflow of multi-source data streams, using a rodent risk prediction model to perform online inference of rodent risk, and generating a spatiotemporal probability distribution of rodent risk within a preset future time step, specifically includes: Establish a real-time data stream processing pipeline to receive current waste collection vehicle trajectory data, waste collection data, and traffic flow data at preset fixed time intervals; Perform feature extraction for each data batch and recalculate the cleaning delay time, load anomaly index and cleaning efficiency decay coefficient for each grid. Update the feature vector values ​​of the corresponding nodes in the spatiotemporal graph, and input the updated spatiotemporal graph into the trained rodent risk prediction model. The rodent infestation risk prediction model outputs the rodent infestation risk probability value for each spatial grid cell within a preset future time step. The discrete risk probability values ​​are converted into a continuous risk distribution surface, generating a smooth risk gradient change map.

[0013] The technical effects and advantages of the big data-based method for predicting sanitary rodent pest risks in this invention are as follows: This approach achieves dynamic modeling and accurate prediction of rodent infestation risk based on the fusion of multi-source heterogeneous data, offering advantages such as high efficiency, strong applicability, and intuitive and reliable prediction results. By incorporating data from garbage collection plans, vehicle trajectories, traffic flow, and rodent infestation reports, and performing spatial gridding and temporal processing, it can meticulously depict the collection efficiency status of different areas at different time periods. This allows potential risk factors such as collection delays, abnormal loads, and efficiency degradation to be expressed quantitatively and used as model input, significantly improving the explanatory power for risk factors. Furthermore, it utilizes spatiotemporal graph convolutional networks to jointly model spatial dependencies and temporal evolution patterns.

[0014] Compared to traditional methods that rely on static statistics or single monitoring data, this invention can continuously update node characteristics and generate the probability distribution of rodent infestation risk for future periods in a real-time data stream environment. This provides urban public health management departments with continuous and dynamic risk monitoring results, facilitating the accurate formulation of prevention and control measures and the optimization of waste disposal resource scheduling, thereby improving the scientific nature and timeliness of urban vector-borne disease control. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of a big data-based method for predicting rodent pest risks according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1 Figure 1This invention presents a method for predicting sanitary rodent pest risks based on big data, which includes the following steps: S1. Obtain regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data, and perform spatial grid mapping on the above multi-source data; S2. Based on the date sequence of the rodent infestation cycle, calculate the garbage collection delay time and load anomaly index, and integrate traffic flow data to generate the garbage collection efficiency attenuation coefficient. S3. Construct a spatiotemporal graph structure for the spatial grid, and use the cleaning delay time, load anomaly index, and efficiency decay coefficient as dynamic node features; S4. Train a rodent risk prediction model by inputting the spatiotemporal graph structure of the spatial grid and the rodent risk label into the graph neural network. S5. Based on the real-time inflow of multi-source data streams, dynamically update the feature vectors of each node in the spatiotemporal graph, use the rodent risk prediction model to infer rodent risk online, and generate the spatiotemporal probability distribution of rodent risk within a preset time step in the future.

[0018] In step S1, regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data are acquired, and the above multi-source data are spatially gridded and mapped.

[0019] Historical rodent infestation report data is acquired, organized, and located. This data typically originates from monitoring records of urban sanitation agencies, containing specific locations, placement of rodent traps, number of rodents caught in the area, reporting time, and event descriptions. To ensure spatial usability, the geographic location information in the reports is uniformly converted to a standard coordinate system. Using latitude and longitude coordinates, the processed rodent infestation report data is mapped to pre-divided regional spatial grid cells. Each grid cell represents a fixed geographical area within the city, for example, a 100m x 100m or 500m x 500m area, with the specific dimensions determined based on the urban geographical environment and data distribution. In this way, all rodent infestation report data is categorized into their respective spatial grids, and an independent historical rodent infestation event database is established for each grid. This database includes the frequency, temporal distribution, and density characteristics of rodent infestation events within that grid, thereby achieving spatialized accumulation of historical rodent infestation data.

[0020] The system accesses regional waste collection data, including planned service time windows and collection volume information for each collection point. Collection points are centralized collection locations for urban waste, and each point is planned with corresponding collection times and estimated collection volumes. After collecting and organizing this data, the planned service times for each collection point are digitally stored to ensure that the time period information and waste processing volume for each collection point are clearly defined within the spatial grid. Simultaneously, historical actual operating trajectory data of collection vehicles is collected. This trajectory data is typically provided by vehicle positioning equipment and includes the vehicle's temporal location points during operation. Using map matching methods, the trajectory points are compared with the urban road network to ensure that the trajectory points accurately fall on road segments. Next, the vehicle trajectories are associated with the collection point locations. When a vehicle's trajectory stops at a collection point and generates a work record, it is marked as an actual collection action.

[0021] In step S2, based on the trajectory data of the garbage collection vehicles and the garbage collection data corresponding to the date sequence of the rodent infestation cycle, the collection delay time and load anomaly index are calculated, and the traffic flow data is integrated to generate the collection efficiency attenuation coefficient.

[0022] The time window preceding a rodent infestation event is extracted from the historical rodent infestation event database and marked as a rodent infestation cycle. Specifically, when a rodent infestation report is received within a spatial grid cell, the report date is extended forward by a certain number of days to form the observation window preceding the infestation. The default observation window length is 30 days, close to the rodent breeding cycle, thus forming a complete rodent infestation cycle. This cycle covers potential garbage collection and environmental changes that may precede a rodent infestation outbreak. Subsequently, the rodent infestation cycle is divided into a continuous date sequence, ensuring that each date is clearly identified. This date sequence serves as the foundational time frame for subsequent extraction of garbage collection and traffic features.

[0023] Within a defined rodent infestation cycle, the planned collection time windows for each collection point are extracted, and combined with historical trajectory data of collection vehicles, the collection delay time for each collection point is calculated. The planned collection time windows are pre-set by the sanitation department, typically including specific start and end times, such as 8:00 to 10:00 daily. The delay time is calculated by comparing the arrival time of collection vehicles at each collection point with the planned time. When a vehicle fails to complete garbage collection within the planned window, the difference is recorded as a positive delay time; for example, if collection is completed 30 minutes later than the planned time, the delay time for that collection point on the corresponding date is 30 minutes. For cases where collection is completed before the planned time, the difference is recorded as a negative delay time to reflect the early collection status.

[0024] Historical waste collection data for each collection point was acquired, and a baseline model for waste collection volume was established. Historical data was stored daily or weekly, reflecting the waste output and collection levels of each collection point under normal conditions. The baseline model was established by calculating the long-term statistical mean and fluctuation range. Using the daily collection volume of the most recent three months as a sample, the average value was calculated, and the standard fluctuation range was recorded as the baseline capacity for that collection point. During rodent infestation cycles, the actual collection volume was compared with the baseline capacity, and the degree of deviation was calculated. When the actual collection volume significantly exceeded the baseline capacity, it was marked as overload; when it significantly fell below the baseline capacity, it was marked as underload. The magnitude of the deviation served as an indicator of load anomalies; for example, a deviation exceeding 20% ​​was defined as abnormal.

[0025] Based on all date sequences corresponding to historical rodent infestation cycles, data on all waste collection delays and abnormal load indicators within that timeframe are extracted and matched with the operational trajectories of waste collection vehicles. The operational trajectories include the road segment numbers traversed by each vehicle; traffic flow data for that road segment within the corresponding time period is obtained by combining these road segment numbers. Traffic flow data is retrieved from the urban traffic monitoring database and primarily includes three dimensions: average speed of the road segment, traffic flow density, and a comprehensively calculated congestion index. The congestion index uses the designed free-flow speed of the road as a reference, statistically calculating the actual average speed of the road segment within a specified time slice. The ratio of the actual speed to the free-flow speed is calculated, and the difference between this ratio and 1 is used as the congestion index. To ensure data consistency, all traffic flow features are statistically analyzed on a time-slice basis, with a complete set of speed, density, and congestion index data generated every five minutes by default, and aligned with the records of waste collection vehicles passing through that road segment within that time slice.

[0026] After data alignment, a regression model was constructed with traffic flow characteristics as independent variables and waste collection delay duration and load anomaly indicators as dependent variables. Specifically, the delay duration and load anomaly values ​​of each waste collection point within a rodent infestation cycle were linked to the traffic flow characteristics of the road segment where the corresponding waste collection vehicle's trajectory was located, forming a sample dataset. Each data record in the sample dataset contains a set of traffic flow characteristics and corresponding waste collection delay and load anomaly values. Subsequently, a linear regression model was established to calculate the influence coefficients of traffic flow characteristics on delay and load. During the regression process, to avoid bias caused by inconsistent feature dimensions, speed, density, and congestion index were normalized to ensure that each feature falls within the same numerical range before regression fitting. The coefficients obtained through the model solution quantify the degree of influence of different traffic flow characteristics on waste collection delay and load anomaly. For example, the coefficient of average driving speed on delay duration may be negative, indicating that increasing speed reduces delay.

[0027] After obtaining the impact coefficients, they are normalized to ensure that the delay impact coefficient and the load impact coefficient can be fused at the same scale. The normalization method maps the coefficient values ​​to the range of 0 to 1, ensuring comparability between different indicators. Subsequently, the delay impact coefficient and the load impact coefficient are weighted and fused according to preset weights to obtain the final waste collection efficiency attenuation coefficient. For example, if delay has a greater impact on overall efficiency, the delay weight can be set to 0.6 and the load weight to 0.4. The specific weight values ​​are flexibly set based on the road conditions and waste collection volume of the region, and a comprehensive coefficient is obtained through weighted calculation. This waste collection efficiency attenuation coefficient reflects the degree of decline in waste collection capacity under specific traffic conditions and is numerically bound to the corresponding spatial grid nodes, providing a basis for the node feature input of the subsequent spatiotemporal graph model.

[0028] In S3, a spatiotemporal graph structure of a spatial grid is constructed, and the cleaning delay time, load anomaly index, and efficiency decay coefficient are used as dynamic node features.

[0029] Data such as collection delay duration, load anomaly indicators, and collection efficiency decay coefficient within each grid cell are compiled and stored as feature vectors for that node. These indicators are continuously recorded in time slice order, forming a dynamic feature sequence of the node in the time dimension. Undirected edges are established between nodes at the spatial level. Spatial adjacency is defined by the geometric position of the spatial grid. When two grid cells share a common edge or a common vertex geographically, they are considered adjacent, and an undirected edge is established between them. Such undirected edges can reflect the spatial proximity and potential interactions between regions. For example, when two adjacent grids are directly adjacent in space, their collection efficiency and waste accumulation status are likely to be correlated, and therefore must be represented by undirected edges. When constructing the graph structure, each grid cell's four boundaries are checked one by one to see if they contact the boundaries of surrounding cells. If they do, a connection is established between the two cells. This process ensures that the spatiotemporal graph maintains structural integrity in the spatial dimension, allowing each node to reflect its spatial association with surrounding nodes through undirected edges.

[0030] In the network topology dimension, directed connections are established by combining road network information between collection points. The road network topology comes from urban road data and includes the directional attributes and topological relationships of roads. By matching the trajectories of collection vehicles with the road network, the actual travel paths of collection vehicles between collection points can be obtained. When a collection vehicle departs from collection point A and travels to collection point B, a directed edge is established between the corresponding two spatial grid cells, with the direction from A to B. This clearly represents the dynamic flow relationship of the waste collection process. In practice, the trajectory sequence of collection vehicles is traversed, the pairs of collection points traversed by the trajectory are identified, and they are mapped to the corresponding spatial grid nodes, generating directed edges one by one. In this way, each collection route is represented as a directed connection in the spatiotemporal graph, enabling the model to capture the temporal and fluid nature of the collection path, thus ensuring that the spatiotemporal graph structure not only reflects static spatial adjacency relationships but also truly reflects the dynamic process of collection activities.

[0031] In step S4, the spatiotemporal graph structure of the spatial grid and the rodent risk identification input graph neural network are used to train a rodent risk prediction model.

[0032] The model design employs a spatiotemporal graph convolutional network as the core computational unit, simultaneously characterizing the complex dependencies in both spatial and temporal dimensions during training. In the spatial dimension, graph convolution operations are used to process the graph structure with spatial grid cells as nodes. This graph structure, already constructed through the preceding S3 steps, includes dynamic features such as waste collection delay duration, load anomaly indicators, and waste collection efficiency decay coefficients. To achieve effective convolution operations on the graph structure, the spatial graph convolutional layer uses Chebyshev polynomial approximation of spectral graph convolution. This method approximates the Laplacian operator of the graph using multi-order polynomials, enabling effective propagation and aggregation of node features among neighboring nodes. Specifically, in each convolution operation, the features of the target node not only depend on its current value but also incorporate the feature information of its neighboring nodes in the spatial structure, thereby capturing the potential dependencies between neighboring grid cells sharing a boundary or waste collection path nodes connected by the road network. For example, when the waste collection delay indicator of a certain grid cell increases abnormally, neighboring grid cells may also be indirectly affected. This mutual relationship is propagated and learned through the graph convolutional layer, thus achieving the expression of spatial dependencies.

[0033] Simultaneously, in the time dimension, gated recurrent units are employed to handle the evolution of dynamic node features over time. Each node has a corresponding feature vector within different time slices, containing the aforementioned data on collection delays, load anomalies, and efficiency degradation. The gated recurrent units model long-term dependencies and short-term changes in the time series through update and reset gate operations, enabling the model to automatically extract temporal patterns related to the evolution of rodent infestation risk during multi-day or multi-period training. For example, when a region experiences collection delays accompanied by load anomalies for several consecutive days, the model will capture this persistent trend through the time-dimensional recurrent units and reflect it in subsequent risk predictions. By combining spatial convolution and temporal recursion, the spatiotemporal graph convolutional network can simultaneously model spatial adjacency relationships and temporal evolution processes, thus forming a prediction framework that considers both spatiotemporal features.

[0034] For each established regional spatial grid unit, rodent infestation reporting data is extracted. The reporting data includes information on the number of rodents caught by traps deployed in each grid, along with the corresponding number of traps deployed. To ensure data integrity, a statistical analysis of the number of rodents caught and the number of traps is performed for each grid unit. Specifically, within the same time window, the total number of rodents caught by all traps is summed, and the number of traps within that time window is recorded. Then, the total number of rodents caught is divided by the number of traps to obtain the rodent density value for that grid unit within that time window. For example, if a grid has 10 traps deployed and a total of 30 rodents are caught, the rodent density value for that grid during that time period is 3. This rodent density value accurately reflects the intensity of rodent activity per unit deployment quantity, thus avoiding bias caused by relying solely on absolute catch quantities. The results are compared with a pre-set rodent infestation risk threshold. This threshold is determined based on statistical patterns and the frequency of resident complaints in the area; for example, a rodent density of 2 represents a high-risk state for rodent infestation. If the rat density value of a certain grid exceeds the threshold, the grid cell is defined as a rat infestation risk cell, and a rat infestation risk label is added to the corresponding spatial grid data structure.

[0035] Rodent risk labels are used as supervisory signals input into the model. To simplify computation during training, the rodent risk labels are converted into a binary label sequence, where 1 represents a risk unit and 0 represents a non-risk unit. This sequence corresponds to the time series of node features, ensuring that the input data for each time slice has clear supervisory labels. During training, a masked binary cross-entropy loss function is used to measure the difference between the model's prediction and the true risk labels. The mask is used to filter out nodes and time slices with missing labels or incomplete data, ensuring that the training process remains stable and effective even with incomplete data. In each iteration, the loss function calculates the deviation between the predicted probability and the actual label, and the error signal is propagated layer by layer through the backpropagation algorithm to update all parameters in the spatial convolutional layer and the temporal recurrent unit. After multiple iterations of training, the model's parameters gradually converge, ultimately enabling it to accurately learn the spatiotemporal correlation between dynamic features such as waste collection delays, abnormal loads, and traffic impacts and the occurrence of rodent risks.

[0036] In step S5, based on the real-time inflow of multi-source data streams, the feature vectors of each node in the spatiotemporal graph are dynamically updated, and the rodent risk prediction model is used to perform online inference of rodent risk, generating the spatiotemporal probability distribution of rodent risk within a preset time step in the future.

[0037] A real-time data stream processing pipeline is constructed, which collects and processes data in batches at preset fixed time intervals, with a default interval of 30 minutes. At each fixed time interval, it receives current collection vehicle trajectory data, garbage collection data, and traffic flow data. Collection vehicle trajectory data includes continuously collected location information and timestamps during vehicle operation; the trajectory point sequence allows determination of whether the vehicle has completed arrival and service at the designated collection point within the predetermined time window. Garbage collection data includes the actual collection volume at each collection point within the current time window, and the degree of difference between this and the historical baseline collection volume. Traffic flow data covers indicators such as average speed, traffic density, and congestion level of the road segment where the vehicle is located, reflecting the vehicle's operational efficiency under current traffic conditions. For each data batch, a rigorous feature extraction operation is performed, transforming the raw trajectory and traffic data into a parameter set for subsequent calculations. The time series length of this parameter set is consistent with the rodent infestation cycle length. After feature extraction, based on a comparison of the trajectory data and the planned collection time window, the collection delay time for each grid corresponding to each collection point is recalculated, i.e., the difference between the actual arrival time and the planned time is recorded. The load anomaly index is regenerated using the difference between the actual waste collection volume and the historical baseline waste collection volume. The index value is then normalized to a comparable range. Finally, combined with real-time traffic flow data, the waste collection efficiency attenuation coefficient is updated according to a predetermined regression coefficient or mapping model, thus forming a complete feature vector for each grid cell at the current moment.

[0038] After completing the above calculations, the updated cleaning delay duration, load anomaly indicators, and cleaning efficiency decay coefficient are rewritten into the feature fields of the corresponding nodes in the spatiotemporal graph. Each node in the spatiotemporal graph represents a spatial grid unit, and the value of the node feature vector is updated with the data input at each time interval. The updated spatiotemporal graph is input as a whole into the trained rodent risk prediction model. This model calculates the temporal and spatial dependencies of the input data based on graph convolution and temporal recursive unit structures, and finally outputs the rodent risk probability value of each spatial grid unit in the future preset time step. These probability values ​​are predictions for the future and are usually represented by floating-point numbers between 0 and 1 to indicate the risk level. Spatial interpolation and continuousization processing are performed on the probability values ​​of each spatial grid. The discrete probability values ​​are connected into a continuous surface curve through the interpolation algorithm, thereby generating a smooth risk gradient change map. This change map can clearly reflect the distribution trend of rodent risk in the entire area within a set time range (not exceeding 24 hours) in the future, providing intuitive data basis for subsequent management and decision-making.

[0039] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0040] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0041] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0042] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0043] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0044] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0045] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0046] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0047] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0048] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting sanitary rodent pest risks based on big data, characterized in that, Includes the following steps: S1. Obtain regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data, and perform spatial grid mapping on the above multi-source data; S2. Based on the date sequence of the rodent infestation cycle, calculate the garbage collection delay time and load anomaly index, and integrate traffic flow data to generate the garbage collection efficiency attenuation coefficient. S3. Construct a spatiotemporal graph structure for the spatial grid, and use the cleaning delay time, load anomaly index, and efficiency decay coefficient as dynamic node features; S4. Train a rodent risk prediction model by inputting the spatiotemporal graph structure of the spatial grid and the rodent risk label into the graph neural network. S5. Based on the real-time inflow of multi-source data streams, dynamically update the feature vectors of each node in the spatiotemporal graph, use the rodent risk prediction model to infer rodent risk online, and generate the spatiotemporal probability distribution of rodent risk within a preset time step in the future.

2. The method for predicting sanitary rodent pest risk based on big data according to claim 1, characterized in that, In step S1, acquiring regional garbage collection data, collection vehicle trajectory data, and rodent infestation report data, and performing spatial gridding mapping on the above multi-source data specifically includes: Acquire historical rodent infestation report data, map the rodent infestation report data to the corresponding spatial grid cells according to geographical location, and establish a historical rodent infestation event database for each spatial grid; Waste collection data for the accessed area, including the planned service time window and collection volume for each collection point; Collect historical trajectory data of waste collection vehicles and map the vehicle trajectories and waste collection point locations into spatial grid cells through map matching.

3. The method for predicting sanitary rodent pest risk based on big data according to claim 1, characterized in that, In step S2, based on the date sequence of the rodent infestation cycle, the garbage collection vehicle trajectory data and garbage collection data are used to calculate the collection delay time and load anomaly index, and traffic flow data is integrated to generate a collection efficiency attenuation coefficient. Specifically, this includes: Extract the observation time window before the occurrence of rodent infestation events from the historical rodent infestation event database and mark it as a rodent infestation cycle, then convert the rodent infestation cycle into the corresponding date sequence; Extract the planned removal time window for each removal point within the rodent infestation cycle, and calculate the removal delay time for each removal point by combining the trajectory of the removal vehicles. Obtain historical waste collection data for each waste collection point and establish a waste collection benchmark model to calculate the deviation between the current waste collection volume and the historical benchmark and obtain load anomaly indicators. Obtain traffic flow data for the road segments corresponding to the trajectories of waste collection vehicles, establish a mapping relationship between traffic flow data and the degree of waste collection efficiency attenuation, and output the waste collection efficiency attenuation coefficient.

4. The method for predicting sanitary rodent pest risk based on big data according to claim 3, characterized in that, The establishment of a mapping relationship between traffic flow data and the degree of waste disposal efficiency attenuation, and the output of the waste disposal efficiency attenuation coefficient, specifically includes: Based on the historical data of cleaning delay and load anomaly indicators within the date sequence corresponding to the historical rodent infestation cycle, a regression model is constructed by combining the traffic flow data of the road segments where the corresponding cleaning vehicles are located. The traffic flow data includes average driving speed of road segments, traffic flow density, and congestion index characteristics. The influence coefficients between traffic flow characteristics, waste removal delay time, and abnormal load indicators are calculated by fitting regression. After normalizing the influence coefficients, the delay influence coefficients and load influence coefficients are weighted and fused according to preset weights to obtain the waste removal efficiency attenuation coefficient.

5. The method for predicting sanitary rodent pest risk based on big data according to claim 1, characterized in that, In step S3, the spatiotemporal graph structure of the spatial grid is constructed, and the cleaning delay time, load anomaly index, and efficiency decay coefficient are used as dynamic node features, specifically including: A spatiotemporal graph structure with spatial grid cells as nodes is constructed. The node feature vectors include the cleaning delay duration, load anomaly index, and cleaning efficiency attenuation coefficient. Based on the spatial adjacency relationship of spatial grid cells, undirected connection edges are established between nodes to connect adjacent spatial grid cells that share common edges; Directed connections between nodes are established based on the road network topology between each waste collection point, with the direction of the edges consistent with the travel direction of the waste collection vehicles.

6. The method for predicting sanitary rodent pest risk based on big data according to claim 1, characterized in that, In step S4, training the rodent risk prediction model by combining the spatiotemporal graph structure of the spatial grid with the rodent risk identification input graph neural network specifically includes: A spatiotemporal graph convolutional network is used as the core computing unit. The spatiotemporal graph convolutional network includes graph convolution operations in the spatial dimension and gated recurrent units in the temporal dimension. Spatial graph convolutional layers use Chebyshev polynomial approximation of spectral graph convolution to capture the spatial dependencies between dynamic nodes in the spatiotemporal graph; The time-gated cyclic unit layer processes the time-series changes of dynamic node features in the spatiotemporal graph and learns the temporal evolution law; The rodent risk identifiers are converted into binary label sequences, and the rodent risk prediction model is trained using a masked binary cross-entropy loss function.

7. The method for predicting sanitary rodent pest risk based on big data according to claim 6, characterized in that, The labeling method for the rodent risk markers is as follows: Extract the ratio of the total number of rats caught to the number of rat traps deployed in each spatial grid cell from the rat infestation report data, and calculate the rat density in each spatial grid cell. Spatial grid cells with rodent density exceeding a set rodent risk threshold are defined as rodent risk cells and marked with a rodent risk label.

8. The method for predicting sanitary rodent pest risk based on big data according to claim 1, characterized in that, In step S5, based on the real-time inflow of multi-source data streams, the feature vectors of each node in the spatiotemporal graph are dynamically updated. A rodent infestation risk prediction model is used to perform online inference of rodent infestation risk, generating the spatiotemporal probability distribution of rodent infestation risk within a future preset time step. Specifically, this includes: Establish a real-time data stream processing pipeline to receive current waste collection vehicle trajectory data, waste collection data, and traffic flow data at preset fixed time intervals; Perform feature extraction for each data batch and recalculate the cleaning delay time, load anomaly index and cleaning efficiency decay coefficient for each grid. Update the feature vector values ​​of the corresponding nodes in the spatiotemporal graph, and input the updated spatiotemporal graph into the trained rodent risk prediction model; The rodent infestation risk prediction model outputs the rodent infestation risk probability value for each spatial grid cell within a preset future time step. The discrete risk probability values ​​are converted into a continuous risk distribution surface, generating a smooth risk gradient change map.

Citation Information

Patent Citations

  • Attention mechanism-based monkey pox epidemic situation prediction method based on tensor space-time diagram convolution

    CN117711636A

  • Garbage collection full-scene data monitoring system

    CN120297663A

  • Intelligent monitoring and early warning device for mouse-borne diseases based on multi-source data fusion

    CN120496880A

  • Urban inland inundation risk multi-level prediction method and device based on space-time diagram learning, storage medium and computer program product

    CN120542667A

  • Estimation of crop pest risk and / or crop disease risk at sub-farm level

    US20210350295A1