Water body microorganism risk analysis method based on reinforcement learning

Through the water body microbial risk analysis method based on reinforcement learning, a dynamic water body map is constructed and the diffusion path sequence is optimized, which solves the problem of singularity and unreal-time risk assessment in the existing technology, and achieves high-efficiency emergency modeling and response to the diffusion of water body pollution.

CN120220788AActive Publication Date: 2025-06-27CHINESE RES ACAD OF ENVIRONMENTAL SCI

Patent Information

Application Number
CN202510694609.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing water pathogen microbial risk assessment technology has the problems of single detection methods, simple evaluation model, poor spatial adaptability and lack of dynamic updates, and it is difficult to capture the pollution spread process in real time and effectively respond to emergencies.

Method used

A water body microbial risk analysis method based on reinforcement learning is proposed. By collecting and preprocessing observation data, a dynamic water body diagram is constructed and dynamic graph representation learning is performed, a spatiotemporal risk heat matrix is ​​generated, and a Bayesian reward function and the maximum posterior inverse reinforcement learning model is combined to optimize the diffusion path sequence, and global search and local refinement are performed through artificial bee colony optimization algorithm.

Benefits of technology

It significantly improves the system's emergency modeling capabilities in sudden pollution scenarios, shortens the delay in pollution spread response, and improves the physical interpretability of the model and the ability to support public health interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220788A_ABST
    Figure CN120220788A_ABST
Patent Text Reader

Abstract

The invention discloses a water microorganism risk analysis method based on reinforcement learning. The method comprises the following steps: S1, generating a standardized observation data set; s2, constructing a dynamic water body map based on the standardized observation data set; s3, executing dynamic graph representation learning for the dynamic water body graph structure, and generating a space-time risk heat matrix on the basis of the node vector representation matrix; s4, constructing prior distribution of a Bayesian reward function, initializing the reward function and establishing a reward function parameter set; s5, performing incremental learning on the award function parameter set by adopting a maximum posterior reverse reinforcement learning model, and outputting a first propagation potential field; s6, obtaining an optimal diffusion path sequence; and S7, performing path reconstruction error analysis on the optimal diffusion path sequence and the standardized observation data set to obtain a dynamic feedback index. According to the method, pollution diffusion response delay caused by emergencies in actual measurement is greatly shortened, and the emergency modeling capability of the system in an emergent pollution scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water body detection, and particularly to a method for analyzing the risk of water body microorganisms based on reinforcement learning. Background Art

[0002] At present, the existing water body pathogenic microorganism risk assessment technologies generally have problems such as single detection means, simple assessment models, poor spatial adaptability, and lack of dynamic updates. Traditional risk identification methods mostly rely on on-site fixed-point monitoring, with a single bacterial index of Escherichia coli as the main detection object, making it difficult to comprehensively cover various types of pathogens such as viruses and parasites; at the same time, the existing detection technologies have limitations such as long cycle, low sensitivity, and poor on-site adaptability, resulting in the inability to capture the pollution diffusion process in real time and hindering the emergency response efficiency.

[0003] At the modeling level, most existing risk assessment models use empirical formulas or static statistical methods, which are only applicable under specific water areas and fixed parameters, and it is difficult to adapt to water body systems with strong nonlinear characteristics. In addition, mainstream methods often fail to comprehensively consider multi-dimensional environmental driving factors such as pollution source distribution, hydrodynamic processes, meteorological disturbances, and land use changes. The risk driving mechanism is not clearly described, the accuracy of path prediction results is limited, and there is a lack of in-depth characterization ability for the risk propagation mechanism. Spatial modeling means also mostly use manual division or rule-based zoning methods, and it is impossible to extract high-quality spatio-temporal propagation clues from dynamic hydrological evolution characteristics.

[0004] In summary, it is urgent to introduce a new modeling strategy for spatio-temporal evolution mechanisms, integrate multi-source information, enhance the intelligence of path generation, and have the ability of adaptive update to address the modeling challenges of the propagation process of pathogenic microorganisms under complex hydrodynamic conditions. Summary of the Invention

[0005] An object of the present invention is to propose a method for analyzing the risk of water body microorganisms based on reinforcement learning. In actual measurements of the present invention, the response delay caused by sudden events to pollution diffusion is greatly shortened, significantly improving the emergency modeling ability of the system in sudden pollution scenarios.

[0006] A method for analyzing the risk of water body microorganisms based on reinforcement learning according to an embodiment of the present invention includes the following steps: S1. Collect the original observation data set in the target water area and perform preprocessing to generate a standardized observation data set; S2. Construct a dynamic water body map based on the standardized observation data set; S3. Perform dynamic graph representation learning on the dynamic water body map structure to obtain a node vector representation matrix, and generate a spatio-temporal risk heat matrix based on the node vector representation matrix; S4. Using the standardized observation dataset, the node vector characterization matrix, and the spatio-temporal risk heat matrix, construct the prior distribution of the Bayesian reward function, initialize the reward function, and establish the reward function parameter set; S5. Taking the standardized observation dataset as the input of the behavioral result on the dynamic water body graph structure, use the maximum a posteriori inverse reinforcement learning model to perform incremental learning on the reward function parameter set, and output the first propagation potential field; S6. Using the first propagation potential field as the fitness evaluation basis, call the artificial bee colony optimization algorithm to perform global search on the dynamic water body graph structure, generate the candidate diffusion path set, and based on the candidate diffusion path set and the first propagation potential field, perform the onlooker bee local refinement and scout bee random perturbation operations to obtain the optimal diffusion path sequence; S7. Conduct path reconstruction error analysis on the optimal diffusion path sequence and the standardized observation dataset to obtain the dynamic feedback index, and form a joint convergence iterative loop. When the joint convergence iterative loop meets the preset convergence threshold, output the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index.

[0007] Optionally, the original observation dataset includes the concentration of pathogenic microorganisms, water flow velocity, water temperature, pH value, water flow direction, physical connectivity attributes, and boundary contact attributes.

[0008] Optionally, S2 includes the following steps: S21. Construct a spatial node set. Each spatial node is used to represent a monitoring section, a tributary inlet, a potential pollution source, or a hydraulic control point in the target water area. All spatial nodes together form the spatial node set, and the spatial node set covers the entire water body research scope. The node numbers range from 1 to , where represents the total number of spatial nodes; S22. For each spatial node, extract the concentration of pathogenic microorganisms, water temperature, and pH value of the spatial node at each time step. Use the concentration of pathogenic microorganisms, water temperature, and pH value as the node attributes of the spatial node at the corresponding time step. The concentration of pathogenic microorganisms, water temperature, and pH value together form the node attribute set of the spatial node; S23. Based on the water flow direction and water flow velocity recorded in the standardized observation dataset, establish a directed edge relationship between spatial nodes. If there is a physical connection between two spatial nodes and the water flow direction points to the physical connection path, then establish a directed edge from the upstream node to the downstream node. The directed edge set consists of all edges that meet the conditions; S24. For each established directed edge, extract the water flow velocity corresponding to the directed edge at each time step. The water flow velocity value is used as the edge weight of the directed edge at the time step and construct the edge weight set. The edge weight is used to represent the dynamic intensity of pollutant propagation in this channel of the water body; S25. For each established directed edge, extract the water flow direction attribute, physical connectivity attribute, and boundary contact attribute of the edge from the standardized observation dataset, and jointly form the edge attribute set of the edge; S26. The spatial node set , the directed edge set , the node attribute set , the edge weight set , and the edge attribute set are uniformly encapsulated to construct a dynamic water body graph structure .

[0009] Optionally, the S3 includes the following steps: S31. Based on the dynamic water body graph structure , extract the node attribute set of each spatial node in the spatial node set in sequence from the initial time step to the final time step , and the edge weight set of each edge in the directed edge set , to form a temporal input sequence for dynamic graph representation learning; S32. For the temporal input sequence, construct a node vector representation model based on the dynamic graph convolutional network. The node feature update process of the node vector representation model is defined as: ; where represents the node feature vector corresponding to the spatial node at the time step in the -th layer network, represents the non-linear activation function, represents the weight matrix of the -th layer network, represents the bias term of the -th layer network, represents the neighbor node aggregation function, represents the set of spatial nodes adjacent to the node ; S33. Use the node vector representation model to perform forward propagation training on the dynamic graph temporal input sequence from to , to obtain the node feature vectors of each spatial node at each time step , to form a node vector representation matrix . The node vector representation matrix is used to express the spatio-temporal dynamic characteristics of spatial nodes; Use the node vector representation matrix as the input to construct a spatio-temporal risk heat matrix generation model; S35. Traverse all spatial nodes in the spatial node set according to the spatio-temporal risk heat matrix generation model, obtain the risk heat values of all nodes at each time step, and form a complete spatio-temporal risk heat matrix .

[0010] Optionally, the S4 includes the following steps: S41. Construct a physically-driven prior space for the Bayesian reward function. The physically-driven prior space takes the standardized observation data set as the input basis and combines the node vector representation matrix and the spatio-temporal risk heat matrix; S42. Define the Bayesian reward function as the conditional probability distribution of the spatial node and the time step . Each Bayesian reward function value follows a Gaussian process distribution with the prior mean function and the covariance function as parameters; S43. Construct the prior mean function of the Bayesian reward function. The prior mean function consists of three weighted parts. The first part is the relative residual term of the pathogenic microorganism concentration, the second part is the water temperature proportion term, and the third part is the pH deviation term; S44. Construct the covariance function of the Bayesian reward function. The covariance function takes the Euclidean distance between the node and the node at different times of the node vectors and the node vector as the input feature, and is mapped through an exponential decay function controlled by a scaling coefficient and a characteristic scale parameter to output the covariance value representing the correlation of the reward value; S45. Combine the prior mean function and the covariance function to form a complete prior distribution set of the Bayesian reward function. The prior distribution set covers all spatial nodes and all time steps, and is used to provide the initial structural relationship of the reward function in the spatial and temporal dimensions; S46. Based on the prior distribution set of the Bayesian reward function, initialize the reward function parameter set . Each element in the sample set represents the initial reward estimate value of the node at the time step , and all the estimated values form the initial sample set.

[0011] Optionally, S5 includes the following steps: S51. Using the pathogenic microorganism concentration distribution in the standardized observation dataset as the expert behavior sample input, based on the dynamic water body graph structure, node vector representation matrix, and Bayesian reward function parameter set, construct a maximum a posteriori inverse reinforcement learning model that integrates a double-layer attention mechanism and dynamic Bayesian update; S52. In the maximum a posteriori inverse reinforcement learning model, use the node attribute set of the spatial node set as the input to construct a node-level attention mechanism, which is used to learn the state weight of the spatial node at the time step ; S53. Based on the state weight output by the node-level attention mechanism, construct a time-level attention mechanism, which is used to capture the time-level attention weight of the pathogenic microorganism diffusion path changing with the time step ; S54. Based on the state weight output by the node-level attention mechanism and the time-level attention weight , reconstruct the objective optimization function of the maximum a posteriori inverse reinforcement learning : ; where is the likelihood probability that node exhibits the pathogenic microorganism concentration under the reward function , and is the prior distribution of the Bayesian reward function; S55. Through the dynamic Bayesian update method, perform iterative updates on the objective optimization function of the maximum a posteriori inverse reinforcement learning, and update the reward function parameter set ; S56. After the reward function parameter set completes iterative updates and reaches the predetermined convergence threshold, output the optimized reward function set, and generate the first propagation potential field based on the optimized reward function set: ; where is the optimized optimal diffusion strategy, is the discount factor, and represents the expected value calculation operator for the future propagation process under the optimal diffusion strategy . ​​​​​​

[0012] Optionally, S6 includes the following steps: S61. Calculate the path fitness value of each diffusion path. The calculation method of the path fitness value is to sum the propagation potential energy values of all spatial nodes on the diffusion path at the corresponding time step multiplied by their risk weight coefficients in the path in turn. The propagation potential energy values are obtained from the first propagation potential field; S62. Initialize the bee colony individual set of the artificial bee colony optimization algorithm, set the number of scout bees, the number of follower bees, and the number of scout bees, and generate an initial path set with the same number. Each path individual is composed of a group of spatial nodes; S63. For each path generated by the scout bee, perform a path update operation according to the gradient direction of the propagation potential field. The path update operation is to insert new spatial nodes or adjust the path direction along the direction of increasing propagation potential energy on the original path; S64. For each path optimized by the follower bee, perform a local refinement operation based on the path with the highest current propagation potential field fitness; S65. The scout bee is responsible for executing the fully random path search mechanism. The random path is constructed by randomly selecting connected spatial nodes on the dynamic water body graph structure. If the cumulative value of the propagation potential energy of the random path is higher than the cumulative value of the propagation potential energy of the worst path in the current path set by more than the minimum improvement threshold, then replace the worst path in the original set with the diffusion path; S66. Repeat the global guiding search of the scout bee, the local refinement search of the follower bee, and the random search mechanism of the scout bee. Record the path with the highest current fitness in each iteration and use it as the reference optimal solution of the path population. When the maximum number of iterations is reached or the improvement amplitude of the path fitness value is less than the preset termination threshold, terminate the artificial bee colony optimization process; S67. After the artificial bee colony optimization process is terminated, output the final optimal diffusion path. The optimal diffusion path is the path with the highest cumulative value of the propagation potential energy among all current path individuals. The diffusion path represents the most likely propagation path of pathogenic microorganisms from the potential source to the high-risk area in the dynamic water body graph structure.

[0013] Optionally, S7 includes the following steps: S71. Perform path reconstruction error analysis on the optimal diffusion path sequence output by the artificial bee colony optimization algorithm and the pathogenic microorganism concentration distribution in the standardized observation dataset; S72. Construct a dynamic feedback index set based on the path reconstruction error analysis results. The dynamic feedback index set includes spatial node matching error, temporal concentration fitting error, and propagation direction consistency error; S73. Take the set of dynamic feedback metrics as the evaluation basis for joint optimization updates, and dynamically update the set of reward function parameters using the feedback information of spatial node matching error and temporal concentration fitting error; S74. Re-enter the set of reward function parameters and the results after updating the artificial bee colony optimization control parameters into the maximum a posteriori inverse reinforcement learning model and the artificial bee colony optimization algorithm, and continue to perform reward function training and diffusion path search to form a joint convergence iteration loop; S75. Output the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity metrics: Final diffusion path sequence: Represents the optimal path trajectory of the waterborne pathogenic microorganisms spreading downstream from the starting high-risk source node. The node sequence satisfies the flow direction constraint and has the maximum cumulative potential energy in the propagation potential field; Dynamic risk area distribution map: Calculate the risk index value of each spatial node at each time step based on the final diffusion path sequence and the propagation potential field. The risk index value is used to divide the risk level areas, which are divided into high-risk areas with a risk index value ≥ 0.8, medium-risk areas with 0.5 ≤ risk index value < 0.8, and low-risk areas with a risk index value < 0.5; Reward function sensitivity metric: Based on the partial derivative response between the finally converged set of reward function samples and the physical and chemical metrics in the standardized observation dataset, output the sensitivity score of each driving factor to the change of the reward function. Factors with a sensitivity score greater than 0.3 are defined as key propagation driving factors; S77. Standardize the format and perform layer mapping on the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity metrics.

[0014] Optionally, the spatial node matching error: Measures the degree of coincidence between the node set in the optimal diffusion path sequence and the observed high-concentration node set; Temporal concentration fitting error: Measures the degree of difference between the propagation potential field values of each node in the diffusion path at the corresponding time step and the concentration values in the standardized observation dataset; Propagation direction consistency error: Measures the degree of consistency between the direction between path nodes and the flow direction recorded in the edge attribute set in the dynamic water body map structure.

[0015] The beneficial effects of the present invention are: (1) The present invention introduces a double-layer attention mechanism on the basis of the traditional inverse reinforcement learning structure: the node-level attention mechanism can dynamically adjust the node weights according to the transmission potential of pathogenic microorganisms at different spatial nodes; the time-level attention mechanism dynamically adjusts the weight distribution of each time step in the learning process based on the time evolution trend of the diffusion path, enabling the model to more effectively focus on key transmission nodes and time periods, thereby learning a reward function that better conforms to the water flow direction and pathogenic dynamic mechanism. This not only improves the policy convergence efficiency but also enhances the physical interpretability of the model, providing direct support for subsequent public health interventions.

[0016] (2) The present invention improves the traditional artificial bee colony optimization strategy. A transmission potential field gradient guidance mechanism is introduced in the leading bee stage to achieve the construction of a globally guiding path. In the follower bee stage, Gaussian perturbation is used to perform path refinement optimization, and a structural consistency constraint is introduced in the scout bee stage to enhance the physical rationality of path generation. At the same time, with the transmission potential energy as the core of fitness evaluation, multiple rounds of iterative search for the optimal path sequence are carried out in the dynamic water body graph structure, effectively avoiding the common problem of falling into local optima, and improving the structural fidelity of the optimal path in real complex watershed scenarios.

[0017] (3) The present invention establishes a dynamic convergence mechanism for joint update of the reward function and path search, improving the adaptability and real-time performance of the model to water body condition changes. By introducing a path reconstruction error feedback mechanism, the joint adaptive update of the reward function parameters and the control parameters of the optimization algorithm is realized. In each iteration, dynamic feedback indicators such as spatial node matching error and propagation direction consistency error are used to adjust the inverse reinforcement learning model and the artificial bee colony optimization strategy, forming a "behavior-structure-update" closed-loop learning process. In the face of scenarios with frequent hydrodynamic disturbances and dynamically changing hydrological parameters, it has the ability of rapid convergence and robust optimization. In actual measurements, the response delay of pollution diffusion caused by emergencies is greatly shortened, significantly improving the emergency modeling ability of the system in sudden pollution scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a flowchart of a method for analyzing the risk of water body microorganisms based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0020] Refer to Figure 1, A method for analyzing the risk of water - body microorganisms based on reinforcement learning, comprising the following steps: S1. Collect the original observation data set in the target water area, perform time synchronization, missing value repair, anomaly elimination, and dimension unification processing on the original observation data set, and generate a standardized observation data set according to the water - area space grid mapping rule; S2. Construct a dynamic water - body graph based on the standardized observation data set. In the structure of the dynamic water - body graph, define the monitoring section, tributary inlet, potential pollution source, and hydraulic control point as graph nodes, write the pathogenic microorganism concentration, water temperature, and pH value in the standardized observation data set into the node attributes, write the water flow velocity in the standardized observation data set into the edge weights, and define the flow direction, connectivity relationship, and boundary attributes between nodes as graph edge attributes; S3. Perform dynamic graph representation learning on the structure of the dynamic water - body graph, use the standardized observation data set containing the complete sequences of node attributes and edge weights at all time steps for training to obtain a node vector representation matrix, and generate a spatio - temporal risk heat matrix based on the node vector representation matrix; S4. Utilize the standardized observation data set, node vector representation matrix, and spatio - temporal risk heat matrix to construct the prior distribution of the Bayesian reward function, initialize the reward function, and establish a reward - function parameter set; S5. On the structure of the dynamic water - body graph, take the standardized observation data set as the input of the behavioral result, and use the maximum a posteriori inverse reinforcement learning model to perform incremental learning on the reward - function parameter set, and output the first propagation potential field; S6. Take the first propagation potential field as the fitness evaluation basis, call the artificial bee colony optimization algorithm and set the number of leading bees, following bees, and scout bees, perform global search on the structure of the dynamic water - body graph to generate a set of candidate diffusion paths, and based on the set of candidate diffusion paths and the first propagation potential field, perform the local refinement of following bees and the random perturbation operation of scout bees to obtain the optimal diffusion path sequence; S7. Perform path reconstruction error analysis on the optimal diffusion path sequence and the standardized observation data set to obtain a dynamic feedback index, and simultaneously update the reward - function parameter set and the artificial bee colony optimization algorithm control parameters with the dynamic feedback index to form a joint convergence iterative loop. When the joint convergence iterative loop meets the preset convergence threshold, output the final diffusion path sequence, dynamic risk - area distribution map, and reward - function sensitivity index, and complete the visualization display on the structure of the dynamic water - body graph.

[0021] In this embodiment, the original observation data set includes pathogenic microorganism concentration, water flow velocity, water temperature, pH value, water flow direction, physical connectivity attribute, and boundary contact attribute.

[0022] The boundary contact attribute is used to identify whether an edge in the graph is connected to the boundary area of the water body. The boundary area includes: water body inflow, outflow, sewage outlet, upstream confluence inlet, and downstream discharge sluice. These boundary points are usually key areas where pathogenic microorganisms enter or leave the target water area.

[0023] The physical connectivity attribute is used to determine whether there is a real physical water channel between two spatial nodes, that is, to determine whether these two nodes are connected by a river channel, water canal, or water pipe hydraulic structure system in the water body.

[0024] In this embodiment, S2 includes the following steps: S21. Construct a set of spatial nodes. Each spatial node is used to represent a monitoring section, tributary inlet, potential pollution source, or hydraulic control point in the target water area. All spatial nodes together form a set of spatial nodes, which covers the entire water body research scope. The node numbers range from 1 to , where represents the total number of spatial nodes; S22. For each spatial node, extract the concentration of pathogenic microorganisms, water temperature and pH value at each time step. The concentration of pathogenic microorganisms, water temperature, and pH value are used as the node attributes of the spatial node at the corresponding time step. The concentration of pathogenic microorganisms, water temperature, and pH value together form the node attribute set of the spatial node, which is used to characterize the physical and chemical state of the water body at the time step. The unit of the concentration of pathogenic microorganisms is CFU / m³, and the unit of water temperature is °C; S23. Based on the water flow direction and water flow velocity recorded in the standardized observation dataset, establish a directed edge relationship between spatial nodes. If there is a physical connection between two spatial nodes and the water flow direction points to the physical connection path, then establish a directed edge from the upstream node to the downstream node. The set of directed edges consists of all edges that meet the conditions; S24. For each established directed edge, extract the water flow velocity corresponding to the directed edge at each time step. The water flow velocity value is used as the edge weight of the directed edge at the time step and construct an edge weight set. The edge weight is used to represent the dynamic intensity of the propagation of pollutants in the water body through this channel; S25. For each established directed edge, extract the water flow direction attribute, physical connectivity attribute, and boundary contact attribute of the edge from the standardized observation dataset, and jointly form the edge attribute set of the edge. Among them, the water flow direction attribute is used to mark whether the directed edge conforms to the actual water flow direction of the water body, the physical connectivity attribute is used to mark whether the two nodes are truly connected, and the boundary contact attribute is used to mark whether the directed edge is connected to the water body boundary. The value of the water flow direction attribute is downstream or upstream, and the values of the physical connectivity attribute and the boundary contact attribute are boolean values; S26. Package the set of spatial nodes , the set of directed edges , the set of node attributes , the set of edge weights and the set of edge attributes for unified encapsulation to construct a dynamic water body graph structure . The dynamic water body graph structure is constructed in time series, and a complete graph instance is generated at each time step.

[0025] In this embodiment, S3 includes the following steps: S31. Based on the dynamic water body graph structure , extract the node attribute sets of each spatial node in the set of spatial nodes from the initial time step to the final time step in sequence, as well as the edge weight sets of each edge in the set of directed edges , forming a time series input sequence for dynamic graph representation learning; S32. For the time series input sequence, construct a node vector representation model based on the dynamic graph convolutional network. The node feature update process of the node vector representation model is defined as: ; Among them, represents the node feature vector corresponding to the spatial node at the time step in the -th layer network, represents the non-linear activation function, represents the weight matrix of the -th layer network, represents the bias term of the -th layer network, represents the neighbor node aggregation function, represents the set of spatial nodes adjacent to the node ; S33. Use the node vector representation model to perform forward propagation training on the dynamic graph time series input sequence from to to obtain the node feature vectors of each spatial node at each time step, forming a node vector representation matrix . The node vector representation matrix is used to express the spatio-temporal dynamic characteristics of spatial nodes; S34. Use the node vector representation matrix ​​​Build a spatio-temporal risk heat matrix generation model for the input: ; Among them, represents the risk heat value of the spatial node at the time step , and the value range is between 0 and 1; represents a multi-layer perceptron with the parameter set as the parameter, is a normalization function to ensure that the sum of the risk heat values of the spatial node at each time step is 1; S35. Traverse all the spatial nodes in the spatial node set to obtain the risk heat values of all nodes at each time step, and form a complete spatio-temporal risk heat matrix , which is used to quantitatively characterize the diffusion risk intensity of waterborne pathogenic microorganisms at different spatial nodes and time steps.

[0026] In this embodiment, the S4 includes the following steps: S41. Build a physically driven prior space for the Bayesian reward function. The physically driven prior space takes the pathogenic microorganism concentration, water temperature, and pH value in the standardized observation dataset as the input basis, and combines the node vector representation matrix and the spatio-temporal risk heat matrix; The construction process of the physically driven prior space of the Bayesian reward function is as follows: Take the water body physical and chemical information recorded in the standardized observation dataset as the prior information input basis, mainly including three indicators: pathogenic microorganism concentration, water temperature, and pH value. Among them, the pathogenic microorganism concentration is used to measure the actual presence intensity of pathogens at spatial nodes, the water temperature is used to reflect the support degree of the environment for the survival activity of microorganisms, and the pH value is used to indicate the chemical suitability of microorganisms to survive in the water body. Align the pathogenic microorganism concentration, water temperature, and pH value in space and time according to the spatial nodes and time steps to form an environmental driving vector for each node at each time step.

[0027] On this basis, fuse the node environmental driving vector with the node vector representation matrix. The node vector representation matrix is output by the dynamic graph representation learning model and is used to characterize the structural role and temporal evolution characteristics of spatial nodes in the propagation network. The fusion process is to splice the node physical and chemical state information with its potential diffusion ability representation in the dynamic graph as a high-dimensional feature input, so as to form a physical semantic feature set for reward function modeling.

[0028] Combine the physical semantic feature set with the spatio-temporal risk heat matrix. The spatio-temporal risk heat matrix provides the probability weight that the current node is judged as a high-risk propagation channel in the propagation graph, and is used to adjust the attention degree to different nodes and time steps in the prior modeling process. By jointly modeling the node physicochemical state features, graph structure representation features and risk heat values, a Bayesian reward function physical-driven prior space covering all spatial nodes and time steps is formed.

[0029] S42. Define the Bayesian reward function as the conditional probability distribution of spatial nodes and time steps . Each Bayesian reward function value follows a Gaussian process distribution with the prior mean function and covariance function as parameters. The prior mean function is used to characterize the reward trend, and the covariance function is used to describe the correlation degree between spatial nodes and between time steps; S43. Construct the prior mean function of the Bayesian reward function. The prior mean function consists of three weighted parts. The first part is the relative residual term of the pathogenic microorganism concentration, and the calculation method is the ratio difference between the pathogenic microorganism concentration of the target node at time step and the standardized maximum concentration value; the second part is the water temperature proportion term, and the calculation method is the ratio of the water temperature of this node at time step to the maximum water temperature value ; the third part is the pH deviation term, and the calculation method is the absolute difference between the pH of the node at time step and the neutral value 7; S44. Construct the covariance function of the Bayesian reward function. The covariance function takes the Euclidean distance between the node vectors of node and the node vector of node at different times as the input feature. The Euclidean distance is used to measure the similarity of nodes in the representation space, and finally it is mapped through an exponential decay function controlled by a scaling coefficient and a feature scale parameter to output the covariance value representing the reward value correlation; The exponential decay function takes the Euclidean distance between the node vectors of node and node at the time step as the core input. The Euclidean distance is obtained by subtracting the vectors of two nodes in the node vector representation matrix element by element, squaring and summing them, and then taking the square root, which measures the ecological environment and propagation state differences between two nodes in the high-dimensional feature space; To convert distance into covariance values, a scaling coefficient and a characteristic scale parameter are introduced as control factors. The scaling coefficient is used to control the overall amplitude of the covariance function, and the characteristic scale parameter is used to adjust the sensitivity of the function to the differences between nodes. In actual calculations, the squared Euclidean distance value is divided by twice the square of the characteristic scale parameter, and then the negative value is taken and mapped with the natural exponential function to obtain the final covariance value.

[0030] S45. The prior mean function and the covariance function are jointly used to form the complete prior distribution set of the Bayesian reward function , and the prior distribution set covers all spatial nodes and all time steps, and is used to provide the initial structural relationship of the reward function in the spatial and temporal dimensions; S46. Based on the prior distribution set of the Bayesian reward function , the parameter set of the reward function is initialized . Each element in the sample set represents the initial reward estimate of node at time step , and all the estimates form the initial sample set.

[0031] In this embodiment, S5 includes the following steps: S51. Using the pathogenic microorganism concentration distribution in the standardized observation dataset as the input of the expert behavior sample, based on the dynamic water body graph structure, the node vector representation matrix, and the parameter set of the Bayesian reward function, a maximum a posteriori inverse reinforcement learning model integrating a double-layer attention mechanism and dynamic Bayesian update is constructed; S52. In the maximum a posteriori inverse reinforcement learning model, using the node attribute set of the spatial node set as the input, a node-level attention mechanism is constructed. The node-level attention mechanism is used to learn the state weight of the spatial node at time step : ; where represents the learnable weight vector, represents the vector concatenation operation, and LeakyReLU is the non-linear activation function; In traditional inverse reinforcement learning, each spatial node is usually regarded as having equal weight in policy learning, and the importance differences of nodes in the process of pathogenic microorganism transmission are not distinguished. Especially in scenarios with complex water body structures or highly uneven pollutant source distributions, it is easy to cause overfitting of policy learning to low-risk nodes.

[0032] The node-level attention mechanism introduced in S52 calculates the weight coefficient of each node at each time step by integrating the spatio-temporal representation features of the node and the current reward value, realizing risk focusing on key nodes, making the reinforcement learning model automatically tend to information-intensive areas, propagation hub areas or areas of drastic changes during the training process, thereby improving the fitting efficiency of the reward function, shortening the model convergence time, and enhancing the physical interpretability of node weights.

[0033] S53. Based on the state weights output by the node-level attention mechanism , a time-level attention mechanism is constructed. The time-level attention mechanism is used to capture the time-level attention weights of the diffusion path of pathogenic microorganisms changing with time steps : ; where is a multi-layer perceptron with parameter set as parameters; Traditional time series reinforcement learning methods often process each time step in an equal-weight manner, ignoring the fact that the propagation behavior in pollution events has stages, lags and suddenness, thus affecting the generalization ability of policy learning and the accuracy of time-effect judgment.

[0034] The time-level attention mechanism introduced in S53 further converges the dynamic evolution trend across nodes on the basis of node attention, generating the state weights of each time step , dynamically emphasizing the key time windows of propagation, thereby effectively enhancing the response sensitivity of the model in the critical stages of pollution acceleration, turning point and diffusion, and improving the discrimination ability and emergency adaptation ability of the policy in the time dimension.

[0035] S54. Based on the state weights output by the node-level attention mechanism and the time-level attention weights , the objective optimization function of maximum a posteriori inverse reinforcement learning is reconstructed : ; where is the likelihood probability that node exhibits the concentration of pathogenic microorganisms under the reward function , is the prior distribution of the Bayesian reward function; Based on the traditional maximum a posteriori objective function, the attention weights provided by S52 and S53 are introduced to achieve joint weighted optimization of space and time, thereby reconstructing the loss function. This reconstruction not only enhances the model's selective focusing ability on data but also guides the model to "focus on fitting high-weight regions and reduce ineffective gradient propagation" during the optimization process. The finally obtained reward function is more globally representative and physically consistent.

[0036] S55. Through the dynamic Bayesian update method, execute the objective optimization function of maximum a posteriori inverse reinforcement learning for iterative update to update the set of reward function parameters , and the update rule is defined as: ; where is the learning rate, represents the gradient of the optimization function with respect to the reward function parameters; S56. After the iterative update of the set of reward function parameters is completed and the predetermined convergence threshold is reached, output the set of optimized reward functions, and generate the first propagation potential field based on the set of optimized reward functions : ; where is the optimized optimal diffusion strategy, is the discount factor, represents the expected value calculation operator for the future propagation process under the optimal diffusion strategy , and the first propagation potential field is used to accurately describe the optimal propagation path and spatio-temporal risk evolution trend of waterborne pathogenic microorganisms on the dynamic water body graph structure.

[0037] In this embodiment, the S6 includes the following steps: S61. Calculate the path fitness value of each diffusion path. The path fitness value is used to measure the strength of the propagation ability of the diffusion path in the propagation potential field. The calculation method of the path fitness value is to sum the propagation potential energy values of all spatial nodes on the diffusion path at the corresponding time steps after multiplying them by their risk weight coefficients in the path. The propagation potential energy values are obtained from the first propagation potential field, and the risk weight coefficients are used to reflect the importance of the propagation of pathogenic microorganisms at different positions on the path. The diffusion path fitness value is used as the evaluation index for artificial bee colony optimization; S62. Initialize the bee swarm individual set of the artificial bee colony optimization algorithm, set the number of leading bees, the number of following bees, and the number of scout bees, and generate an initial path set with the same number. Each path individual consists of a group of spatial nodes. The adjacent spatial nodes in the path must meet the requirements of edge connectivity, flow direction consistency, and boundary contact in the dynamic water body graph structure. All path sets together constitute the initial solution space for the search of artificial bee colony optimization; S63. For each path generated by the leading bee, perform a path update operation according to the gradient direction of the propagation potential field. The path update operation makes the path tend to the high propagation potential energy region by inserting new spatial nodes or adjusting the path direction along the direction of increasing propagation potential energy on the basis of the original path. The step size is controlled by the global step size coefficient, and the gradient direction is determined by the potential energy difference between adjacent nodes in the propagation potential field; S64. For each path optimized by the following bee, perform a local refinement operation based on the path with the highest fitness in the current propagation potential field. The local refinement operation includes performing operations such as inserting points, deleting points, or replacing path segments in the path, and introducing fine-grained structural changes on the basis of the original path by adding Gaussian perturbations. The mean of the Gaussian perturbation is zero, the variance is controlled by the set characteristic scale parameter, and the perturbation intensity is controlled by the local perturbation coefficient, which is used to enhance the local exploration ability of the path; S65. The scout bee is responsible for implementing a fully random path search mechanism. The random path is constructed by randomly selecting connected spatial nodes on the dynamic water body graph structure. If the cumulative value of the propagation potential energy of the random path is higher than the cumulative value of the propagation potential energy of the worst path in the current path set by more than the minimum improvement threshold, then replace the worst path in the original set with the diffusion path to avoid falling into local optimum and increase the global search ability; S66. Repeat the global guiding search of the leading bee, the local refinement search of the following bee, and the random search mechanism of the scout bee. Record the path with the highest fitness in each round of iteration and use it as the reference optimal solution of the path population. When the maximum number of iterations is reached or the improvement amplitude of the path fitness value is less than the preset termination threshold, terminate the artificial bee colony optimization process; S67. After the artificial bee colony optimization process is terminated, output the final optimal diffusion path. The optimal diffusion path is the path with the highest cumulative value of the propagation potential energy among all current path individuals. The diffusion path represents the most likely propagation path of pathogenic microorganisms from the potential source to the high-risk area in the dynamic water body graph structure.

[0038] In this embodiment, the S7 includes the following steps: S71. Perform path reconstruction error analysis on the optimal diffusion path sequence output by the artificial bee colony optimization algorithm and the pathogenic microorganism concentration distribution in the standardized observation dataset. The path reconstruction error analysis is used to evaluate the fitting degree of the diffusion path generated by the model to the actual observation in terms of both the spatial node distribution and the temporal concentration evolution; S72. Construct a dynamic feedback index set based on the path reconstruction error analysis results. The dynamic feedback index set includes spatial node matching error, temporal concentration fitting error, and propagation direction consistency error; S73. Use the dynamic feedback index set as the evaluation basis for joint optimization and update. Dynamically update the reward function parameter set using the feedback information of spatial node matching error and temporal concentration fitting error. The update method is to perform posterior adjustment on the sampling distribution of the current reward function, regenerate the reward function sample set, and adjust the control parameters of the artificial bee colony optimization algorithm using the feedback information of propagation direction consistency error. The control parameters include the perturbation intensity of the leading bee, the local search range of the following bee, and the re-initialization probability of the scout bee; S74. Re-input the updated results of the reward function parameter set and the artificial bee colony optimization control parameters into the maximum a posteriori inverse reinforcement learning model and the artificial bee colony optimization algorithm, and continue to perform reward function training and diffusion path search to form a joint convergence iteration loop; S75. Determine whether the joint convergence iteration loop meets the preset convergence threshold. If the change amplitude of the reconstruction error of the optimal diffusion path sequence is lower than the error threshold in several consecutive rounds, or the overall model evaluation index tends to be stable, it is determined to converge, terminate the iteration, and output the final result; S76. Output the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index: Final diffusion path sequence: Represents the optimal path trajectory of the waterborne pathogenic microorganisms spreading downstream from the starting high-risk source node. The node sequence satisfies the flow direction constraint and has the maximum cumulative potential energy in the propagation potential field; Dynamic risk area distribution map: Calculate the risk index value of each spatial node at each time step according to the final diffusion path sequence and the propagation potential field. The risk index value is used to divide the risk level area, which is divided into a high-risk area with a risk index value ≥ 0.8, a medium-risk area with 0.5 ≤ risk index value < 0.8, and a low-risk area with a risk index value < 0.5; Reward function sensitivity index: Based on the partial derivative response between the finally converged reward function sample set and the physical and chemical indexes in the standardized observation dataset, output the sensitivity score of each driving factor to the change of the reward function. The factor with a sensitivity score greater than 0.3 is defined as the key propagation driving factor; S77. Standardize the format and layer mapping of the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index, and perform visual display on the dynamic water body map structure, and support the export of a risk tracking report with pictures and texts and model traceability analysis materials.

[0039] In this embodiment, the spatial node matching error: measures the coincidence degree between the node set in the optimal diffusion path sequence and the observed high-concentration node set; Temporal concentration fitting error: Measures the degree of difference between the propagation potential field values of each node in the diffusion path at the corresponding time step and the concentration values in the standardized observation dataset. Propagation direction consistency error: Measures the degree of consistency between the directions among path nodes and the flow directions recorded in the edge attribute set of the dynamic water body map structure.

[0040] Example 1: During a certain online water quality monitoring process, the system detected continuous anomalies in the concentration values of pathogenic microorganisms at the downstream water intake point. Specifically, the concentration value recorded by the monitoring device with the sensor number "WS-305" increased rapidly from a stable 112 CFU / m³ to 412 CFU / m³ within the time period from 07:30 to 08:30. Moreover, the water temperature changed abnormally during the rising process, increasing from the normal 21.3°C to 24.6°C. The abnormal data exceeded the internally set warning threshold for the concentration change rate (the hourly increase value exceeded 200 CFU / m³), and the system triggered a "primary pollution warning" signal, requiring a source tracing analysis of the upstream area.

[0041] The system automatically invoked the method for inverting the diffusion path of pathogenic microorganisms proposed in the present invention, imported historical observation data for the past 12 hours, with a total data volume of 7,680 records, covering 72 water body section nodes. Each node contains 5 standardized indicators: pathogenic concentration, water temperature, pH value, flow rate, and flow direction. The system automatically constructed a dynamic water body map structure with 72 nodes, 213 edges, 24 frames for the time step, and an interval of 30 minutes for each frame.

[0042] In the node representation learning stage, the vector representation values in the area from node "P-22" to "P-27" showed a significant aggregation trend. After being normalized and encoded by the graph representation learning model, the risk heat value in this area exceeded 0.92 for 3 consecutive frames, becoming a spatio-temporal heat focus area.

[0043] In the risk-driven modeling stage, the prior mean part of the reward function significantly depends on three dimensions: the standardized concentration residual distribution, the amplitude of abnormal water temperature fluctuations, and the degree of deviation of the pH value from neutrality. Taking node "P-25" as an example, its concentration value at the 08:00 time frame is 393 CFU / m³, the corresponding standardized residual is 0.77, the water temperature is 24.4°C (deviating from the mean by 2.3°C), and the pH value is 8.1 (deviating from the neutral value by +1.1), resulting in a prior reward value of 0.86 for it, and the spatial similarity with its neighboring nodes in the covariance function reaching 0.94.

[0044] Subsequently, the system executes the maximum a posteriori inverse reinforcement learning process with a fusion attention mechanism. After the 13th iteration, a high-reward path is formed among nodes "P-20" to "P-28". The maximum value of the corresponding propagation potential field appears at node "P-24", with a value of 0.97. The path is input into the artificial bee colony optimization module. The initial bee colony contains 60 path individuals. After 12 rounds of search, the optimal diffusion path "P-18 → P-20 → P-23 → P-25 → P-27 → P-29" is finally selected.

[0045] The system automatically performs error reconstruction evaluation on the optimal path. The fitting error value is 0.074, which is better than the average value (0.13) in the system's historical records. At the same time, the system records the cumulative risk integral of this path in the propagation potential field as 5.74 (the threshold is 3.5). Based on this result, the system marks the area between path "P-18 to P-29" as a "high-risk diffusion zone", and generates an automatic risk level report with a risk level of "Level II: Moderate pollution warning", accompanied by a dynamic risk heat map, a key index deviation map, and a diffusion direction trend map.

[0046] The report is synchronously pushed to the management terminal through the interface. The management staff receives the alarm at 08:45 and initiates an investigation of the upstream river mouth according to the platform's recommendation. On-site staff found that industrial wastewater had been discharged recently at "River mouth number - R12". Water sample tests showed that the concentration of Escherichia coli was 678 CFU / m³, which was highly consistent with the starting concentration in the system's inversion path. Subsequently, the river mouth was temporarily blocked, and the source tracing control measures were completed at 09:20.

[0047] To verify the practicality of the method of the present invention, the research team compared the source tracing performance of the system using the traditional inverse numerical simulation method. Using the same data volume, the traditional method identified the path "P-17 → P-22 → P-26 → P-30" in this scenario, with a path reconstruction error of 0.162 and a fitting time of 14.8 minutes. While the method of the present invention has a coincidence rate of 94.2% between the identified path and the actual pollution point, the fitting time is only 3.7 minutes, and the early warning time is 11 minutes, effectively avoiding the entry of water from the high-pollution area into the drinking water treatment system.

[0048] The system further evaluated the sensitivity of the reward function and conducted an attribution analysis on the influence weights of water temperature, pH value, and concentration at each node in the path. The results showed that among the changes in the reward value of node "P-25", the weight of the water temperature factor was 0.41, the concentration residual was 0.36, and the pH value was 0.23, indicating that the driving effect of the high-temperature environment on the diffusion of pathogen activity was more significant in this scenario.

[0049] In summary, this embodiment demonstrates how the method of the present invention efficiently constructs the water body diffusion path, dynamically corrects the propagation potential, quickly locates the source, and assists decision-makers in completing effective intervention by integrating the reinforcement learning mechanism and the swarm intelligence search method in the scenario of sudden pollution events, sparse monitoring points, and complex water body structures, fully verifying the engineering adaptability and policy intelligence of the method.

[0050] The present invention introduces a two-layer attention mechanism on the basis of the traditional inverse reinforcement learning structure: the node-level attention mechanism can dynamically adjust the node weights according to the propagation potential of pathogenic microorganisms at different spatial nodes; the time-level attention mechanism dynamically adjusts the weight distribution of each time step in the learning process based on the time evolution trend of the diffusion path, enabling the model to more effectively focus on key propagation nodes and time periods, thereby learning a reward function that better conforms to the water flow direction and pathogenic dynamic mechanism, not only improving the policy convergence efficiency but also enhancing the physical interpretability of the model, providing direct support for subsequent public health interventions.

[0051] The present invention improves the traditional artificial bee colony optimization strategy. In the leading bee stage, a propagation potential field gradient guidance mechanism is introduced to construct a globally guided path. In the follower bee stage, Gaussian perturbation is used to refine and optimize the path. In the scout bee stage, a structural consistency constraint is introduced to enhance the physical rationality of path generation. At the same time, with the propagation potential energy as the core of fitness evaluation, the optimal path sequence is searched iteratively in the dynamic water body graph structure, effectively avoiding the common problem of falling into local optima, and improving the structural fidelity of the optimal path in the real complex watershed scenario.

[0052] The present invention establishes a dynamic convergence mechanism for joint update of the reward function and path search, improving the adaptability and real-time performance of the model to water body condition changes. By introducing a path reconstruction error feedback mechanism, the joint adaptive update of the reward function parameters and the control parameters of the optimization algorithm is realized. In each iteration, the inverse reinforcement learning model and the artificial bee colony optimization strategy are adjusted using dynamic feedback indicators such as spatial node matching error and propagation direction consistency error, forming a "behavior-structure-update" closed-loop learning process. In the scenario of frequent hydrodynamic disturbances and dynamic changes in hydrological parameters, it has the ability of fast convergence and robust optimization. In actual measurements, the response delay of pollution diffusion caused by sudden events is significantly shortened, significantly improving the emergency modeling ability of the system in the scenario of sudden pollution.

[0053] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.

Claims

1. A method for analyzing the risk of waterborne microorganisms based on reinforcement learning, characterized in that, It includes the following steps: S1. Collect the original observation data set in the target water area and perform preprocessing to generate a standardized observation data set; S2. Construct a dynamic water body map based on the standardized observation data set; S3. Perform dynamic graph representation learning on the structure of the dynamic water body map to obtain a node vector characterization matrix, and generate a spatio-temporal risk heat matrix based on the node vector characterization matrix; S4. Use the standardized observation data set, the node vector characterization matrix, and the spatio-temporal risk heat matrix to construct a prior distribution of the Bayesian reward function, initialize the reward function, and establish a reward function parameter set; S5. On the structure of the dynamic water body map, take the standardized observation data set as the input of the action result, and use the maximum a posteriori inverse reinforcement learning model to perform incremental learning on the reward function parameter set, and output the first propagation potential field; S6. Take the first propagation potential field as the fitness evaluation basis, call the artificial bee colony optimization algorithm to perform global search on the structure of the dynamic water body map, generate a candidate diffusion path set, and based on the candidate diffusion path set and the first propagation potential field, perform the onlooker bee local refinement and scout bee random perturbation operations to obtain the optimal diffusion path sequence; S7. Perform path reconstruction error analysis on the optimal diffusion path sequence and the standardized observation data set to obtain a dynamic feedback index, and form a joint convergence iteration loop. When the loop ends, output the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index.

2. The method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 1, wherein, The original observation data set includes pathogenic microorganism concentration, water flow velocity, water temperature, pH value, water flow direction, physical connectivity attribute, and boundary contact attribute.

3. The method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 2, wherein, The S2 includes the following steps: S21. Construct a set of spatial nodes. Each spatial node is used to represent a monitoring section, a tributary inlet, a potential pollution source, or a hydraulic control point in the target water area. All spatial nodes together form a set of spatial nodes, and the set of spatial nodes covers the entire water body research scope. The node numbers range from 1 to , where represents the total number of spatial nodes; S22. For each spatial node, extract the pathogenic microorganism concentration, water temperature, and pH value of the spatial node at each time step. Take the pathogenic microorganism concentration, water temperature, and pH value as the node attributes of the spatial node at the corresponding time step. The pathogenic microorganism concentration, water temperature, and pH value together constitute the node attribute set of the spatial node; S23. Based on the water flow direction and water flow velocity recorded in the standardized observation data set, establish a directed edge relationship between spatial nodes. If there is a physical connection between two spatial nodes and the water flow direction points to the physical connection path, then establish a directed edge from the upstream node to the downstream node. The directed edge set consists of all edges that meet the conditions; S24. For each established directed edge, extract the corresponding water flow velocity at each time step. The water flow velocity value is used as the edge weight of the directed edge at the time step and construct an edge weight set. The edge weight is used to represent the dynamic intensity of the pollutant propagation in the water body through this channel; S25. For each established directed edge, extract the water flow direction attribute, physical connectivity attribute, and boundary contact attribute of the edge from the standardized observation data set, and jointly form the edge attribute set of the edge; S26. Unify and encapsulate the set of spatial nodes , the set of directed edges , the set of node attributes , the set of edge weights and the set of edge attributes to construct a dynamic water body graph structure .

4. The method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 3, characterized in that, The S3 includes the following steps: S31. Based on dynamic water body graph structure , from the initial time step To the final time step Extract spatial node sets sequentially The node attribute set of each spatial node in and the directed edge set The set of edge weights for each edge in , forming a dynamic graph to represent the temporal input sequence of learning; S32. For the time series input sequence, construct a node vector characterization model based on the dynamic graph convolutional network. The node feature update process of the node vector characterization model is defined as: ; Among them, represents the node feature vector of the spatial node in the time step corresponding to the represents the non-linear activation function, represents the weight matrix of the layer network, represents the neighbor node aggregation function, represents the set of spatial nodes adjacent to the node S33. Using node vectors to represent the model arrive The dynamic graph temporal input sequence is forward propagated for training, and each spatial node is obtained At each time step The node feature vector of , forming a node vector representation matrix ,The node vector representation matrix is ​​used to express the spatiotemporal dynamic characteristics of spatial nodes; Using the node vector representation matrix as input, construct a spatio-temporal risk heat matrix generation model; S35. Traverse all spatial nodes in the set of spatial nodes according to the spatio-temporal risk heat matrix generation model to obtain the risk heat values of all nodes at each time step, and form a complete spatio-temporal risk heat matrix .

5. The method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 1, characterized in that, The S4 includes the following steps: S41. Construct a physically-driven prior space for the Bayesian reward function. The physically-driven prior space takes the standardized observed dataset as the input basis and combines the node vector representation matrix and the spatio-temporal risk heat matrix; S42. Define the Bayesian reward function as the conditional probability distribution of the spatial nodes and the time step , and each Bayesian reward function value obeys a Gaussian process distribution with the prior mean function and the covariance function as parameters; S43. Construct the prior mean function of the Bayesian reward function , the prior mean function consists of three weighted parts. The first part is the relative residual term of the pathogenic microorganism concentration, the second part is the water temperature proportion term, and the third part is the pH deviation term; S44. Construct the covariance function of the Bayesian reward function. The covariance function takes the Euclidean distance between the node and the node at different times, with the node vectors and the node vector as input features, and is mapped through an exponential decay function controlled by a scaling coefficient and a feature scale parameter, and outputs a covariance value representing the correlation of the reward value; S45. Combine the prior mean function with the covariance function to jointly form a complete set of prior distributions of the Bayesian reward function . The set of prior distributions covers all spatial nodes and all time steps and is used to provide the initial structural relationship of the reward function in the spatial and temporal dimensions; S46. Set of prior distributions based on Bayesian reward function , initialize the set of reward function parameters , for each element in the sample set represents a node at time step of the initial reward estimate, and all the estimates form the initial sample set.

6. The method for analyzing the risk of water - body microorganisms based on reinforcement learning according to claim 5, wherein, The S5 includes the following steps: S51. Use the pathogen concentration distribution in the standardized observed dataset as the input of the expert behavior sample. Based on the dynamic water body graph structure, the node vector representation matrix, and the Bayesian reward function parameter set, construct a maximum a posteriori inverse reinforcement learning model that integrates a double-layer attention mechanism and dynamic Bayesian update; S52. In the maximum a posteriori inverse reinforcement learning model, using the node attribute set of the spatial node set as input, a node-level attention mechanism is constructed, and the node-level attention mechanism is used to learn the state weight of the spatial node at the time step ; The state weight ; S53. State weights output based on the node-level attention mechanism , a time-level attention mechanism is constructed, and the time-level attention mechanism is used to capture the time-level attention weights of the diffusion path of pathogenic microorganisms changing with time steps ; State weights output based on the node-level attention mechanism and the time-level attention weights to reconstruct the objective optimization function of the maximum a posteriori inverse reinforcement learning : ; Among them, is the node showing the likelihood probability of the pathogenic microorganism concentration under the reward function ; is the prior distribution of the Bayesian reward function. S55. Execute the objective optimization function of maximum a posteriori inverse reinforcement learning through the dynamic Bayesian update method for iterative update to update the reward function parameter set ; S56. In the reward function parameter set After completing the iterative update and reaching the predetermined convergence threshold, output the optimized reward function set, and generate the first propagation potential field based on the optimized reward function set : ; Among them, is the optimized optimal diffusion strategy, is the discount factor, represents the expected value calculation operator for the future propagation process under the optimal diffusion strategy.

7. A method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 6, characterized in that, The S6 includes the following steps: S61. Calculate the path fitness value of each diffusion path. The calculation method of the path fitness value is to sum the propagation potential values of all spatial nodes on the diffusion path at the corresponding time step multiplied by their risk weight coefficients in the path in turn. The propagation potential value is obtained from the first propagation potential field; S62. Initialize the swarm individual set of the artificial bee colony optimization algorithm, set the number of leading bees, the number of following bees, and the number of scout bees, and generate an initial path set with the same number. Each path individual consists of a group of spatial nodes; S63. For each path generated by the leading bee, perform a path update operation according to the gradient direction of the propagation potential field. The path update operation is to insert new spatial nodes or adjust the path direction along the direction of increasing propagation potential on the basis of the original path; S64. For each path optimized by the following bee, perform a local refinement operation based on the path with the highest fitness in the current propagation potential field; S65. The scout bee is responsible for executing the fully random path search mechanism. The random path is constructed by randomly selecting connected spatial nodes on the dynamic water body graph structure. If the cumulative value of the propagation potential of the random path is higher than the cumulative value of the propagation potential of the worst path in the current path set by more than the minimum improvement threshold, then replace the worst path in the original set with the diffusion path; S66. Repeat the global guiding search of the leading bee, the local refinement search of the following bee, and the random search mechanism of the scout bee. Record the path with the highest fitness in each round of iteration and use it as the reference optimal solution of the path population. When the maximum number of iterations is reached or the improvement amplitude of the path fitness value is less than the preset termination threshold, terminate the artificial bee colony optimization process; S67. After the artificial bee colony optimization process is terminated, output the final optimal diffusion path. The optimal diffusion path is the path with the highest cumulative value of the propagation potential among all current path individuals. The diffusion path represents the most likely propagation path of the pathogen from the potential source to the high-risk area in the dynamic water body graph structure.

8. A method for analyzing the risk of waterborne microorganisms based on reinforcement learning according to claim 7, characterized in that, The S7 includes the following steps: S71. Conduct path reconstruction error analysis on the optimal diffusion path sequence output by the artificial bee colony optimization algorithm and the pathogen concentration distribution in the standardized observed dataset; S72. Construct a dynamic feedback index set based on the path reconstruction error analysis results. The dynamic feedback index set includes spatial node matching error, temporal concentration fitting error, and propagation direction consistency error; S73. Take the set of dynamic feedback indicators as the evaluation basis for joint optimization update, and dynamically update the set of reward function parameters using the feedback information of spatial node matching error and temporal concentration fitting error; S74. Re-input the set of reward function parameters and the results after updating the artificial bee colony optimization control parameters into the maximum a posteriori inverse reinforcement learning model and the artificial bee colony optimization algorithm, and continue to perform reward function training and diffusion path search to form a joint convergence iteration loop; S75. Output the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index: Final diffusion path sequence: It represents the optimal path trajectory of the waterborne pathogenic microorganisms diffusing from the starting high-risk source node downstream. The node sequence satisfies the flow direction constraint and has the maximum cumulative potential energy in the propagation potential field; Dynamic risk area distribution map: Calculate the risk index value of each spatial node at each time step according to the final diffusion path sequence and the propagation potential field. The risk index value is used to divide the risk level area, which is divided into a high-risk area with a risk index value ≥ 0.8, a medium-risk area with 0.5 ≤ risk index value < 0.8, and a low-risk area with a risk index value < 0.5; Reward function sensitivity index: Based on the partial derivative response between the finally converged reward function sample set and the physical and chemical indicators in the standardized observation dataset, output the sensitivity score of each driving factor to the change of the reward function. The factor with a sensitivity score greater than 0.3 is defined as the key propagation driving factor; S76. Standardize the format and perform layer mapping on the final diffusion path sequence, the dynamic risk area distribution map, and the reward function sensitivity index.

9. The method for analyzing the risk of water microorganisms based on reinforcement learning according to claim 8, wherein, The spatial node matching error: Measures the coincidence degree between the node set in the optimal diffusion path sequence and the observed high-concentration node set; Temporal concentration fitting error: Measures the difference degree between the propagation potential field value of each node in the diffusion path at the corresponding time step and the concentration value in the standardized observation dataset; Propagation direction consistency error: Measures the consistency degree between the direction between path nodes and the flow direction recorded in the edge attribute set in the dynamic water body map structure.

Citation Information

Patent Citations

  • Detection method of abnormal event of multi-variable water quality parameter time sequence data

    CN106872657A

  • Bayesian migration learning method in sudden water quality pollution environment

    CN111667104A

  • Sewage treatment process abnormal working condition monitoring method based on width slow feature neural network with incremental learning capability

    CN116842420A

  • Environment variable driven perchlorate point source pollution treatment method and system

    CN119005757A

  • Multi-scale space-time diagram convolution blue-green algae forecasting method and device based on hydrodynamic perception

    CN119026085A

Cited By

  • Remote calibration method and system for intelligent water station

    CN120562471A

  • Camera calibration method based on reward feedback sampling and fixed reference

    CN121236182A

  • Camera calibration method based on reward feedback sampling and fixed benchmark

    CN121236182B