A data fusion-based water quality anomaly tracing method and system for a water supply network

By aligning and weighting the water quality parameter data of the water supply network in time, and constructing a node connection matrix in combination with the topology, the problem of inaccurate data analysis in the source tracing of water quality anomalies in the water supply network is solved, and efficient pollution source location and early warning response are achieved.

CN120741805BActive Publication Date: 2025-12-05XIANGYUAN WATER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511192574.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-12-05
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing technologies for tracing the source of water quality anomalies in water supply networks fail to effectively utilize the time differences and parameter weights of distributed monitoring point data, resulting in a lack of a unified time benchmark for the data, which affects the accuracy of the analysis. Furthermore, they fail to deeply construct the relationships between network nodes, making it difficult to accurately track the propagation path of anomalies and locate the source of pollution.

Method used

By aligning water quality parameter data from distributed monitoring points in the water supply network with real-time parameters, assigning and correcting weight coefficients, constructing a fusion feature vector and node connection matrix, and combining topological data to trace anomalies, locate pollution sources, and generate early warning signals.

Benefits of technology

It has achieved high accuracy and rapid response in tracing water quality anomalies in water supply networks, significantly improving the accuracy of pollution source location and the timeliness of early warning response, and enhancing water quality safety assurance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120741805B_ABST
    Figure CN120741805B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of water quality detection, and discloses a water quality anomaly tracing method and system based on data fusion for a water supply network, which comprises the following steps: aligning water quality parameter data of distributed monitoring points in the water supply network at all times to obtain a standardized water quality parameter sequence; according to the measurement accuracy of parameters in the standardized water quality parameter sequence, the weight coefficients of the parameters are distributed, and the weight coefficients are corrected according to the hydraulic retention time in the water supply network to obtain correction coefficients of the water quality parameter data; the standardized water quality parameter sequence is weighted and fused based on the correction coefficients to obtain a fusion feature vector; a node connection relationship matrix is constructed according to topological structure data of the fusion feature vector; the fusion feature vector is subjected to anomaly tracing according to the node connection relationship matrix to obtain an anomaly result; a pollution source in the water supply network is located according to the anomaly result, and a water quality early warning signal is generated; and the application can improve the accuracy of water quality anomaly tracing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water quality testing technology, and in particular to a method and system for tracing the source of water quality anomalies in water supply networks based on data fusion. Background Technology

[0002] In the field of tracing the source of water quality anomalies in water supply networks, existing technologies are not precise enough in processing water quality parameter data collected from distributed monitoring points. They often ignore the differences in the time dimension of data from different monitoring points, resulting in a lack of a unified time benchmark and affecting the accuracy of subsequent analysis. At the same time, the weighting of various water quality parameters is relatively simplistic, failing to fully consider the differences in parameter measurement accuracy and the impact of hydraulic residence time in the water supply network on parameter validity. This leads to discrepancies between the data fusion results and the actual water quality conditions of the network, making it difficult to accurately reflect anomaly characteristics.

[0003] Existing technologies, when constructing relationships between pipeline nodes, do not make in-depth use of the topology and fail to effectively link the fused feature vectors with the actual connection status of the pipeline network. This makes it difficult to accurately trace the propagation path of anomalies during anomaly tracing. Furthermore, when locating the initial anomaly triggering node, there is a lack of effective reference to historical normal operating conditions, and the analysis of anomaly deviation is not comprehensive enough. This results in low accuracy in locating pollution sources, and the generation of early warning signals is insufficient to meet the needs of rapid response and precise handling in practical applications. Summary of the Invention

[0004] This invention provides a method and system for tracing the source of water quality anomalies in water supply networks based on data fusion, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, this invention provides a method for tracing the source of water quality anomalies in water supply networks based on data fusion, comprising:

[0006] S1. Align the water quality parameter data of the distributed monitoring points in the water supply network with time to obtain the standardized water quality parameter sequence of the water supply network;

[0007] S2. Based on the measurement accuracy of the parameters in the standardized water quality parameter sequence, assign weight coefficients to the parameters, and correct the weight coefficients based on the hydraulic residence time in the water supply network to obtain correction coefficients for the water quality parameter data.

[0008] S3. Based on the correction coefficient, the standardized water quality parameter sequence is weighted and fused to obtain the fused feature vector of the water supply network;

[0009] S4. Construct the node connection matrix of the water supply network based on the topological structure data of the fused feature vectors;

[0010] S5. Based on the node connection relationship matrix, perform anomaly tracing on the fused feature vector to obtain the anomaly results of the water supply network;

[0011] S6. Locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal.

[0012] In a preferred embodiment, the step of time-aligning the water quality parameter data from distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence for the water supply network includes:

[0013] Receive water quality parameter data packets uploaded by various distributed monitoring points and extract the timestamp information from the data packets;

[0014] The offset between the timestamp information and the standard time reference is calculated to generate time calibration parameters.

[0015] The water quality parameter data packets are time-series rearranged according to the time calibration parameters to eliminate transmission delay differences;

[0016] The time-series rearranged water quality parameter data are spatially correlated according to the location of pipeline nodes to generate the standardized water quality parameter sequence.

[0017] In a preferred embodiment, the weighting coefficient is calculated using the following formula:

[0018] ;

[0019] In the formula, The weighting coefficients are... This refers to the difference between the measured accuracy value of a parameter in the standardized water quality parameter sequence and the accuracy reference standard. This serves as a reference value for the accuracy of parameters in the standardized water quality parameter sequence.

[0020] In a preferred embodiment, the step of correcting the weighting coefficients based on the hydraulic retention time in the water supply network to obtain correction coefficients for the water quality parameter data includes:

[0021] Obtain the topology data of the water supply network and extract the hydraulic residence time parameters corresponding to the distributed monitoring points;

[0022] The hydraulic residence time parameter is converted into a time-dependent decay factor;

[0023] The weighting coefficients are dynamically compensated and adjusted based on the time-dependent decay factor to generate the correction coefficients.

[0024] In a preferred embodiment, the step of weighting and fusing the standardized water quality parameter sequence based on the correction coefficient to obtain the fused feature vector of the water supply network includes:

[0025] The standardized water quality parameter sequence is split into multidimensional data vector groups according to parameter type;

[0026] A diagonal weight matrix is ​​constructed using the correction coefficients, and feature fusion is performed on the multidimensional data vector group based on the diagonal weight matrix. The calculation formula for the feature fusion is as follows: ;

[0027] In the formula, The result of the feature fusion, The total number of types of the parameters. The type ordinal number of the parameter. For the first The class parameters are a diagonal matrix composed of corresponding correction coefficients. For the multidimensional data vector group, the first... Standardized data for class parameters;

[0028] The results are then subjected to feature dimension compression to obtain the fused feature vector of the water supply network.

[0029] In a preferred embodiment, constructing the node connection matrix of the water supply network based on the topological data of the fused feature vectors includes:

[0030] Extract the topological connection feature parameters of the network node identifiers from the fused feature vector;

[0031] A directed connection rule is established based on the fluid transmission direction of the nodes in the water supply network, and the topology connection characteristic parameters are converted into connection strength values ​​based on the directed connection rule.

[0032] A two-dimensional relational grid is constructed according to a preset node sorting. The connection strength value is mapped to the corresponding grid position of the two-dimensional relational grid, and zero-value parameters are filled between grid positions of nodes that are not directly connected in the two-dimensional relational grid to obtain the node connection relation moment of the water supply network.

[0033] In a preferred embodiment, the step of tracing the anomalies in the fused feature vector based on the node connection matrix to obtain the anomaly results of the water supply network includes:

[0034] The fused feature vector is mapped to the corresponding node position of the node connection relationship matrix to obtain the node anomaly feature distribution map of the water supply network;

[0035] By traversing the node connection matrix along the reverse path of the water flow in the water supply network, the abnormal feature propagation chain of the water supply network is obtained.

[0036] Based on the gradient changes of the abnormal features of the nodes in the abnormal feature propagation chain, the initial abnormal triggering node of the water supply network is located.

[0037] Based on the characteristic correlation strength between the initial anomaly triggering node and the monitoring points in the middle and lower reaches of the water supply network, anomaly results of the water supply network are generated.

[0038] In a preferred embodiment, locating the initial anomalous trigger node of the water supply network based on the gradient changes of node anomalous features in the anomalous feature propagation chain includes:

[0039] Obtain the baseline pattern of feature distribution under historical normal operating conditions;

[0040] Based on the difference between the abnormal feature distribution map and the feature distribution baseline pattern, an abnormal deviation heat map of the water supply network is generated.

[0041] Identify the region of maximum deviation in the abnormal deviation heatmap and associate the region of maximum deviation with the set of upstream nodes in the node connection matrix;

[0042] The connectivity continuity of the upstream node set is verified based on the reverse flow path of the water supply network to obtain the initial abnormal triggering node of the water supply network.

[0043] In a preferred embodiment, the step of locating the pollution source in the water supply network based on the abnormal result and generating a water quality early warning signal includes:

[0044] Extract the spatial coordinates of the initial anomaly triggering node from the anomaly results;

[0045] Based on the node connection relationship matrix, the pollution propagation path is traced to determine the candidate areas of pollution sources in the water supply network;

[0046] Based on the historical pollution event feature database and the spatial coordinates, the candidate areas are identified by risk level to obtain the pollution source location instructions for the water supply network;

[0047] Based on the results of the risk level identification, a preset early warning response strategy is matched to activate the multi-level water quality early warning signal of the water supply network.

[0048] To address the aforementioned problems, this invention also provides a data fusion-based system for tracing the source of water quality anomalies in water supply networks, the system comprising:

[0049] The water quality data extraction module is used to perform time-alignment of water quality parameter data from distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence for the water supply network.

[0050] The weight coefficient correction module is used to assign weight coefficients to the parameters according to the measurement accuracy of the parameters in the standardized water quality parameter sequence, and to correct the weight coefficients according to the hydraulic residence time in the water supply network, so as to obtain the correction coefficients of the water quality parameter data.

[0051] The feature fusion module is used to perform weighted fusion of the standardized water quality parameter sequence based on the correction coefficient to obtain the fused feature vector of the water supply network;

[0052] The topology connection module is used to construct the node connection relationship matrix of the water supply network based on the topology data of the fused feature vector;

[0053] An anomaly tracing module is used to trace anomalies in the fused feature vector based on the node connection relationship matrix to obtain the anomaly results of the water supply network.

[0054] The pollution early warning module is used to locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] 1. This invention achieves precise weighted fusion of standardized water quality parameter sequences by real-time alignment of water quality parameter data from distributed monitoring points in the water supply network, combining the parameter measurement accuracy with weight coefficients, and correcting the weight coefficients based on hydraulic residence time. This effectively enhances the ability of the fused feature vector to characterize the water quality status of the network, providing a high-quality data foundation for subsequent anomaly tracing, thereby improving the accuracy of water quality anomaly tracing.

[0057] 2. This invention constructs a node connection matrix by fusing topological data of feature vectors, and uses this matrix to trace anomalies in the fused feature vectors. This enables precise location of the initial anomaly triggering node, clear tracing of pollution propagation paths, accurate identification of pollution sources in the water supply network, and generation of water quality early warning signals. This significantly enhances the accuracy of pollution source location and the timeliness of early warning response, providing strong support for ensuring the safety of water quality in the water supply network. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating a method for tracing the source of water quality anomalies in a water supply network based on data fusion, as provided in an embodiment of the present invention.

[0059] Figure 2A functional module diagram of a water quality anomaly tracing system for a water supply network based on data fusion, provided in an embodiment of the present invention;

[0060] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0061] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0062] This application provides a method for tracing the source of water quality anomalies in a water supply network based on data fusion. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for tracing the source of water quality anomalies in a water supply network based on data fusion can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0063] Reference Figure 1 The diagram shown is a flowchart illustrating a method for tracing the source of water quality anomalies in a water supply network based on data fusion, according to an embodiment of the present invention. In this embodiment, the method for tracing the source of water quality anomalies in a water supply network based on data fusion includes:

[0064] S1. Align the water quality parameter data of the distributed monitoring points in the water supply network with time to obtain the standardized water quality parameter sequence of the water supply network;

[0065] In this embodiment of the invention, the step of time-aligning the water quality parameter data of distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence for the water supply network includes:

[0066] Receive water quality parameter data packets uploaded by various distributed monitoring points and extract the timestamp information from the data packets;

[0067] The offset between the timestamp information and the standard time reference is calculated to generate time calibration parameters.

[0068] The water quality parameter data packets are time-series rearranged according to the time calibration parameters to eliminate transmission delay differences;

[0069] The time-series rearranged water quality parameter data are spatially correlated according to the location of pipeline nodes to generate the standardized water quality parameter sequence.

[0070] Specifically, when receiving water quality parameter data packets uploaded by various distributed monitoring points, the source identifier of each data packet is first determined by the source code contained in the data packet transmission protocol. This identifier is composed of the monitoring point number and the unique serial number of the device, ensuring a unique mapping with the corresponding distributed monitoring point. Then, the parsing tool of the data receiving terminal is used to open the data packets one by one. The parsing tool automatically scans the metadata area of ​​the data packets. If no time record is found in the metadata, the information column labeled "collection time" in the data fields is searched to extract the time record accurate to milliseconds as timestamp information. After extraction, the timestamp information and the source identifier are bound and stored through the association field of the database. The storage is in the form of a table, with each row corresponding to one record, including the source identifier, timestamp information and data packet storage path, ensuring that each timestamp can be traced back to the monitoring point that uploaded the data through the source identifier.

[0071] Furthermore, each extracted timestamp is compared with a standard time reference. The standard time reference directly uses the UTC time signal output by the GPS receiver, which is updated every second and accurate to the nanosecond level. The current time of the standard time is obtained in real time through the time reading module, and the timestamp information to be compared is retrieved. Both are converted to the same time format, i.e., year-month-day-hour-minute-second-millisecond. Then, using the current time of the standard time as the reference, the current time of the standard time is subtracted from the time corresponding to the timestamp information. The result is the offset. If the time corresponding to the timestamp information is earlier than the current time of the standard time, the offset is negative; otherwise, it is positive. Time calibration parameters are generated based on the specific value of the offset and the corresponding monitoring point identifier. The time calibration parameters are saved in the form of a text file. Each parameter includes the monitoring point source identifier, the offset value, and the parameter generation time for accurate retrieval in subsequent steps.

[0072] Furthermore, based on the generated time calibration parameters, the corresponding water quality parameter data packets are matched using the monitoring point source identifier in the parameters. For each successfully matched data packet, the time field of the stored data record within the data packet is opened. According to the offset in the time calibration parameters corresponding to the data packet, each record in the time field is adjusted. If the offset is positive, the corresponding duration is added to the original time; if the offset is negative, the corresponding duration is reduced from the original time. The adjusted time field is uniformly referenced to the UTC time provided by the Global Positioning System. After the adjustment is completed, all water quality parameter data packets are rearranged in ascending order of the adjusted time field values ​​to eliminate the transmission delay differences caused by different monitoring points due to different transmission path lengths and network congestion levels, so that all data are in a unified time dimension.

[0073] Furthermore, the time-series rearranged water quality parameter data are categorized according to the specific locations of their corresponding distributed monitoring points within the pipeline network. Each pipeline node location has a preset geographical coordinate range. By comparing the installation coordinates of the monitoring points with the node coordinate range, the data is assigned to the corresponding node category, ensuring that the dataset under each pipeline node location comes from monitoring points within that location range. Subsequently, the pipeline network topology diagram is invoked, which clearly marks the connecting pipes, distances, and water flow directions between each node. Based on the information in the diagram, the water quality parameter data of adjacent node locations are associated. For example, upstream nodes and downstream nodes are associated through connecting pipes. At the same time, the water flow transmission time between nodes is recorded during the association process to clarify the spatial transmission relationship between data from different nodes. Finally, all associated node data are integrated into a structured table. The table uses the unique code of the pipeline node location as an index, and each index contains the water quality parameter records of that node at different time points, forming a complete standardized water quality parameter sequence.

[0074] In general, in the complex operating environment of water supply networks, distributed monitoring points collect a large amount of water quality parameter data. However, due to differences in the operating conditions, transmission paths, and clock accuracy of the equipment at each monitoring point, these data often have inconsistent timestamps in their initial state. For example, some older equipment may experience time shifts due to clock module aging, resulting in data acquired by different monitoring points at the same "nominal time" not actually corresponding to the water quality conditions in the network at the same moment. While some newly deployed equipment possesses high-precision clocks, data transmission may still be delayed due to network fluctuations, leading to time misalignments.

[0075] In summary, time alignment of these water quality parameter data is of paramount importance. On one hand, it ensures that data from all monitoring points are on the same time horizon, providing a reliable time series basis for subsequent data analysis. For example, when studying the changing trends of water quality parameters over time, accurate time alignment can prevent biases in trend analysis caused by time misalignment, thus making the analysis results more accurately reflect the actual changes in water quality within the pipe network. On the other hand, time alignment facilitates multi-parameter collaborative analysis. Complex interrelationships exist between different water quality parameters in the water supply network; for instance, changes in pH value may affect the solubility of certain heavy metal ions.

[0076] In summary, time alignment enables precise temporal matching of data from these different parameters, facilitating in-depth analysis of the intrinsic relationships between parameters. This provides more accurate data support for key applications such as tracing the source of water quality anomalies and early warning of sudden water quality changes, greatly enhancing the scientific rigor and effectiveness of water quality monitoring and management in water supply networks.

[0077] S2. Based on the measurement accuracy of the parameters in the standardized water quality parameter sequence, assign weight coefficients to the parameters, and correct the weight coefficients based on the hydraulic residence time in the water supply network to obtain correction coefficients for the water quality parameter data.

[0078] In this embodiment of the invention, the formula for calculating the weighting coefficient is as follows: ;

[0079] In the formula, The weighting coefficients are... This refers to the difference between the measured accuracy value of a parameter in the standardized water quality parameter sequence and the accuracy reference standard. This serves as a reference value for the accuracy of parameters in the standardized water quality parameter sequence.

[0080] The step of correcting the weighting coefficients based on the hydraulic retention time in the water supply network to obtain correction coefficients for the water quality parameter data includes:

[0081] Obtain the topology data of the water supply network and extract the hydraulic residence time parameters corresponding to the distributed monitoring points;

[0082] The hydraulic residence time parameter is converted into a time-dependent decay factor;

[0083] The weighting coefficients are dynamically compensated and adjusted based on the time-dependent decay factor to generate the correction coefficients.

[0084] Specifically, the weighting coefficients in the formula are obtained by processing the standardized water quality parameter sequence. The measurement accuracy value of the parameter in the standardized water quality parameter sequence is the actual measurement accuracy data of each parameter in the sequence, and the accuracy reference value is a standard reference value set for evaluating the measurement accuracy. The difference between the two and the accuracy reference value together serve as the basic data for calculating the weighting coefficients.

[0085] Furthermore, the significance of this formula is to calculate the weighting coefficient. By considering the difference between the measurement accuracy value of a parameter in the standardized water quality parameter sequence and the accuracy reference value, the importance of the parameter in the correlation analysis or calculation is determined. The smaller the difference, the larger the weighting coefficient, indicating that the measurement accuracy of the parameter is closer to the reference reference and that it has a higher proportion in subsequent processing.

[0086] Furthermore, judging from the trend of the formula, when the absolute value of the difference between the measured accuracy value and the accuracy reference value of the parameter in the standardized water quality parameter sequence increases, the denominator increases accordingly, and the value of the whole fraction, i.e., the weight coefficient, decreases; when the absolute value of the difference decreases, the denominator decreases accordingly, and the weight coefficient increases. That is, the weight coefficient and the absolute value of the difference between the measured accuracy value and the accuracy reference value of the parameter show an inverse trend.

[0087] Specifically, when acquiring the topology data of the water supply network, the network management system first retrieves a vector graphic file containing pipeline routes, pipe diameters, node locations, valve distribution, and connection relationships. This file is stored in CAD format, with each node location labeled with a corresponding distributed monitoring point number. The file is then opened using a graphic analysis tool, which automatically identifies information such as node coordinates, pipe lengths, and water flow directions to form a structured topology data table. Next, node records matching the distributed monitoring point numbers are selected from the data table. Based on the location of these nodes in the topology and combined with pipeline flow velocity data, the time required for water to flow from that node to the next associated node is calculated. This time is the hydraulic residence time parameter corresponding to the distributed monitoring point. After the calculation is completed, the hydraulic residence time parameter is bound to the monitoring point number and stored in a dedicated data table.

[0088] Furthermore, when converting the hydraulic residence time parameter into a time-dependent decay factor, a time decay benchmark value is first determined. This benchmark value is the standard hydraulic residence time specified in the water supply network design code. Then, the hydraulic residence time parameter corresponding to each distributed monitoring point is extracted and compared with the time decay benchmark value. If the hydraulic residence time parameter is greater than the time decay benchmark value, it indicates that the water flow residence time at this node is too long, and the water quality timeliness decreases significantly. In this case, a value less than 1 is generated as the time-dependent decay factor according to the proportion of residence time exceeding the benchmark value. The larger the proportion of exceeding the benchmark value, the smaller the value. If the hydraulic residence time parameter is less than or equal to the time decay benchmark value, it indicates that the water flow residence time is within a reasonable range, and the water quality timeliness is not significantly affected. In this case, the time-dependent decay factor is directly set to 1. After the conversion, each hydraulic residence time parameter corresponds to a unique time-dependent decay factor.

[0089] Furthermore, when dynamically compensating and adjusting the weight coefficients based on the time-sensitivity decay factor, the weight coefficients to be processed are first retrieved from the data storage module. These weight coefficients are values ​​corresponding to each distributed monitoring point obtained through prior calculation. Then, the time-sensitivity decay factor belonging to the same monitoring point as the weight coefficient is found. The weight coefficient and the time-sensitivity decay factor are multiplied together, and the result is the value after compensation and adjustment. If the time-sensitivity decay factor is less than 1, the weight coefficient will decrease accordingly after multiplication, indicating that the water quality parameter of the monitoring point has decreased in importance in the overall analysis due to the decline in timeliness. If the time-sensitivity decay factor is equal to 1, the weight coefficient remains unchanged after multiplication, indicating that the water quality parameter of the monitoring point has good timeliness and its importance has not been affected. After all weight coefficients have been adjusted, these compensated and adjusted values ​​together constitute the correction coefficient.

[0090] In summary, from the perspective of parameter measurement accuracy, the accuracy of measuring equipment varies for different water quality parameters, leading to differences in data reliability. By using the weighting coefficient calculation formula, higher initial weights can be assigned to parameters with higher measurement accuracy (i.e., smaller differences from the accuracy reference), ensuring that the fusion process focuses more on reliable data and reduces the interference of low-accuracy data on the results, thus improving the quality of data fusion from the source. Furthermore, the introduction of hydraulic residence time for correction further optimizes the weighting coefficients.

[0091] In general, hydraulic residence time directly affects the timeliness of water quality parameters in water supply networks—the longer the residence time, the lower the effectiveness of the parameters in reflecting the current water quality status. By converting hydraulic residence time into a timeliness decay factor and dynamically compensating for and adjusting the initial weighting coefficients, the corrected coefficients can better reflect the actual changes in water quality parameters within the network. For example, the weight of parameters corresponding to water bodies with longer residence times in the network will be appropriately reduced to avoid distortion of the fusion results due to excessive influence from outdated data.

[0092] In summary, the final correction coefficients not only reflect the differences in measurement reliability of different parameters, but also take into account the impact of the hydraulic characteristics of the pipeline network on the timeliness of the parameters. This provides an accurate and realistic weighting basis for the subsequent weighted fusion of standardized water quality parameter sequences, enabling the fused feature vectors to more realistically and comprehensively reflect the water quality status of the water supply network, and laying a high-quality data foundation for subsequent anomaly tracing and pollution source location.

[0093] S3. Based on the correction coefficient, the standardized water quality parameter sequence is weighted and fused to obtain the fused feature vector of the water supply network;

[0094] In this embodiment of the invention, the step of weighting and fusing the standardized water quality parameter sequence based on the correction coefficient to obtain the fused feature vector of the water supply network includes:

[0095] The standardized water quality parameter sequence is split into multidimensional data vector groups according to parameter type;

[0096] A diagonal weight matrix is ​​constructed using the correction coefficients, and feature fusion is performed on the multidimensional data vector group based on the diagonal weight matrix. The calculation formula for the feature fusion is as follows:

[0097] ;

[0098] In the formula, The result of the feature fusion, The total number of types of the parameters. The type ordinal number of the parameter. For the first The class parameters are a diagonal matrix composed of corresponding correction coefficients. For the multidimensional data vector group, the first... Standardized data for class parameters;

[0099] The results are then subjected to feature dimension compression to obtain the fused feature vector of the water supply network.

[0100] Specifically, when splitting the standardized water quality parameter sequence into a multidimensional data vector group according to parameter type, the parameter types included in the standardized water quality parameter sequence are first identified, such as pH value, turbidity, and residual chlorine content. Each parameter type is treated as an independent dimension. Then, all data in the standardized water quality parameter sequence are traversed, and the datasets of the same parameter type are extracted according to the parameter type label corresponding to the data. The datasets are arranged in the order of data collection time to form a vector. For example, all pH value data are arranged in time to form a pH value vector, all turbidity data are arranged in time to form a turbidity vector, and so on. Finally, a multidimensional data vector group composed of vectors of different parameter types is obtained. Each vector corresponds to a parameter type, and the data within the vector maintains the original temporal correlation.

[0101] Furthermore, when constructing a diagonal weight matrix using the correction coefficients and performing feature fusion on the multidimensional data vector group based on the diagonal weight matrix, the number of vectors contained in the multidimensional data vector group is first determined, which is the dimension of the diagonal weight matrix. Then, the correction coefficients are sequentially filled into the main diagonal of the matrix according to the order of the vectors in the multidimensional data vector group, and all positions outside the main diagonal of the matrix are filled with 0, forming a diagonal weight matrix. Next, each vector in the multidimensional data vector group is taken out, and each data in the vector is multiplied by the correction coefficient at the corresponding position in the diagonal weight matrix to obtain the data of each vector after weight adjustment. Then, all the weight-adjusted vectors are superimposed according to the corresponding time point, that is, the data of different parameter types at the same time point are added to form a new vector. This new vector is the result of feature fusion.

[0102] Furthermore, when compressing the feature dimensions of the results to obtain the fused feature vector of the water supply network, the length of the vector after feature fusion is first analyzed. This length is equal to the total number of data collections. Then, the number of key feature points to be retained is determined. These key feature points include the maximum value, minimum value, start point, end point, and the point with the largest data change rate in the data sequence. The values ​​corresponding to these key feature points are accurately extracted from the fused results and arranged in chronological order to form a new vector. If the number of key feature points is less than the original vector length, then the dimension compression is completed. This new vector composed of the values ​​of key feature points is the fused feature vector of the water supply network.

[0103] Specifically, the parameter sources of the feature fusion results are clear and explicit. The total number of parameter types is determined by the number of vectors contained in the multidimensional data vector group split from the standardized water quality parameter sequence. Each vector corresponds to a parameter type, and the specific number is thus calculated. The parameter type ordinal number is a sequential number assigned to distinguish different parameter types, starting from 1 and increasing sequentially, corresponding one-to-one with the arrangement order of vectors in the multidimensional data vector group. The diagonal matrix formed by the correction coefficients corresponding to the class parameters originates from the correction coefficients generated after dynamic compensation adjustment of the weight coefficients based on the time-dependent decay factor, and is filled into the rank diagonal of the diagonal matrix according to the order of parameter types in the multidimensional data vector group; the first parameter in the multidimensional data vector group... The standardized data vectors for class parameters are derived from the multidimensional data vector group obtained after splitting the standardized water quality parameter sequences according to parameter type, corresponding to the first parameter. A vector of class parameters, which consists of data of the same parameter type arranged in chronological order.

[0104] Furthermore, the significance of this formula lies in realizing the feature fusion of multidimensional data vector groups. By multiplying the standardized data vector of each type of parameter in the multidimensional data vector group with the diagonal matrix formed by the corresponding correction coefficient, the vector of each type of parameter after weight adjustment is obtained. Then, all the weight-adjusted vectors are added together to form a new vector. This new vector is the result of feature fusion. It integrates the information of various parameters and reflects the difference in importance of different parameters through the correction coefficient, so that the fused result can better reflect the overall water quality characteristics of the water supply network.

[0105] Furthermore, judging from the trend of the formula, when the correction coefficients in the diagonal matrix formed by the correction coefficients of a certain type of parameter increase, the vector value obtained by multiplying the standardized data vector of that type of parameter with the diagonal matrix will increase accordingly. In the process of adding all vectors, the influence of that type of parameter on the feature fusion result will be enhanced. Conversely, when the correction coefficients corresponding to a certain type of parameter decrease, the vector value obtained by multiplying its standardized data vector will decrease, and the influence on the feature fusion result will also weaken. If all correction coefficients increase or decrease at the same ratio, the overall value of the feature fusion result will increase or decrease at the same ratio, and the influence ratio of different parameter types on the result remains unchanged.

[0106] In summary, from a data integration perspective, the standardized water quality parameter sequence includes multidimensional water quality parameters (such as pH, turbidity, and residual chlorine) from multiple distributed monitoring points in the water supply network. These parameters each have their own emphasis when reflecting water quality conditions. By constructing a diagonal weight matrix through correction coefficients and performing weighted fusion on the multidimensional data vector groups, the value of different parameters can be integrated according to their actual contribution. Parameters with high measurement accuracy and strong timeliness (after hydraulic residence time correction) occupy a higher weight in the fusion result, thereby highlighting the influence of key information and weakening the interference of low-value data, making the fusion result more focused on the core characteristics of the network water quality.

[0107] In summary, from the perspective of feature optimization, the weighted fusion data, after feature dimension compression, effectively reduces data redundancy by forming a fused feature vector. Water quality parameter data from water supply networks often exhibit high dimensionality and strong coupling, making direct analysis potentially increasing computational complexity and increasing susceptibility to noise. The fused feature vector, by retaining key information and eliminating invalid redundancy, transforms scattered parameter data into a structured feature set, simplifying subsequent analysis and enhancing the correlation between features and the actual water quality status of the network.

[0108] In summary, the fused feature vectors comprehensively reflect the correlation between time and space dimensions: on the one hand, based on a standardized sequence aligned to time, it ensures data consistency across time; on the other hand, by incorporating the influence of network spatial characteristics such as hydraulic residence time through weight correction, the feature vectors more realistically reflect the distribution and variation patterns of water quality parameters within the network. This highly integrated feature vector provides a unified and high-quality analytical framework for subsequently constructing node connection matrices and tracing anomalies, significantly improving the accuracy and efficiency of water quality anomaly feature identification.

[0109] S4. Construct the node connection matrix of the water supply network based on the topological structure data of the fused feature vectors;

[0110] In this embodiment of the invention, constructing the node connection matrix of the water supply network based on the topological data of the fused feature vector includes:

[0111] Extract the topological connection feature parameters of the network node identifiers from the fused feature vector;

[0112] A directed connection rule is established based on the fluid transmission direction of the nodes in the water supply network, and the topology connection characteristic parameters are converted into connection strength values ​​based on the directed connection rule.

[0113] A two-dimensional relational grid is constructed according to a preset node sorting. The connection strength value is mapped to the corresponding grid position of the two-dimensional relational grid, and zero-value parameters are filled between grid positions of nodes that are not directly connected in the two-dimensional relational grid to obtain the node connection relation moment of the water supply network.

[0114] Specifically, when extracting the topology connection feature parameters of the network node identifiers in the fused feature vector, the field storing the network node identifiers is first found in the data structure of the fused feature vector. This field contains the unique code of each node. Then, the adjacent nodes corresponding to each node identifier are determined through the network topology diagram. The number of connections between each node and its adjacent nodes is counted. The length and diameter of the connecting pipes between nodes are measured. The position coordinates of the nodes in the network are recorded. This information is summarized and organized to form topology connection feature parameters containing node identifiers, a list of adjacent nodes, the number of connections, pipe length, pipe diameter, and position coordinates. Each parameter is bound to and stored with the corresponding network node identifier.

[0115] Furthermore, when establishing directed connection rules based on the fluid transmission direction of nodes in the water supply network, and converting the topology connection characteristic parameters into connection strength values ​​based on the directed connection rules, the fluid transmission direction between each node is first determined through the hydraulic calculation report of the water supply network, i.e., from which node the water flows to which node. Based on this, directed connection rules are established: if node A transmits fluid to node B, then the connection from A to B is a valid directed connection, and the connection from B to A is an invalid connection. The strength of a valid directed connection is directly proportional to the pipe diameter and inversely proportional to the pipe length. The strength of an invalid connection is zero. Then, by referring to the adjacent node list, pipe length, and pipe diameter in the topology connection characteristic parameters, calculations are performed for each valid directed connection. First, the basic strength value is determined based on the pipe diameter; the larger the diameter, the higher the basic strength value. Then, the basic strength value is divided by the pipe length, and the result is the connection strength value of the valid directed connection. Invalid connections are directly marked with a connection strength value of zero.

[0116] Further, a two-dimensional relationship grid is constructed according to a preset node sorting. The connection strength values ​​are mapped to the corresponding grid positions in the two-dimensional relationship grid, and zero-value parameters are filled between grid positions of nodes that are not directly connected in the two-dimensional relationship grid. When obtaining the node connection relationship matrix of the water supply network, the nodes are first sorted according to the numerical size of the network node identifiers to form a preset order from the first node to the last node. The number of rows and columns of the two-dimensional relationship grid is determined according to the total number of nodes, and the number of rows and columns are equal to the total number of nodes. The rows and columns of the grid correspond to the nodes arranged in the preset order. Then, the connection strength values ​​of each node are traversed to find the row position of the node in the preset order and the column position of its adjacent nodes in the preset order. The corresponding connection strength values ​​are filled into the grid position where the row and column intersect in the two-dimensional relationship grid. After completing the mapping of all connection strength values, all unfilled grid positions in the two-dimensional relationship grid are checked. There are no direct connections between the nodes corresponding to these positions, and zero-value parameters are filled into these positions. The final two-dimensional relationship grid containing connection strength values ​​and zero-value parameters is the node connection relationship matrix of the water supply network.

[0117] In summary, from the perspective of topological feature transformation, although the fused feature vector integrates multi-dimensional water quality information, it needs to be combined with the actual physical structure of the pipeline network to accurately locate the source of anomalies. By extracting the topological connection feature parameters of the pipeline node identifiers in the fused feature vector, water quality characteristics can be correlated with physical attributes such as the spatial location and connection method of the nodes. Based on the fluid transmission direction of nodes in the water supply pipeline network, directed connection rules are established, and the topological connection feature parameters are transformed into connection strength values. This allows for the quantification of water flow propagation relationships between different nodes—such as the influence intensity of upstream nodes on downstream nodes and the connectivity between nodes—transforming the originally ambiguous pipeline connection relationships into a computable numerical matrix.

[0118] In summary, from the perspective of the structural value of matrix construction, a two-dimensional relational grid is constructed by sorting nodes according to a preset order, and the connection strength values ​​are mapped to the corresponding grid positions. Nodes not directly connected are filled with zero values, resulting in a node connection relation matrix with a clear mathematical structure and physical meaning. This matrix not only fully preserves the topological information of the pipeline network (such as the upstream and downstream relationships of nodes and connection paths), but also reflects the degree of mutual influence between nodes through numerical connection strength, providing a standardized input carrier for subsequent anomaly tracing based on matrix operations.

[0119] In summary, the construction of this structured matrix allows for a direct correlation between water quality anomaly information in the fused feature vectors and the physical propagation path of the pipe network. Subsequent analysis can quickly locate the propagation chain of anomaly features through matrix traversal, clarifying the path and intensity of the anomaly's spread from the initial node to other nodes. Simultaneously, the matrix format facilitates efficient computer processing, improving the computational efficiency of anomaly tracing and ensuring accurate capture of spatial propagation patterns even in complex pipe network structures, providing crucial spatial correlation evidence for ultimately locating the pollution source.

[0120] S5. Based on the node connection relationship matrix, perform anomaly tracing on the fused feature vector to obtain the anomaly results of the water supply network;

[0121] In this embodiment of the invention, the step of performing anomaly tracing on the fused feature vector based on the node connection relationship matrix to obtain the anomaly results of the water supply network includes:

[0122] The fused feature vector is mapped to the corresponding node position of the node connection relationship matrix to obtain the node anomaly feature distribution map of the water supply network;

[0123] By traversing the node connection matrix along the reverse path of the water flow in the water supply network, the abnormal feature propagation chain of the water supply network is obtained.

[0124] Based on the gradient changes of the abnormal features of the nodes in the abnormal feature propagation chain, the initial abnormal triggering node of the water supply network is located.

[0125] Based on the characteristic correlation strength between the initial anomaly triggering node and the monitoring points in the middle and lower reaches of the water supply network, anomaly results of the water supply network are generated.

[0126] The step of locating the initial anomalous trigger node of the water supply network based on the gradient changes of the anomalous features of nodes in the anomalous feature propagation chain includes:

[0127] Obtain the baseline pattern of feature distribution under historical normal operating conditions;

[0128] Based on the difference between the abnormal feature distribution map and the feature distribution baseline pattern, an abnormal deviation heat map of the water supply network is generated.

[0129] Identify the region of maximum deviation in the abnormal deviation heatmap and associate the region of maximum deviation with the set of upstream nodes in the node connection matrix;

[0130] The connectivity continuity of the upstream node set is verified based on the reverse flow path of the water supply network to obtain the initial abnormal triggering node of the water supply network.

[0131] Specifically, when mapping the fused feature vector to the corresponding node position in the node connection relationship matrix to obtain the node anomaly feature distribution map of the water supply network, firstly, the feature value corresponding to each node is extracted from the fused feature vector. These feature values ​​reflect the water quality anomaly related information of the node. Then, the position index of each node in the node connection relationship matrix is ​​checked. This index corresponds one-to-one with the node identifier. Subsequently, the feature value of each node is accurately filled into the node position at the intersection of the row and column corresponding to that node in the node connection relationship matrix. After completing the mapping of all node feature values, the matrix is ​​visualized. Different colors are used to indicate the size of the feature values. The larger the feature value, the darker the color. This visually presents the distribution of the anomaly features of each node, forming the node anomaly feature distribution map of the water supply network.

[0132] Furthermore, when traversing the node connection matrix along the reverse water flow path of the water supply network to obtain the abnormal feature propagation chain of the water supply network, the forward transmission path of the water flow is first determined through the topology diagram of the water supply network, that is, the direction from the water source node to the terminal node. The reverse water flow path is the opposite direction from the terminal node to the water source node. Then, abnormal nodes with feature values ​​exceeding a preset threshold are found from the node abnormal feature distribution diagram. Starting from the abnormal node, the upstream node directly connected to it is found in the node connection matrix according to the reverse water flow path. The upstream node is recorded. Then, the upstream node found is used as the new starting point to continue to find its corresponding upstream node along the reverse water flow path. This process is repeated until the water source node is traced back. All nodes involved in the search process are arranged in traversal order. The resulting sequence is the abnormal feature propagation chain of the water supply network.

[0133] Furthermore, based on the gradient changes of the abnormal features of nodes in the abnormal feature propagation chain, when locating the initial abnormal triggering node of the water supply network, the feature values ​​of each node in the abnormal feature propagation chain in the node abnormal feature distribution map are first extracted. These feature values ​​are recorded sequentially according to the arrangement order of the nodes in the propagation chain. Then, the difference between the feature values ​​of two adjacent nodes is calculated. If the feature value of the next node (the node closer to the water source) is smaller than the feature value of the previous node, and this decreasing trend continues until the feature value of a certain node no longer decreases but begins to increase, then the node whose feature value no longer decreases is the source of the abnormal feature, i.e., the initial abnormal triggering node. If the feature value in the propagation chain continues to increase from the starting point, then the starting point of the propagation chain is the initial abnormal triggering node.

[0134] Furthermore, when generating the abnormal results of the water supply network based on the characteristic correlation strength between the initial abnormal triggering node and the mid-to-downstream monitoring points of the water supply network, firstly, all monitoring points downstream of the initial abnormal triggering node are determined. These monitoring points constitute the mid-to-downstream monitoring point set. Then, the overlap between the characteristic value of the initial abnormal triggering node and the characteristic value of each mid-to-downstream monitoring point is calculated. The higher the overlap, the stronger the characteristic correlation strength between the two. Subsequently, the number of mid-to-downstream monitoring points with characteristic correlation strength exceeding the set standard and the distribution range of these monitoring points in the network are counted. Combining the location of the initial abnormal triggering node, the size of the characteristic value, and the associated mid-to-downstream monitoring points, a report containing the abnormal starting location, the scope of influence, and the abnormal characteristic manifestations is formed. This report is the abnormal result of the water supply network.

[0135] Specifically, when obtaining the feature distribution benchmark pattern under historical normal operating conditions, the system first retrieves all operation records of the water supply network marked as normal operating conditions in the past year from the database. These records contain the feature values ​​of each node and the corresponding timestamps. Then, it filters out the record segments that have no abnormal alarms for three consecutive months. The feature values ​​of each node during this period are averaged in chronological order to obtain the typical feature values ​​of each node under normal operating conditions. Then, according to the node arrangement order of the node connection relationship matrix, these typical feature values ​​are filled into the corresponding node positions to form a fixed feature distribution template. This template is the feature distribution benchmark pattern under historical normal operating conditions. When storing the template, the time range and data volume of the record selection are included.

[0136] Furthermore, when generating the heat map of abnormal deviation of the water supply network based on the difference between the abnormal feature distribution map and the feature distribution benchmark pattern, the feature value of each node in the abnormal feature distribution map is first compared with the typical feature value of the corresponding node in the feature distribution benchmark pattern. The difference between the feature value in the abnormal feature distribution map and the typical feature value in the feature distribution benchmark pattern is subtracted from the feature value in the abnormal feature distribution map. The result is the difference of the node. A positive difference indicates that the feature value is higher than the normal benchmark, and a negative difference indicates that it is lower than the normal benchmark. Then, the levels are divided according to the absolute value of the difference. The larger the absolute value, the higher the level. Then, a corresponding color is assigned to each level. The highest level is represented by dark red, the lowest level by light red, and the difference is represented by white when it is zero. Finally, the color corresponding to each node is filled into the same grid position as the node connection relationship matrix to form a heat map of abnormal deviation of the water supply network that intuitively shows the degree of deviation of each node from the normal state.

[0137] Furthermore, when identifying the region of maximum deviation in the abnormal deviation heatmap and associating it with the upstream node set in the node connection matrix, first observe the darkest area in the abnormal deviation heatmap, which is the region of maximum deviation. Determine all node identifiers contained in this region using grid coordinates. Then, examine the node connection matrix to find the rows corresponding to these nodes. The nodes corresponding to the columns with non-zero values ​​in each row are the upstream nodes of that node. Collect all the identifiers of these upstream nodes, remove duplicate identifiers, and form a set. This set is the upstream node set in the node connection matrix associated with the region of maximum deviation. At the same time, record the connection relationship between each upstream node and the nodes in the region of maximum deviation.

[0138] Furthermore, when verifying the continuity of the upstream node set based on the reverse flow path of the water supply network to obtain the initial anomalous triggering node of the water supply network, the specific direction of the reverse flow path is first obtained from the topology diagram of the water supply network, that is, the sequential connection order from the terminal node to the water source node. Then, a node is randomly selected from the upstream node set as the starting point, and the upstream node directly connected to it is searched along the reverse flow path. It is checked whether the upstream node is also in the node set. If it is, the tracing continues forward. If it is not, the starting point node is marked as a breakpoint. This process is repeated to check all nodes in the upstream node set. Finally, a subsequence of nodes that can be continuously traced along the reverse flow path and always belong to the upstream node set is selected. The node at the front of this subsequence that cannot be traced forward is the initial anomalous triggering node of the water supply network. If there are multiple such nodes, the node most directly associated with the difference component of the maximum deviation area is selected as the final result.

[0139] In summary, from the perspective of spatial mapping of abnormal features, mapping the fused feature vector to the corresponding node positions in the node connection matrix can generate a node abnormal feature distribution map, directly linking abstract water quality abnormal features with specific pipeline node locations. This mapping transforms the water quality abnormal information contained in the fused feature vector into a visualized spatial distribution, intuitively showing which nodes are abnormal and to what extent, providing a clear spatial reference for source tracing analysis.

[0140] In summary, tracing the path of anomaly propagation by reversing the water flow along the node connection matrix of the water supply network can construct an anomaly propagation chain. The quantified connection strength values ​​in the node connection matrix clearly reflect the direction and degree of fluid transmission between nodes. Based on this, reverse tracing can accurately reconstruct the process of anomaly propagation from downstream monitoring points to upstream sources, clarify the diffusion path of anomalies in the network and the associated impact of each node, and avoid deviations in the tracing direction caused by ignoring the network topology.

[0141] In summary, regarding the accurate localization of initial anomaly nodes, a heatmap of anomaly deviation is generated based on the gradient changes of anomaly features in the anomaly feature propagation chain, combined with the feature distribution benchmark pattern under historical normal operating conditions. This heatmap quickly identifies the region of maximum deviation and associates it with the upstream node set. The continuity of the node set is then verified through the reverse flow path, ultimately determining the initial anomaly trigger node. This process fully utilizes the topological logic of node connections in the matrix and the gradient change patterns of anomaly features, eliminating interference from non-initial nodes and making the localization of initial anomaly nodes more accurate and reliable.

[0142] In summary, the abnormal results obtained through the above steps not only clarified the spatial distribution and propagation path of the abnormality in the pipeline network, but also accurately identified the initial abnormality triggering node, providing a direct and reliable basis for subsequent pollution source location and early warning signal generation, and significantly improving the accuracy and efficiency of water quality abnormality tracing.

[0143] S6. Locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal.

[0144] In this embodiment of the invention, the step of locating the pollution source in the water supply network based on the abnormal result and generating a water quality early warning signal includes:

[0145] Extract the spatial coordinates of the initial anomaly triggering node from the anomaly results;

[0146] Based on the node connection relationship matrix, the pollution propagation path is traced to determine the candidate areas of pollution sources in the water supply network;

[0147] Based on the historical pollution event feature database and the spatial coordinates, the candidate areas are identified by risk level to obtain the pollution source location instructions for the water supply network;

[0148] Based on the results of the risk level identification, a preset early warning response strategy is matched to activate the multi-level water quality early warning signal of the water supply network.

[0149] Specifically, when extracting the spatial coordinates of the initial anomaly triggering node from the anomaly results, first find the unique identifier marked as the initial anomaly triggering node from the anomaly results, then open the node attribute database of the water supply network. Each record in this database contains information such as node identifier, latitude and longitude coordinates, and installation height. The node identifier is used to perform precise matching in the database. After finding the corresponding record, the latitude and longitude coordinate data is extracted from it. These data are the spatial coordinates of the initial anomaly triggering node. After extraction, the spatial coordinates and node identifier are recorded together in a dedicated positioning file.

[0150] Furthermore, when tracing the pollution propagation path based on the node connection relationship matrix and determining the candidate areas of pollution sources in the water supply network, the row corresponding to the initial abnormal triggering node is first found in the node connection relationship matrix. The nodes corresponding to the columns containing all non-zero values ​​in this row are the downstream nodes directly connected to the initial abnormal triggering node. These downstream nodes are taken as the first-level nodes of pollution propagation. Then, the row corresponding to each first-level node in the matrix is ​​checked in turn, and their downstream nodes are found as second-level nodes. This process is repeated step by step until all nodes that may be affected by pollution propagation are covered. Then, based on the location of these nodes in the geographical distribution map of the water supply network, a geographical range containing all nodes is delineated. This range is the candidate area of ​​pollution sources in the water supply network, and the area is marked with a dashed line on the geographical distribution map.

[0151] Furthermore, based on the historical pollution event feature database and the spatial coordinates, the candidate areas are identified by risk level. When the pollution source location command for the water supply network is obtained, the historical pollution event feature database is first opened. This database stores information such as the location coordinates, pollution type, and impact range of past pollution events. The distance between the spatial coordinates of the initial abnormal triggering node and the location coordinates of each pollution event in the database is calculated. Historical events within 500 meters are selected, and the number and severity of pollution in the area during these events are counted. The more times and the more severe the pollution, the higher the risk level. Subsequently, the risk level of the candidate areas is divided into high, medium, and low levels according to the statistical results, and the corresponding level is marked next to the geographical marker of the candidate areas. Finally, the information including the candidate area range, risk level, and spatial coordinates of the initial abnormal triggering node is integrated into a single command, which is the pollution source location command for the water supply network.

[0152] Furthermore, when activating the multi-level water quality early warning signal of the water supply network by matching the preset early warning response strategy based on the risk level identification results, the preset early warning response strategy document is first consulted. This document clearly specifies the early warning signal types and response measures corresponding to different risk levels. If the risk level of the candidate area is high, a red early warning signal is matched, and the strategy includes immediately closing the relevant valves in the area and activating the emergency water supply plan; if it is medium, a yellow early warning signal is matched, and the strategy includes increasing the frequency of water quality monitoring in the area and notifying surrounding users to pay attention to water safety; if it is low, a blue early warning signal is matched, and the strategy includes increasing the number of daily inspections and closely monitoring water quality changes. After determining the corresponding early warning signal, an activation command is sent to the early warning devices at all levels through the control system of the water supply network. A red early warning activates the audible and visual alarm devices in all areas, a yellow early warning activates the alarm devices in the area and adjacent areas, and a blue early warning only activates the display alarm in the control center, thereby activating the multi-level water quality early warning signal of the water supply network.

[0153] In summary, from the perspective of the accuracy of pollution source location, extracting the spatial coordinates of the initial anomaly triggering nodes from the anomaly results can directly pinpoint the core area of ​​the anomaly source. Tracing the pollution propagation path based on the node connection matrix can further clarify the scope and direction of pollution spread, thereby narrowing down the candidate areas of the pollution source. Combining the historical pollution event feature database to label the candidate areas with risk levels allows for the quantitative assessment of the pollution probability of the candidate areas using experience such as pollution patterns and diffusion laws from historical data. The final pollution source location command has clear spatial orientation and risk level classification, significantly reducing the cost of blind investigation and improving the efficiency and accuracy of pollution source location.

[0154] In summary, considering the timeliness and targeted nature of the early warning response, matching pre-set early warning response strategies with risk level identification results can activate multi-level water quality early warning signals. Different risk levels correspond to different early warning intensities and response plans. For example, high-risk levels can trigger emergency water outages and comprehensive investigations, while low-risk levels can initiate enhanced monitoring and localized disinfection plans. This tiered early warning mechanism avoids the limitations of a single early warning model, ensuring that the early warning signal matches the actual degree of pollution hazard. This enables relevant departments to quickly take targeted measures to control the spread of pollution in a timely manner, minimizing the impact of water quality anomalies on residents' drinking water safety, and providing dynamic and efficient protection for the safety of water supply networks.

[0155] like Figure 2 The diagram shown is a functional block diagram of a water quality anomaly tracing system for water supply networks based on data fusion, provided by an embodiment of the present invention.

[0156] The water quality anomaly tracing system 100 for water supply networks based on data fusion described in this invention can be installed in an electronic device. Depending on the functions implemented, the water quality anomaly tracing system 100 may include a water quality data extraction module 101, a weight coefficient correction module 102, a feature fusion module 103, a topology connection module 104, an anomaly tracing module 105, and a pollution early warning module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0157] In this embodiment, the functions of each module / unit are as follows:

[0158] The water quality data extraction module 101 is used to perform time alignment on the water quality parameter data of distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence of the water supply network.

[0159] The weight coefficient correction module 102 is used to allocate weight coefficients for the parameters according to the measurement accuracy of the parameters in the standardized water quality parameter sequence, and to correct the weight coefficients according to the hydraulic residence time in the water supply network, so as to obtain the correction coefficients of the water quality parameter data.

[0160] The feature fusion module 103 is used to perform weighted fusion of the standardized water quality parameter sequence based on the correction coefficient to obtain the fused feature vector of the water supply network.

[0161] The topology connection module 104 is used to construct the node connection relationship matrix of the water supply network based on the topology data of the fused feature vector;

[0162] The anomaly tracing module 105 is used to perform anomaly tracing on the fused feature vector based on the node connection relationship matrix to obtain the anomaly results of the water supply network.

[0163] The pollution early warning module 106 is used to locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal. In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of modules is only a logical functional division, and other division methods may exist in actual implementation.

[0164] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0165] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0166] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0167] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for tracing the source of water quality anomalies in a water supply network based on data fusion, characterized in that, The method includes: S1. Align the water quality parameter data of the distributed monitoring points in the water supply network with time to obtain the standardized water quality parameter sequence of the water supply network; S2. Based on the measurement accuracy of the parameters in the standardized water quality parameter sequence, assign weight coefficients to the parameters, and correct the weight coefficients according to the hydraulic retention time in the water supply network to obtain correction coefficients for the water quality parameter data, including: The formula for calculating the weighting coefficient is as follows: ; In the formula, The weighting coefficients are... This refers to the difference between the measured accuracy value of a parameter in the standardized water quality parameter sequence and the accuracy reference standard. This serves as a reference value for the accuracy of parameters in the standardized water quality parameter sequence; Obtain the topology data of the water supply network and extract the hydraulic residence time parameters corresponding to the distributed monitoring points; The ratio of the predetermined time decay benchmark value to the hydraulic residence time parameter is output as the time-dependent decay factor of the distributed monitoring point, including: The standard hydraulic residence time specified in the water supply network design code shall be used as the time decay benchmark value; Extract the hydraulic residence time parameter corresponding to each distributed monitoring point and compare it with the time decay benchmark value; If the hydraulic residence time parameter is greater than the time decay benchmark value, a value less than 1 is generated as the time decay factor. If the hydraulic residence time parameter is less than or equal to the time decay reference value, the time decay factor is directly set to 1; The weighting coefficients are dynamically compensated and adjusted based on the time-dependent decay factor to generate the correction coefficients; S3. The standardized water quality parameter sequence is split into multi-dimensional data vector groups according to parameter type; The multidimensional data vector group is weighted and fused based on the correction coefficient to obtain the fused feature vector of the water supply network. S4. Construct the node connection matrix of the water supply network based on the topological structure data of the fused feature vectors; S5. Based on the node connection relationship matrix, perform anomaly tracing on the fused feature vector to obtain the anomaly results of the water supply network; S6. Locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal.

2. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 1, characterized in that, The step of aligning the water quality parameter data from distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence for the water supply network includes: Receive water quality parameter data packets uploaded by various distributed monitoring points and extract the timestamp information from the data packets; The offset between the timestamp information and the standard time reference is calculated to generate time calibration parameters. The water quality parameter data packets are time-series rearranged according to the time calibration parameters to eliminate transmission delay differences; The time-series rearranged water quality parameter data are spatially correlated according to the location of pipeline nodes to generate the standardized water quality parameter sequence.

3. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 1, characterized in that, The weighted fusion of the standardized water quality parameter sequence based on the correction coefficient to obtain the fused feature vector of the water supply network includes: A diagonal weight matrix is ​​constructed using the aforementioned correction coefficients, and feature fusion is performed on the multidimensional data vector group based on the diagonal weight matrix. The calculation formula for the feature fusion is as follows: ; In the formula, The result of the feature fusion, The total number of types of the parameters. The type ordinal number of the parameter. For the first The class parameters are a diagonal matrix composed of corresponding correction coefficients. For the multidimensional data vector group, the first... Standardized data for class parameters; The results are then subjected to feature dimension compression to obtain the fused feature vector of the water supply network.

4. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 1, characterized in that, The step of constructing the node connection matrix of the water supply network based on the topological data of the fused feature vectors includes: Extract the topological connection feature parameters of the network node identifiers from the fused feature vector; A directed connection rule is established based on the fluid transmission direction of the nodes in the water supply network, and the topology connection characteristic parameters are converted into connection strength values ​​based on the directed connection rule. A two-dimensional relational grid is constructed according to a preset node sorting. The connection strength value is mapped to the corresponding grid position of the two-dimensional relational grid, and zero-value parameters are filled between grid positions of nodes that are not directly connected in the two-dimensional relational grid to obtain the node connection relation moment of the water supply network.

5. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 1, characterized in that, The step of tracing the anomalies in the fused feature vector based on the node connection matrix to obtain the anomaly results of the water supply network includes: The fused feature vector is mapped to the corresponding node position of the node connection relationship matrix to obtain the node anomaly feature distribution map of the water supply network; By traversing the node connection matrix along the reverse path of the water flow in the water supply network, the abnormal feature propagation chain of the water supply network is obtained. Based on the gradient changes of the abnormal features of the nodes in the abnormal feature propagation chain, the initial abnormal triggering node of the water supply network is located. Based on the characteristic correlation strength between the initial anomaly triggering node and the monitoring points in the middle and lower reaches of the water supply network, anomaly results of the water supply network are generated.

6. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 5, characterized in that, The step of locating the initial anomalous trigger node of the water supply network based on the gradient changes of the anomalous features of nodes in the anomalous feature propagation chain includes: Obtain the baseline pattern of feature distribution under historical normal operating conditions; Based on the difference between the abnormal feature distribution map and the feature distribution baseline pattern, an abnormal deviation heat map of the water supply network is generated. Identify the region of maximum deviation in the abnormal deviation heatmap and associate the region of maximum deviation with the set of upstream nodes in the node connection matrix; The connectivity continuity of the upstream node set is verified based on the reverse flow path of the water supply network to obtain the initial abnormal triggering node of the water supply network.

7. The method for tracing the source of water quality anomalies in a water supply network based on data fusion as described in claim 6, characterized in that, The step of locating the pollution source in the water supply network based on the abnormal results and generating a water quality early warning signal includes: Extract the spatial coordinates of the initial anomaly triggering node from the anomaly results; Based on the node connection relationship matrix, the pollution propagation path is traced to determine the candidate areas of pollution sources in the water supply network; Based on the historical pollution event feature database and the spatial coordinates, the candidate areas are identified by risk level to obtain the pollution source location instructions for the water supply network; Based on the results of the risk level identification, a preset early warning response strategy is matched to activate the multi-level water quality early warning signal of the water supply network.

8. A water quality anomaly tracing system for a water supply network based on data fusion, used to implement the water quality anomaly tracing method for a water supply network based on data fusion as described in claim 1, the system comprising: The water quality data extraction module is used to perform time-alignment of water quality parameter data from distributed monitoring points in the water supply network to obtain a standardized water quality parameter sequence for the water supply network. The weighting coefficient correction module is used to assign weighting coefficients to the parameters based on the measurement accuracy of the parameters in the standardized water quality parameter sequence, and to correct the weighting coefficients based on the hydraulic retention time in the water supply network, thereby obtaining correction coefficients for the water quality parameter data, including: The formula for calculating the weighting coefficient is as follows: ; In the formula, The weighting coefficients are... This refers to the difference between the measured accuracy value of a parameter in the standardized water quality parameter sequence and the accuracy reference standard. This serves as a reference value for the accuracy of parameters in the standardized water quality parameter sequence; Obtain the topology data of the water supply network and extract the hydraulic residence time parameters corresponding to the distributed monitoring points; The ratio of the predetermined time decay benchmark value to the hydraulic residence time parameter is output as the time decay factor of the distributed monitoring point. The weighting coefficients are dynamically compensated and adjusted based on the time-dependent decay factor to generate the correction coefficients; The feature fusion module is used to split the standardized water quality parameter sequence into multi-dimensional data vector groups according to parameter type; The multidimensional data vector group is weighted and fused based on the correction coefficient to obtain the fused feature vector of the water supply network. The topology connection module is used to construct the node connection relationship matrix of the water supply network based on the topology data of the fused feature vector; An anomaly tracing module is used to trace anomalies in the fused feature vector based on the node connection relationship matrix to obtain the anomaly results of the water supply network. The pollution early warning module is used to locate the pollution source in the water supply network based on the abnormal results and generate a water quality early warning signal.

Citation Information

Patent Citations

  • Pollution traceability analysis method and system based on water quality sampling data

    CN120524875A

  • A water hammer protection device for water supply pipe network

    CN221034597U