Water quality trend prediction system based on big data
By constructing a concentration sequence shift module, a matching similarity extraction module, a lag interval derivation module, and a ratio trend construction module, the problem of insufficient dynamic modeling in water quality trend prediction in existing technologies is solved. This enables the identification of concentration transmission characteristics and trend prediction in complex environments, thereby improving prediction accuracy and scheduling efficiency.
Patent Information
- Application Number
- CN202510554349.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing water quality trend prediction technologies are unable to perform continuous dynamic modeling of concentration propagation processes, lack the ability to judge the correlation of differences in concentration sequences over time, resulting in insufficient trend deconstruction capabilities in complex environments, difficulty in effectively judging trend inflection points or reversals, and lack of identification of structural evolutionary relationships among multiple variables.
By constructing a concentration sequence offset module, a matching similarity extraction module, a lag interval derivation module, and a ratio trend construction module, the continuous time series of pollutant concentrations at each cross-section in the water conveyance path is obtained. An error sliding window is established and the difference is accumulated to identify the stable time period of concentration transmission. The ratio of concentration decrease rate per unit distance to flow velocity is calculated to generate a concentration decay ratio change sequence and construct a downstream predicted concentration trend curve.
It enhances the ability to identify the dynamic transmission path of pollution signals, strengthens the ability to predict complex water quality fluctuations in multiple stages and the support effect of scheduling strategies, and improves the accuracy and efficiency of early warning.
Smart Images

Figure CN120470334B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of water quality analysis, and particularly relates to a water quality trend prediction system based on big data. BACKGROUND
[0002] As an important non-engineering measure to ensure water safety, water quality analysis and prediction model has been a popular research field in the industry for a long time. The current mainstream water quality analysis and prediction technologies include mathematical statistics method, numerical simulation method, and deep learning method. The mathematical statistics method is simple in principle and easy to implement, but the prediction accuracy is limited due to the non-linear problem of water quality change; the numerical simulation method is based on the principles of hydrodynamics and conservation of mass to construct a water quality model, analyze the migration, transformation and degradation process of pollutants in water body, and has high prediction accuracy, but the model has many parameters, large amount of calculation, and needs a large amount of measured data for calibration and verification; the deep learning method is to learn from historical data to establish the mapping relationship between environmental quantity and pollutant concentration change, which can solve complex nonlinear problems, but the interpretability is relatively poor, and the prediction accuracy will be reduced for extreme cases beyond the historical experience category. In order to cope with sudden water pollution incidents, it is necessary to further study the prediction algorithm with higher timeliness, stable and relatively high prediction accuracy, to automatically calibrate the model parameters according to the real-time collected monitoring data such as water depth, flow rate, pollutant concentration, etc. in engineering application, and dynamically predict the evolution trend of water pollution incidents to provide technical support for command and decision-making.
[0003] The technical field of water quality analysis includes detection and analysis of physical, chemical and biological indicators in water body, aiming to evaluate the water quality state and its change trend, and the core content is to collect water samples, quantitatively or qualitatively analyze the components, and judge whether there are harmful substances in the water body and whether the water body meets the corresponding use standard. The indicators covered by water quality analysis usually include pH value, dissolved oxygen, turbidity, conductivity, ammonia nitrogen, total phosphorus, total nitrogen, heavy metals, etc. According to different detection methods, it can be divided into specific methods such as optical analysis, electrochemical analysis, and chromatographic analysis. Water quality analysis is widely used in environmental monitoring, water resource management, water treatment engineering and other scenes, and is a basic technical link in environmental protection and public health protection.
[0004] The water quality trend prediction system based on big data refers to a technical scheme for analyzing and judging the water quality change trend by using historical water quality data and related environmental data, combining data modeling and prediction analysis methods, collecting water quality monitoring data of nodes such as water sources along the line, water channels, reservoirs and supporting regulation facilities, and establishing a water quality database containing time series, geographical position and monitoring indexes. Subsequently, according to the change law of various pollutants in water, multivariate statistical analysis and time series modeling method are used to process the monitoring data, so as to complete the water quality trend analysis and prediction of key sections and key nodes. Based on historical monitoring data, combined with the characteristics of water transfer process and regional environmental factors, trend modeling and dynamic deduction of key water quality parameters are carried out.
[0005] In the existing water quality trend prediction process, based on static parameter detection, it is difficult to carry out continuous dynamic modeling in the concentration propagation process, and there is a blind area in identifying the response period when dealing with water quality mutation. The single index evaluation method is often used, ignoring the structural evolution relationship between multivariate, resulting in insufficient trend disintegration ability in complex environment. The monitoring data focuses on the change of numerical value itself, lacks dynamic correlation matching mechanism between water quality trend and hydraulic transport characteristics, and it is difficult to reflect the response structure of water quality trend to flow velocity and other hydrodynamic factors. The traditional prediction mode mainly depends on time extension, lacks trend direction division, resulting in that the prediction curve lacks change logic support, and it is difficult to realize effective judgment before the trend inflection point or reversal. In the face of multiple sections and complex nodes in the water transfer line, the existing method is difficult to construct the concentration conduction chain of the whole path, and does not have the ability of segment trend analysis and multi-time lag interval positioning. In the application scene with increasing risk identification demand, this kind of technical mode is easy to cause prediction delay, misjudgment or omission of key trend stage, reduce the regulation efficiency and early warning reliability. SUMMARY
[0006] The purpose of the present application is to solve the problems existing in the prior art, and to provide a water quality trend prediction system based on big data.
[0007] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows: the water quality trend prediction system based on big data comprises:
[0008] The concentration sequence offset module obtains the continuous time sequence of the pollutant concentration of each section in the water transfer path, obtains the difference sequence of the offset concentration value and the upstream concentration value and carries out error calculation, establishes an error sliding window and carries out difference accumulation in the sliding process of the error sliding window, counts and matches the error index, and establishes a concentration trend offset sample set;
[0009] The matching similarity extraction module extracts the concentration trend offset sample set, screens all time window positions with error levels lower than a concentration difference value identification threshold, extracts overlapping regions according to frequency of occurrence in a time period, and obtains a concentration conduction stable time period;
[0010] The lag interval derivation module identifies an offset value range based on the offset step sequence in the concentration conduction stable time period, matches the corresponding offset value interval with the concentration conduction time window, analyzes a probabilistic conduction time length range, and obtains a stable time lag offset interval.
[0011] The ratio trend construction module obtains concentration values and spatial intervals of sampling points in each section of the water delivery path, calculates a unit distance concentration decrement rate, combines a flow rate sensor reading, calculates a unit distance concentration decrement rate and flow rate value ratio, records a continuous change section identified according to a ratio change direction, and generates a concentration attenuation ratio change sequence.
[0012] As a further scheme of the present application, the concentration trend offset sample set includes a time offset parameter, an error accumulation feature, and a concentration difference value window index, the concentration conduction stable time period includes a time period grouping number, a frequency sorting result, and an overlapping interval label, the stable time lag offset interval includes an offset value interval range, a probabilistic response interval, and an accumulated frequency statistical result, and the concentration attenuation ratio change sequence includes a concentration decrement ratio sequence, a ratio change direction identifier, and a continuous trend section number.
[0013] As a further scheme of the present application, the concentration sequence offset module includes:
[0014] The concentration difference value generation submodule sets a time offset step based on a continuous time sequence of pollutant concentrations at each section of the water delivery path, performs a step-by-step reverse sliding operation on a downstream concentration sequence, obtains a difference value sequence of downstream offset concentration values and upstream concentration values at the same time window at each time offset step, and obtains a concentration difference value sequence corresponding to all offset steps after sequence synchronization processing.
[0015] The error sliding window construction submodule sets a time window length based on the concentration difference value sequence and a difference value item at each position, divides the difference value sequence into sub-sections according to the time window, and uses the formula:
[0016] ;
[0017] Calculate the average deviation value of each time window position , construct a time sliding window with a fixed window length, cover the error value sequence with the window, and obtain an error sliding window sequence, wherein, represents the upstream and downstream concentration difference value at time in the th window, and Representing the downstream and upstream respectively in the first stage The first window Concentration value at time, This represents the path response factor at that moment. The time step scale at that moment. This represents the total number of time steps within the window.
[0018] The matching error accumulation submodule performs an error term accumulation operation based on the error sliding window sequence. According to the total error at each sliding window position, it compares the cumulative error of the sliding window at all time offset steps with the error benchmark value, selects the minimum cumulative error value, marks the corresponding time offset step size, organizes the corresponding concentration sequence, error value sequence and time index according to the offset step size, and establishes a concentration trend offset sample set.
[0019] As a further aspect of the present invention, the matching similarity extraction module includes:
[0020] The error screening submodule, based on the concentration trend offset sample set, filters out positions in all time windows where the error level is lower than the concentration difference identification threshold, and extracts the sequence number and time period information corresponding to each time window at different time steps to obtain the sequence time mapping set.
[0021] The time grouping submodule groups all time periods according to the start and end boundaries of the time periods based on the sequence time mapping set, counts the frequency of each time period in all sequence numbers, and sorts them according to the repetition of time periods to obtain the frequency sorted time set.
[0022] The overlap extraction submodule compares the time intervals in pairs based on the frequency sorting time set to determine whether there is a boundary intersection. If there is an intersection, it records the intersection interval range of the corresponding time interval, integrates all intersection intervals, and establishes a stable concentration conduction time period.
[0023] As a further aspect of the present invention, the lag interval derivation module includes:
[0024] The step length frequency statistics submodule calculates the frequency of each offset step length in all time periods based on the offset step length sequence in the concentration conduction stable time period, obtains the total number of occurrences corresponding to each step length value, and performs numbering and frequency alignment processing on all offset step lengths to generate a step length frequency distribution set.
[0025] The lag interval identification submodule, based on the step size frequency distribution set, sorts all offset step size values by frequency, progressively accumulates the frequency values of the preceding values, calculates the accumulated frequency, and compares it with the set time lag threshold using the following formula:
[0026] ;
[0027] Determine the set of all step sizes that satisfy the condition that the cumulative frequency is greater than a threshold. Generate a range of hysteresis offset values, where, Indicates the first Each offset step value Indicates the first Frequency of occurrence of each step size Indicates the first Frequency of occurrence of each step size This represents the total number of step size types. Indicates the time lag threshold;
[0028] The response mapping matching submodule extracts the corresponding concentration conduction time window position based on the hysteresis offset value range, analyzes the mapping relationship between each offset value and the upstream and downstream concentration sequences corresponding to the time period, determines the duration of the response corresponding to the upstream concentration change in the downstream concentration sequence, integrates all response time period ranges, and establishes a stable time hysteresis offset interval.
[0029] As a further aspect of the present invention, the ratio trend construction module includes:
[0030] The decay rate calculation submodule obtains the concentration values and spatial spacing of each sampling point in the water conveyance path, calculates the unit distance concentration decay rate between adjacent sampling points in turn, records the concentration change and corresponding spatial interval of each pair of sampling points in different time periods, and obtains the decay rate sequence.
[0031] The ratio generation submodule, based on the decrease rate sequence and the corresponding flow rate sensor readings for each time period, uses the following formula:
[0032] ;
[0033] Calculate the first Concentration to flow rate ratio over a time period The information on the direction of change in the ratio is summarized to generate a concentration-flow-rate ratio sequence, in which... This represents the concentration decrease rate per unit distance in the corresponding segment. The interval between the sampling points of the corresponding segment. This is the flow velocity reading for the corresponding section. This represents the flow flux in the corresponding sampling area. The sampling time interval span of the corresponding segment. The sampling frequency for the corresponding flow velocity segment;
[0034] The trend extraction submodule analyzes the continuous change direction in the ratio sequence based on the concentration-flow-rate ratio sequence, identifies continuous segments with consistent directions as change segments within the trend segment, numbers each segment, locates its start and end points and marks its change direction, and establishes a concentration decay ratio change sequence.
[0035] As a further scheme of the present application, the system further comprises a trend data output module;
[0036] The trend data output module extracts a time range in which the stable time lag offset interval and the concentration decay ratio change sequence coincide, and judges whether there is a trend increasing section in the range, and if so, marks it as a future fluctuation potential conduction interval, and recursively extends the concentration values in each prediction window to construct a downstream predicted concentration trend curve.
[0037] The downstream predicted concentration trend curve comprises the coinciding time range, the trend increasing section label, the potential conduction interval range, and the predicted concentration curve.
[0038] As a further scheme of the present application, the trend data output module comprises:
[0039] The interval intersection submodule compares the time intervals of the stable time lag offset interval and the concentration decay ratio change sequence, identifies the intersection region of the respective start and end times, extracts the continuous time interval information of the coinciding part, and establishes a time overlap section set.
[0040] The trend identification submodule judges whether there is a continuous increasing section in the ratio change direction in each time interval based on the time overlap section set, judges the ratio direction in sections, extracts and numbers the ratio increasing trend sections, and establishes a predicted trend section set.
[0041] The concentration recursion submodule recursively extends the downstream concentration values in each trend increasing section according to the predicted trend section set, performs concentration deduction processing in time sequence according to the current concentration value and the ratio increasing trend, and draws a downstream predicted concentration trend curve.
[0042] Compared with the prior art, the present application has the following advantages and positive effects:
[0043] In the application, by constructing the time difference sliding mechanism between the upstream and downstream concentration sequences, the time sequence offset law in the water quality conduction characteristics can be captured without relying on the fixed sampling frequency, the identification ability of the dynamic conduction path of the pollution signal is expanded, the coupling analysis of the difference error sliding window and the time step improves the distinguishability of the concentration response period, the stable conduction interval has stronger statistical basis, the offset step sequence frequency is aggregated to form a lag interval, the high probability concentration response area is located in the spatial and temporal multi-dimensional scene, the prediction ability of the pollution change influence range is strengthened, the concentration decreasing rate and the flow rate ratio are constructed, the material migration trend and the hydrodynamic factor are dynamically coupled, the attenuation trend is no longer limited to a single spatial function relationship, but has a concentration evolution trajectory reflecting the flow characteristics, the identification of the change direction of the ratio trend makes the trend sequence have the ability to extract adaptive change segments, provides structured basis for critical judgment of trend enhancement and decline, finally through the extraction of the trend increasing segment in the overlapping time period and the concentration value extension, a multi-dimensional condition screening system for trend prediction is formed, the accuracy of early warning of potential risk signals is improved, and the multi-segment prediction ability and scheduling strategy support effect of complex water quality fluctuations are effectively enhanced. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The system flowchart of the application is shown in the figure;
[0045] Figure 2 The concentration sequence offset module flowchart of the application is shown in the figure;
[0046] Figure 3 The matching similar extraction module flowchart of the application is shown in the figure;
[0047] Figure 4 The lag interval derivation module flowchart of the application is shown in the figure;
[0048] Figure 5 The ratio trend construction module flowchart of the application is shown in the figure;
[0049] Figure 6 The trend data output module flowchart of the application is shown in the figure. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical scheme and advantages of the application clearer and more apparent, the application will be further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.
[0051] In the description of the present application, it should be understood that the terms "length", "width", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.
[0052] Please refer to Figure 1 The water quality trend prediction system based on big data comprises:
[0053] The concentration sequence offset module acquires the continuous time sequence of the pollutant concentration of each section in the water conveyance path, sets a time offset step, performs a step-by-step reverse sliding operation on the downstream concentration sequence, acquires the difference value sequence of the offset concentration value and the upstream concentration value under the same time window, performs error calculation on the difference value sequence, establishes an error sliding window according to the time position of the error value, accumulates the difference value during the sliding process of the error sliding window, counts the matching error index, and establishes a concentration trend offset sample set;
[0054] The matching similarity extraction module filters all time window positions with an error level lower than a concentration difference identification threshold according to the concentration trend offset sample set, extracts the sequence number and time period of the corresponding window under each time step, groups and clusters the time period, extracts the overlapping area according to the frequency of occurrence in the time period, and obtains a concentration conduction stable time period;
[0055] The lag interval derivation module counts the occurrence frequency corresponding to each step based on the offset step sequence in the concentration conduction stable time period, identifies the offset value range with a cumulative occurrence frequency exceeding a time lag threshold, matches the corresponding offset value interval with the concentration conduction time window, analyzes the probabilistic conduction time length range of the upstream concentration change on the downstream response, and obtains a stable time lag offset interval;
[0056] The ratio trend construction module acquires the concentration values and spatial intervals of each sampling point in the water conveyance path, calculates the unit distance concentration decrement rate between each pair of sampling points, combines the flow rate sensor readings corresponding to each period, calculates the unit distance concentration decrement rate and flow rate value ratio of each time period, records the ratio change direction of all time periods, identifies the continuous change section in the trend section, and generates a concentration attenuation ratio change sequence;
[0057] The trend data output module extracts a time range in which the stable time lag offset interval and the concentration attenuation ratio change sequence coincide, and determines whether there is a trend increasing section in the range. If there is, it is marked as a potential conduction interval of future fluctuations, and the concentration values in each prediction window are recursively extended to construct a downstream predicted concentration trend curve.
[0058] The concentration trend offset sample set includes a time offset parameter, an error accumulation feature, and a concentration difference value window index. The concentration conduction stable time period includes a time period grouping number, a frequency sorting result, and an overlapping interval label. The stable time lag offset interval includes an offset value interval range, a probabilistic response interval, and a cumulative frequency statistical result. The concentration attenuation ratio change sequence includes a concentration decreasing ratio sequence, a ratio change direction identifier, and a continuous trend section number. The downstream predicted concentration trend curve includes a coincident time range, a trend increasing section label, a potential conduction interval range, and a predicted concentration curve.
[0059] Please refer to Figure 2 The concentration sequence offset module includes:
[0060] The concentration difference value generation submodule sets a time offset step based on the continuous time sequence of pollutant concentrations at each section in the water conveyance path, performs a step-by-step reverse sliding operation on the downstream concentration sequence, obtains the difference value sequence of the downstream offset concentration value and the upstream concentration value at the same time window for each time offset step, and after sequence synchronization processing, obtains the concentration difference value sequence corresponding to all offset steps;
[0061] Based on the continuous time series of pollutant concentration of each section in the water delivery path, the water delivery path is first segmented by grid, a concentration monitoring point is set at each section, and time series data is recorded. For example, assuming that three sections A, B, and C are set in a certain river basin, respectively at positions 0 km, 5 km, and 10 km, and the monitoring time series length is 24 hours, the pollutant concentration value is recorded once an hour, after the sequence is obtained, the time offset step is set to 1 hour, indicating that each downstream sequence is shifted forward by 1 time unit, during the sliding process, the downstream concentration value at each offset step is synchronized and aligned to form the upstream and downstream concentration comparison set in the same time window, for example, for offset step 1, the concentration value of point C at hour 2 is compared with the concentration value of point B at hour 1, on this basis, step by step sliding is performed to construct a concentration difference value set corresponding to multiple offset steps, the difference value is calculated by directly taking the algebraic difference of the concentrations of the two sections, for example, if the concentration of point B at hour 1 is 2.3 mg / L and the concentration of point C at hour 2 is 2.9 mg / L, then the difference value at this offset step is 0.6 mg / L, and so on to form a complete difference value sequence, attention should be paid to the time alignment and length uniformity of the data during processing, for example, if the downstream data is missing at a certain time, the position should be excluded from the calculation to ensure that errors are not introduced, finally, the sliding generation is completed under all offset steps, and the concentration difference value sequence corresponding to each offset step is obtained to form the concentration difference value sequence.
[0062] The error sliding window construction submodule is based on the concentration difference value sequence, based on the difference value at each position, sets the time window length, divides the difference value sequence into sub-sections according to the time window, uses the formula:
[0063] ;
[0064] Calculate the average deviation value at each time window position , construct a time sliding window with a fixed window length, cover the error value sequence with the window, get the error sliding window sequence, where represents the upstream and downstream concentration difference value at time in the th window, and represent the concentration values of the downstream and upstream at time in the th window, is the path response factor at this time, is the time step scale at this time, is the total number of time steps in the window;
[0065] The time step scale refers to the actual time interval corresponding to each time step in the discretization calculation process, which determines the time granularity of each simulation step during model running and has an important influence on the stability and accuracy of the model results. In models based on actual monitoring data, the time step scale is usually directly determined by the data sampling frequency. For example, if the data is collected every second, the time step scale is one second; if it is collected every minute, the time step scale is one minute, ensuring that the time step is consistent with the observation.
[0066] According to the concentration difference value sequence, the sliding window length is set to 3 hours, the difference value sequence is segmented, each sub-section contains 3 adjacent time point difference items, for example, the first sliding window covers the difference value from the 1st to the 3rd hour, the second sliding window covers the difference value from the 2nd to the 4th hour, and so on, in each sliding window, the , downstream concentration , upstream concentration , path response factor and time step scale corresponding values are extracted, the path response factor value is set based on the flow velocity change rate between sections and the reflection degree of transport lag, specifically, the value is derived from the normalized weighted result of the flow velocity change coefficient and the section reaction delay time in the hydrodynamic monitoring data, assuming that the flow velocity of B-C section is 1.2 m / s at the 1st hour, 0.95 m / s at the 2nd hour, and 1.1 m / s at the 3rd hour within 3 hours of monitoring, and the corresponding pollutant transport delay time is 20 minutes, 30 minutes, and 25 minutes respectively, according to the above index, the ratio of flow fluctuation rate and lag time is calculated and normalized by time, then the path response factors of the 1st, 2nd, and 3rd hours are 0.8, 0.6, and 0.7 respectively, and the time step scale is set to one hour, i.e. , the concentration values are , , and so on, the specific calculation is performed by substituting the data in Table 1 below:
[0067] Table 1 Error sliding window parameter table
[0068]
[0069] Substitute the formula:
[0070] ;
[0071] The calculation is as follows:
[0072] ;
[0073] ;
[0074] ;
[0075] ;
[0076] ;
[0077] ;
[0078] ;
[0079] The results show that the average error under the sliding window position is 0.94 mg / L, which is recorded as an item in the error value sequence, and the error value of all sliding window positions is constructed to obtain the error sliding window sequence.
[0080] The operation logic of the formula aims to comprehensively evaluate the influence of upstream and downstream concentration changes and path characteristics on the error in each time window. First, the difference term represents the difference between the upstream and downstream concentrations at time in the th sliding window, which is directly used as the basis for error evaluation; second, the term reflects the trend of coordinated changes between concentrations by taking the square root of the product of upstream and downstream concentrations. If both concentrations rise or fall simultaneously, the product will be larger, and the root result will increase, reflecting the convergence in the transport process; third, the term represents the ratio of the path response factor to the time step size, which is used to quantify the time-varying influence of the path transport conditions, where characterizes the degree of concentration response of the path at that time, reflects the time span at that time, and the larger the ratio, the more intense the path response, the larger the error correction amplitude, and the entire formula reflects the average error level formed by the concentration difference under each offset step after considering the path factors and transport changes through the addition and subtraction between the difference term, the coordinated change term, and the path response term, and the summation and averaging within the window, thereby providing a quantitative basis for subsequent sliding window matching and error judgment.
[0081] The matching error accumulation submodule performs error accumulation operations based on the error sliding window sequence. According to the total error under each sliding window position, the cumulative error of all time offset steps is compared with the error reference value, the minimum cumulative error value is selected, the corresponding time offset step is marked, and the corresponding concentration sequence, error value sequence, and time index are organized according to the offset step to establish a concentration trend offset sample set;
[0082] Based on the error sliding window sequence, the sum of error values within the sliding window range corresponding to each offset step is calculated one by one. For example, in offset step 1, the sliding window range is from the 1st to the 5th time point, with corresponding error values of [0.94, 0.82, 1.01, 0.90, 0.95] and a cumulative error of 5.62 mg / L. This value is compared with the set error benchmark value, which is set according to the historical stable fluctuation range of pollutants. Assuming that the benchmark is set to 6.0 mg / L based on the annual average fluctuation value of the measured data, then offset step 1 is a valid sample. Other offset steps are evaluated sequentially, and all offset steps with a cumulative error of less than 6.0 mg / L are selected. The corresponding concentration sequence fragment, error value sequence fragment, and offset step number are recorded together to construct a concentration trend offset sample. Each sample structure is {offset step number, concentration difference vector, cumulative error value}. The sample set is organized in a two-dimensional structure to form a concentration trend offset sample set.
[0083] Please see Figure 3 The similarity extraction module includes:
[0084] The error screening submodule, based on the concentration trend offset sample set, filters out positions in all time windows where the error level is lower than the concentration difference identification threshold, and extracts the sequence number and time period information corresponding to each time window at different time steps to obtain the sequence time mapping set.
[0085] Based on the concentration trend offset sample set, the positions with error level lower than the concentration difference identification threshold in all time window positions are screened, which first needs to traverse each time window in the sample set and extract the error value of the window at each offset step to determine whether it is less than the identification threshold 0.12 mg / L (set according to the average fluctuation range of the standard deviation σ of the concentration of pollutants (such as ammonia nitrogen and total phosphorus) in the coastal area in multiple measured sections, the 1σ range of ammonia nitrogen concentration fluctuation in a typical section within 24 hours is usually not more than 0.10 mg / L, if the 95% confidence level is taken as the limit, the concentration fluctuation interval is ±1.96σ, and the corresponding value is 0.196 mg / L. Combined with the field background noise and error control requirements, 0.12 mg / L is selected as the offset judgment standard for removing disturbance type concentration anomalies. This value fluctuates with the background pollution level of the river, the flow velocity of the section and the time window span. In the section with high flow velocity and low concentration, the threshold value tends to decrease. In the low flow velocity and high background concentration area, the threshold value is appropriately floated. Therefore, 0.12 mg / L as the judgment reference value under the fluctuation background of medium water quality has reasonableness), if the error value is 0.08 mg / L, it is determined that the sample item meets the condition, and then the index number of the time window and its corresponding time period start and end point are recorded, for example, the start time is the 15th minute, the end time is the 45th minute, and the corresponding time period is [15, 45]. For different time steps, such as 5 minutes, 10 minutes or 15 minutes, the time period [15, 45] can be expressed as interval [10, 50] under 10 minute step. In practice, the following data records may be obtained: window number W1, corresponding error value 0.08 mg / L, time period [15, 45], step type 10 minutes, corresponding sequence number S3, sequence numbers S1 to S5 are generated under different step lengths, in order to ensure the integrity of the results, a mapping matrix between the time period and the sequence number is constructed, where the row represents the time period, the column represents the sequence number under different time steps, and the corresponding cell is the error value. When screening, the rows and columns of the cells that meet the conditions are extracted to obtain the sequence number and time period information set that meets the threshold requirement, which is the sequence time mapping set. In order to ensure the accuracy of screening, the concentration difference identification threshold needs to be explained, which can be based on the stable concentration fluctuation range of the same type of water body in the history measurement. The concentration difference under the 95% confidence level is calculated as ±0.12 mg / L. Therefore, the value is set as the identification threshold. If the error value in the sample exceeds ±0.12 mg / L, it is considered as an unstable section and is removed. For example, the following concentration difference value set {0.05, 0.08, 0.13, 0.07, 0.15} is screened, only 0.05, 0.08 and 0.07 are retained, and the rest is removed, so as to obtain the sequence time mapping set.
[0086] The time grouping submodule groups all time periods according to the start and end boundaries of the time periods according to the sequence time mapping set, counts the frequency of each time period in all sequence numbers, sorts according to the time period repetition degree, and obtains the frequency-sorted time set;
[0087] Based on the sequence time mapping set, all time periods are grouped according to the start and end boundaries of the time periods. In operation, the time period set needs to be sorted in ascending order according to the start time, and then it is determined whether there is time overlap or continuity between adjacent time periods. If there is a relationship between two time periods, for example, time period A is [10, 30] and time period B is [30, 50], they can be classified as the same group and form a combined segment [10, 50], and then the combined segment is counted in the frequency statistics. If a certain time period appears in three different sequence numbers, its frequency is 3. According to this logic, the frequency of all combined segments is counted to form a mapping relationship between the time period and its frequency. For example, if the time period [20, 40] appears in sequence numbers S2, S3 and S4, the frequency is 3, and if [50, 70] only appears in S1, the frequency is 1. After all the statistics are completed, all time period combinations are sorted in descending order according to the frequency value, such as the segment with a frequency of 5 is sorted first, and the segment with a frequency of 3 is sorted second, to form an ordered frequency-sorted time period set. To supplement the frequency information of each time period, a set of simulation data is shown in the following table, which lists the time period, frequency and covered sequence number:
[0088] Table 2 Frequency-sorted time period table
[0089] Time period (minutes) Frequency (times) Cover sequence number [10,40] 4 S1, S2, S3, S4 [20,50] 3 S2, S3, S4 [30,60] 2 S1, S3 [45,75] 1 S2
[0090] As shown in Table 2, by grouping and counting the sequence time mapping set, the frequency of each time period appearing in multiple sequences can be obtained, and the [10, 40] segment appears the most frequently, with a frequency of 4. In subsequent sorting, this segment is extracted first to obtain the frequency-sorted time set.
[0091] The overlap extraction submodule compares the time periods based on the frequency-sorted time set, judges whether there is an intersection between the boundaries, records the intersection interval range of the corresponding time period if there is an intersection, integrates all intersection intervals, and establishes the concentration conduction stable time period.
[0092] According to the frequency ranking time set, the time periods ranked in the front are compared with each other in two-way intersection interval, and whether there is boundary intersection is judged. The specific judgment method is to check whether the starting time of time period A is earlier than the ending time of time period B, and whether the ending time of A is later than the starting time of B. If both satisfy the condition, there is an overlapping interval, and the intersection part needs to be calculated and its start and end boundaries are recorded. For example, time period A is [10, 40] and B is [30, 60], and the intersection is [30, 40]. The intersection can be derived from the following formula: the starting point = max(A starting, B starting) = 30, the ending point = min(A ending, B ending) = 40, so the intersection interval is [30, 40]. The intersection interval is recorded and added to the overlapping interval set. If there are continuous overlapping conditions in multiple time periods, further interval merging operation is performed to form a stable section. For example, [10, 40] and [30, 60] overlap to form [30, 40], and then overlap with [35, 55] to get [35, 40], and the final stable transmission section is [35, 40]. Finally, all intersection parts are merged to obtain a non-repeating time interval set, which is the time period set with stable transmission characteristics in the concentration time sequence, that is, the concentration transmission stable time period. The stable section reflects the time period when the concentration change between multiple offset steps and multiple sample sections remains small and highly repetitive, which can be used as a time reference benchmark for subsequent pollutant transport mechanism or trend modeling.
[0093] Please refer to Figure 4 , the lag interval derivation module comprises:
[0094] The step frequency statistical sub-module is based on the offset step sequence in the concentration transmission stable time period, and the total number of occurrences corresponding to each step value is obtained by counting the frequency of each offset step in all time periods. All offset steps are numbered and frequency aligned to generate a step frequency distribution set.
[0095] Based on the offset step sequence in the concentration conduction stable period, first, the offset step recorded in each stable period is counted segment by segment, for example, set a total of 5 concentration conduction stable time periods, which are [120, 130], [150, 165], [180, 190], [200, 215] and [230, 240], respectively, and the offset step sequences recorded in these segments are as follows: [3, 4, 4, 5, 6], [4, 5, 6, 6, 7], [3, 3, 4, 5], [4, 5, 6, 6], [3, 4, 5], all the offset steps are summarized and counted to form the offset step and frequency control data, and the frequency value is obtained by normalizing the number of occurrences; for example, offset step 3 occurs 5 times, step 4 occurs 6 times, step 5 occurs 6 times, step 6 occurs 5 times, and step 7 occurs 1 time, a total of 23 times, so the frequency of step 3 is 5 / 23≈0.217, and the frequencies of other steps are obtained in the same way, see Table 3.
[0096] Table 3 Step frequency distribution table
[0097] Offset step Number of occurrences Frequency 3 5 0.217 4 6 0.261 5 6 0.261 6 5 0.217 7 1 0.043
[0098] As shown in Table 3, the table summarizes the statistical number and frequency of different offset steps in all stable time periods, and all frequency values are normalized values, satisfying the condition that the sum is 1. Threshold setting is not involved in this sub-module, so there is no need to introduce a judgment interval, and no external parameters need to be calculated. The result is the step frequency distribution set, which will be used as the input parameter for the next step analysis.
[0099] The hysteresis interval identification submodule is based on the step frequency distribution set, and all offset step values are sorted by frequency, and the frequency values of the previous items are accumulated step by step, and the cumulative frequency is calculated and compared with the set time hysteresis threshold, using the formula:
[0100] ;
[0101] Determine all step sets that satisfy the cumulative frequency greater than the threshold , generate the hysteresis offset value interval, where, represents the th offset step value, represents the th step frequency, represents , is the total number of step types, represents the time hysteresis threshold;
[0102] According to the step frequency distribution set, the frequency values in Table 1 are arranged in ascending order of offset step, and the cumulative frequency judgment operation is performed, that is, the frequency is accumulated from step 3 in turn until the cumulative frequency first reaches or exceeds the time lag threshold, which is 0.75 in this case. The threshold is set to extract the interval range with the largest cumulative impact to reflect the time lag response. The value is set according to the cumulative frequency statistical results of all offset steps in the concentration conduction stable time period in the early stage. Among them, by analyzing the cumulative frequency change trend of the interval corresponding to the step with the smallest concentration difference and stable error distribution, when the cumulative frequency accounts for 75% of the total frequency, the error value variance presents a rapid downward trend and is stably maintained below 0.02, indicating that the frequency ratio forms the main contribution range in the concentration conduction path. Therefore, τ = 0.75 is determined as the reasonable threshold value under this system. The specific setting of the time lag threshold τ depends on three key parameters: first, the change amplitude of the concentration difference sequence before and after the lag interval, specifically the concentration fluctuation amplitude decreases by more than 50%; second, the error sum of squares (SSE) decreases slowly when the cumulative frequency reaches 75%, with a decrease ratio of less than 5%; third, the response amplitude of the path response factor decreases before and after the threshold, and the contribution rate of the first 75% frequency to the system response is not less than 90%. Through the interactive analysis of these three parameters, τ = 0.75 is determined as the reasonable limit of the transmission lag response in the system. In addition, this value will fluctuate with different concentration time series, path length and sampling step number. Especially when the path length is lengthened or the concentration change rate is accelerated, the τ value usually presents an upward trend, while in short path or slow change system, the τ value tends to the lower limit of 0.70. According to the test results of the comprehensive model, the τ value fluctuates between [0.70, 0.80], and the default value is 0.75, but when the step sample exceeds 30 groups or the measurement period is greater than 100, the τ value can be adjusted appropriately. The current value is determined based on the sequence group number of 5 groups and the conduction structure characteristics of the offset step sample of 23 groups; calculate according to the formula:
[0103] Item 1 (step 3): cumulative frequency = 0.217;
[0104] Item 2 (step 4): cumulative frequency = 0.217 + 0.261 = 0.478;
[0105] Item 3 (step 5): cumulative frequency = 0.478 + 0.261 = 0.739;
[0106] Item 4 (step 6): cumulative frequency = 0.739 + 0.217 = 0.956;
[0107] When accumulated to the 4th item, the cumulative frequency exceeds 0.75, so steps 3 to 6 are taken as intervals meeting the condition, that is, the offset value interval [3, 6] is extracted as the lag offset value interval, which will be used in subsequent mapping of the concentration conduction time window, indicating that the offset step in this interval is the dominant delay factor of the conduction process. The results show that the offset steps are concentrated between [3, 6], which have a dominant effect on the system response.
[0108] The operation logic of the formula is to determine which interval in the offset step set has a dominant contribution to the concentration conduction by step-by-step accumulation of frequency values. The formula uses The operator accumulates the frequency of each step in order to obtain the cumulative value of all step frequencies before the current offset step position, which reflects the overall influence intensity in this step range; the denominator is the sum of all step frequencies, that is, the normalized reference total amount, so that the cumulative frequency value becomes a relative proportion, which varies between 0 and 1; this ratio is compared with the lag threshold , which essentially determines whether a certain offset step range has accumulated to the system's acceptable delay interpretation threshold, so the formula as a whole establishes a screening judgment relationship between the step frequency distribution and the threshold by summing and forming a ratio, maintains the monotonicity and physical interpretability of the judgment standard, and ensures the stability and reliability of the results.
[0109] The response mapping matching submodule extracts the corresponding concentration conduction time window position according to the lag offset value interval, analyzes the upstream and downstream concentration sequence mapping relationship corresponding to each offset value and time period, judges the duration of the corresponding response of the upstream concentration change in the downstream concentration sequence, integrates all response time period ranges, and establishes a stable time lag offset interval;
[0110] Based on the lag offset value interval, all time period intervals with step lengths of 3 to 6 are obtained, for example, if the time positions of the offset step length of 3 in time period 1 are 125-127, the time positions of the offset step length of 4 in time period 2 are 152-155, the time positions of the offset step length of 5 in time period 3 are 182-186, and the time positions of the offset step length of 6 in time period 4 are 205-208, these time positions are recorded and compared with the upstream and downstream concentration change sequences to determine the delay degree of the upstream concentration change peak position (for example, the upstream concentration changes at t=120, 150, 180, and 200 seconds) between the start time of the downstream concentration response. If the upstream concentration changes at t=120 and the downstream response detects the corresponding increase at t=125, it is considered that the conduction delay is 5 seconds. After repeating such operations in each sample, a plurality of sets of concentration response delay lengths are obtained. By integrating and summarizing the delay length values, it is concluded that the common conduction length range is [5, 8] seconds. This range will correspond to the lag offset value interval [3, 6] to be linked and marked, and a stable time lag offset interval is established as a time index for concentration response conduction, indicating that there is a unified conduction logic within the identified offset interval, and the statistically significant delay effect can be effectively characterized.
[0111] Please refer to Figure 5 The ratio trend construction module includes:
[0112] The decreasing rate calculation submodule obtains the concentration values and spatial intervals of the sampling points in each section of the water conveyance path, sequentially calculates the concentration decreasing rate per unit distance between adjacent sampling points, records the concentration changes and corresponding spatial intervals of each pair of sampling points in different time periods, and obtains a decreasing rate sequence.
[0113] To obtain the concentration values and spatial intervals of the sampling points in each section of the water conveyance path, a plurality of sampling points need to be arranged along the water conveyance path, for example, sampling points A, B, and C are arranged at the starting point, the middle section, and the end, respectively. Assuming that the distance between A and B is 80 meters and the distance between B and C is 120 meters, at t1, the sampling concentrations of A, B, and C are 8.6 mg / L, 7.2 mg / L, and 6.1 mg / L, respectively. After recording the spatial intervals, the concentration changes of adjacent sampling point pairs are calculated, wherein the concentration change between A and B is 8.6−7.2=1.4 mg / L, and the concentration decreasing rate per unit distance is 1.4÷80=0.0175 mg / L·m⁻¹. The concentration change between B and C is 7.2−6.1=1.1 mg / L, and the decreasing rate is 1.1÷120=0.0092 mg / L·m⁻¹. Further extending to multiple time periods, the concentration values at t2 and t3 are collected to repeat the calculation, forming a data set containing the concentration decreasing rates between each pair of sampling points in different time periods. All the decreasing rate information is recorded and labeled for the time sequence to form a concentration attenuation data table for each pair of sampling points in each time period, as shown in the following table:
[0114] Table 4. Concentration Decrease Rate per Unit Distance at Sampling Points
[0115]
[0116] As shown in Table 4, the decrease rate for each pair of sampling points is based on the same formula. Perform the calculation, where The concentration at the sampling point. The distance between two sampling points is used as the basis for constructing subsequent ratios, resulting in a decrease rate sequence.
[0117] The ratio generation submodule, based on the decrease rate sequence and the corresponding flow rate sensor readings for each time period, uses the following formula:
[0118] ;
[0119] Calculate the first Concentration to flow rate ratio over a time period The information on the direction of change in the ratio is summarized to generate a concentration-flow-rate ratio sequence, in which... This represents the concentration decrease rate per unit distance in the corresponding segment. The interval between the sampling points of the corresponding segment. This is the flow velocity reading for the corresponding section. This represents the flow flux in the corresponding sampling area. The sampling time interval span of the corresponding segment. The sampling frequency for the corresponding flow velocity segment;
[0120] The flow flux of the sampling area is set based on the flow rate variation characteristics determined by the combined flow velocity and cross-sectional area of the fluid within the sampling area. The flux magnitude is dynamically adjusted according to the channel cross-sectional size and flow velocity fluctuations, and is usually used to reflect the contribution of a local area to the overall flow intensity. The sampling time interval is set based on the typical response time range during fluid transmission, and is determined by balancing the sampling frequency and data stability. The interval value varies with the data acquisition cycle, fluid change rate, and model synchronization requirements to ensure the representativeness and completeness of the sampling results within each time period. The flow velocity sampling frequency is set based on the dynamic characteristics of flow velocity changes and the sensor's response capability. The frequency directly affects the accuracy of capturing flow details, and its value is adjusted according to the flow velocity fluctuation amplitude, data timeliness requirements, and equipment sampling performance to ensure a reasonable match between sampling density and system load.
[0121] After obtaining the decline rate sequence, the flow velocity sensor readings for each time period are paired, and the flow velocity information corresponding to the sampling segment is collected. For example, at time t1, the flow velocity corresponding to segment AB is 0.42 m / s and segment BC is 0.39 m / s. The flow velocity readings and decline rate values for the same time period are recorded, and the ratio is calculated using a formula.
[0122] Wherein take t1, A-B segment as an example, known mg / L·m⁻¹, m, m / s, m³ / s (corresponding to the flow of the measurement area), s (as a single cycle time span), (sampling frequency), then put into the formula as follows:
[0123] ;
[0124] Continue to calculate the rest of the time period and sampling segment, get multiple ratio results, build a complete ratio change sequence, and record the ratio increase or decrease trend of each time period, for example, the ratio increases from 0.0486 to 0.0532 from t1 to t2, which indicates an increasing trend, mark each trend direction as +1 (up) or -1 (down), and establish the concentration flow rate ratio sequence.
[0125] The operation logic of the formula is to fuse multiple physical quantities to measure the sensitivity change of concentration decay to flow rate disturbance in the water delivery path, wherein the decrement rate represents the concentration loss per unit distance, multiplied by the distance reflects the total amount of actual concentration change in this time period, which is used as the numerator part of the ratio to quantify the decay amplitude, while the denominator part is composed of the flow rate value , flow flux , sampling time span , sampling frequency , reflects the direct influence of fluid motion, Through square root processing of the regional flow, the stretching effect of the ratio value caused by the sharp difference of the flow between different measurement areas is alleviated, represents the data collection time length of each sampling point, the larger the ratio, the longer the data coverage time, which increases the system stability, but also reduces the instantaneous concentration response sensitivity, so it is subtracted as an adjustment term, and the denominator is combined as a whole to form a dynamic flow rate disturbance factor of concentration change, which is used to balance the decay amplification or weakening caused by flow rate variation, so that the ratio Comprehensively reflects the change trend of concentration decrement under the current hydraulic condition.
[0126] The trend extraction submodule is based on the concentration flow rate ratio sequence, analyzes the continuous change direction in the ratio sequence, identifies the continuous segment with consistent direction as the change section in the trend segment, numbers, locates and marks the change direction of each section, and establishes the concentration decay ratio change sequence;
[0127] After obtaining the concentration flow rate ratio value sequence, the trend of the ratio value change is analyzed along the time axis, for example, the corresponding ratio value directions at t1, t2 and t3 are +1, +1 and -1, respectively, indicating that the trend is rising from t1 to t2 and falling from t2 to t3, the time period combination with consistent direction is identified as a continuous trend section, and a continuous trend section information set is constructed, and the minimum continuous duration is set to two sampling periods in the identification process, if the continuous same direction section is less than two periods, it is not included in the trend section statistics, if the condition is met, the section is marked as a trend section, for example, t1 and t2 are rising sections, t2-t3 is not counted, t4-t6 is a falling section, and is recorded as a falling trend section, the trend extraction of all time periods is completed and numbered, and finally a sequence of all trend sections is formed, arranged according to the time period, and the concentration decay ratio value change sequence is obtained. The advantage of the formula is that by introducing flow flux, sampling frequency, time span and other multi-parameter coupling terms, the concentration gradient change under nonlinear flow disturbance is standardized, and the stability of the ratio sequence to continuous trend identification is improved.
[0128] See Figure 6 , the trend data output module comprises:
[0129] The interval intersection submodule compares the time intervals of the stable time lag offset interval and the concentration decay ratio value change sequence, identifies the intersection area of the start and end time of each other, extracts the continuous time period information of the overlapping part, and establishes a time overlap section set;
[0130] According to the stable time lag offset interval and the concentration decay ratio value change sequence, the time intersection of the two types of time periods is processed, first, each section in the stable time lag offset interval is represented as a pair of start and end time stamps, for example, {t1, t2}, and the start and end time of each trend section in the concentration decay ratio value change sequence is represented as a trend section structure, the comparison method adopts a double-pointer traversal strategy, and the offset interval and the ratio trend interval are arranged in ascending order of time, and the overlap relationship of the start and end time of each other is compared pair by pair, if t1^A≤t2^B and t2^A≥t1^B, it is marked as an overlapping section, if there is a partial intersection, the time stamp range of the intersection section is reconstructed and stored, in the actual scene, assuming that the offset interval is {8:00, 10:30}, and the ratio section is {9:00, 11:00}, the intersection of the two is {9:00, 10:30}, and the intersection section is extracted for subsequent judgment; if multiple overlapping sections exist, they are numbered and recorded respectively, to prepare data for subsequent trend identification, the time stamp index needs to be maintained in each matching process, so as to establish an accurate time overlap mapping relationship, in the actual example, the concentration and ratio data collected from the monitoring system are one group every 5 minutes, and the table is recorded as follows:
[0131] Table 5 Time overlap section table
[0132]
[0133] As shown in Table 5, the time overlap section can be directly constructed by time range comparison, and the constructed time overlap section set is used as the input basis for subsequent trend increasing identification.
[0134] The trend identification submodule judges whether there is a continuous increasing section in the change direction of the ratio in each time section based on the time overlap section set, judges the section of the ratio direction, extracts and numbers the ratio increasing trend section, and establishes a prediction trend section set;
[0135] According to the time overlap section set, it is judged whether there is a ratio continuously increasing paragraph in each overlap section. The specific operation steps are as follows: first, the concentration and flow rate ratio sequence in each time section is extracted, and the corresponding time index is constructed according to the sampling frequency. In this embodiment, the sampling frequency is 5 minutes once, and the corresponding time section length is 1 hour, so 12 groups of data should be included. The ratio sequence B={0.78, 0.81, 0.86, 0.89, 0.85, 0.87, 0.90, 0.94, 0.91, 0.92, 0.95, 0.98} is constructed. The difference between adjacent values is calculated, and the length of the continuous increasing section is judged. If there is a continuous increasing section with a length greater than 3, it is considered that there is a trend increasing segment, and the marking process is continued. During the execution process, a micro-increase judgment threshold ε should be set for the increasing judgment operation to avoid the increasing caused by the small fluctuation. The setting basis of the threshold is the difference order between the flow rate sampling accuracy and the concentration change stability in the actual water delivery path. Specifically, in this embodiment, the flow rate sampling accuracy is 0.01 m / s, and the concentration change is often disturbed by background noise and sensor response error when it is less than 0.005 mg / L. Therefore, a threshold higher than the basic error term but lower than the mutation critical value is needed for micro-increase identification. Finally, ε=0.02 is set, which is between the concentration sampling resolution (0.005 mg / L) and the typical fluctuation upper limit (0.03 mg / L), and can produce interval adaptation ability through the difference of concentration baseline level in different concentration interval, such as in the medium concentration interval 0.3~0.7 mg / L, the micro-increase amplitude above 0.02 mg / L can represent the trend change, and below this value it is treated as background fluctuation. Therefore, the value of ε, the sampling interval Δt, the flow rate stable interval σv and the concentration sampling stable interval σc together determine the change trend, which fluctuates with the sampling accuracy and the concentration difference distribution. In this example, the first to fourth groups of data constitute an increasing trend section, the fifth group is not counted, and the sixth to eighth groups form an increasing segment again. These sections are numbered respectively to obtain the numbering sequence S1={1-4}, S2={6-8}. Finally, the mapping relationship between these continuous sections and the original time index is connected to generate the prediction trend section set. In the evaluation process, a minimum length threshold of the increasing trend section is also set based on the total section length, for example, 15 minutes. If it is lower than the value, the section is removed to avoid misjudgment caused by short-term disturbance. After this process, the prediction trend section set that meets all the conditions can be obtained.
[0136] The concentration recursive sub-module performs hour-by-hour extension processing on the downstream concentration values in each trend increasing section according to the prediction trend section set. According to the current value of the concentration value and the ratio increasing trend, the concentration deduction processing is performed in time sequence, and the downstream predicted concentration trend curve is drawn.
[0137] Based on the predicted trend section set, the downstream concentration values in each trend increasing section are processed by time extension, and the specific operation steps are as follows: firstly, the starting concentration value C0 in the trend section is extracted, and the corresponding ratio change rate r is calculated, then the concentration value C1=C0+r*Delta t at the next time is calculated according to the current sampling time interval Delta t, and the complete trend sequence is constructed by recursive calculation, in the actual example, the starting time concentration is 0.42mg / L, Delta t is 5 minutes, and r is 0.008mg / L*min-1, then the first step concentration is C1=0.42+0.008*5=0.46mg / L, the second step is 0.46+0.008*5=0.50mg / L, and so on, until the end of the complete trend section, all the calculated concentration values are generated into a trend sequence according to the time index, and the starting and ending time are recorded, so as to construct the downstream predicted concentration trend curve, and in order to ensure the stability of the data, it is necessary to judge whether the ratio is in the reasonable growth interval during the extension process, generally, the maximum concentration growth rate threshold is set to 0.01mg / L*min-1, if it exceeds the value, the recursive calculation is interrupted, and the fluctuation section is avoided from being amplified too much, and finally the complete downstream predicted concentration trend curve is drawn.
[0138] The above is only the preferred embodiment of the present application, and does not limit the present application in other forms, any skilled person in the art can change or modify the above disclosed technical content into equivalent embodiments applied to other fields, but any simple modification, equivalent change and modification made on the basis of the technical essence of the present application to the above embodiments without departing from the technical solution content of the present application still belongs to the protection scope of the present application technical solution.
Claims
1. A water quality trend prediction system based on big data, characterized by, The system comprises: The concentration sequence offset module obtains the continuous time sequence of pollutant concentration of each section in the water delivery path, obtains the difference sequence of the offset concentration value and the upstream concentration value, and performs error calculation, establishes an error sliding window, and accumulates the difference in the sliding process of the error sliding window, counts the matching error index, and establishes a concentration trend offset sample set; The matching similar extraction module screens all time window positions with error level lower than the concentration difference identification threshold according to the concentration trend offset sample set, sorts and extracts the overlapping area according to the frequency in the time period, and obtains the concentration conduction stable time period; The lag interval derivation module identifies the offset value range based on the offset step sequence in the concentration conduction stable time period, matches the corresponding offset value interval with the concentration conduction time window, analyzes the probabilistic conduction time length range, and obtains the stable time lag offset interval; The ratio trend construction module obtains the concentration value and the spatial interval of each sampling point in the water delivery path, calculates the unit distance concentration decrement rate, combines the flow rate sensor reading, calculates the unit distance concentration decrement rate and the flow rate value ratio, records the change direction of the ratio to identify the continuous change section, and generates a concentration attenuation ratio change sequence.
2. The big data based water quality trend prediction system as claimed in claim 1, wherein, The concentration trend offset sample set comprises a time offset parameter, an error accumulation feature, and a concentration difference value window index, the concentration conduction stable time period comprises a time period grouping number, a frequency sorting result, and an overlapping interval label, the stable time lag offset interval comprises an offset value interval range, a probabilistic response interval, and a cumulative frequency statistical result, and the concentration attenuation ratio change sequence comprises a concentration decrement ratio sequence, a ratio change direction identifier, and a continuous trend section number.
3. The big data based water quality trend prediction system as claimed in claim 1, wherein, The concentration sequence offset module comprises: The concentration difference value generation submodule sets a time offset step based on the continuous time sequence of pollutant concentration of each section in the water delivery path, performs a step-by-step reverse sliding operation on the downstream concentration sequence, obtains the difference sequence of the downstream offset concentration value and the upstream concentration value at the same time window at each time offset step, performs sequence synchronization processing, and obtains the concentration difference value sequence corresponding to all offset steps; The error sliding window construction submodule sets a time window length based on the concentration difference value sequence and the difference value item at each position, divides the difference value sequence into sub-sections according to the time window, and adopts the formula: ; Calculate the average deviation value of each time window position , construct a time sliding window with fixed window length, cover the error value sequence with the window, and obtain an error sliding window sequence, wherein, represents the upstream and downstream concentration difference at time t in the first window, is the path response factor at this time, is the time step scale at this time, is the total number of time steps in the window; The matching error accumulation submodule performs accumulation operation on the error items based on the error sliding window sequence, compares the sliding window cumulative error and the error reference value at each sliding window position, screens the minimum cumulative error value, marks the corresponding time offset step, organizes the corresponding concentration sequence, error value sequence, and time index according to the offset step, and establishes the concentration trend offset sample set.
4. The big data based water quality trend prediction system as claimed in claim 1, wherein, The matching similar extraction module comprises: The error screening submodule screens the positions with error level lower than the concentration difference identification threshold in all time window positions based on the concentration trend offset sample set, extracts the sequence number and time period information corresponding to each time window at different time steps, and obtains a sequence time mapping set. The time grouping submodule groups all time periods according to the start and end boundaries of the time periods according to the sequence time mapping set, counts the frequency of each time period in all sequence numbers, sorts the time periods according to the time period repetition degree, and obtains a frequency-sorted time set; The overlap extraction submodule judges whether there is an intersection of boundaries based on the frequency-sorted time set, records the intersection interval range of the corresponding time period if there is an intersection, integrates all intersection intervals, and establishes a concentration conduction stable time period.
5. The big data based water quality trend prediction system as claimed in claim 1, wherein, The lag interval derivation module comprises: The step frequency statistical submodule counts the frequency of each offset step in all time periods based on the offset step sequence in the concentration conduction stable time period, obtains the total number of occurrences corresponding to each step value, and aligns the numbers and frequencies of all offset steps to generate a step frequency distribution set; The lag interval identification submodule sorts all offset step values according to the frequency based on the step frequency distribution set, accumulates the frequency values of the previous items in the order step by step, calculates the cumulative frequency, compares it with the set time lag threshold value, and uses the formula: ; determining all sets of step lengths that satisfy a cumulative frequency greater than a threshold value generating a hysteresis offset value interval, wherein, represents the th offset step value, represents the frequency of occurrence of the th step length, represents the frequency of occurrence of the th step length, is the total number of step length classes, represents a time hysteresis threshold value; The response mapping matching submodule extracts the corresponding concentration conduction time window position according to the lag offset value interval, analyzes the upstream and downstream concentration sequence mapping relationship corresponding to each offset value and time period, judges the duration of the corresponding response of the upstream concentration change in the downstream concentration sequence, integrates all response time period ranges, and establishes a stable time lag offset interval.
6. The big data based water quality trend prediction system as claimed in claim 1, wherein, The ratio trend construction module comprises: The decreasing rate calculation submodule obtains the concentration values and spatial intervals of each segment of the water delivery path, calculates the concentration decreasing rate per unit distance between adjacent sampling points in sequence, and records the concentration change and corresponding spatial interval of each pair of sampling points in different time periods to obtain a decreasing rate sequence; The ratio generation submodule uses the formula: ; Computing the concentration-flow rate ratio value under the time period Concentration-flow rate ratio value under the time period , aggregating the ratio direction change information to generate a concentration-flow rate ratio sequence, wherein, is the concentration decrement rate per unit distance of the corresponding section, is the sampling point interval of the corresponding section, is the flow rate reading of the corresponding section, is the sampling area flux of the corresponding section, is the sampling time interval span of the corresponding section, is the flow rate sampling frequency of the corresponding section; The trend extraction submodule analyzes the continuous change direction in the concentration flow rate ratio sequence based on the concentration flow rate ratio sequence, identifies the continuous segments with consistent directions as the change segments in the trend segment, numbers, positions, and marks the directions of each segment, and establishes a concentration attenuation ratio change sequence.
7. The big data based water quality trend prediction system as claimed in claim 1, wherein, The system further comprises a trend data output module; The trend data output module extracts the time range of the overlap of the stable time lag offset interval and the concentration attenuation ratio change sequence, judges whether there is a trend increasing segment in the range, marks it as a potential conduction interval in the future if there is, and recursively extends the concentration values in each prediction window to construct a downstream predicted concentration trend curve; The downstream predicted concentration trend curve comprises the overlapping time range, the trend increasing segment label, the potential conduction interval range, and the predicted concentration curve.
8. The big data based water quality trend prediction system as claimed in claim 7, wherein, The trend data output module comprises: The interval intersection submodule compares the time intervals of the stable time lag offset interval and the concentration attenuation ratio change sequence, identifies the intersection area of the start and end times of each other, extracts the continuous time period information of the overlapping part, and establishes a time overlap segment set. The trend identification submodule judges whether there is a continuous increasing section in the change direction of the ratio in each time section based on the time overlap section set, judges the section of the ratio direction, extracts and numbers the ratio increasing trend section, and establishes a prediction trend section set; The concentration recursion submodule performs hourly extension processing on the downstream concentration value in each trend increasing section according to the prediction trend section set, performs concentration deduction processing in time sequence according to the current value of the concentration value and the ratio increasing trend, and draws a downstream predicted concentration trend curve.
Citation Information
Patent Citations
River water quality real-time monitoring platform
CN118052450A
Method and system for evaluating water quality change trend
CN118247673A