Data acquisition method and system based on multi-source water service sensors
By analyzing water data and influencing factors and dynamically adjusting the sensor collection frequency, the problems of resource waste and monitoring lag in traditional methods are solved, and efficient and intelligent management of the water system is achieved.
Patent Information
- Application Number
- CN202510874079.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Existing multi-source water sensor data acquisition methods and systems generate redundant data when the water system is operating stably, wasting resources. However, they are unable to capture key data in a timely manner under abnormal circumstances, resulting in monitoring lags and affecting the efficiency and safety of water management.
By analyzing historical water data and influencing factor data, the sensor collection frequency is dynamically adjusted, the abnormal probability is predicted using cycle and trend rules, and the correlation of influencing factors is combined to optimize the collection frequency to adapt to the dynamic changes of the water system.
It can increase the collection frequency when the probability of abnormality is high, obtain key data in time, reduce resource waste, improve the response capability and data accuracy of water management, and avoid water accidents.
Smart Images

Figure CN120373916B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more particularly to a method and system for collecting data based on multi-source water service sensors. Background Art
[0002] In modern water management, ensuring the efficient use of water resources, water quality safety, and stable operation of water supply systems is crucial. Multi-source water sensor data collection is a key means of obtaining water information. The accuracy, timeliness, and completeness of its data play a decisive role in water management decisions. By deploying various sensors, such as water level sensors, flow sensors, and water quality sensors, at key nodes such as water sources, water supply networks, and sewage treatment plants, important information such as water level changes, water flow rates, and water quality parameters can be monitored in real time, providing data support for water management.
[0003] However, the existing multi-source water sensor data acquisition methods and systems have certain limitations in practical applications. Traditional sensor data acquisition often adopts a fixed acquisition frequency, which fails to fully take into account the dynamic changes in the operating status of the water system. During periods when the water system is operating relatively stably and the probability of water anomalies is low, excessively high acquisition frequencies will lead to the generation of a large amount of redundant data, which not only increases the burden of data storage and transmission, but also consumes too much sensor power and network bandwidth resources, reducing the operating efficiency of the system. On the other hand, at critical moments when the water system may face abnormal conditions, such as heavy rain causing rapid rise in water levels or sudden water pollution, and the probability of water anomalies is high, a fixed low acquisition frequency cannot capture changes in key data in a timely manner, resulting in delayed monitoring of abnormal conditions and difficulty in taking effective response measures in a timely manner. This may cause serious water accidents and have adverse effects on residents' lives and the ecological environment.
[0004] Therefore, there is an urgent need to develop a multi-source water sensor data collection method and system that can dynamically adjust sensor collection frequency based on the probability of water anomalies. By intelligently setting sensor collection frequencies in different situations—reducing the frequency when the probability of anomalies is low to optimize resource utilization, and increasing it when the probability of anomalies is high to ensure data timeliness and accuracy—this approach will significantly improve the sophistication of water management and the ability to respond to emergencies, effectively avoiding resource waste and potential water risks, and driving water management towards more efficient and intelligent development. Summary of the Invention
[0005] In order to solve the problem of how to adaptively adjust the collection frequency according to the abnormal situation of water service data, the present invention proposes a data collection method and system based on multi-source water service sensors.
[0006] In a first aspect, the present invention provides a method for collecting data based on multi-source water service sensors, comprising:
[0007] Obtain each type of historical water affairs data at each historical moment and each type of historical influencing factor data at the alignment moment;
[0008] Splitting the time series sequence composed of the historical water service data into a period component and a trend component; calculating the correlation between the characteristic descriptors of all period segments of the period component and the abnormal probability as period correlation, obtaining the trend correlation, respectively obtaining abnormal probability prediction values for future time periods based on the periodic law and the trend law, using the periodic correlation and the trend correlation as weights of the corresponding abnormal probability prediction values, performing a weighted summation on the two abnormal probability prediction values to obtain a comprehensive abnormal probability prediction value for the future time period, wherein the future time period includes a plurality of future sub-periods;
[0009] The information gain between historical water affairs data and each historical influencing factor data is calculated as the abnormal correlation of each historical influencing factor data. The abnormal possibility of water affairs data in the first future sub-period is obtained based on the abnormal correlation and the co-occurrence of abnormal data in each influencing factor data and historical water affairs data. The comprehensive abnormal probability prediction value of the first future sub-period is corrected according to the abnormal possibility of water affairs data to obtain the final abnormal probability prediction value. The collection frequency of the first future sub-period is set according to the final abnormal probability prediction value. The water affairs data collection is controlled based on the collection frequency.
[0010] Preferably, the method for obtaining the feature descriptor of the periodic segment includes:
[0011] The product of the mean and variance of all historical water service data in a periodic segment is used as the feature descriptor of the periodic segment.
[0012] Preferably, the method for obtaining the abnormality probability includes:
[0013] Performing anomaly detection on the historical water service data to obtain abnormal data in the historical water service data, and recording the abnormal data in the historical water service data as abnormal water service data;
[0014] The abnormal probability of the periodic segment is obtained by dividing the number of abnormal water service data in the periodic segment by the total number of data in the periodic segment.
[0015] Preferably, the method for obtaining the trend correlation includes:
[0016] The trend component is segmented to obtain several trend segments;
[0017] Get the feature descriptor of the trend segment;
[0018] Get the abnormal probability of the trend segment;
[0019] The correlation between the characteristic descriptors of all trend segments in the trend component and the abnormal probability is taken as the trend relevance.
[0020] Preferably, the obtaining of abnormal probability prediction values for future time periods based on periodic laws and trend laws respectively includes:
[0021] The current cycle segment is matched with the previous cycle segments to obtain a matching value, and the cycle segment with a matching value greater than a preset threshold is used as an alternative reference segment; the next cycle segment of the alternative reference segment with the shortest interval with the current moment is used as the target cycle segment; the abnormal probability in the target cycle segment is used as the abnormal probability prediction value of the future time period based on the periodic law;
[0022] The least squares method is used to fit polynomials to the abnormal probabilities of all trend segments, and the abnormal probabilities of future periods are fitted using the fitted polynomials and recorded as the predicted values of the abnormal probabilities of future periods based on the trend law.
[0023] Preferably, the method for obtaining the abnormal correlation of each historical influencing factor data includes:
[0024] The abnormal information entropy of historical water data is calculated based on the probability of abnormal water data and the probability of non-abnormal water data; the influencing factor data are divided into several category layers; the probability of abnormal water data occurring and the probability of no abnormal water data occurring in the aligned moments of all moments of the influencing factor data in any category layer are recorded as the probability of abnormal water data and the probability of non-abnormal water data under the category layer, and the information entropy of the category layer is calculated based on the probability of abnormal water data and the probability of non-abnormal water data under the category layer; the probability of influencing factor data in each category layer is obtained and recorded as the probability of each category layer, and the information gain is calculated based on the probability of each category layer and the information entropy of each category layer as the abnormal correlation of each influencing factor data.
[0025] Preferably, obtaining the possibility of abnormality of water service data in the first future sub-period includes:
[0026] The category layer in which the predicted value of each influencing factor data at the aligned moments of each moment in the first future sub-period of the future period is located is recorded as the analysis category layer; the probability of abnormal water affairs data under the analysis category layer is obtained as the individual co-occurrence of each influencing factor data at each moment and the abnormal data in the historical water affairs data; the average of the individual co-occurrences of all moments in the first future sub-period of the future period is taken as the comprehensive co-occurrence of each influencing factor data predicted value and the abnormal data in the historical water affairs data; the proportion of the abnormal correlation of each abnormal factor data is used as the weight, and the comprehensive co-occurrence of the predicted values of all kinds of influencing factor data is weightedly summed to obtain the abnormal possibility of water affairs data in the first future sub-period.
[0027] Preferably, the step of correcting the comprehensive abnormal probability prediction value of the first future sub-period according to the abnormal possibility of the water service data to obtain the final abnormal probability prediction value includes:
[0028] The abnormal probability of water service data is multiplied by the comprehensive abnormal probability prediction value of the first future sub-period in the future period and then normalized to obtain the final abnormal probability prediction value of the first future sub-period.
[0029] Preferably, setting the collection frequency of the first future sub-period according to the final abnormality probability prediction value includes:
[0030] The final abnormality probability prediction value of the first future sub-period is multiplied by the preset collection frequency to obtain the collection frequency of the first future sub-period.
[0031] In a second aspect, the present invention provides a data acquisition system based on multi-source water affairs sensors, comprising: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned data acquisition method based on multi-source water affairs sensors is implemented.
[0032] By adopting the above technical solution, the above-mentioned multi-source water sensor data acquisition method is generated into a computer program and stored in the memory to be loaded and executed by the processor, so that a terminal device is made based on the memory and the processor for easy use.
[0033] The present invention has the following beneficial effects:
[0034] The present invention adjusts the frequency of collecting water service data according to the occurrence of anomalies in the water service data, so that when the probability of anomalies is high, a higher frequency of collecting water service data is used to collect water service data, thereby enabling the timely acquisition of abnormal information and providing a basis for taking timely corresponding measures;
[0035] Furthermore, considering that the probability of abnormal occurrence of water service data has certain regularity, the probability of abnormal occurrence in future time periods can be relatively accurately predicted by analyzing the change pattern information of historical water service data. The probability of abnormal occurrence in future time periods obtained through historical patterns can better reflect the long-term regularity information.
[0036] Furthermore, considering that whether water data anomalies occur will be affected by other factors, the correlation between other factors and the occurrence of water data anomalies is analyzed to relatively accurately predict the possibility of water data anomalies in future periods. Predicting the possibility of water data anomalies by the correlation between other factors and the occurrence of water data anomalies can better reflect real-time information.
[0037] Furthermore, the probability of abnormal occurrence in future time periods is used to correct the probability of abnormal occurrence, thereby combining real-time information with long-term regular information to more accurately predict the occurrence of abnormalities in future time periods. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a flowchart of the steps of the multi-source water sensor data collection method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0039] See also Figure 1 , which shows a flowchart of a method for collecting data based on multi-source water service sensors according to an embodiment of the present invention, the method includes the following steps:
[0040] S1: Obtain each type of historical water affairs data at each historical moment and each type of historical influencing factor data at the alignment moment.
[0041] Preferably, as an example, obtaining each type of historical water service data at each historical moment and each type of historical influencing factor data at the alignment moment includes:
[0042] Each type of water service data obtained at each historical moment is recorded as each type of historical water service data.
[0043] Collect data of each influencing factor at each historical moment and record it as each historical influencing factor data, match the historical water affairs data at all historical moments with the historical influencing factor data, obtain the time correspondence between the historical water affairs data and the historical influencing factor data when the matching value is the largest, and perform time alignment on the historical water affairs data and the historical influencing factor data according to the time correspondence.
[0044] The types of historical water data include but are not limited to the following: river water level, groundwater level, pH value, water dissolved oxygen, and conductivity.
[0045] The types of influencing factor data include but are not limited to the following: rainfall, temperature, aquatic organism content, and industrial emissions.
[0046] S2: Split the time series sequence composed of the historical water affairs data into a periodic component and a trend component; calculate the correlation between the characteristic descriptors of all periodic segments of the periodic component and the abnormal probability and record it as periodic correlation, obtain trend correlation, and obtain abnormal probability prediction values of future time periods based on periodic laws and trend laws respectively, use periodic correlation and trend correlation as weights of corresponding abnormal probability prediction values respectively, perform weighted summation on the two abnormal probability prediction values to obtain a comprehensive abnormal probability prediction value of the future time period, where the future time period includes several future sub-periods.
[0047] It's important to note that anomalies in water data exhibit certain regularities. For example, during the rainy summer months, river water levels are more likely to exceed set heights, exhibiting anomalies. Therefore, anomalies in water data exhibit certain regularities. Therefore, the probability of future anomalies can be predicted based on these regularities.
[0048] S20: Split the time series consisting of the historical water affairs data into a period component and a trend component.
[0049] It should be noted that because historical water service data contains multiple patterns of periodicity and trend, analyzing multiple patterns together will affect the accuracy of the analysis. Therefore, each pattern needs to be analyzed separately. First, the pattern is separated.
[0050] Preferably, as an example, splitting the time series consisting of the historical water service data into a period component and a trend component includes:
[0051] The seasonal decomposition of time series (STL) method is used to split the time series consisting of historical water service data into periodic components and trend components.
[0052] It should be noted that the method of splitting the time series consisting of historical water service data into period components and trend components using the seasonal decomposition method is an existing technology and will not be described in detail here.
[0053] S21: Calculate the correlation between the feature descriptors of all periodic segments of the periodic component and the abnormal probability, record it as periodic correlation, obtain trend correlation, obtain abnormal probability prediction values for future time periods based on periodic rules and trend rules respectively, use periodic correlation and trend correlation as weights of corresponding abnormal probability prediction values respectively, perform weighted summation on the two abnormal probability prediction values, and obtain the comprehensive abnormal probability prediction value for the future time period.
[0054] It should be noted that abnormal probability prediction values can be obtained according to both cycle and trend rules. Therefore, it is necessary to effectively combine the abnormal probability prediction values obtained by this rule to obtain relatively accurate abnormal probability prediction values.
[0055] It should be further explained that, due to the different correlations between the two patterns, cycle and trend, and the probability of anomaly occurrence, if the correlation between cycle and the probability of anomaly data is greater, the predicted value under the cycle pattern should be used more frequently when predicting the probability of anomaly data occurrence. If the correlation between trend and the probability of anomaly data occurrence is greater, the predicted value under the trend pattern should be used more frequently when predicting the probability of anomaly data occurrence. Therefore, based on this, the anomaly probability prediction values obtained from the two patterns can be effectively combined.
[0056] Preferably, as an example, the correlation between the characteristic descriptors of all periodic segments of the periodic component and the abnormal probability is recorded as periodic correlation, trend correlation is obtained, and abnormal probability prediction values of future time periods based on periodic laws and trend laws are obtained respectively. The periodic correlation and trend correlation are used as weights of the corresponding abnormal probability prediction values, and the two abnormal probability prediction values are weighted and summed to obtain a comprehensive abnormal probability prediction value for the future time period, including:
[0057] The correlation between the sequence composed of the feature descriptors of all periodic segments of the periodic component and the sequence composed of the abnormal probability is recorded as periodic correlation.
[0058] The correlation between the sequence composed of the feature descriptors of all trend segments of the trend component and the sequence composed of the abnormal probability is recorded as trend correlation.
[0059] Obtain anomaly probability predictions for future periods based on trend patterns and cycle patterns, respectively. Using cycle correlation and trend correlation as weights for the corresponding anomaly probability predictions, take the weighted sum of the two anomaly probability predictions to obtain a comprehensive anomaly probability prediction for the future period.
[0060] It can be understood that the feature descriptor of the periodic segment reflects the information of the periodic segment. If the abnormal probability is highly correlated with the information of the periodic segment, it means that the abnormal probability prediction is more dependent on the periodic information. Therefore, more reference should be made to the periodic information when making the abnormal probability prediction. Therefore, the weight of the abnormal probability prediction value predicted based on the periodic information should be set larger; the feature descriptor of the trend segment reflects the information of the trend segment. If the abnormal probability is highly correlated with the information of the trend segment, it means that the abnormal probability prediction is more dependent on the periodic information. Therefore, more reference should be made to the trend information when making the abnormal probability prediction. Therefore, the weight of the abnormal probability prediction value predicted based on the trend information should be set larger.
[0061] The above embodiments involve period segments, trend segments, characteristic descriptors of period segments and trend segments, as well as abnormal probabilities, future time periods, and abnormal probability prediction values of future time periods based on trend laws, and abnormal probability prediction values of future time periods based on periodic laws. The following needs to explain the method for determining the characteristic descriptors of period segments, trend segments, period segments and trend segments, as well as abnormal probabilities, future time periods, and abnormal probability prediction values of future time periods based on trend laws, and abnormal probability prediction values of future time periods based on periodic laws.
[0062] First, the method of obtaining period segments and trend segments is introduced.
[0063] Preferably, as an example, the method for obtaining the period segment and the trend segment includes:
[0064] Perform Fourier transform on the periodic component to obtain several frequency components. Use the amplitude of the frequency component as the weight, perform weighted summation on the frequencies of all frequency components to obtain the weighted frequency value, take the inverse of the weighted frequency value to obtain the period length, and evenly divide the periodic component into several periodic segments of L, where L represents the period length.
[0065] A window of preset size is obtained with each data in the trend component as the center, and the change rate of each data in the window of each data is calculated. The Euclidean distance between the change rate of all data in the window of each data and the change rate of all data in the window of the previous data is calculated and recorded as the change difference between each data and the previous data. The opposite of the transformation difference is used as the exponent and the natural number is used as the base to perform power calculation to obtain the change similarity between each data and the previous data. All data in the trend component are clustered according to the change similarity, and the data segment composed of the data in each category obtained by clustering is used as the trend segment.
[0066] It can be understood that the amplitude reflects the proportion of the frequency component in the periodic component. The weighted frequency value obtained by weighting the frequency using the amplitude as the weight can reflect the overall frequency of the periodic component. Therefore, segmenting the period length based on the weighted frequency can effectively separate each period.
[0067] In addition, the rate of change of data reflects the changing trend of the data. The data in the trend component are clustered by the similarity of the rate of change, so that the data with similar changing trends are divided into one segment, and the data with different changing trends are divided into different segments, so that each trend segment reflects relatively single trend information.
[0068] Then the feature descriptors of period segments and trend segments and the method of obtaining abnormal probability are introduced.
[0069] Preferably, as an example, the method for obtaining the characteristic descriptors and abnormal probabilities of the period segment and the trend segment includes:
[0070] The product of the mean and variance of all historical water service data in the cycle segment is used as the feature descriptor of the cycle segment; the LOF algorithm is used to perform anomaly detection on the historical water service data to obtain the abnormal data in the historical water service data, and the abnormal data in the historical water service data is recorded as abnormal water service data. The number of abnormal water service data in the cycle segment is divided by the total number of data in the cycle segment to obtain the abnormal probability of the cycle segment.
[0071] The product of the mean and variance of all historical water service data in the cycle segment is used as the feature descriptor of the cycle segment; the LOF algorithm is used to perform anomaly detection on the historical water service data to obtain the abnormal data in the historical water service data, and the abnormal data in the historical water service data is recorded as abnormal water service data. The number of abnormal water service data in the cycle segment is divided by the total number of data in the cycle segment to obtain the abnormal probability of the cycle segment.
[0072] The product of the mean and variance of all historical water service data in the trend segment is used as the feature descriptor of the trend segment; the abnormal probability of the trend segment is obtained by dividing the number of abnormal water service data in the trend segment by the total number of data in the trend segment.
[0073] It can be understood that the feature descriptor contains not only value information but also data change information, so the feature descriptor can more comprehensively reflect the information of each segment, thereby providing a basis for accurate correlation analysis.
[0074] Finally, the method of obtaining the abnormal probability prediction value of the future time period and the future time period based on the trend law and the abnormal probability prediction value of the future time period based on the periodic law is introduced.
[0075] Preferably, as an example, a method for obtaining a future time period and an abnormality probability prediction value of a future time period based on a trend law and an abnormality probability prediction value of a future time period based on a cycle law includes:
[0076] The trend length is the average of all trend segment lengths, rounded upwards. The length of the future period is the average of the trend length and the cycle length, rounded upwards. The future period is the period consisting of M consecutive moments starting at the next moment after the current moment. M represents the length of the future period.
[0077] The current cycle segment is matched with the previous cycle segments to obtain a matching value, and the cycle segment with a matching value greater than a preset threshold is used as an alternative reference segment; the next cycle segment of the alternative reference segment with the shortest interval with the current moment is used as the target cycle segment; the abnormal probability in the target cycle segment is used as the abnormal probability prediction value of the future time period based on the periodic law;
[0078] The least squares method is used to fit polynomials to the abnormal probabilities of all trend segments, and the abnormal probabilities of future periods are fitted using the fitted polynomials and recorded as the predicted values of the abnormal probabilities of future periods based on the trend law.
[0079] It can be understood that the more similar the period segment is to the current moment and the shorter the time interval is to the current moment, the more similar it is to the periodic variation law of the current moment. Therefore, the period segment most similar to the periodic variation law of the current moment can be obtained through the similarity with the periodic variation law of the current moment and the interval length. Therefore, the next period segment of the period segment most similar to the periodic variation law of the current moment can better reflect the periodic variation information of the future time period.
[0080] S3: Calculate the information gain between historical water affairs data and each historical influencing factor data as the abnormal correlation of each historical influencing factor data, and obtain the abnormal possibility of water affairs data in the first future sub-period based on the abnormal correlation and the co-occurrence of abnormal data in each influencing factor data and historical water affairs data; correct the comprehensive abnormal probability prediction value of the first future sub-period based on the abnormal possibility of water affairs data to obtain the final abnormal probability prediction value; set the collection frequency of the first future sub-period based on the final abnormal probability prediction value; and control the water affairs data collection based on the collection frequency.
[0081] S30: Calculate the information gain of the historical water affairs data and each historical influencing factor data as the abnormal correlation of each historical influencing factor data, and obtain the abnormal possibility of the water affairs data in the first future sub-period based on the abnormal correlation and the co-occurrence of each influencing factor data with the abnormal data in the historical water affairs data.
[0082] It should be noted that the anomaly probability derived from historical patterns cannot accurately reflect the real-time anomaly situation. Therefore, the anomaly probability derived solely from historical patterns cannot accurately reflect the anomaly situation at each moment. Because water supply anomalies can be affected by other factors, such as rainfall causing river levels to rise, the real-time situation of water supply anomalies can be analyzed by analyzing the impact of other factors on water supply anomalies.
[0083] Preferably, as an example, the information gain between the historical water service data and each historical influencing factor data is calculated as the abnormal correlation of each historical influencing factor data, and the abnormal possibility of the water service data in the first future sub-period is obtained according to the abnormal correlation and the co-occurrence of each influencing factor data with the abnormal data in the historical water service data, including:
[0084] First, get the exception relevance:
[0085] ,in, represents the probability of occurrence of abnormal water service data, represents the probability of occurrence of non-abnormal water service data, represents the logarithmic function with base 2, and S0 represents the abnormal information entropy.
[0086]
[0087] The probability of abnormal water service data appearing and the probability of no abnormal water service data appearing in the aligned moments of all moments of the influencing factor data in any category layer are recorded as the probability of abnormal water service data appearing and the probability of non-abnormal water service data appearing under the category layer. represents the probability of occurrence of abnormal water service data under the i-th category layer, represents the probability of occurrence of non-abnormal water service data under the i-th category layer, represents the information entropy of the i-th category layer.
[0088] , represents the number of category layers, represents information gain.
[0089] Information gain is taken as the abnormal correlation of each historical influencing factor data.
[0090] It can be understood that information gain reflects the determination of abnormal water service data under the condition of the determination of influencing factor data. In other words, the influence of influencing factor data on water service data anomalies. The larger the value, the greater the influence of the influencing factor data on the abnormal water service data, and therefore the greater the anomaly correlation of the influencing factor data.
[0091] Then, the abnormal possibility of water service data in the first future sub-period is obtained based on the abnormal correlation and the co-occurrence of abnormal data in each influencing factor data and historical water service data.
[0092] The category layer in which the predicted value of each influencing factor data at the aligned moments of each moment in the first future sub-period of the future period is located is recorded as the analysis category layer; the probability of abnormal water affairs data under the analysis category layer is obtained as the individual co-occurrence of each influencing factor data at each moment and the abnormal data in the historical water affairs data; the average of the individual co-occurrences of all moments in the first future sub-period of the future period is taken as the comprehensive co-occurrence of each influencing factor data predicted value and the abnormal data in the historical water affairs data; the proportion of the abnormal correlation of each abnormal factor data is used as the weight, and the comprehensive co-occurrence of the predicted values of all kinds of influencing factor data is weightedly summed to obtain the abnormal possibility of water affairs data in the first future sub-period.
[0093] It can be understood that individual co-occurrence reflects the probability of abnormal water service data under the condition that the predicted value of the influencing factor data occurs. Abnormal correlation reflects the influence of the influencing factor data on the determination of abnormal water service data. By using the abnormal influence as a weight, the individual co-occurrence of all influencing factor data is combined to comprehensively determine the probability of abnormal water service data in the future sub-period.
[0094] The above embodiments involve the predicted values of each influencing factor data at the category layer and the alignment time. The following describes a method for determining the predicted values of each influencing factor data at the category layer and the alignment time.
[0095] First, determine the category layer.
[0096] Preferably, as an example, the division of the category layer includes:
[0097] The classification layer is divided according to the division method of the field of each influencing factor data. Taking precipitation as an example, rainfall in the range of (0,10] is considered light rain, rainfall in the range of (10,24.9) is moderate rain, rainfall in the range of (25,49.9) is considered greater than, rainfall in the range of (50,99.9) is considered heavy rain, and rainfall greater than 100 is considered heavy rain. The category layers are light rain, moderate rain, heavy rain, heavy rain, and heavy rain.
[0098] Then determine the predicted value of each influencing factor data at the alignment moment.
[0099] Preferably, as an example, the method for determining the predicted value of each influencing factor data at the alignment moment includes:
[0100] If the alignment time is before the current time, the actual value of each influencing factor data at the alignment time is used as the predicted value of each influencing factor data at the alignment time.
[0101] If the alignment moment is after the current moment, the prediction method of the field where each influencing factor data is located is used to obtain the predicted value of each influencing factor data. Taking rainfall as an example, the rainfall at the alignment moment predicted in the existing weather forecast is used as the rainfall prediction value.
[0102] It should be added that the methods for obtaining future sub-periods include:
[0103] The future period is evenly divided into K sub-periods, which are recorded as future sub-periods, where K represents the preset number of divisions.
[0104] It should be noted that the future time period spans a long time, and the occurrence of anomalies changes in real time. Therefore, it is not accurate enough to use a fixed collection frequency in a long time period.
[0105] S31: According to the abnormal possibility of the water service data, the comprehensive abnormal probability prediction value of the first future sub-period is corrected to obtain a final abnormal probability prediction value.
[0106] Preferably, as an example, the comprehensive abnormal probability prediction value of the first future sub-period is corrected according to the abnormal possibility of the water service data to obtain the final abnormal probability prediction value, including:
[0107] The abnormal probability of water service data is multiplied by the comprehensive abnormal probability prediction value of the first future sub-period in the future period and then normalized to obtain the final abnormal probability prediction value of the first future sub-period.
[0108] S32: Setting the collection frequency of the first future sub-period according to the final abnormality probability prediction value.
[0109] Preferably, as an example, setting the collection frequency of the first future sub-period according to the final abnormality probability prediction value includes:
[0110] The final abnormality probability prediction value of the first future sub-period is multiplied by the preset collection frequency to obtain the collection frequency of the first future sub-period.
[0111] S33: Perform water data collection control based on collection frequency.
[0112] An embodiment of the present invention also discloses a multi-source water affairs sensor data acquisition system, including a processor and a memory, wherein the memory stores computer program instructions. When the computer program instructions are executed by the processor, the multi-source water affairs sensor data acquisition method according to the present invention is implemented.
[0113] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.
[0114] In the present invention, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory, dynamic random access memory, static random access memory, enhanced dynamic random access memory, high bandwidth memory, hybrid memory cube, etc., or any other medium that can be used to store the required information and can be accessed by an application, module, or both. Any such computer storage medium may be part of, accessible to, or connectable to the device.
Claims
1. A data collection method based on multi-source water service sensors, characterized in that: include: Obtain each type of historical water affairs data at each historical moment and each type of historical influencing factor data at the alignment moment; Splitting the time series sequence composed of the historical water service data into a period component and a trend component; calculating the correlation between the characteristic descriptors of all period segments of the period component and the abnormal probability as period correlation, obtaining the trend correlation, respectively obtaining abnormal probability prediction values for future time periods based on the periodic law and the trend law, using the periodic correlation and the trend correlation as weights of the corresponding abnormal probability prediction values, performing a weighted summation on the two abnormal probability prediction values to obtain a comprehensive abnormal probability prediction value for the future time period, wherein the future time period includes a plurality of future sub-periods; Calculating the information gain between the historical water affairs data and each historical influencing factor data as the abnormal correlation of each historical influencing factor data, wherein the method for obtaining the abnormal correlation includes: The abnormal information entropy of historical water service data is calculated based on the probability of abnormal water service data and the probability of non-abnormal water service data; the influencing factor data is divided into several category layers; the probability of abnormal water service data appearing and the probability of no abnormal water service data appearing in the aligned moments of all moments of the influencing factor data in any category layer are recorded as the probability of abnormal water service data and the probability of non-abnormal water service data under the category layer, and the information entropy of the category layer is calculated based on the probability of abnormal water service data and the probability of non-abnormal water service data under the category layer; the probability of the influencing factor data in each category layer is obtained and recorded as the probability of each category layer, and the information gain is calculated based on the probability of each category layer and the information entropy of each category layer as the abnormal correlation of each influencing factor data; The abnormal probability of water service data in the first future sub-period is obtained based on the abnormal correlation and the co-occurrence of abnormal data in each influencing factor data and historical water service data; the comprehensive abnormal probability prediction value of the first future sub-period is corrected according to the abnormal probability of water service data to obtain the final abnormal probability prediction value; the collection frequency of the first future sub-period is set according to the final abnormal probability prediction value; and water service data collection is controlled based on the collection frequency.
2. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method for obtaining the feature descriptor of the periodic segment includes: The product of the mean and variance of all historical water service data in a periodic segment is used as the feature descriptor of the periodic segment.
3. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method for obtaining the abnormality probability includes: Performing anomaly detection on the historical water service data to obtain abnormal data in the historical water service data, and recording the abnormal data in the historical water service data as abnormal water service data; The abnormal probability of the periodic segment is obtained by dividing the number of abnormal water service data in the periodic segment by the total number of data in the periodic segment.
4. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method for obtaining the trend correlation includes: The trend component is segmented to obtain several trend segments; Get the feature descriptor of the trend segment; Get the abnormal probability of the trend segment; The correlation between the characteristic descriptors of all trend segments in the trend component and the abnormal probability is taken as the trend relevance.
5. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method of respectively obtaining the abnormal probability prediction values for future time periods based on the periodic law and the trend law includes: The current cycle segment is matched with the previous cycle segments to obtain a matching value, and the cycle segment with a matching value greater than a preset threshold is used as an alternative reference segment; the next cycle segment of the alternative reference segment with the shortest interval with the current moment is used as the target cycle segment; the abnormal probability in the target cycle segment is used as the abnormal probability prediction value of the future time period based on the periodic law; The least squares method is used to fit polynomials to the abnormal probabilities of all trend segments, and the abnormal probabilities of future periods are fitted using the fitted polynomials and recorded as the predicted values of the abnormal probabilities of future periods based on the trend law.
6. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method of obtaining the abnormal possibility of water service data in the first future sub-period includes: The category layer in which the predicted value of each influencing factor data at the aligned moments of each moment in the first future sub-period of the future period is located is recorded as the analysis category layer; the probability of abnormal water affairs data under the analysis category layer is obtained as the individual co-occurrence of each influencing factor data at each moment and the abnormal data in the historical water affairs data; the average of the individual co-occurrences of all moments in the first future sub-period of the future period is taken as the comprehensive co-occurrence of each influencing factor data predicted value and the abnormal data in the historical water affairs data; the proportion of the abnormal correlation of each abnormal factor data is used as the weight, and the comprehensive co-occurrence of the predicted values of all kinds of influencing factor data is weightedly summed to obtain the abnormal possibility of water affairs data in the first future sub-period.
7. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The method of correcting the comprehensive abnormal probability prediction value of the first future sub-period according to the abnormal possibility of the water service data to obtain the final abnormal probability prediction value includes: The abnormal probability of water service data is multiplied by the comprehensive abnormal probability prediction value of the first future sub-period in the future period and then normalized to obtain the final abnormal probability prediction value of the first future sub-period.
8. The method for collecting data based on multi-source water service sensors according to claim 1, characterized in that: The step of setting the collection frequency of the first future sub-period according to the final abnormal probability prediction value includes: The final abnormality probability prediction value of the first future sub-period is multiplied by the preset collection frequency to obtain the collection frequency of the first future sub-period.
9. Based on the multi-source water sensor data acquisition system, it is characterized by: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the multi-source water affairs sensor data collection method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Power grid host dynamic threshold setting method based on FARIMA-LSTM prediction
CN113435725A
Charging shed photovoltaic power generation reserve prediction method based on multi-data fusion
CN116011686A