Water affair data management method and system
By formulating rules for identifying missing values, filling and identifying abnormal data in water business business, and automated processing with Transformer model, the problems of high manpower investment, difficulty in handling missing values and difficult outlier values in data preprocessing in the water industry are solved, and data management efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202411779679.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
The existing data preprocessing model of the water industry has problems such as high manpower investment, difficulty in processing missing values, and difficulty in identifying outliers, resulting in low data preprocessing efficiency and accuracy.
Formulate rules for identifying missing data, filling and identifying abnormal data in water business, combine the Transformer model to automatically identify and process data missing and abnormal data, and improve identification efficiency and accuracy through a combination of shallow and deep analysis.
Through automated data missing and abnormal data identification processing, the proportion of manual verification is significantly reduced, the efficiency and accuracy of data management are improved, and the data preprocessing efficiency and accuracy are improved.
Smart Images

Figure CN120105032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart water technology, and in particular to a water data management method and system. Background Art
[0002] With the rise of the concepts of smart cities and digital twin water networks, the water industry is gradually entering a new era of management driven by big data. As a key hub connecting data sources and business applications, the data preprocessing technology of the water data center is particularly important. Data preprocessing technology aims to clean, convert, integrate and control the quality of massive, multi-source and heterogeneous water data through a series of automated or semi-automated steps, providing a solid foundation for data analysis and decision-making.
[0003] At present, most of the data reported by the water industry through monitoring and sensing equipment is checked and verified by front-line staff, and a small amount is processed by programs. The common processing methods currently include: over-range data identification, data missing item identification, etc. The above existing processing mode has the following main disadvantages:
[0004] 1) Human resources input: Tasks such as manual data cleaning and feature engineering can be very time-consuming and require a large number of people, resulting in high human resource costs.
[0005] 2) Difficulty in handling missing values: How to effectively and reasonably handle missing values is a challenge. Simple deletion or filling strategies may introduce bias, while complex prediction models may overfit or be inaccurate.
[0006] 3) Difficulty in identifying outliers: Humans can quickly identify sudden increases or decreases reported by devices based on experience, but programs are difficult to process automatically and are prone to misjudgment or omission, affecting data accuracy.
[0007] How to improve the efficiency and accuracy of data preprocessing is a technical problem that needs to be solved urgently. Summary of the invention
[0008] In this regard, the present invention provides a water affairs data management method, system, electronic equipment, computer storage medium and computer program product to solve the above technical problems.
[0009] The present invention discloses a water affairs data management method, which is applied to a water affairs data middle platform, and the method comprises the following steps:
[0010] Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data;
[0011] Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result;
[0012] Performing missing data filling processing on the missing data identification result based on the second rule;
[0013] An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
[0014] Optionally, using the first rule to identify missing data on the water service data to be preprocessed to obtain missing data identification results includes:
[0015] Analyze and obtain the remote water meter data collection frequency and the user water fee settlement cycle from the first rule, calculate the first target number of remote water meter data groups according to the remote water meter data collection frequency, and calculate the second target number of user water fee settlement data groups according to the user water fee settlement cycle;
[0016] Determine whether the first group number of remote water meter data in the water service data is the same as the first target group number, and determine whether the second group number of user water fee settlement data in the water service data is the same as the second target group number, to obtain the missing data identification result.
[0017] Optionally, performing missing data filling processing on the missing data identification result based on the second rule includes:
[0018] The following missing data filling methods are obtained from the second rule: mode filling, mean / median filling, and model-based prediction filling;
[0019] At least one of the missing data filling methods mentioned above is used to perform missing data filling processing on the missing data identification result.
[0020] Optionally, using the third rule to identify abnormal data on the water service data to be preprocessed to obtain a first abnormal data identification result includes:
[0021] The following abnormal data identification methods are obtained from the third rule: identification based on fixed values, identification based on relative values, and identification based on statistical distribution;
[0022] At least one of the above abnormal data identification methods is used to perform abnormal data identification on the water service data to be preprocessed to obtain the first abnormal data identification result.
[0023] Optionally, performing abnormality evaluation on the first abnormal data identification result to obtain an abnormality evaluation value includes:
[0024] Extracting the amount of abnormal data, the divergence of abnormal data, the time distribution characteristics of abnormal data, and the spatial distribution characteristics of abnormal data from the first abnormal data identification result;
[0025] The amount of abnormal data, the divergence of the abnormal data, the time distribution characteristics of the abnormal data, and the spatial distribution characteristics of the abnormal data are input into an abnormality evaluation model based on the BERT large model, and the abnormality evaluation model outputs the abnormality evaluation value.
[0026] Optionally, the abnormal threshold is determined by:
[0027] The first confirmation times and the second confirmation times of the missing data identification result and the first abnormal data identification result are obtained by statistical calculation based on historical data, and the abnormal threshold is determined based on the first confirmation times and the second confirmation times:
[0028]
[0029] Among them, nor is the abnormal threshold, nor1 is the minimum abnormal threshold, n1 is the first confirmation number, n2 is the second confirmation number; a and b are adjustment constant values, and a>b.
[0030] The present invention also discloses a water affairs data management system, which is applied to a water affairs data middle station. The system includes a processor and a memory. The processor calls and executes a computer code in the memory to implement the following steps:
[0031] Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data;
[0032] Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result;
[0033] Performing missing data filling processing on the missing data identification result based on the second rule;
[0034] An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
[0035] The present invention also discloses an electronic device, comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement any of the above methods.
[0036] The present invention also discloses a computer storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement any of the above methods.
[0037] The present invention also discloses a computer program product, wherein the computer program product includes a computer program, and the computer program is executed by a processor of an electronic device to implement any of the above methods.
[0038] The solution of the present invention can realize automatic identification of data missing and data anomalies based on preset rules, greatly reducing the proportion of manual data verification, thereby improving the efficiency of data management. At the same time, the present invention adopts a combination of shallow and deep analysis for the identification of data anomalies. First, shallow analysis is used to perform a rapid abnormality analysis on the pre-processed water affairs data based on the third rule. If the abnormality evaluation value of the abnormal data identification result is too high, it means that the accuracy of the shallow analysis is insufficient. At this time, the abnormal data deep identification model is used to perform a deeper secondary abnormal data identification on the pre-processed water affairs data, so as to finally obtain a more accurate second abnormal data identification result. This method achieves an organic balance between the efficiency and accuracy of abnormal data identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 It is a flow chart of a water affairs data management method disclosed in an embodiment of the present invention;
[0041] Figure 2 It is a structural diagram of a water affairs data management system disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0042] The following is a description of the implementation of the present application by specific specific embodiments. People familiar with the technology can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0043] In addition, the technical features involved in the different embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0044] like Figure 1 As shown, an embodiment of the present invention discloses a water affairs data management method, which is applied to a water affairs data middle station, and the method comprises the following steps:
[0045] Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data;
[0046] Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result;
[0047] Performing missing data filling processing on the missing data identification result based on the second rule;
[0048] An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
[0049] The present invention pre-establishes a first rule, a second rule, and a third rule, wherein the first rule is used to identify missing water service data, the second rule is used to fill missing water service data, and the third rule is used to identify abnormal water service data. After receiving the water service data to be pre-processed reported by the monitoring and sensing device, the water service data center first identifies missing data for the pre-processed water service data based on the first rule, and uses the third rule to identify abnormal data for the pre-processed water service data, thereby obtaining missing data identification results and first abnormal data identification results, respectively. For the identified missing data points, the second rule is used to fill in the missing data.
[0050] In addition, for the first abnormal data identification result, the present invention first performs an abnormality evaluation on the first abnormal data identification result to obtain an abnormality evaluation value. If the obtained abnormality evaluation value is higher than the abnormal threshold, it indicates that the first abnormal data identification result itself is abnormal and its confidence is not high. At this time, the Transformer-based abnormal data deep recognition model is further used to perform a deeper secondary abnormal data recognition on the preprocessed water data to obtain a more accurate second abnormal data recognition result.
[0051] Therefore, the scheme of the present invention can realize automatic identification of data missing and data anomalies based on preset rules, greatly reducing the proportion of manual data verification, thereby improving the efficiency of data management. At the same time, the present invention adopts a combination of shallow and deep analysis for the identification of data anomalies. First, shallow analysis is used to perform a rapid abnormality analysis on the pre-processed water affairs data based on the third rule. If the abnormality evaluation value of the abnormal data identification result is too high, it means that the accuracy of the shallow analysis is insufficient. At this time, the abnormal data deep identification model is used to perform a deeper secondary abnormal data identification on the pre-processed water affairs data, so as to finally obtain a more accurate second abnormal data identification result. This method achieves an organic balance between the efficiency and accuracy of abnormal data identification.
[0052] It should be noted that, for the specific structure of the abnormal data deep recognition model, the existing structure can be adopted and will not be elaborated here.
[0053] Optionally, using the first rule to identify missing data on the water service data to be preprocessed to obtain missing data identification results includes:
[0054] Analyze and obtain the remote water meter data collection frequency and the user water fee settlement cycle from the first rule, calculate the first target number of remote water meter data groups according to the remote water meter data collection frequency, and calculate the second target number of user water fee settlement data groups according to the user water fee settlement cycle;
[0055] Determine whether the first group number of remote water meter data in the water service data is the same as the first target group number, and determine whether the second group number of user water fee settlement data in the water service data is the same as the second target group number, to obtain the missing data identification result.
[0056] In this embodiment, according to the reporting frequency of the data source, a data missing identification task is formulated, and the amount of data in the database of the water service data center is scanned regularly. For example, the frequency of remote water meter data collection is one group every 30 minutes, and there should be 48 groups of data per day; the user water fee settlement data is collected once a month, and there should be only one group per month. The first group number of remote water meter data and the second group number of user water fee settlement data are respectively parsed from the water service data to be preprocessed, and by comparing them with the first target group number and the second target group number, it is possible to quickly identify whether the business data is missing. The remote water meter data and the user water fee settlement data are sorted in chronological order, and according to the data interval calculated from other remote water meter data and user water fee settlement data, the missing points can be quickly determined.
[0057] Optionally, performing missing data filling processing on the missing data identification result based on the second rule includes:
[0058] The following missing data filling methods are obtained from the second rule: mode filling, mean / median filling, and model-based prediction filling;
[0059] At least one of the missing data filling methods mentioned above is used to perform missing data filling processing on the missing data identification result.
[0060] In this embodiment, the present invention can be set to fill missing data by mode filling, mean / median filling, model-based prediction filling, etc., and each filling method is specifically described as follows:
[0061] 1) Majority filling
[0062] For discrete variables, the most frequent value (mode) of the variable can be used to fill missing values. This method is simple and easy to implement, but it may introduce bias, especially when the mode does not represent the overall data distribution.
[0063] 2) Mean / Median Filling
[0064] For continuous variables, missing values can be filled with the mean or median of the column. The mean is more affected by extreme values, while the median is more robust. This method is suitable for cases where the data is normally distributed or approximately normally distributed.
[0065] Optionally, the model-based predictive filling includes regression analysis or classification model or MissForest algorithm.
[0066] In this embodiment, the model-based predictive filling may be a regression analysis or a MissForest algorithm, which is described in detail as follows:
[0067] Regression analysis: For numerical data, models such as linear regression and polynomial regression can be used to predict missing values based on other complete variables.
[0068] Classification models: For categorical variables, models such as logistic regression and decision trees can be used to predict the missing categories.
[0069] MissForest algorithm: This is an iterative prediction method based on random forests, which is suitable for mixed type data and can handle missing values of numerical and categorical data.
[0070] Optionally, using the third rule to identify abnormal data on the water service data to be preprocessed to obtain a first abnormal data identification result includes:
[0071] The following abnormal data identification methods are obtained from the third rule: identification based on fixed values, identification based on relative values, and identification based on statistical distribution;
[0072] At least one of the above abnormal data identification methods is used to perform abnormal data identification on the water service data to be preprocessed to obtain the first abnormal data identification result.
[0073] In this embodiment, the present invention sets the recognition of abnormal data to be performed by using methods such as recognition based on fixed values, recognition based on relative values, and recognition based on statistical distribution. Each method is described as follows:
[0074] 1) Identification based on fixed values
[0075] Set one or several fixed thresholds, and any data points exceeding these thresholds are considered anomalies. This method is simple and direct, but may not be accurate enough because it does not take into account the distribution and variation of the data.
[0076] 2) Relative numerical identification
[0077] Year-over-year and quarter-over-quarter comparisons: Calculate the percentage of data compared to the same or previous period to identify data points that deviate significantly from normal patterns.
[0078] Percentiles: Using a specific percentile (such as 1% or 99%) as a threshold, data falling outside these limits is considered anomaly.
[0079] 3) Identification based on statistical distribution
[0080] Three times standard deviation method: Under the normal distribution assumption, approximately 99.7% of the data points are within the range of ±3 standard deviations from the mean. Data points outside this range are considered outliers.
[0081] Box plot analysis: Detect outliers by identifying the interquartile range (IQR) and upper and lower boundaries of the data. Points outside the upper and lower boundaries are considered outliers.
[0082] Optionally, performing abnormality evaluation on the first abnormal data identification result to obtain an abnormality evaluation value includes:
[0083] Extracting the amount of abnormal data, the divergence of abnormal data, the time distribution characteristics of abnormal data, and the spatial distribution characteristics of abnormal data from the first abnormal data identification result;
[0084] The amount of abnormal data, the divergence of the abnormal data, the time distribution characteristics of the abnormal data, and the spatial distribution characteristics of the abnormal data are input into an abnormality evaluation model based on the BERT large model, and the abnormality evaluation model outputs the abnormality evaluation value.
[0085] In this embodiment, the present invention constructs an abnormality assessment model based on the BERT large model, which is constructed by fine-tuning the BERT large model using small sample level training data. Since fine-tuning the BERT large model to obtain a local vertical model belongs to mature prior art, the present invention will not repeat it here. The present invention uses the constructed abnormality assessment model to perform abnormality assessment on the first abnormal data recognition result itself. The details are as follows:
[0086] First, the amount of abnormal data, the divergence of abnormal data, the time distribution characteristics of abnormal data, and the spatial distribution characteristics of abnormal data are extracted from the first abnormal data identification result. Among them, the amount of abnormal data refers to the total amount of all data identified as abnormal contained in the first abnormal data identification result; the divergence of abnormal data refers to the mean of the difference between all data identified as abnormal and the mean / median of similar data in the water service data to be preprocessed; the time distribution characteristics of abnormal data refer to the maximum global time span of all abnormal data and the maximum local time span of each time aggregation area (some similar abnormal data are aggregated at the upload time); the spatial distribution characteristics of abnormal data refer to the maximum global distance of the geographical distance between the monitoring and sensing devices corresponding to all abnormal data (i.e., the monitoring and sensing devices that upload abnormal data) and the maximum local distance of each geographical aggregation area (the uploaded data of multiple geographically close monitoring and sensing devices in the area have abnormal data). The abnormality evaluation value can be obtained by using the above-mentioned constructed abnormality evaluation model to comprehensively evaluate these abnormal characteristic data.
[0087] Among them, the small sample-level training data used in the above-mentioned fine-tuning training process is actually the historical abnormal data recognition results, so the trained abnormality evaluation model can grasp the general law of the above-mentioned abnormal feature data of the abnormal data recognition results. The abnormality evaluation value is used to characterize the gap between the first abnormal data recognition result represented by the above-mentioned abnormal feature data and the general law. Obviously, the larger the gap, the higher the abnormality evaluation value. If the abnormality evaluation value is higher than the abnormal threshold, it is determined that the credibility of the first abnormal data recognition result is insufficient. At this time, it is necessary to use the Transformer-based abnormal data deep recognition model to perform secondary abnormal data recognition on the pre-processed water data; otherwise, there is no need to use the Transformer-based abnormal data deep recognition model to perform secondary abnormal data recognition on the pre-processed water data.
[0088] Optionally, the abnormal threshold is determined by:
[0089] The first confirmation times and the second confirmation times of the missing data identification result and the first abnormal data identification result are obtained by statistical calculation based on historical data, and the abnormal threshold is determined based on the first confirmation times and the second confirmation times:
[0090]
[0091] Among them, nor is the abnormal threshold, nor1 is the minimum abnormal threshold, n1 is the first confirmation number, n2 is the second confirmation number; a and b are adjustment constant values, and a>b.
[0092] In this embodiment, the abnormal threshold in the present invention is dynamically adjusted, which is specifically determined according to the first confirmation number of the missing data identification result and the second confirmation number of the first abnormal data identification result, wherein the first confirmation number and the second confirmation number are respectively the total number of recognized records of the missing data identification result and the first abnormal data identification result during the manual review contained in the historical record data, and the total number of recognized records reflects the accuracy of the aforementioned first rule and the third rule in the missing data identification and abnormal data identification. Therefore, when the accuracy is higher, the higher the abnormal threshold is set, that is, the threshold for using the Transformer-based abnormal data deep identification model to perform secondary abnormal data identification on the pre-processed water affairs data is increased, so as to achieve the premise of ensuring that the accuracy of abnormal data identification is high enough, and appropriately improve the efficiency of abnormal data identification (the abnormal data deep identification model requires more complex calculations, which takes longer than the abnormal data identification based on the third rule); conversely, the lower the abnormal threshold is set, that is, the threshold for using the Transformer-based abnormal data deep identification model to perform secondary abnormal data identification on the pre-processed water affairs data is lowered, so as to ensure the accuracy of abnormal data identification.
[0093] Among them, in addition to using the second confirmation number of the first abnormal data identification result to determine the above-mentioned abnormal threshold, the present invention also uses the first confirmation number of the missing data identification result. The reason for this setting is that the aforementioned first rule and third rule are both artificially set based on experience values, and they have a certain similarity in the rationality of the setting. Therefore, the present invention uses the first confirmation number of the missing data identification result to assist in determining the abnormal threshold, but sets the adjustment constant value a>b, so as to reduce the weight of the first confirmation number, that is, mainly relying on the second confirmation number to determine the abnormal threshold.
[0094] like Figure 2 As shown, an embodiment of the present invention further discloses a water affairs data management system, which is applied to a water affairs data middle station. The system includes a processor and a memory. The processor calls and executes a computer code in the memory to implement the following steps:
[0095] Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data;
[0096] Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result;
[0097] Performing missing data filling processing on the missing data identification result based on the second rule;
[0098] An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
[0099] An embodiment of the present invention further discloses an electronic device, comprising: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method described in the above embodiment.
[0100] An embodiment of the present invention further discloses a computer storage medium, wherein the computer storage medium stores a computer program, and the computer program is executed by a processor to implement the method described in the above embodiment.
[0101] An embodiment of the present invention further discloses a computer program product, wherein the computer program product includes a computer program, and the computer program is executed by a processor of an electronic device to implement any of the above methods.
[0102] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0104] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0105] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.
[0106] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0107] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0108] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A water affairs data management method, applied to a water affairs data middle platform, characterized in that: The method comprises the following steps: Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data; Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result; Performing missing data filling processing on the missing data identification result based on the second rule; An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
2. A water affairs data management method according to claim 1, characterized in that: Using the first rule to identify missing data on the water affairs data to be preprocessed, and obtaining missing data identification results, includes: Analyze and obtain the remote water meter data collection frequency and the user water fee settlement cycle from the first rule, calculate the first target number of remote water meter data groups according to the remote water meter data collection frequency, and calculate the second target number of user water fee settlement data groups according to the user water fee settlement cycle; Determine whether the first group number of remote water meter data in the water service data is the same as the first target group number, and determine whether the second group number of user water fee settlement data in the water service data is the same as the second target group number, to obtain the missing data identification result.
3. A water affairs data management method according to claim 2, characterized in that: Performing missing data filling processing on the missing data identification result based on the second rule includes: The following missing data filling methods are obtained from the second rule: mode filling, mean / median filling, and model-based prediction filling; At least one of the missing data filling methods mentioned above is used to perform missing data filling processing on the missing data identification result.
4. A water affairs data management method according to claim 3, characterized in that: Using the third rule to identify abnormal data on the water service data to be preprocessed to obtain a first abnormal data identification result includes: The following abnormal data identification methods are obtained from the third rule: identification based on fixed values, identification based on relative values, and identification based on statistical distribution; Use at least one of the above abnormal data identification methods to perform abnormal data identification on the water service data to be preprocessed to obtain the first abnormal data identification result.
5. A water affairs data management method according to claim 4, characterized in that: Performing an abnormality evaluation on the first abnormal data identification result to obtain an abnormality evaluation value includes: Extracting the amount of abnormal data, the divergence of abnormal data, the time distribution characteristics of abnormal data, and the spatial distribution characteristics of abnormal data from the first abnormal data identification result; The amount of abnormal data, the divergence of the abnormal data, the time distribution characteristics of the abnormal data, and the spatial distribution characteristics of the abnormal data are input into an abnormality evaluation model based on the BERT large model, and the abnormality evaluation model outputs the abnormality evaluation value.
6. A water affairs data management method according to claim 5, characterized in that: The abnormal threshold is determined by: The first confirmation times and the second confirmation times of the missing data identification result and the first abnormal data identification result are obtained by statistical calculation based on historical data, and the abnormal threshold is determined based on the first confirmation times and the second confirmation times: Among them, nor is the abnormal threshold, nor1 is the minimum abnormal threshold, n1 is the first confirmation number, n2 is the second confirmation number; a and b are adjustment constant values, and a>b.
7. A water affairs data management system, applied to a water affairs data middle station, the system comprising a processor and a memory, characterized in that: The processor calls and executes the computer code in the memory to implement the following steps: Formulate a first rule for identifying missing water service data, a second rule for filling missing water service data, and a third rule for identifying abnormal water service data; Receiving water service data to be preprocessed reported by the monitoring sensing device, and using the first rule and the third rule to perform data missing identification and abnormal data identification on the water service data to be preprocessed, respectively, to obtain a missing data identification result and a first abnormal data identification result; Performing missing data filling processing on the missing data identification result based on the second rule; An abnormality degree evaluation is performed on the first abnormal data identification result to obtain an abnormality degree evaluation value. If the abnormality degree evaluation value is higher than the abnormal threshold, a Transformer-based abnormal data deep recognition model is used to perform secondary abnormal data recognition on the preprocessed water service data to obtain a second abnormal data recognition result.
8. An electronic device comprising: At least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 6.
9. A computer storage medium storing a computer program, characterized in that: The computer program is executed by a processor to implement the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that: The computer program is executed by a processor of an electronic device to implement the method according to any one of claims 1 to 6.
Citation Information
Cited By
Intelligent water meter recognition method and system and related equipment
CN120745463A