Water affair data exception identification method, server and storage medium
By implementing technical means such as multi-source access, DML and DDL extraction, task scheduling, monitoring layer analysis and data blood relationship analysis on the water data management platform, the problem of low accuracy of outliers recognition in water data is solved, and higher identification accuracy and better troubleshooting support is achieved.
Patent Information
- Application Number
- CN202510581095.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has low accuracy when identifying outliers in water data, especially when the data distribution is unstable.
By deploying a water data management platform on the data processing server, outliers in water data are identified and visualized by using technical means such as multi-source access, DML and DDL extraction, task scheduling, monitoring layer analysis and data blood relationship analysis.
It significantly improves the accuracy of abnormal identification of water data, adapts to the unstable data distribution of various links of the water system, reduces misjudgment and misjudgment problems caused by data fluctuations, and provides support for troubleshooting and decision-making optimization through abnormal topology diagrams.
Smart Images

Figure CN120086544A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of water service data processing, and in particular, to a method for identifying abnormal water service data, a server, and a storage medium. Background Art
[0002] The water service system is an important part of urban infrastructure, responsible for key links such as water resource collection, treatment, distribution, and discharge. With the acceleration of urbanization and the continuous growth of water resource demand, the complexity and scale of the water service system are also increasing. Modern water service systems rely on a large number of sensors, monitoring devices, and data management systems to collect and process massive amounts of water service data in real time, including information such as water quality, water volume, water pressure, and equipment status. These data are important bases for the operation, scheduling, and decision-making of the water service system, and are directly related to water supply safety, resource optimization, and environmental protection. However, due to the diverse data sources involved in the water service system (such as sensors, databases, external systems, etc.), and the fact that the data collection and transmission processes may be affected by factors such as equipment failures, network delays, and human errors, abnormal values often appear in water service data. These abnormal data may manifest as data missing, data errors, abnormal data fluctuations, etc. Since the management and decision-making of the water service system rely on data analysis and prediction, abnormal data may lead to deviations in analysis results, affecting the accuracy and scientific nature of decision-making.
[0003] Currently, statistical principles are mainly used to identify abnormal water service data. By analyzing the distribution characteristics of the data, abnormal values that deviate from the normal distribution are identified. For example, statistical quantities such as the mean and variance are used for anomaly detection. However, this method is applicable to scenarios where the data distribution is stable. Since the water service system involves multiple links such as water sources, water treatment, water supply networks, and drainage systems, the data characteristics and change laws of each link are different, resulting in an unstable overall data distribution. Therefore, for the case of unstable data distribution, the current method of using statistical principles to identify abnormal water service data has a low accuracy rate. Summary of the Invention
[0004] Embodiments of the present application provide a method for identifying abnormal water service data, a server, and a storage medium, which are used to effectively improve the accuracy rate of identifying abnormal water service data.
[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions: In a first aspect, a method for identifying abnormal water service data is provided, which is applied to a data processing server. The data processing server is deployed with a water service data management platform. The water service data management platform includes a data access layer, a data exchange layer, a data processing layer, and a monitoring layer. The method includes: In response to receiving water service data from different data sources, multi-source access to the water service data from different data sources is performed through the data access layer; Perform DML extraction and DDL extraction on water service data through the data exchange layer to achieve data exchange between different data sources; Perform task scheduling on the water service data after data exchange through the data processing layer according to a preset workflow; Obtain task execution status data and hardware resource usage data through the monitoring layer; Based on the task execution status data and hardware resource usage data, identify abnormal water service data from the water service data after task scheduling; Conduct data lineage analysis on the abnormal water service data to determine the abnormal topology graph of the abnormal water service data and visualize the abnormal topology graph.
[0006] In a possible implementation manner of the first aspect, perform multi-source access to the water service data of different data sources through the data access layer, including: Identify the data types of different data sources, where the data types include structured data, semi-structured data, and unstructured data; Determine the corresponding data access plug-ins according to the data types; Access the water service data of the data source through the data access plug-ins in a passive reception or active pulling manner; Store the water service data in the original database according to a preset access output distribution strategy.
[0007] In another possible implementation manner of the first aspect, the data source includes a relational database. Perform DML extraction and DDL extraction on the water service data through the data exchange layer to achieve data exchange between different data sources, including: Summarize the water service data of different data sources through the data exchange layer, where the water service data of the relational database is obtained through full-volume collection or incremental collection; Establish a mapping relationship of heterogeneous systems through the visualization configuration interface of the data exchange layer, and perform DML extraction and DDL extraction on the water service data to achieve data exchange between different data sources, where each data source corresponds to a data operating system, and the heterogeneous systems are different data operating systems.
[0008] In another possible implementation manner of the first aspect, the data processing layer includes a workflow management module and a task scheduling module. Perform data scheduling on the water service data after data exchange through the data processing layer according to a preset workflow, including: Configure the workflow through the workflow management module, where the workflow includes data cleaning tasks, data conversion tasks, data quality detection tasks, and data analysis tasks; The task scheduling module determines the dependencies and priorities among the data cleaning task, data transformation task, data quality detection task, and data analysis task, and schedules the water service data after data exchange according to the dependencies and priorities.
[0009] In another possible implementation of the first aspect, the monitoring layer includes a task status monitoring module and a hardware resource monitoring module. Task execution status data and hardware resource usage data are obtained through the monitoring layer, including: The task status monitoring module monitors the task execution status data of each task in the workflow in real time. The task execution status data includes the waiting status, running status, successfully completed status, and failed status; The hardware resource monitoring module obtains the hardware resource usage data. The hardware resource usage data includes CPU usage rate, memory occupancy, and disk space.
[0010] In another possible implementation of the first aspect, based on the task execution status data and the hardware resource usage data, abnormal water service data is identified from the water service data after task scheduling, including: Obtain technical metadata and business metadata. Technical metadata is used to describe the technical attributes of data, and business metadata is used to describe the business meaning of data; Establish an association relationship between the technical metadata and the business metadata to obtain a water service data resource catalog; Based on the water service data resource catalog, analyze the task execution status data to identify abnormal task execution; Based on the water service data resource catalog, analyze the hardware resource usage data to identify abnormal resource usage; According to the abnormal task execution and abnormal resource usage, determine the abnormal water service data from the water service data after task scheduling.
[0011] In another possible implementation of the first aspect, establishing an association relationship between the technical metadata and the business metadata to obtain a water service data resource catalog includes: Extract the data structure information, storage location information, and data format information in the technical metadata; Extract the data definition information, business rule information, and data quality standard information in the business metadata; Establish a mapping relationship between the technical metadata and the business metadata; Generate a water service data resource catalog according to the mapping relationship.
[0012] The analysis of the task execution status data based on the water service data resource catalog to identify abnormal task execution includes: Obtain the historical execution status data of each task in the workflow; Based on the historical execution status data, establish a normal behavior model for task execution; Compare the current task execution status data with the normal behavior model to obtain the task status and task execution time; When the task status is a failure status or the task execution time exceeds the corresponding second preset threshold, it is determined that the task execution is abnormal.
[0013] The analysis of the hardware resource usage data based on the water service data resource catalog to identify resource usage anomalies includes: Obtain the historical hardware resource usage data of the system operation; Based on the historical hardware resource usage data, establish a normal range model for resource usage; Compare the current hardware resource usage data with the normal range model; When any one of the CPU usage rate, memory occupancy, or disk space exceeds the corresponding second preset threshold, it is determined that the resource usage is abnormal.
[0014] In another possible implementation manner of the first aspect, perform data lineage analysis on the abnormal water service data to determine the abnormal topology graph of the abnormal water service data, including: Perform data lineage analysis on the abnormal water service data to determine the lineage relationship of the abnormal water service data, where the lineage relationship includes data processing rules and data sources; Determine the influence range of the abnormal water service data on other water service data, and the influence range includes the affected tables and fields; Generate an abnormal topology graph in the preset data map, where the abnormal topology graph includes the lineage relationship and the influence range.
[0015] In a second aspect, the present application provides a data processing server deployed with a water service data management platform, including: A memory configured to store instructions; and A processor configured to call the instructions from the memory and be able to implement the above-mentioned water service data anomaly identification method when executing the instructions.
[0016] In a third aspect, the present application provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to cause a machine to execute the above-mentioned water service data anomaly identification method.
[0017] Through the above technical solution, the data exchange layer realizes seamless data exchange between different data sources through DML extraction and DDL extraction technologies, solves the problem of data heterogeneity, and ensures the consistency and integrity of data. The data processing layer performs task scheduling on the exchanged water service data according to the preset workflow, optimizes the data processing process, and improves the data processing efficiency. On this basis, through the task execution status data and the hardware resource usage data, combined with the preset anomaly recognition algorithm, anomaly data is accurately recognized in the water service data after task scheduling, significantly improving the accuracy of anomaly data recognition. Compared with the traditional statistical method, this technical solution can adapt to the characteristics of unstable data distribution in all links of the water service system, avoiding misjudgment and missed judgment problems caused by data fluctuations. In addition, this technical solution also generates an anomaly topology diagram through data lineage analysis technology, intuitively showing the source, propagation path, and influence range of anomaly data, providing strong support for fault troubleshooting and decision-making optimization of the water service system. The visualization of the anomaly topology diagram enables operation and maintenance personnel to quickly locate the root cause of the problem, take targeted measures, and reduce the impact of anomaly data on the operation of the water service system. In summary, the accuracy of anomaly recognition in the water service system is significantly improved.
[0018] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings
[0019] Figure 1 It is a schematic flowchart of a method for identifying anomalies in water service data provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a water service data management platform provided by an embodiment of the present application; Figure 3 It is a schematic diagram of a data source provided by an embodiment of the present application. Detailed Description of the Embodiments
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present application, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0022] In addition, if the descriptions such as "first" and "second" are involved in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.
[0023] Figure 1 Schematically shows a flowchart of a method for identifying abnormal water service data according to an embodiment of the present application. As Figure 1 shown, the embodiment of the present application provides a method for identifying abnormal water service data, which is applied to a data processing server. The data processing server is deployed with a water service data management platform. The water service data management platform includes a data access layer, a data exchange layer, a data processing layer, and a monitoring layer. The method may include the following steps.
[0024] S110. In response to receiving water service data from different data sources, perform multi-source access to the water service data from different data sources through the data access layer; S120. Perform DML extraction and DDL extraction on the water service data through the data exchange layer to achieve data exchange between different data sources; S130. Perform task scheduling on the water service data after data exchange according to a preset workflow through the data processing layer; S140. Obtain task execution status data and hardware resource usage data through the monitoring layer; S150. Based on the task execution status data and the hardware resource usage data, identify abnormal water service data in the water service data after task scheduling; S160. Perform data lineage analysis on the abnormal water service data to determine the abnormal topology graph of the abnormal water service data and visualize the abnormal topology graph.
[0025] Figure 2 Is a schematic structural diagram of a water service data management platform provided by an embodiment of the present application. As Figure 2As shown in the figure, the data access layer is used to connect to water service data from different sources (such as sensors, databases, APIs, etc.); the data exchange layer is used to standardize the data format and achieve cross-data source conversion (DML for data operation processing, DDL for table structure processing); the data processing layer is used to define the data processing process and manage the execution order and resource allocation of data processing tasks; the monitoring layer is used to track the task execution progress and success rate and monitor hardware metrics such as CPU, memory, and storage.
[0026] After receiving water service data from different data sources, the data access layer first performs multi-source access processing on this data. Water service data may come from a variety of different systems, such as sensor networks, databases, external data interfaces, etc. The data formats, transmission protocols, and storage methods of the above data sources are all different. The data access layer realizes data access by identifying the data types of each data source (such as structured data, semi-structured data, and unstructured data) and calling the corresponding data access plug-ins. For example, for structured data (such as data in a relational database), the data access layer will use an SQL query plug-in to extract the data; for semi-structured data (such as data in XML or JSON format), the data access layer will use the corresponding parser to parse the data; for unstructured data (such as text files or image data), the data access layer will use a preset processing module to extract useful information.
[0027] The data access layer also supports two data access methods: passive reception (push) and active pulling (pull) to ensure that data can be transmitted in real time or regularly. The data exchange layer is used to perform DML (Data Manipulation Language) and DDL (Data Definition Language) extraction on water service data to achieve data exchange between different data sources. DML extraction mainly targets data insert, update, and delete operations (such as insert, update, delete statements) to extract the changed parts of the data; DDL extraction targets the database structure definition (such as tablespaces, users, views, indexes, etc.) to extract the metadata information of the database.
[0028] The data exchange layer establishes a mapping relationship between different data sources through a visual configuration interface to ensure that data can be transmitted seamlessly between different systems. For example, when transmitting data from a relational database (such as Oracle) to a NoSQL database (such as MongoDB), the data exchange layer will convert the SQL query results into JSON format and store them according to the storage structure of the target database. The data exchange layer also supports two modes: full-volume collection and incremental collection. Full-volume collection is used for the first data migration, and incremental collection is used for subsequent data synchronization.
[0029] The data processing layer schedules tasks for the water service data after data exchange according to a preset workflow. The workflow consists of multiple tasks, including data cleaning, data transformation, data quality detection, and data analysis, etc. The data processing layer configures the execution order and dependency relationships of each task through the workflow management module, and determines the priority and execution time of the tasks through the task scheduling module. For example, the data cleaning task needs to be executed before the data transformation task to ensure the accuracy and consistency of the data; the data analysis task needs to be executed after all data processing tasks are completed to generate the final analysis results. The task scheduling module can also adopt various scheduling mechanisms such as scheduled execution and execution after dependent tasks are completed to ensure that tasks can be completed on time.
[0030] In addition, the task scheduling module can also set up timeout alarms and retry mechanisms to handle possible exceptions during task execution.
[0031] In this embodiment, the monitoring layer is responsible for obtaining task execution status data and hardware resource usage data to monitor the running status of the system in real time. The task execution status data includes the waiting status, running status, successful completion status, and failure status of the tasks, etc. These data are collected and displayed in real time through the task status monitoring module. For example, when a certain task is in the running state for a long time, the monitoring layer will issue a timeout alarm to prompt the user to intervene; when the task execution fails, the monitoring layer will record the failure reason and provide a retry option. The hardware resource usage data includes CPU usage rate, memory occupancy, disk space, etc. These data are collected and displayed in real time through the hardware resource monitoring module. For example, when the CPU usage rate exceeds the preset threshold, the monitoring layer will issue a resource shortage alarm to prompt the user to optimize resources. The monitoring layer also provides a historical data query function, and users can query the task execution history and hardware resource usage according to conditions such as time range and task name.
[0032] Technical metadata describes the technical attributes of data (such as data structure, storage location, etc.), and business metadata describes the business meaning of data (such as data definition, business rules, etc.). By establishing the association relationship between technical metadata and business metadata, a water service data resource catalog is generated. Then, the task execution status data is analyzed to identify abnormal task execution. For example, when the task execution time exceeds the preset threshold or the task status is failure, it is determined that the task execution is abnormal. At the same time, the hardware resource usage data is analyzed to identify abnormal resource usage. For example, when the CPU usage rate, memory occupancy, or disk space exceeds the preset threshold, it is determined that the resource usage is abnormal. Finally, based on the abnormal task execution and abnormal resource usage, abnormal data is determined in the water service data after task scheduling. For example, when a certain data processing task fails due to insufficient resources, the data processed by this task is marked as abnormal data to improve the accuracy of abnormal data identification.
[0033] Perform data lineage analysis on abnormal water service data to determine the data lineage of the abnormal data (such as data processing rules and data sources) and its influence scope on other data (such as affected tables and fields). Data lineage analysis determines the source and propagation path of abnormal data by tracing the data processing path. For example, when a certain water quality monitoring data is marked as abnormal, then trace the collection, transmission, processing, and analysis process of this data to determine the source of the abnormal data (such as sensor failure or data processing error) and its influence on other data (such as related water quality analysis results). Then, generate an abnormal topology map in the preset data map. The abnormal topology map graphically displays the data lineage and influence scope of the abnormal data. For example, the abnormal topology map can show the source node, propagation path, and affected nodes of the abnormal data, helping users intuitively understand the distribution and influence of the abnormal data. Finally, visualize the abnormal topology map, and users can view the detailed information of the abnormal topology map through the interactive interface and take corresponding measures for processing.
[0034] In this embodiment, the data exchange layer realizes seamless data exchange between different data sources through DML extraction and DDL extraction technologies, solves the data heterogeneity problem, and ensures data consistency and integrity. The data processing layer performs task scheduling on the exchanged water service data according to the preset workflow, optimizes the data processing process, and improves the data processing efficiency. On this basis, through the task execution status data and hardware resource usage data, combined with the preset abnormal recognition algorithm, abnormal data is accurately recognized in the water service data after task scheduling, significantly improving the accuracy of abnormal data recognition. Compared with traditional statistical methods, this technical solution can adapt to the characteristics of unstable data distribution in all links of the water service system, avoiding misjudgment and missed judgment problems caused by data fluctuations. In addition, this technical solution also generates an abnormal topology map through data lineage analysis technology, intuitively showing the source, propagation path, and influence scope of the abnormal data, providing strong support for fault troubleshooting and decision-making optimization of the water service system. The visualization of the abnormal topology map enables operation and maintenance personnel to quickly locate the root cause of the problem, take targeted measures, and reduce the impact of abnormal data on the operation of the water service system. In summary, the accuracy of abnormal recognition in the water service system is significantly improved.
[0035] In one implementation manner of this embodiment, multi-source access to water service data from different data sources is performed through the data access layer, including the following steps: S210. Identify the data types of different data sources, where the data types include structured data, semi-structured data, and unstructured data; S220. Determine the corresponding data access plug-ins according to the data types; S230. Access the water service data of the data source in a passive reception or active pulling manner through the data access plug-ins; S240. Store the water service data into the original database according to the preset access-output distribution strategy.
[0036] During the data access process, it is first necessary to identify the data types of different data sources in order to select appropriate data access methods. Figure 3 A schematic diagram of a data source provided for an embodiment of this application is as Figure 3 shown. The data source types of different data sources may be different. For example, for data sources named water plant scheduling, meter reading system, and Internet of Things system, their corresponding data source type is MYSQL, while for the data source named tag library, its corresponding data source type is CLICKHOUSE. Data types are mainly divided into structured data, semi-structured data, and unstructured data. Structured data refers to data with a fixed format and clear field definitions, usually stored in relational databases such as SQL Server, MySQL, etc. Structured data has clear field types and clear data relationships, and can be directly queried and processed through SQL statements. Semi-structured data refers to data with a certain format but not completely fixed, such as data in XML, JSON, etc. Semi-structured data has a flexible field structure and usually contains tags or markers to define the hierarchical relationship of the data. Unstructured data refers to data without a fixed format, such as text files, images, videos, etc. Unstructured data has various formats and requires specific parsing methods to extract useful information. Methods for identifying data types include analyzing the metadata of the data source, file extensions, data content characteristics, etc. For example, for relational databases, the table structure and field types can be obtained by querying the metadata table of the database; for file data, the data type can be determined by the file extension (such as.xml,.json); for unstructured data, the file type can be determined by analyzing the first few bytes of the file content.
[0037] After identifying the data type, it is necessary to determine the corresponding data access plugin according to the data type. The data access plugin is used to process specific types of data and parse and extract data. For structured data, in this embodiment, an SQL query plugin is used to extract data from a relational database by executing SQL statements. For example, for a MySQL database, a JDBC plugin can be used to connect to the database and execute a SELECT statement to obtain data. For semi-structured data, XML or JSON parsing plugins can be used to extract data by parsing tags or markers. For example, for JSON-formatted data, a JSON parsing plugin can be used to parse the data into key-value pairs and extract the required fields. For unstructured data, text parsing plugins or image processing plugins are usually used to extract useful information through specific algorithms. For example, for text files, natural language processing plugins can be used to extract keywords and entities; for image files, image recognition plugins can be used to extract image features. The selection of the data access plugin depends not only on the data type but also on the transmission protocol and storage method of the data source. For example, for data transmitted through the HTTP protocol, an HTTP plugin is required to receive the data; for data stored in a distributed file system, an HDFS plugin is required to read the data.
[0038] The data access plugin supports two data access methods: passive reception (push) and active pulling (pull) to adapt to different data sources and business requirements. The passive reception method means that the data source actively transmits data to the data access layer, and the data access layer receives the data by listening on a port or an API interface. For example, for real-time sensor data, the sensor will regularly send the data to a specified port of the data access layer, and the data access layer receives and stores the data by listening on the port. The active pulling method means that the data access layer actively requests data from the data source, usually used to obtain data regularly or on demand. For example, for historical data stored in a relational database, the data access layer will regularly execute SQL query statements to extract data from the database and store it. The data access plugin also supports multiple transmission protocols, such as HTTP, FTP, MQTT, etc., to adapt to different data sources and network environments. For example, for data transmitted through the HTTP protocol, the data access plugin will send an HTTP request and receive the response data; for data transmitted through the MQTT protocol, the data access plugin will subscribe to a specified topic and receive messages. The data access plugin also supports data compression and encrypted transmission to improve the efficiency and security of data transmission. For example, for large-volume data transmission, the data access plugin can use the GZIP compression algorithm to compress the data; for the transmission of sensitive data, the data access plugin can use the SSL / TLS protocol to encrypt the data.
[0039] After accessing water utility data, it is necessary to store the data in the original database according to the preset access output distribution strategy. The access output distribution strategy defines the storage location, storage format, and distribution rules of the data to ensure that the data can be correctly stored and subsequently processed. The storage location can be a relational database, a NoSQL database, a distributed file system, etc., depending on the data type and business requirements. For example, for structured data, it is usually stored in a relational database; for semi-structured data, it is usually stored in a NoSQL database; for unstructured data, it is usually stored in a distributed file system. The storage format can be a table structure, a document structure, a file format, etc., depending on the data type and storage location. For example, for structured data, it is usually stored in a table structure; for semi-structured data, it is usually stored in a document structure; for unstructured data, it is usually stored in a file format. The distribution rules define the storage path and naming rules of the data to ensure that the data can be effectively managed and retrieved. For example, for real-time sensor data, files can be named and stored according to the timestamp and sensor ID; for historical data, files can be named and stored according to the date and data type. The access output distribution strategy also supports data partitioning and indexing to improve the query efficiency of the data. For example, for large-scale data storage, it can be partitioned and stored according to time or region; for frequently queried data, indexes can be created to speed up the query speed. By storing the data in the original database according to the preset access output distribution strategy, it can be ensured that the data can be stored and managed efficiently and orderly, providing a reliable data foundation for subsequent data analysis and applications.
[0040] In this embodiment, by identifying the data type and selecting the appropriate data access plug-in, it is ensured that the data can be parsed and extracted efficiently and accurately. Then, by supporting two data access methods, passive reception and active pulling, it flexibly adapts to different data sources and business requirements, ensuring that the data can be transmitted to the system efficiently and securely. Finally, the data is stored in the original database according to the preset access output distribution strategy, ensuring that the data can be stored and managed efficiently and orderly. Through this multi-source access mechanism, the data access layer can integrate water utility data from different data sources, providing a unified data foundation for subsequent data processing and analysis, significantly improving the efficiency and reliability of data processing, and providing strong support for the operation and decision-making of the water utility system.
[0041] In one implementation of this embodiment, the data source includes a relational database. The water utility data is extracted by DML and DDL through the data exchange layer to achieve data exchange between different data sources, including the following steps: S310. Aggregate the water utility data from different data sources through the data exchange layer, where the water utility data in the relational database is obtained through full-volume collection or incremental collection; S320. Establish the mapping relationship of heterogeneous systems through the visualization configuration interface of the data exchange layer, and perform DML extraction and DDL extraction on water service data to achieve data exchange between different data sources. Each data source corresponds to a data operating system, and the heterogeneous systems are different data operating systems.
[0042] The data exchange layer first aggregates the water service data from different data sources. The data in the relational database can be obtained through full-scale collection or incremental collection. Full-scale collection means that when migrating or synchronizing data for the first time, all the data in the database is extracted at one time, which is suitable for scenarios with a small amount of data or requiring full synchronization. For example, when initializing the water service data management platform, historical data can be fully imported into the system through full-scale collection. Incremental collection means that in subsequent data synchronization, only the changed data in the database is extracted, which is suitable for scenarios with a large amount of data or requiring real-time synchronization. For example, during daily operation, newly added or modified data can be imported into the system in real time through incremental collection. Incremental collection is usually achieved through timestamps, logs, or triggers. For example, by recording the last modification timestamp of each piece of data, only the data that has changed within the specified time range is extracted; or by parsing the transaction logs of the database, newly added, modified, or deleted data is extracted. The data exchange layer also supports multiple data collection strategies, such as scheduled collection, event-triggered collection, etc., to adapt to different business requirements. For example, scheduled collection can be set to occur every day at midnight, or collection can be triggered when a specific event (such as data update) occurs. Through the combination of full-scale collection and incremental collection, the data exchange layer can efficiently aggregate the water service data from different data sources, ensuring the integrity and real-time nature of the data.
[0043] After aggregating water service data, the data exchange layer establishes mapping relationships between heterogeneous systems through a visual configuration interface and performs DML (Data Manipulation Language) extraction and DDL (Data Definition Language) extraction on the water service data to achieve data exchange between different data sources. Heterogeneous systems refer to data sources using different data operating systems, such as Oracle, MySQL, SQL Server, etc. The visual configuration interface allows users to define mapping relationships between different data sources by dragging and connecting nodes. For example, a table in an Oracle database can be mapped to a table in a MySQL database, and the corresponding relationships between fields can be specified. DML extraction mainly targets the operations of adding, deleting, and modifying data, extracting the changed parts of the data. For example, by parsing SQL statements, newly added or modified data records can be extracted. DDL extraction mainly targets the structure definition of the database, extracting the metadata information of the database. For example, by parsing DDL statements, table structures and field definitions can be extracted. The data exchange layer also supports various data conversion rules, such as field type conversion, data format conversion, etc., to ensure that data can be seamlessly transmitted between different systems. For example, a DATE type field in an Oracle database can be converted to a DATETIME type field in a MySQL database; or XML format data in a SQL Server database can be converted to JSON format data. By establishing mapping relationships for heterogeneous systems and performing DML extraction and DDL extraction on water service data, the data exchange layer can efficiently achieve data exchange between different data sources, ensuring data consistency and integrity.
[0044] In this embodiment, through the methods of full - volume collection and incremental collection, the integrity and real - time nature of the data can be ensured. Then, by establishing mapping relationships between heterogeneous systems through a visual configuration interface and performing DML extraction and DDL extraction on water service data, it can be ensured that data can be seamlessly transmitted between different systems. Through the data exchange mechanism, the data exchange layer can effectively solve the problem of data heterogeneity, ensure data consistency and integrity, provide a reliable data basis for subsequent data processing and analysis, and significantly improve the efficiency and reliability of data processing.
[0045] In one implementation of this embodiment, the data processing layer includes a workflow management module and a task scheduling module. The data processing layer performs data scheduling on the water service data after data exchange according to a preset workflow, including the following steps: S410. Configure the workflow through the workflow management module. The workflow includes data cleaning tasks, data conversion tasks, data quality detection tasks, and data analysis tasks; S420. Determine the dependencies and priorities among the data cleaning task, data transformation task, data quality detection task, and data analysis task through the task scheduling module, and perform task scheduling on the water service data after data exchange according to the dependencies and priorities.
[0046] The workflow management module is responsible for configuring the workflow of data processing. The workflow consists of multiple tasks, including data cleaning tasks, data transformation tasks, data quality detection tasks, and data analysis tasks. The data cleaning task refers to preprocessing the original data to remove noise, fill in missing values, correct errors, etc. For example, for the sensor data in the water service data, the data cleaning task can include removing outliers (such as data outside the reasonable range), filling in missing values (such as using interpolation to fill in missing time series data), and correcting errors (such as correcting data deviations caused by sensor drift). The data transformation task refers to converting data from one format or structure to another to meet the needs of subsequent processing and analysis. For example, for water service data from different data sources, the data transformation task can include converting XML-format data to JSON format, or converting table data in a relational database to a matrix format suitable for input to machine learning algorithms. The data quality detection task refers to checking the accuracy, integrity, and consistency of the data to ensure that the data quality meets the requirements. For example, the data quality detection task can include checking the integrity of data fields (such as ensuring that all required fields have values), the consistency of data (such as ensuring that the data in different data sources for the same entity is consistent), and the accuracy of data (such as verifying the reasonableness of data through business rules). The data analysis task refers to performing statistical analysis, pattern recognition, and predictive analysis on the cleaned and transformed data to extract valuable information. For example, the data analysis task can include performing time series analysis on the water service data to predict future water quality change trends; or performing clustering analysis to identify different water quality patterns. The workflow management module allows users to define the execution order and parameter configuration of each task by dragging and connecting nodes through a visual interface. For example, users can define a workflow that first performs the data cleaning task, then the data transformation task, then the data quality detection task, and finally the data analysis task. By configuring the workflow, the workflow management module can ensure the integrity and logic of the data processing process, providing basic support for subsequent task scheduling.
[0047] The task scheduling module is used to determine the dependency relationships and priorities among various tasks in the workflow, and perform task scheduling on the water service data after data exchange based on these dependency relationships and priorities. The dependency relationship refers to the execution sequence relationship among tasks, that is, some tasks must be executed after other tasks are completed. For example, the data cleaning task must be executed before the data transformation task because only after data cleaning can the accuracy and consistency of the data be ensured; the data analysis task must be executed after the data quality detection task because only when the data quality meets the requirements can effective analysis be carried out. The priority refers to the importance and urgency of tasks, that is, some tasks need to be executed preferentially. For example, for real-time water service data, the data cleaning task and the data transformation task need to be executed preferentially to ensure the real-time and availability of the data; for historical water service data, the data analysis task can be executed later to make full use of computing resources. The task scheduling module generates a task scheduling plan by analyzing the dependency relationships and priorities of various tasks in the workflow. For example, a task scheduling plan can be generated to first execute the high-priority data cleaning task and data transformation task, then execute the data quality detection task, and finally execute the data analysis task.
[0048] In this embodiment, the task scheduling module also supports multiple scheduling strategies, such as timed scheduling, event-triggered scheduling, etc., to adapt to different business requirements. For example, timed scheduling can be set to occur every day at midnight, or scheduling can be triggered when a specific event (such as data update) occurs. The task scheduling module also provides task monitoring and alarm functions, monitors the execution status of tasks in real time, and issues alarms when tasks fail to execute or time out. For example, when a certain task is in a running state for a long time, the task scheduling module will issue a timeout alarm to prompt the user to intervene; when a task fails to execute, the task scheduling module will record the reason for the failure and provide a retry option. Through the task scheduling module, the data processing layer can efficiently manage complex data processing processes, ensure that tasks can be completed on time, and improve the efficiency and reliability of data processing.
[0049] The data processing layer of this embodiment can efficiently manage complex data processing processes, ensuring that data cleaning, data transformation, data quality detection, and data analysis tasks can be completed on time. First, by configuring the workflow through the workflow management module, the integrity and logic of the data processing process can be ensured. Then, through the task scheduling module, the dependency relationships and priorities among tasks are determined, a task scheduling plan is generated, and the execution status of tasks is monitored in real time. Through the task scheduling mechanism, the data processing layer can efficiently manage complex data processing processes, improving the efficiency and reliability of data processing.
[0050] In one implementation of this embodiment, the monitoring layer includes a task status monitoring module and a hardware resource monitoring module. Obtaining task execution status data and hardware resource usage data through the monitoring layer includes the following steps: S510. Real-time monitor the task execution status data of each task in the workflow through the task status monitoring module. The task execution status data includes the waiting status, running status, successfully completed status, and failure status; S520. Obtain the hardware resource usage data through the hardware resource monitoring module. The hardware resource usage data includes CPU usage rate, memory occupancy, and disk space.
[0051] The task status monitoring module is responsible for real-time monitoring of the execution status of each task in the workflow to ensure that tasks can be executed as planned. The task execution status data includes the waiting status, running status, successfully completed status, and failure status. The waiting status indicates that the task has been submitted but has not started execution. The running status indicates that the task is being executed, and at this time, the progress and execution log of the task can be updated in real time. The successfully completed status indicates that the task has been successfully executed, and the execution result and output data of the task can be recorded. The failure status indicates that an error occurred during the task execution, and the error reason can be recorded and a retry option can be provided.
[0052] The task status monitoring module obtains the task execution status data in real time through polling or event triggering. For example, the execution status of the task can be polled once per second, or an event notification can be triggered when the task status changes. The task status monitoring module also provides a real-time log viewing function. Users can refresh the execution log of the task in real time on the monitoring interface, which is convenient for troubleshooting and problem location. For example, when a certain task fails, the user can view the execution log to understand the specific reason for the task failure and take corresponding measures to handle it. By real-time monitoring the execution status of the task, the task status monitoring module can timely detect abnormal situations in the task execution and take corresponding measures to handle them, ensuring that the task can be completed on time and improving the efficiency and reliability of data processing.
[0053] The hardware resource monitoring module is responsible for real-time monitoring of the usage of hardware resources in the cluster to ensure the efficient operation of the system. The hardware resource usage data includes CPU usage rate, memory occupancy, and disk space. The CPU usage rate represents the utilization rate of the CPU, usually expressed as a percentage. A high CPU usage rate may cause the system response to slow down or task execution to be delayed. Memory occupancy represents the usage of system memory, usually expressed as a percentage or a specific value. High memory occupancy may lead to insufficient system memory and affect task execution. Disk space represents the available space on the disk, usually expressed as a percentage or a specific value. Insufficient disk space may cause data not to be stored or task execution to fail. The hardware resource monitoring module generates a hardware resource usage report by collecting and statistically analyzing the hardware resource usage data of each node in the cluster. For example, it is possible to set to collect CPU usage rate, memory occupancy, and disk space data once per minute and generate corresponding line charts or pie charts to facilitate users' resource analysis and adjustment.
[0054] In this embodiment, the hardware resource monitoring module also provides real-time monitoring charts to graphically display the real-time changes in hardware resources. For example, the real-time change trend of the CPU usage rate can be displayed through a line chart, or the real-time distribution of memory occupancy can be displayed through a pie chart. By real-time monitoring the usage of hardware resources, the hardware resource monitoring module can promptly detect resource bottlenecks and abnormal situations and take corresponding measures for adjustment and optimization to ensure the stability and efficiency of the system.
[0055] In this implementation, the task status monitoring module is used to real-time monitor the execution status of tasks, promptly detect abnormal situations during task execution, and take corresponding measures for handling to ensure that tasks can be completed on time. Then, the hardware resource monitoring module real-time monitors the usage of hardware resources, promptly detects resource bottlenecks and abnormal situations, and takes corresponding measures for adjustment and optimization to ensure the high efficiency of anomaly identification.
[0056] In one implementation of this embodiment, based on the task execution status data and the hardware resource usage data, abnormal water service data is identified from the water service data after task scheduling, including the following steps: S610. Obtain technical metadata and business metadata. The technical metadata is used to describe the technical attributes of the data, and the business metadata is used to describe the business meaning of the data; S620. Establish an association relationship between the technical metadata and the business metadata to obtain a water service data resource catalog; S630. Analyze the task execution status data based on the water service data resource catalog to identify abnormal task execution; S640. Analyze the hardware resource usage data based on the water service data resource catalog to identify abnormal resource usage; S650. Determine the abnormal water service data in the water service data after task scheduling based on task execution exceptions and resource usage exceptions.
[0057] In the process of identifying abnormal water service data, it is first necessary to obtain technical metadata and business metadata. Technical metadata refers to information that describes the technical attributes of data, including the storage location of data, data structure, data format, data source, etc. For example, for sensor data in water service data, technical metadata may include the ID of the sensor, data collection time, data storage location (such as database table name or file path), data format (such as CSV, JSON), etc. Business metadata refers to information that describes the business meaning of data, including data definitions, business rules, data quality standards, etc. For example, for water quality data in water service data, business metadata may include the names of water quality parameters (such as pH value, dissolved oxygen), units of parameters (such as mg / L), reasonable ranges of parameters (such as the pH value should be between 6.5 and 8.5), etc. Methods for obtaining technical metadata and business metadata include querying the metadata tables of the database, parsing the metadata part of the data file, calling API interfaces to obtain metadata, etc. For example, the table structure and field types can be obtained by querying the system tables of a relational database (such as ALL_TABLES and ALL_TAB_COLUMNS in Oracle); the business meaning of the data can be obtained by parsing the metadata part in the JSON file.
[0058] After obtaining technical metadata and business metadata, it is necessary to establish an association relationship between the two to generate a water service data resource catalog. The water service data resource catalog is a structured data catalog that contains the technical attributes and business meaning of the data, facilitating users to search for and use the data. Methods for establishing the association relationship include mapping tables, rule engines, machine learning models, etc. For example, the field names in the technical metadata can be associated with the parameter names in the business metadata through a mapping table; the data types in the technical metadata can be associated with the units in the business metadata through a rule engine; the data structure in the technical metadata can be associated with the business rules in the business metadata through a machine learning model. The generation process of the water service data resource catalog includes steps such as data cleaning, data transformation, and data integration. For example, redundant information in the technical metadata can be cleaned, inconsistent formats in the business metadata can be transformed, and the association relationships in the technical metadata and business metadata can be integrated. By generating the water service data resource catalog, a unified data basis can be provided for subsequent abnormal data identification, ensuring that the data can be parsed and analyzed in a consistent manner.
[0059] Based on the water service data resource catalog, analyze the task execution status data to identify abnormal task execution. The task execution status data includes the waiting status, running status, successful completion status, and failure status of the task. In this embodiment, the steps of analyzing the task execution status data include statistical analysis, pattern recognition, and machine learning. Specifically, calculate metrics such as the average execution time, success rate, and failure rate of the task through statistical analysis; determine abnormal patterns during task execution (such as overly long task execution time, overly high task failure rate) through pattern recognition; and finally predict abnormal situations during task execution (such as the probability of task execution failure) through a machine learning model.
[0060] Based on the water service data resource catalog, analyze the hardware resource usage data to identify abnormal resource usage. The hardware resource usage data includes CPU usage rate, memory occupancy, disk space, etc. The hardware resource usage data can be analyzed through steps such as statistical analysis, trend analysis, and anomaly detection. Specifically, calculate metrics such as the average usage rate and peak usage rate of the hardware resources through statistical analysis; determine abnormal trends in hardware resource usage (such as continuous increase in CPU usage rate, continuous increase in memory occupancy) through trend analysis; and identify abnormal situations in hardware resource usage (such as CPU usage rate exceeding the preset threshold, memory occupancy exceeding the preset threshold) through anomaly detection algorithms.
[0061] Finally, based on the abnormal task execution and abnormal resource usage, determine the abnormal water service data in the water service data after task scheduling. The steps of determining the abnormal water service data can include rule matching, statistical analysis, machine learning, etc. Specifically, the abnormal task execution and abnormal resource usage can be associated with the water service data through rule matching; calculate abnormal metrics of the water service data (such as data deviation, data fluctuation) through statistical analysis, such as: (1) Zero water volume rule: Rule description: The water meter reading is 0 for a long time (possibly damaged, out of service, or stolen). Configuration parameters: Continuous days threshold: Generally set to 0 readings for 7 consecutive days. Exclusion conditions: User-reported service suspension, newly installed meter (generally in the first 10 days).
[0062] (2) Sudden increase / sudden decrease rule: Rule description: The daily water consumption exceeds the historical normal range.
[0063] Configuration parameters: Sudden increase threshold: For example, when the daily water volume > 1 times the average of the past 30 days.
[0064] Sudden decrease threshold: The daily water volume < 50% of the average of the past 30 days.
[0065] (3) Constant value rule: Rule description: The water meter reading has not changed for several consecutive days (the meter may be stuck).
[0066] Configuration parameter: The readings are exactly the same for N consecutive days, usually 5 days.
[0067] (4) Night minimum flow rule: Rule description: Abnormal water consumption at night (there may be a pipeline leak).
[0068] Configuration parameter: Time period: 00:00 - 05:00.
[0069] Flow threshold: Night flow > 30% of the average daily flow.
[0070] (5) Seasonal fluctuation anomaly: Rule description: Water consumption does not conform to the seasonal pattern (e.g., water consumption increases instead in winter).
[0071] Configuration parameter: Historical year-on-year threshold: Current month's water consumption > last year's same period ± 30%.
[0072] (6) Household average comparison rule: Rule description: The water consumption of a single household is significantly higher than that of similar users (e.g., the water consumption of a single elderly person living alone is the same as that of a family).
[0073] Configuration parameter: Grouping of similar users: Divided by housing type and region.
[0074] Deviation threshold: Water volume > 1.5 times the standard deviation of the average of users in the same group.
[0075] By predicting the anomaly of water service data (such as the probability of data anomaly) through a machine learning model, the abnormal water service data can be determined, such as: 1) Time series prediction Rule description: Predict water consumption based on ARIMA or LSTM, and the actual value exceeding the confidence interval is regarded as abnormal.
[0076] Configuration parameter: Confidence interval: Such as 90% confidence band.
[0077] Training data period: At least 6 months of historical data.
[0078] (2) Clustering anomaly detection Rule description: Mark outliers through K-Means or DBSCAN clustering.
[0079] Configuration parameter: Feature dimensions: Average daily water volume, volatility, time period distribution, etc.
[0080] This embodiment can identify abnormal water service data in the water service data after task scheduling based on task execution status data and hardware resource usage data, effectively improving the accuracy of abnormal water service data identification. Establishing an association relationship between technical metadata and business metadata to generate a water service data resource catalog can provide a unified data basis for abnormal data identification. Based on the water service data resource catalog, analyze the task execution status data and hardware resource usage data to identify task execution anomalies and resource usage anomalies. Finally, based on task execution anomalies and resource usage anomalies, determine abnormal water service data in the water service data after task scheduling, which can significantly improve the accuracy rate of abnormal data identification.
[0081] In one implementation of this embodiment, establishing an association relationship between technical metadata and business metadata to obtain a water service data resource catalog includes the following steps: S710. Extract data structure information, storage location information, and data format information in the technical metadata; S720. Extract data definition information, business rule information, and data quality standard information in the business metadata; S730. Establish a mapping relationship between the technical metadata and the business metadata; S740. Generate a water service data resource catalog according to the mapping relationship.
[0082] Based on the water service data resource catalog, analyze the task execution status data to identify task execution anomalies, including: S11. Obtain the historical execution status data of each task in the workflow; S12. Based on the historical execution status data, establish a normal behavior model for task execution; S13. Compare the current task execution status data with the normal behavior model to obtain the task status and task execution time; S14. When the task status is a failure status or the task execution time exceeds the corresponding second preset threshold, determine that the task execution is abnormal.
[0083] Based on the water service data resource catalog, analyze the hardware resource usage data to identify resource usage anomalies, including: S21. Obtain the historical hardware resource usage data of system operation; S22. Based on the historical hardware resource usage data, establish a normal range model for resource usage; S23. Compare the current hardware resource usage data with the normal range model; S24. When any one of the CPU usage rate, memory occupancy, or disk space exceeds the corresponding second preset threshold, determine that the resource usage is abnormal.
[0084] Before establishing the association relationship between technical metadata and business metadata, it is first necessary to extract key information from technical metadata, including data structure information, storage location information, and data format information. Data structure information describes the organization of data, such as table structure, field type, primary key, foreign key, etc. For example, for water service data in a relational database, data structure information may include table name, field name, field type (such as VARCHAR, INT), field length, etc. Storage location information describes the physical storage location of data, such as database instance, tablespace, file path, etc. For example, for water service data in a distributed file system, storage location information may include file path, file size, file creation time, etc. Data format information describes the storage format of data, such as CSV, JSON, XML, etc. For example, for water service data in JSON format, data format information may include the encoding method of the JSON file (such as UTF-8), the key-value pair structure of the JSON object, etc. Methods for extracting technical metadata include querying the system tables of the database, parsing the metadata section of the data file, calling API interfaces to obtain metadata, etc. For example, the table structure and field type can be obtained by querying the ALL_TABLES and ALL_TAB_COLUMNS tables in the Oracle database; the encoding method and key-value pair structure of the data can be obtained by parsing the beginning part of the JSON file.
[0085] After extracting technical metadata, it is necessary to extract key information from business metadata, including data definition information, business rule information, and data quality standard information. Data definition information describes the business meaning of data, such as the name, unit, value range of the field, etc. For example, for water quality parameters in water service data, data definition information may include parameter name (such as pH value), parameter unit (such as mg / L), reasonable range of the parameter (such as the pH value should be between 6.5 and 8.5), etc. Business rule information describes the business logic of data, such as calculation rules, verification rules of data, etc. For example, for water pressure data in water service data, business rule information may include the calculation formula of water pressure (such as water pressure = water level × acceleration due to gravity), verification rules of water pressure (such as water pressure should be between 0 and 100 kPa), etc. Data quality standard information describes the quality requirements of data, such as integrity, consistency, accuracy of data, etc. For example, for flow data in water service data, data quality standard information may include integrity requirements of data (such as all required fields have values), consistency requirements of data (such as flow data at the same time point should be consistent), accuracy requirements of data (such as flow data should be within a reasonable range), etc.
[0086] The methods for extracting business metadata include querying the metadata table of the business system, parsing business documents, calling API interfaces to obtain metadata, etc. For example, data definition information and business rule information can be obtained by querying the metadata table of the business system; data quality standard information can be obtained by parsing business documents.
[0087] After extracting technical metadata and business metadata, establish a mapping relationship between the two to ensure that technical attributes and business meanings can correspond one by one. Specifically, the mapping relationship can be established through mapping tables, rule engines, machine learning models, etc. For example, the field names in technical metadata can be associated with the parameter names in business metadata through a mapping table; the data types in technical metadata can be associated with the units in business metadata through a rule engine; the data structures in technical metadata can be associated with the business rules in business metadata through a machine learning model.
[0088] Among them, the process of establishing the mapping relationship includes the following steps: cleaning redundant information in technical metadata, converting inconsistent formats in business metadata, and integrating the association relationships in technical metadata and business metadata. By establishing the mapping relationship between technical metadata and business metadata, it can provide basic support for the subsequent generation of the water service data resource catalog, ensuring that data can be parsed and analyzed in a consistent manner.
[0089] After establishing the mapping relationship between technical metadata and business metadata, it is necessary to generate a water service data resource catalog according to the mapping relationship. The water service data resource catalog is a structured data catalog that contains the technical attributes and business meanings of the data, facilitating users to search for and use the data. The methods for generating the water service data resource catalog include steps such as data cleaning, data conversion, and data integration. For example, redundant information in technical metadata can be cleaned, inconsistent formats in business metadata can be converted, and the association relationships in technical metadata and business metadata can be integrated. The generation process of the water service data resource catalog also includes steps such as data classification, data tagging, and data indexing. For example, data can be classified according to attributes such as the theme, type, format, and source of the data; keywords or tags can be added to the data to facilitate user search and filtering of the data; a data index can be created to improve the query efficiency of the data. By generating the water service data resource catalog, it can provide a unified data basis for subsequent identification of abnormal data, ensuring that data can be parsed and analyzed in a consistent manner.
[0090] Before the abnormal execution of the recognition task, it is necessary to obtain the historical execution status data of each task in the workflow. The historical execution status data includes the waiting status, running status, successful completion status, and failure status of the task. The methods for obtaining the historical execution status data include querying the logs of the task scheduling system, calling the API interface to obtain historical data, etc. For example, the execution time, execution result, execution log, etc. of each task can be obtained by querying the logs of the task scheduling system; the historical execution status data of each task can be obtained by calling the API interface. The process of obtaining the historical execution status data also includes steps such as data cleaning, data transformation, and data integration. For example, the noise information in the historical execution status data can be cleaned, inconsistent data formats can be transformed, and the historical execution status data of different tasks can be integrated. By obtaining the historical execution status data of each task in the workflow, it can provide basic support for the establishment of the subsequent normal behavior model and ensure that task execution anomalies can be accurately identified.
[0091] After obtaining the historical execution status data, it is necessary to establish a normal behavior model for task execution based on this data. The normal behavior model describes the execution characteristics of tasks under normal circumstances, such as the execution time, execution result, execution log, etc. of the task. The normal behavior model can calculate indicators such as the average execution time, success rate, and failure rate of tasks through statistical analysis, and discover the normal patterns in task execution (such as the distribution of task execution time, the distribution of task success rate) through pattern recognition; it can also predict the normal behavior of task execution (such as the predicted value of task execution time, the predicted value of task success rate) through machine learning models. The process of establishing the normal behavior model also includes steps such as data preprocessing, feature extraction, model training, and model evaluation. For example, the noise information in the historical execution status data can be preprocessed, features such as task execution time and task dependency can be extracted, a machine learning model can be trained, and the accuracy of the model can be evaluated. By establishing a normal behavior model for task execution, it can provide basic support for the subsequent identification of task execution anomalies and ensure that task execution anomalies can be accurately identified.
[0092] After establishing the normal behavior model, it is necessary to compare the current task execution status data with the normal behavior model to obtain the task status and task execution time. The deviation between the current task execution time and the average execution time in the normal behavior model can be calculated through statistical analysis. In addition, the difference between the current task execution status and the normal pattern in the normal behavior model can be discovered through pattern recognition. In another implementation, the machine learning model can also be used to predict whether the current task execution status is normal to determine whether the task status is the successful completion status or the failure status.
[0093] After comparing the current task execution status data with the normal behavior model, the task status can be compared with the failure status through rule matching, and the task execution time can be compared with the second preset threshold; the deviation between the task execution time and the average execution time in the normal behavior model can be calculated through statistical analysis. The determination process also includes steps such as data preprocessing, feature extraction, model prediction, and result analysis. For example, the noise information in the current task execution status data can be preprocessed, features such as task execution time and task dependency relationship can be extracted, and a machine learning model can be used for prediction to output the prediction result. By determining that the task execution is abnormal, it can provide basic support for subsequent abnormal data processing, ensure that the task can be completed on time, and improve the efficiency and reliability of data processing.
[0094] Before identifying abnormal resource usage, historical hardware resource usage data needs to be obtained. The historical hardware resource usage data includes CPU usage rate, memory occupancy, disk space, etc. The methods for obtaining historical hardware resource usage data include querying the logs of the hardware resource monitoring system, calling API interfaces to obtain historical data, etc. For example, the CPU usage rate, memory occupancy, disk space, etc. of each node can be obtained by querying the logs of the hardware resource monitoring system; the historical hardware resource usage data of each node can be obtained by calling API interfaces.
[0095] After obtaining the historical hardware resource usage data, a normal range model of resource usage needs to be established based on this data. The normal range model describes the usage characteristics of hardware resources under normal circumstances, such as CPU usage rate, memory occupancy, disk space, etc. The methods for establishing the normal range model include statistical analysis, pattern recognition, machine learning, etc. For example, indicators such as the average usage rate and peak usage rate of hardware resources can be calculated through statistical analysis; normal patterns in hardware resource usage (such as the distribution of CPU usage rate, the distribution of memory occupancy) can be discovered through pattern recognition; the normal range of hardware resource usage (such as the predicted value of CPU usage rate, the predicted value of memory occupancy) can be predicted through a machine learning model.
[0096] Among them, the steps for establishing the normal range model include: preprocessing the noise information in the historical hardware resource usage data, extracting features such as CPU usage rate and memory occupancy, training a machine learning model, evaluating the accuracy of the model, and finally obtaining the normal range model. It is not difficult to understand that in this embodiment, the normal range model can be a machine learning model.
[0097] After establishing the normal range model, it is necessary to compare the current hardware resource usage data with the normal range model. Similarly to S13, the deviation between the current CPU usage rate and the average usage rate in the normal range model can be calculated through statistical analysis; the difference between the current hardware resource usage data and the normal mode in the normal range model can be found through pattern recognition; and whether the current hardware resource usage data is normal can be predicted through a machine learning model. By comparing the current hardware resource usage data with the normal range model, it can provide basic support for subsequent determination of abnormal resource usage and ensure that abnormal resource usage can be accurately identified.
[0098] After comparing the current hardware resource usage data with the normal range model, it is necessary to determine abnormal resource usage based on the comparison result. The determination methods include rule matching, statistical analysis, machine learning, etc. For example, the CPU usage rate, memory occupancy, and disk space can be compared with the second preset threshold through rule matching; the deviation between the current hardware resource usage data and the average usage rate in the normal range model can be calculated through statistical analysis; and whether the current hardware resource usage data is normal can be predicted through a machine learning model. The determination process also includes steps such as data preprocessing, feature extraction, model prediction, and result analysis. For example, the noise information in the current hardware resource usage data can be preprocessed, features such as CPU usage rate and memory occupancy can be extracted, and a machine learning model can be used for prediction to output the prediction result.
[0099] This embodiment can generate a water service data resource catalog based on technical metadata and business metadata, providing a unified data basis for abnormal data identification. First, extract the key information in the technical metadata and business metadata, establish the mapping relationship between the two, and generate the water service data resource catalog. Then, based on the water service data resource catalog, analyze the task execution status data and hardware resource usage data to identify abnormal task execution and resource usage. Finally, determine the abnormal water service data based on the abnormal task execution and resource usage, which can significantly improve the accuracy of abnormal data identification.
[0100] In one implementation manner of this embodiment, perform data lineage analysis on the abnormal water service data to determine the abnormal topology graph of the abnormal water service data, including the following steps: S810. Perform data lineage analysis on the abnormal water service data to determine the lineage relationship of the abnormal water service data, where the lineage relationship includes data processing rules and data sources; S820. Determine the influence range of the abnormal water service data on other water service data, where the influence range includes the affected tables and fields; S830. Generate an abnormal topology graph in the preset data map, where the abnormal topology graph includes the lineage relationship and the influence range.
[0101] After identifying abnormal water utility data, it is necessary to conduct data lineage analysis on it to determine its lineage relationship. The lineage relationship refers to the processing path and source of data in the system, including data processing rules and data sources. Data processing rules describe the processing logic of data in the system, such as data cleaning rules, data transformation rules, data quality detection rules, etc. For example, for water quality data in water utility data, data processing rules can include data cleaning rules (such as removing outliers and filling missing values), data transformation rules (such as converting XML format to JSON format), data quality detection rules (such as checking data integrity, consistency, and accuracy), etc. Data sources describe the original sources of data, such as sensors, databases, external systems, etc.
[0102] For example, for water pressure data in water utility data, data sources can include pressure sensors, relational databases, external data interfaces, etc. Methods for data lineage analysis include querying data processing logs, parsing data processing rules, calling API interfaces to obtain data sources, etc. For example, the processing path and processing rules of data can be obtained by querying data processing logs; the cleaning, transformation, detection, etc. logic of data can be obtained by parsing data processing rules; the original source of data can be obtained by calling API interfaces. Through data lineage analysis, the lineage relationship of abnormal water utility data can be determined, providing basic support for subsequent impact scope analysis and abnormal topology map generation, ensuring that abnormal data can be accurately traced and processed.
[0103] After determining the lineage relationship of abnormal water utility data, it is necessary to further determine its impact scope on other water utility data. The impact scope refers to the spread and impact of abnormal data on other data in the system, including affected tables and fields. Affected tables refer to data tables related to abnormal data, such as tables with the same data source, tables with the same processing path, tables with the same business theme, etc. For example, for water quality data in water utility data, affected tables can include water quality monitoring tables, water quality analysis tables, water quality report tables, etc. Affected fields refer to fields related to abnormal data, such as fields with the same data source, fields with the same processing path, fields with the same business theme, etc. For example, for water pressure data in water utility data, affected fields can include pressure value fields, timestamp fields, sensor ID fields, etc. Methods for determining the impact scope include querying data dependencies, parsing data processing rules, calling API interfaces to obtain data relationships, etc. For example, tables and fields related to abnormal data can be obtained by querying data dependencies; the processing path and processing logic of data can be obtained by parsing data processing rules; the relationships and dependencies of data can be obtained by calling API interfaces. By determining the impact scope of abnormal water utility data, basic support can be provided for subsequent abnormal topology map generation, ensuring that the impact of abnormal data can be accurately evaluated and processed.
[0104] After determining the lineage relationship and impact range of abnormal water service data, it is necessary to generate an abnormal topology map in the preset data map. The abnormal topology map is a graphical data map that shows the lineage relationship and impact range of abnormal data, so that users can intuitively understand the distribution and impact of abnormal data. Methods for generating abnormal topology maps include data cleaning, data conversion, data integration, graphical display, etc. For example, you can clean the noise information in the abnormal water service data, convert inconsistent data formats, integrate lineage relationship and impact range data, and use graphical tools to generate abnormal topology maps. The abnormal topology map includes lineage relationship and impact range. The lineage relationship shows the processing path and source of abnormal data, and the impact range shows the impact of abnormal data on other data. For example, the abnormal topology map can display the source node, processing path node, affected node, etc. of the abnormal data, which is convenient for users to troubleshoot and locate problems.
[0105] In this embodiment, the process of generating an abnormal topology map also includes steps such as data classification, data labeling, and data indexing. For example, data can be classified according to attributes such as the subject, type, format, and source of the data; keywords or labels can be added to the data to facilitate users to search and filter data; data indexes can be created to improve data query efficiency. By generating an abnormal topology map, basic support can be provided for subsequent abnormal data processing to ensure that abnormal data can be accurately tracked and processed.
[0106] This implementation method performs data lineage analysis on abnormal water service data to determine its lineage relationship, including data processing rules and data sources. Then, the impact of abnormal water service data on other water service data is determined, including the affected tables and fields. Finally, an abnormal topology map is generated in the preset data map to display the lineage relationship and impact range of abnormal data, which significantly improves the accuracy of abnormal data processing.
[0107] In other embodiments, after analyzing the abnormal data, the abnormal data needs to be processed, and the processing method includes: 1. Automatic repair: (1) Data re-collection: triggering the device to retransmit the lost data; (2) Data correction: Calculate reasonable values based on equipment data (such as calculating missing values based on adjacent water meters).
[0108] 2. Manual processing: (1) Work orders are manually or automatically assigned to responsible persons, and if they are not processed within the time limit, an alarm is escalated; (2) The processing results are fed back to the system to optimize algorithm parameters (such as adjusting threshold sensitivity).
[0109] Common scenarios include:
[0110] An embodiment of the present application further provides a data processing server, which is deployed with a water service data management platform, including: a memory configured to store instructions; and a processor configured to call instructions from the memory and capable of implementing the above-mentioned water service data anomaly recognition method when executing the instructions.
[0111] An embodiment of the present application further provides a machine-readable storage medium, on which instructions are stored, and the instructions are used to cause a machine to execute the above-mentioned water service data anomaly recognition method.
[0112] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the functions in one Figure 1 one flow or multiple flows and / or blocksFigure 1 Steps of functions specified in one or more boxes.
[0116] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0117] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0118] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0119] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0120] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying abnormalities in water service data, characterized in that: Applied to a data processing server, the data processing server is deployed with a water affairs data management platform, the water affairs data management platform includes a data access layer, a data exchange layer, a data processing layer and a monitoring layer, the method includes: In response to receiving water service data from different data sources, performing multi-source access to the water service data from different data sources through a data access layer; DML and DDL extraction of water affairs data is performed through the data exchange layer to achieve data exchange between different data sources; The data processing layer performs task scheduling on the water affairs data after data exchange according to the preset workflow; Obtain task execution status data and hardware resource usage data through the monitoring layer; Based on the task execution status data and hardware resource usage data, abnormal water service data is identified in the water service data after task scheduling; Perform data lineage analysis on abnormal water service data to determine the abnormal topology of the abnormal water service data and visualize the abnormal topology.
2. The method according to claim 1, characterized in that The data access layer provides multi-source access to water data from different data sources, including: Identify the data types of different data sources, including structured data, semi-structured data, and unstructured data; Determine the corresponding data access plug-in according to the data type; Access water data from data sources through data access plug-ins in a passive or active manner; The water affairs data is stored in the original database according to the preset access output distribution strategy.
3. The method according to claim 1, characterized in that The data source includes a relational database. The water service data is extracted by DML and DDL through the data exchange layer to achieve data exchange between different data sources, including: The water affairs data from different data sources are aggregated through the data exchange layer, where the water affairs data from the relational database is obtained through full or incremental collection; The mapping relationship between heterogeneous systems is established through the visual configuration interface of the data exchange layer, and DML and DDL extraction are performed on water service data to realize data exchange between different data sources. Each data source corresponds to a data operating system, and heterogeneous systems are different data operating systems.
4. The method according to claim 1, characterized in that The data processing layer includes a workflow management module and a task scheduling module. The data processing layer schedules the water service data after data exchange according to the preset workflow, including: Configure the workflow through the workflow management module, the workflow includes data cleaning tasks, data conversion tasks, data quality detection tasks and data analysis tasks; The task scheduling module determines the dependencies and priorities among data cleaning tasks, data conversion tasks, data quality detection tasks and data analysis tasks, and performs task scheduling on the water affairs data after data exchange based on the dependencies and priorities.
5. The method according to claim 1, characterized in that The monitoring layer includes a task status monitoring module and a hardware resource monitoring module. The monitoring layer obtains task execution status data and hardware resource usage data, including: The task status monitoring module monitors the task execution status data of each task in the workflow in real time, wherein the task execution status data includes a waiting state, a running state, a successful completion state, and a failed state; The hardware resource usage data is obtained through the hardware resource monitoring module, and the hardware resource usage data includes CPU usage, memory occupancy and disk space.
6. The method according to claim 1, characterized in that Based on the task execution status data and hardware resource usage data, abnormal water service data is identified in the water service data after task scheduling, including: Obtain technical metadata and business metadata. Technical metadata is used to describe the technical attributes of data, and business metadata is used to describe the business meaning of data. Establish an association relationship between technical metadata and business metadata to obtain a water affairs data resource directory; Based on the water affairs data resource directory, analyze the task execution status data to identify task execution anomalies; Based on the water data resource directory, analyze the hardware resource usage data to identify abnormal resource usage; According to the abnormal task execution and abnormal resource usage, abnormal water service data is determined in the water service data after task scheduling.
7. The method according to claim 6, characterized in that By establishing an association between technical metadata and business metadata, a water affairs data resource directory is obtained, including: Extract data structure information, storage location information and data format information from technical metadata; Extract data definition information, business rule information and data quality standard information from business metadata; Establish a mapping relationship between technical metadata and business metadata; Generate a water affairs data resource directory based on the mapping relationship; The task execution status data is analyzed based on the water affairs data resource directory to identify task execution anomalies, including: Obtain the historical execution status data of each task in the workflow; based on the historical execution status data, establish a normal behavior model for task execution; Compare the current task execution status data with the normal behavior model to obtain the task status and task execution time; When the task status is a failed status or the task execution time exceeds the corresponding second preset threshold, determining that the task execution is abnormal; The hardware resource usage data is analyzed based on the water service data resource directory to identify abnormal resource usage, including: Obtain historical hardware resource usage data for system operation; Based on historical hardware resource usage data, establish a normal range model for resource usage; Compare current hardware resource usage data with the normal range model; When any one of the CPU usage, memory occupancy or disk space exceeds the corresponding second preset threshold, it is determined that the resource usage is abnormal.
8. The method according to claim 1, characterized in that Perform data lineage analysis on abnormal water service data to determine the abnormal topology of the abnormal water service data, including: Conduct data lineage analysis on abnormal water service data to determine the lineage relationship of the abnormal water service data, where the lineage relationship includes data processing rules and data sources; Determine the impact of abnormal water service data on other water service data, including the affected tables and fields; In the preset data map, an abnormal topology map is generated, wherein the abnormal topology map includes blood relationship and impact range.
9. A data processing server, characterized in that: A water data management platform is deployed, including: a memory configured to store instructions; and A processor is configured to call the instruction from the memory and implement the water service data anomaly identification method according to any one of claims 1 to 8 when executing the instruction.
10. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores instructions for enabling a machine to execute the method for identifying anomalies in water service data according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data conversion method of heterogeneous data source and middleware
CN111460019A
Multi-source heterogeneous water environment big data management system
CN113010506A
Metadata knowledge graph engine system for rail transit field
CN113392227A
Overall management method and system for multiple tasks in cloud environment
CN117573307A
Multi-source heterogeneous data synchronization method, device and equipment and readable storage medium
CN118445348A