Alarm platform multi-level processing system and method based on data flow
Through a multi-level processing system based on data flow, automatic logging and real-time analysis, the problem of high complexity of logging in the existing technology is solved, troubleshooting speed and system stability are improved, and operation and maintenance efficiency and user satisfaction are improved.
Patent Information
- Application Number
- CN202510338178.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336131A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log management, and more specifically, to a multi-level processing system and method for an alarm platform based on data flow. Background Art
[0002] With the rapid development of information technology, the scale of various software systems, network services, and Internet of Things devices has become increasingly large. The stability and reliability of these systems have become key factors in enterprise operations and user experiences. In this context, as an important tool for monitoring and ensuring the stable operation of the system, the performance and efficiency of a real-time alarm system are directly related to the fault response speed and system recovery ability.
[0003] In existing real-time alarm systems, log recording, as the core link for monitoring and diagnosing the operating status of the system, is of great importance. However, the limitations of traditional log management methods, especially when dealing with complex systems and large-scale data, are particularly prominent. Specifically, mixing various types of information such as normal operations, warnings, and errors in the same log file not only increases the complexity and redundancy of the logs but also brings the following specific inconveniences in the actual operation and maintenance process:
[0004] Low information screening efficiency: When a system fails, the primary task of the operation and maintenance personnel is to quickly locate the root cause of the problem. However, in mixed logs, normal operation logs occupy most of the space, and error information is submerged in them, forcing the operation and maintenance personnel to manually screen and filter a large amount of irrelevant information, greatly reducing the efficiency of information screening. This not only prolongs the fault response time but also may miss key errors due to human negligence, further affecting the fault resolution efficiency.
[0005] Increased log analysis difficulty: Due to the mixed log information and the lack of clear boundaries or classification identifiers between different types of logs, log analysis becomes more difficult. Operation and maintenance personnel need to rely on experience and professional knowledge to pick through the complex log data to find possible abnormal patterns and correlation relationships. This highly difficult analysis work is not only time-consuming and laborious but also error-prone, affecting the accuracy and timeliness of problem diagnosis.
[0006] System performance is affected: In large systems, log recording itself also has a certain impact on system performance. If all logs are written to the same file, it will not only increase the I / O burden on the file system but may also cause the system response speed to decrease due to frequent disk write operations. In addition, when the log file reaches a certain size, archiving or deletion operations need to be performed to free up disk space. These operations may all have additional impacts on system performance.
[0007] Increasing operation and maintenance costs: Due to the many inconveniences of traditional log management methods, operation and maintenance personnel need to invest more time and energy in processing log information. This not only increases labor costs but may also lead to serious consequences such as escalated failures or business interruptions due to untimely processing. At the same time, to meet complex log management requirements, enterprises may also need to purchase more advanced log management tools or software, further increasing operation and maintenance costs.
[0008] There is also a method proposed, which periodically scans OSS (server) logs, analyzes and judges whether there are faults, and sends alarm messages for faults. This method has the following technical problems: The log data volume is large, and obtaining log information regularly has low real-time performance and slow fault troubleshooting speed. Currently, the existing technical solutions have not solved the problem of slow fault troubleshooting speed. Summary of the Invention
[0009] In view of the current technical development needs and deficiencies, the present invention provides a multi-level processing system and method for an alarm platform based on data flow. The aim is to improve the operation and maintenance efficiency of the system and the speed of fault troubleshooting through an automated method. The aim is to implement a dual log separation mechanism to automatically separate and record logs containing error content from normal logs, and retrieve the collection script to locate problems, thereby improving the operation and maintenance efficiency of the system and the speed of fault troubleshooting. By implementing the dual log separation mechanism, logs containing error content are automatically separated and recorded from normal logs, the logs containing error are printed to a separate file, and the relevant collection scripts and relevant tables involved in the error content are retrieved and located, thereby improving the operation and maintenance efficiency of the system and the speed of fault troubleshooting.
[0010] In the first aspect, the present invention provides a multi-level processing system for an alarm platform based on data flow. The technical solution adopted to solve the above technical problems is as follows:
[0011] A multi-level processing system for an alarm platform based on data flow, which includes:
[0012] A data collection layer, which is used to obtain log data generated during the operation of the system from various data sources and multiple platforms;
[0013] A data processing layer, which is used to receive the data from the data collection layer and perform preprocessing operations such as cleaning, format unification, and log separation;
[0014] A data storage layer, which is used to store the data output by the data processing layer in a specified database for subsequent analysis and query;
[0015] A data analysis layer, which is used to perform real-time analysis on the data in the data storage layer to discover anomalies and potential problems;
[0016] The data alarm layer is used to generate alarm information according to predefined rules and policies when anomalies or potential problems are detected in the data analysis layer.
[0017] Optionally, the data collection layer involved obtains log data through the following two methods:
[0018] a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the data collection layer according to preset rules and time intervals;
[0019] b) For monitoring objects without the function of actively pushing data, the data collection layer accesses the monitoring objects through network protocols or directly reads relevant data files from the local file system;
[0020] The data collection layer involved collects data asynchronously and in parallel from multiple platforms, and each collection process is independent of each other and does not interfere with each other.
[0021] Optionally, the data processing layer involved includes:
[0022] The receiving and listening unit is used to design a Python script to monitor the data transmission channel of the data collection layer in real time to ensure the timely and complete reception of data;
[0023] The preliminary integration unit is used to preliminarily integrate the received data and gather data from different data sources and in different formats into the processing queue;
[0024] The data cleaning unit is used to identify and remove useless information and duplicate data by comparing key fields, and at the same time check for outliers to complete data cleaning;
[0025] The conversion processing unit is used to convert the cleaned data into a unified format and encoding standard, and uses the try-except syntax to handle possible errors;
[0026] The log separation unit is used to print logs using the logging module, set the log framework to print at the DEBUG level or above, if data collection or processing fails, set the prefix of the DEBUG information to "error" and append the error message, and print the logs to two files, namely the "system log" and the "ERROR log" respectively.
[0027] Optionally, the data analysis layer involved includes:
[0028] The setting and judgment unit is used to first set thresholds for each indicator and clarify whether the indicator allows null values;
[0029] An acquisition processing unit is used to obtain the required data from the database using an SQL query statement. During the process of obtaining the required data, missing values, outliers, and duplicate record problems in the data are processed to ensure the quality and consistency of the data.
[0030] A format conversion unit is used to perform formatting and aggregation operations on the acquired data for in-depth analysis.
[0031] A real-time analysis unit uses Apache Kafka as a streaming processing framework. Through its consumer, it subscribes to and pulls messages containing real-time monitoring data from the corresponding topic. According to the preset logic, the data is compared with the threshold, and trend and correlation analysis are performed to discover potential problems and anomalies, providing accurate information for the data alarm layer.
[0032] Further optionally, the real-time analysis unit specifically performs the following operations:
[0033] Use Apache Kafka as a streaming processing framework to receive data from various data sources in real time and publish it as messages to different topics.
[0034] Subscribe to data from the corresponding topic through Kafka's consumer. The consumer continuously pulls messages from the topic, and the messages contain real-time monitoring data. For each received message, the real-time analysis unit processes it according to the preset analysis logic to trigger the corresponding alarm mechanism when it is found that the data of a certain metric exceeds the threshold range.
[0035] At the same time, perform trend analysis and correlation analysis on the real-time data.
[0036] The real-time analysis unit timely discovers potential problems and anomalies during the system operation process based on the analysis results, providing accurate information for the data alarm layer.
[0037] In the second aspect, the present invention provides a multi-level processing method for an alarm platform based on data flow. The technical solution adopted to solve the above technical problems is as follows:
[0038] A multi-level processing method for an alarm platform based on data flow includes the following steps:
[0039] S1. Obtain the log data generated during the system operation from various data sources and multiple platforms.
[0040] S2. Receive the log data and perform preprocessing operations of cleaning, format unification, and log separation.
[0041] S3. Store the preprocessed data in a specified database for subsequent analysis and query.
[0042] S4. Perform real-time analysis on the stored data to detect anomalies and potential problems;
[0043] S5. When anomalies or potential problems are detected, generate alarm information according to predetermined rules and strategies.
[0044] Optionally, when performing step S1, obtain log data through the following two methods:
[0045] a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the specified location according to preset rules and time intervals;
[0046] b) For monitoring objects without the function of actively pushing data, access the monitoring objects through network protocols or directly read relevant data files from the local file system;
[0047] When performing step S1, collect data asynchronously and in parallel from multiple platforms, and each collection process is independent of each other and does not interfere with each other.
[0048] Optionally, step S2 specifically includes the following operations:
[0049] Design a Python script to monitor the data transmission channel in real time to ensure timely and complete receipt of data;
[0050] Perform preliminary integration on the received data, and collect data from different data sources and in different formats into the processing queue;
[0051] Identify and remove useless information and duplicate data by comparing key fields, and at the same time check for outliers to complete data cleaning;
[0052] Convert the cleaned data into a unified format and encoding standard, and use the try-except syntax to handle possible errors;
[0053] Use the logging module to print logs, set the log framework to print at DEBUG level or above, if data collection or processing fails, set the DEBUG information prefix to "error" and append the error message, and print the logs to two files: "system log" and "ERROR log" respectively.
[0054] Optionally, step S4 specifically includes the following operations:
[0055] S4.1. Set thresholds for each indicator and clarify whether the indicator allows null values;
[0056] S4.2. Use SQL query statements to obtain the required data from the database. During the process of obtaining the required data, handle the problems of missing values, outliers, and duplicate records in the data to ensure the quality and consistency of the data;
[0057] S4.3. Format and aggregate the acquired data for in-depth analysis;
[0058] S4.4. Use Apache Kafka as a streaming processing framework. Its consumers subscribe to and pull messages containing real-time monitoring data from corresponding topics. According to preset logic, compare the data with thresholds, perform trend and correlation analysis to discover potential problems and anomalies, and provide accurate information for data alarm.
[0059] Further optionally, step S4.4 specifically includes the following operations:
[0060] Use Apache Kafka as a streaming processing framework to receive data from various data sources in real time and publish it as messages to different topics;
[0061] Subscribe to data from corresponding topics through Kafka consumers. The consumers continuously pull messages from the topics, and the messages contain real-time monitoring data. For each received message, process it according to preset analysis logic to trigger corresponding alarm mechanisms when it is found that the data of a certain metric exceeds the threshold range;
[0062] At the same time, perform trend analysis and correlation analysis on real-time data;
[0063] Discover potential problems and anomalies in the system operation process in a timely manner according to the analysis results, and provide accurate information for data alarm.
[0064] The multi-level processing system and method of an alarm platform based on data flow of the present invention has the following beneficial effects compared with the prior art:
[0065] 1. The present invention aims to improve the system operation and maintenance efficiency and the speed of fault troubleshooting through an automated method, effectively reducing the manual operations of operation and maintenance personnel, improving the automation level of system operation and maintenance, and providing strong support for quickly locating and solving system problems;
[0066] 2. The present invention separately records and manages error logs, enabling operation and maintenance personnel to directly access key information related to faults without manually filtering irrelevant logs. At the same time, it can automatically retrieve and locate relevant collection scripts and database tables involved in the error content, further shortening the time for fault location and providing strong support for quickly fixing problems;
[0067] 3. By capturing and recording error logs in a timely manner, the system of the present invention can discover potential problems earlier, provide warning information for operation and maintenance personnel, thereby avoiding small problems from evolving into large-scale failures, and enhancing the overall reliability and stability of the system;
[0068] 4. When facing an emergency failure, the present invention can quickly mobilize resources to prioritize the handling of key issues; while in daily operation and maintenance, it can concentrate on system optimization and preventive maintenance, improving the utilization efficiency of resources and the refinement level of management; by quickly responding to and resolving system failures, the alarm platform can reduce service interruption time, ensuring the continuity and stability of user services. This not only enhances the user's trust and satisfaction with the system, but also strengthens the enterprise's brand image and market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The Figure 1 hierarchical connection diagram of the first embodiment of the present invention;
[0070] The Figure 2 method flowchart of the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] In order to make the technical solutions, the technical problems solved and the technical effects of the present invention more clearly understood, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.
[0072] Embodiment 1:
[0073] Combined with the Figure 1 accompanying drawings, this embodiment proposes a multi-level processing system for an alarm platform based on data flow, which includes:
[0074] A data acquisition layer for obtaining log data generated during the operation of the system from various data sources and multiple platforms;
[0075] A data processing layer for receiving the data from the data acquisition layer and performing preprocessing operations such as cleaning, format unification, and log separation;
[0076] A data storage layer for storing the data output by the data processing layer in a specified database for subsequent analysis and query;
[0077] A data analysis layer for performing real-time analysis on the data in the data storage layer to discover anomalies and potential problems;
[0078] A data alarm layer for generating alarm information according to predetermined rules and strategies when anomalies or potential problems are discovered in the data analysis layer.
[0079] In this embodiment, the data acquisition layer obtains log data through the following two methods:
[0080] a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the data acquisition layer according to preset rules and time intervals. These status information may include various types, such as device performance indicators, operating status, error information, etc.;
[0081] b) For monitoring objects without the function of actively pushing data, the data acquisition layer accesses the monitoring objects through network protocols or directly reads relevant data files from the local file system; wherein, the monitoring objects refer to entities that the system needs to focus on and obtain data from, including but not limited to hardware devices (such as servers, network devices, storage devices), software systems and applications (such as enterprise-level ERP systems, Web applications, database management systems), virtual resources and containers; performance indicators (such as usage rate, temperature, bandwidth, etc.) and status information (such as power on, power off, failure, etc.) of hardware devices are important sources of monitoring data; for software systems and applications, the monitored information may include software performance (such as response time, throughput, etc.), running status (such as whether it is running normally, whether there are errors, etc.) and business-related data (such as the number of user logins, the number of transactions, etc.); for virtual resources and containers, data such as their resource usage (such as CPU allocation, memory occupancy, etc.), running status (such as start, stop, restart, etc.) and application deployment status are monitored;
[0082] The data acquisition layer involved in this case collects data asynchronously and in parallel from multiple platforms, and each acquisition process is independent of each other and does not interfere with each other. Such a design can improve the efficiency of data acquisition, and at the same time can also adapt to the data acquisition speed differences of different platforms and data sources, ensuring the timeliness and integrity of data acquisition.
[0083] In this embodiment, the data processing layer involved includes:
[0084] A receiving and listening unit, which is used to design a Python script to monitor the data transmission channel of the data acquisition layer in real time to ensure the timely and complete reception of data; for example, libraries related to message queues (such as pika for RabbitMQ, kafka-python for Kafka) can be used to implement the listening mechanism;
[0085] A preliminary integration unit, which is used to preliminarily integrate the received data and gather data from different data sources and in different formats into the processing queue; a queue data structure (such as queue.Queue in Python) can be specifically used to store and manage the data to be processed to ensure the orderliness and manageability of the data;
[0086] A data cleaning unit, which is used to identify and remove useless information and duplicate data by comparing key fields (such as ID, timestamp, event type, etc.), and at the same time check for outliers to complete data cleaning; for example, delete duplicate log entries, filter out incomplete or damaged log records, and ensure the reliability of the data through methods such as key field checking, timestamp verification, and data integrity checking;
[0087] The conversion processing unit is used to convert the cleaned data into a unified format and encoding standard, and uses the try-except syntax to handle possible errors; when processing the data format, it may involve parsing and converting data in different formats (such as text, JSON, XML, etc.). For numerical data, normalization and standardization are performed to eliminate the impact of dimensional differences on subsequent analysis.
[0088] The log separation unit is used to print logs using the logging module, set the log framework to print at a level above DEBUG. If data collection or processing fails, the DEBUG information prefix is set to "error" and the error message is appended, and the logs are printed to two files, "system log" and "ERROR log" respectively.
[0089] In this embodiment, the data analysis layer involved includes:
[0090] The setting and judgment unit is used to first set thresholds for each indicator and clarify whether the indicator allows null values; these thresholds will be important bases for subsequent analysis to judge whether the data is abnormal. For example, for the CPU usage rate, the normal range may be set from 0 to 80%, and values outside this range are regarded as abnormal. At the same time, it is clarified whether some indicators can be null. For indicators that cannot be NULL, special processing is required in subsequent analysis.
[0091] The acquisition and processing unit is used to obtain the required data from the database using SQL query statements or database operation languages (such as the query language of MongoDB); during the process of obtaining the required data, missing values, outliers, and duplicate record problems in the data are processed to ensure the quality and consistency of the data; for missing values, appropriate filling methods (such as using the mean, median, or other statistical methods) can be selected according to the characteristics of the data and business logic; for outliers, statistical methods (such as the 3σ principle) can be used for identification and processing; for duplicate records, they are identified and deleted by comparing key fields in the records (such as unique identification ID, timestamp, etc.).
[0092] The format conversion unit is used to format and aggregate the obtained data for in-depth analysis; for example, unify the time format into the standard ISO format, group and aggregate the data according to time periods (such as every hour, every day), and calculate statistical indicators (such as sum, average, maximum, minimum, etc.) within each time period.
[0093] The real-time analysis unit is used to adopt Apache Kafka as a streaming processing framework, subscribe and pull messages containing real-time monitoring data from the corresponding topics through its consumers, compare the data with the thresholds according to the preset logic, and perform trend and correlation analysis to discover potential problems and anomalies, and provide accurate information for the data warning layer.
[0094] The specific operations performed by the real-time analysis unit involved are as follows:
[0095] Apache Kafka is used as a streaming processing framework to receive data from various data sources in real time and publish it as messages to different topics;
[0096] The data is subscribed from the corresponding topic through the Kafka consumer. The consumer continuously pulls messages from the topic, and the messages contain real-time monitoring data. For each received message, the real-time analysis unit processes it according to the preset analysis logic to trigger the corresponding alarm mechanism when it is found that the data of a certain metric exceeds the threshold range;
[0097] At the same time, trend analysis and correlation analysis are performed on the real-time data; for example, by analyzing the change trend of the CPU usage rate over a period of time, predicting possible future performance bottlenecks; or by correlating the relationships between different metrics to discover potential problems; for example, when the network bandwidth usage suddenly increases and the server response time becomes longer at the same time, it may mean that there is a network congestion problem that needs to be processed in a timely manner;
[0098] The real-time analysis unit timely discovers potential problems and abnormal situations during the system operation process based on the analysis results, and provides accurate information for the data alarm layer.
[0099] Embodiment 2:
[0100] Combined with the attached Figure 2 , this embodiment proposes a multi-level processing method for an alarm platform based on data flow, which includes the following steps:
[0101] S1. Obtain the log data generated during the system operation from various data sources and multiple platforms.
[0102] When performing step S1, the log data is obtained through the following two methods:
[0103] a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the specified location according to the preset rules and time intervals;
[0104] b) For monitoring objects that do not have the function of actively pushing data, access the monitoring objects through network protocols or directly read relevant data files from the local file system; wherein, the monitoring objects refer to entities that the system needs to focus on and obtain data, including but not limited to hardware devices (such as servers, network devices, storage devices), software systems and applications (such as enterprise-level ERP systems, Web applications, database management systems), virtual resources and containers; performance indicators (such as utilization rate, temperature, bandwidth, etc.) and status information (such as power on, power off, failure, etc.) of hardware devices are important sources of monitoring data; for software systems and applications, the monitored information may include software performance (such as response time, throughput, etc.), running status (such as whether it is running normally, whether there are errors, etc.) and business-related data (such as the number of user logins, transaction volume, etc.); for virtual resources and containers, monitor data such as their resource usage (such as CPU allocation, memory occupancy, etc.), running status (such as start, stop, restart, etc.) and application deployment status.
[0105] When executing step S1, collect data asynchronously and in parallel from multiple platforms, and each collection process is independent and does not interfere with each other. Such a design can improve the efficiency of data collection, and at the same time can also adapt to the data collection speed differences of different platforms and data sources, ensuring the timeliness and integrity of data collection.
[0106] S2. Receive log data and perform preprocessing operations such as cleaning, format unification, and log separation.
[0107] Step S2 specifically includes the following operations:
[0108] Design a Python script to listen to the data transmission channel in real time to ensure the timely and complete reception of data; for example, relevant libraries related to message queues (such as pika for RabbitMQ, kafka-python for Kafka) can be used to implement the listening mechanism;
[0109] Perform preliminary integration on the received data, and gather data from different data sources and in different formats into the processing queue; a queue data structure (such as queue.Queue in Python) can be specifically used to store and manage the data to be processed, ensuring the orderliness and manageability of the data;
[0110] Identify and remove useless information and duplicate data by comparing key fields (such as ID, timestamp, event type, etc.), and at the same time check for outliers to complete data cleaning; for example, delete duplicate log entries, filter out incomplete or damaged log records, and ensure the reliability of the data through methods such as key field checking, timestamp verification, and data integrity checking;
[0111] Convert the cleaned data into a unified format and encoding standard, and use the try-except syntax to handle possible errors; when processing data formats, it may involve parsing and converting data in different formats (such as text, JSON, XML, etc.). For numerical data, perform normalization and standardization processing to eliminate the impact of dimensional differences on subsequent analysis.
[0112] Use the logging module to print logs. Set the logging framework to print above the DEBUG level. If data collection or processing fails, set the DEBUG information prefix to "error" and append the error message, and print the logs to two files, "system log" and "ERROR log" respectively.
[0113] S3. Store the preprocessed data in a specified database for subsequent analysis and query.
[0114] The specified database can be a relational database (such as MySQL, PostgreSQL). Relational databases are suitable for storing structured data and can use SQL language for complex data queries and transaction processing, which is more suitable for storing log information with a clear relational data table structure.
[0115] The specified database can also be a non-relational database (such as MongoDB, Cassandra). Non-relational databases are suitable for storing a large amount of unstructured or semi-structured log data, can better handle high-concurrency data writing operations, and are especially suitable for storing the original content of log data.
[0116] S4. Perform real-time analysis on the stored data to detect anomalies and potential problems.
[0117] Step S4 specifically includes the following operations:
[0118] S4.1. Set thresholds for each indicator and clarify whether the indicator allows null values; these thresholds will be important bases for subsequent analysis to judge whether the data is abnormal. For example, for CPU usage, the normal range may be set from 0 to 80%, and if it exceeds this range, it is considered abnormal. At the same time, clarify whether some indicators can be null. For indicators that cannot be NULL, special processing is required in subsequent analysis.
[0119] S4.2. Obtain the required data from the database using SQL query statements or database operation languages (such as the query language of MongoDB); during the process of obtaining the required data, handle the missing values, outliers, and duplicate record problems in the data to ensure the quality and consistency of the data; for missing values, appropriate filling methods (such as using the mean, median, or other statistical methods) can be selected according to the characteristics of the data and business logic; for outliers, statistical methods (such as the 3σ principle) can be used for identification and processing; for duplicate records, identify and delete them by comparing the key fields in the records (such as unique identification ID, timestamp, etc.).
[0120] S4.3. Format and aggregate the obtained data for in-depth analysis; for example, unify the time format into the standard ISO format, group and aggregate the data by time period (such as every hour, every day), and calculate the statistical indicators (such as sum, average, maximum, minimum, etc.) within each time period.
[0121] S4.4. Adopt Apache Kafka as the streaming processing framework, and through its consumers, subscribe to and pull messages containing real-time monitoring data from the corresponding topics. According to the preset logic, compare the data with the thresholds, and perform trend and correlation analysis to discover potential problems and anomalies, providing accurate information for data alerts.
[0122] Step S4.4 specifically includes the following operations:
[0123] Adopt Apache Kafka as the streaming processing framework to receive data from various data sources in real time and publish it as messages to different topics;
[0124] Subscribe to data from the corresponding topics through Kafka's consumers. The consumers continuously pull messages from the topics, and the messages contain real-time monitoring data; for each received message, process it according to the preset analysis logic to trigger the corresponding alert mechanism when it is found that the data of a certain indicator exceeds the threshold range;
[0125] At the same time, perform trend analysis and correlation analysis on the real-time data; for example, by analyzing the change trend of CPU usage rate over a period of time, predicting possible future performance bottlenecks; or by correlating and analyzing the relationships between different indicators, discovering potential problems. For example, when the network bandwidth usage rate suddenly increases and the server response time becomes longer at the same time, it may mean that there is network congestion and needs to be processed in a timely manner;
[0126] Discover potential problems and anomalies in the system operation process in a timely manner according to the analysis results, providing accurate information for data alerts.
[0127] S5. When an anomaly or potential problem is detected, generate an alarm message according to the predefined rules and strategies.
[0128] In this step, ① it is necessary to set clear alarm rules, such as triggering an alarm when a certain indicator exceeds a threshold, a specific error code appears, or a specific pattern is met; ② send the alarm message through multiple channels, such as email, SMS, instant messaging tools (such as Slack, Microsoft Teams), mobile application notifications, or call an external alarm service to ensure that relevant personnel can receive the alarm notification in a timely manner; ③ for unhandled alarms, escalation processing can be carried out according to the importance and duration of the alarm. For example, send the alarm message to higher-level management personnel or take more urgent handling measures.
[0129] In summary, by adopting the multi-level processing system and method of an alarm platform based on data flow of the present invention, it aims to improve the system operation and maintenance efficiency and the speed of fault troubleshooting in an automated manner, effectively reduce the manual operations of operation and maintenance personnel, improve the automation level of system operation and maintenance, and provide strong support for quickly locating and solving system problems.
[0130] The above application of specific examples has elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of this technology, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of patent protection of the present invention.
Claims
1. A multi - level processing system for an alarm platform based on data flow, characterized in that, It includes: The data collection layer is used to obtain log data generated during system operation from various data sources and multiple platforms; The data processing layer is used to receive data from the data collection layer and perform preprocessing operations such as cleaning, format unification, and log separation; The data storage layer is used to store the data output by the data processing layer in a specified database for subsequent analysis and query; The data analysis layer is used to analyze the data in the data storage layer in real time to find anomalies and potential problems; The data alarm layer is used to generate alarm information according to predetermined rules and strategies when anomalies or potential problems are found in the data analysis layer.
2. The multi-level processing system of an alarm platform based on data flow according to claim 1, characterized in that The data collection layer obtains log data in the following two ways: a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the data collection layer according to the preset rules and time intervals; b) For monitoring objects that do not have the function of actively pushing data, the data collection layer accesses the monitoring objects through network protocols, or directly reads relevant data files from the local file system; The data collection layer collects data asynchronously and in parallel from multiple platforms, and each collection process is independent of each other and does not interfere with each other.
3. The multi-level processing system of an alarm platform based on data flow according to claim 2, wherein, The data processing layer includes: The receiving monitoring unit is used to design Python scripts to monitor the data transmission channel of the data acquisition layer in real time to ensure timely and complete data reception; A preliminary integration unit is used to perform preliminary integration of the received data and aggregate data from different data sources and in different formats into a processing queue; The data cleaning unit is used to identify and remove useless information and duplicate data by comparing key fields, and to check abnormal values to complete data cleaning; The conversion processing unit is used to convert the cleaned data into a unified format and encoding standard, and use the try-except syntax to handle possible errors; The log separation unit is used to print logs using the logging module. The log framework is set to print at the DEBUG level or above. If data collection or processing fails, the DEBUG information prefix is set to "error" and error information is attached. The logs are printed to the "system log" and "ERROR log" files respectively.
4. A multi-level processing system for an alarm platform based on data flow according to claim 3, characterized in that, The data analysis layer includes: Set up a judgment unit to first set a threshold for each indicator and clarify whether the indicator is allowed to be null; The acquisition processing unit is used to obtain the required data from the database using SQL query statements. In the process of obtaining the required data, it handles the missing values, abnormal values, and duplicate records in the data to ensure the quality and consistency of the data. Format conversion unit, used to format and aggregate acquired data for in-depth analysis; The real-time analysis unit uses Apache Kafka as a stream processing framework. Through its consumers, it subscribes to and pulls messages containing real-time monitoring data from the corresponding topics. According to the preset logic, it compares the data with the threshold, performs trend and correlation analysis, and discovers potential problems and anomalies, providing accurate information for the data alarm layer.
5. A multi-level processing system for an alarm platform based on data flow according to claim 4, characterized in that, The real-time analysis unit specifically performs the following operations: Adopt Apache Kafka as the streaming processing framework to receive data from various data sources in real time and publish it to different topics in the form of messages; Subscribe to data from the corresponding topics through Kafka consumers. The consumers continuously pull messages from the topics. The messages contain real-time monitoring data. For each received message, the real-time analysis unit processes it according to the preset analysis logic to trigger the corresponding alarm mechanism when it is found that the data of a certain metric exceeds the threshold range. At the same time, perform trend analysis and correlation analysis on the real-time data; The real-time analysis unit timely discovers potential problems and abnormal situations during the system operation according to the analysis results and provides accurate information for the data alarm layer.
6. A multi-level processing method for an alarm platform based on data flow, characterized in that It includes the following steps: S1. Obtain the log data generated during the system operation from various data sources and multiple platforms; S2. Receive the log data and perform preprocessing operations such as cleaning, format unification, and log separation; S3. Store the preprocessed data in a specified database for subsequent analysis and query; S4. Perform real-time analysis on the stored data to discover anomalies and potential problems; S5. When anomalies or potential problems are discovered, generate alarm information according to the predetermined rules and strategies.
7. A multi - level processing method for an alarm platform based on data flow according to claim 6, characterized in that, Execute step S1 to obtain log data in the following two ways: a) For monitoring objects with the function of actively pushing data, the monitoring objects send their own status information to the specified location according to the preset rules and time intervals; b) For monitoring objects without the function of actively pushing data, access the monitoring objects through network protocols or directly read the relevant data files from the local file system; When executing step S1, collect data asynchronously and in parallel from multiple platforms, and each collection process is independent of each other and does not interfere with each other.
8. A multi - level processing method for an alarm platform based on data flow according to claim 7, characterized in that, The specific operations of step S2 include the following: Design a python script to monitor the data transmission channel in real time to ensure the timely and complete reception of data; Perform preliminary integration on the received data, and gather data from different data sources and different formats into the processing queue; Identify and remove useless information and duplicate data by comparing key fields, and at the same time check for outliers to complete data cleaning; Convert the cleaned data into a unified format and encoding standard, and use the try-except syntax to handle possible errors; Use the logging module to print logs. Set the log framework to print above the DEBUG level. If data collection or processing fails, set the DEBUG information prefix to "error" and append the error message, and print the logs to two files, namely "system log" and "ERROR log".
9. A multi-level processing method for an alarm platform based on data flow according to claim 8, characterized in that The specific operations of step S4 include the following: S4.
1. Set thresholds for each metric and clarify whether the metric allows null values; S4.
2. Use SQL query statements to obtain the required data from the database. During the process of obtaining the required data, handle the problems of missing values, outliers, and duplicate records in the data to ensure the quality and consistency of the data; S4.
3. Perform formatting and aggregation operations on the obtained data for in-depth analysis; S4.
4. Use Apache Kafka as the streaming processing framework. Its consumers subscribe to and pull messages containing real-time monitoring data from the corresponding topics. According to the preset logic, compare the data with the thresholds, and conduct trend and correlation analysis to discover potential problems and anomalies, providing accurate information for data alerts.
10. A multi-level processing method for an alarm platform based on data flow according to claim 9, characterized in that, The specific operations of step S4.4 are as follows: Use Apache Kafka as the streaming processing framework to receive data from various data sources in real time and publish it as messages to different topics; Subscribe to data from the corresponding topics through Kafka's consumers. The consumers continuously pull messages from the topics, and the messages contain real-time monitoring data. For each received message, process it according to the preset analysis logic to trigger the corresponding alert mechanism when it is found that the data of a certain indicator exceeds the threshold range; At the same time, conduct trend analysis and correlation analysis on the real-time data; Timely discover potential problems and anomalies during the system operation according to the analysis results, providing accurate information for data alerts.
Citation Information
Cited By
Log management system
CN121351425A