Data transfer intelligent monitoring system based on data medium station
By designing an intelligent data flow monitoring system based on the data middle platform, the problems of low data quality and insufficient monitoring of performance indicators in the existing technology are solved, the efficiency and stability of data flow are achieved, and the enterprise's trust and utilization efficiency of data are enhanced.
Patent Information
- Application Number
- CN202510302029.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has problems such as low data quality, inability to detect abnormal situations in time, insufficient performance indicator monitoring and alarm feedback in data flow monitoring, which affects the enterprise's trust and utilization efficiency of data.
Design an intelligent data flow monitoring system based on the data middle platform, including data acquisition module, data preprocessing module, data storage module, data analysis module, data monitoring module and alarm feedback module. Through real-time data acquisition and verification, data preprocessing, system analysis and trend mining, real-time status monitoring and performance indicator tracking, and timely alarm feedback, data quality and system performance are ensured.
It improves data quality and system performance, enhances enterprises' trust in data, improves data utilization efficiency, and helps enterprises have a more competitive advantage in data-driven decision-making.
Smart Images

Figure CN120234322A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent monitoring of data flow, and particularly to an intelligent monitoring system for data flow based on a data middle platform. Background Art
[0002] With the rapid development of information technology, enterprises and organizations are facing the challenge of massive data. Effective data management and flow monitoring can not only improve the accuracy of business decisions, but also enhance operational efficiency and response speed. By real-time monitoring the status of data streams, enterprises can timely discover potential problems, optimize resource allocation, and reduce operational risks, thus providing strong support for the sustainable development of enterprises. In addition, the data monitoring system can help enterprises identify key business trends, thereby providing data support for strategic decisions.
[0003] However, many current enterprises still have technical bottlenecks in data flow monitoring. There is often a lack of effective monitoring mechanisms in the aspects of data collection, preprocessing, and storage, resulting in low data quality and the inability to timely discover abnormal situations in data flow. At the same time, existing monitoring systems also have deficiencies in performance index monitoring and alarm feedback, failing to provide sufficient real-time performance and accuracy. These problems seriously affect the trust and utilization efficiency of enterprises in data. Summary of the Invention
[0004] (I) Technical Problems to be Solved
[0005] In view of the deficiencies of the prior art, the present invention provides an intelligent monitoring system for data flow based on a data middle platform. Through real-time data collection and verification mechanisms, the reliability of data quality is ensured, and abnormal situations caused by data quality problems are reduced. The cleaning and standardization processing of the data preprocessing module improve the consistency and usability of data, thus providing a solid foundation for subsequent analysis. The data analysis module helps enterprises gain insights into key indicators and timely discover potential risks through systematic analysis and trend mining. The real-time status monitoring and performance index tracking of the monitoring module ensure the efficiency and stability of data flow, while the alarm feedback module immediately feeds back information to operation and maintenance personnel when abnormal situations occur, effectively improving the reaction speed of the response mechanism. It not only improves the transparency and real-time performance of data flow, but also enhances the trust of enterprises in data, greatly improving the data utilization efficiency, and helping enterprises gain more competitive advantages in data-driven decision-making, thus solving the above problems.
[0006] (II) Technical Solutions
[0007] To achieve the above object, the present invention provides the following technical solution: An intelligent monitoring system for data flow based on a data middle platform, including a data collection module, a data preprocessing module, a data storage module, a data analysis module, a data monitoring module, and an alarm feedback module;
[0008] The data acquisition module is used to collect raw data from multiple data sources in real time. The data sources include business platforms, sensor devices, and external market data, and after verifying the collected data, it is sent to the data preprocessing module;
[0009] The data preprocessing module receives the output of the data acquisition module and performs data cleaning tasks, including data denoising, removing duplicate data, filling in missing values, correcting data formats, and normalizing the data;
[0010] The data storage module is used to uniformly store the preprocessed cleaned data in a NoSQL database for subsequent data query and analysis;
[0011] The data analysis module obtains data from the data storage module, systematically analyzes the stored data, calculates the data missing ratio, data conversion rate, user growth rate, and conducts key data trend analysis, and generates a data report;
[0012] The data monitoring module, based on the analysis results of the data analysis module, monitors the status of the data stream in real time, continuously analyzes the data transfer speed and the amount of abnormal data, and at the same time monitors system performance indicators, including CPU usage, memory occupancy, and response time, and generates a data monitoring report and sends it to the alarm feedback module;
[0013] The alarm feedback module detects the data monitoring report. When it finds abnormal data transfer and system performance anomalies, it automatically triggers an alarm message and feeds it back to the operation and maintenance personnel.
[0014] Preferably, the formula for data denoising is as follows:
[0015]
[0016] In the formula, Y(t) represents the denoised data value at time point t, X(t) represents the raw data value at time point t, N represents the window size for calculating the moving average, and i represents the subscript index.
[0017] Preferably, the formula for removing duplicate data is as follows:
[0018] D unique = unique(D original )
[0019] In the formula, D unique represents the unique data set after removing duplicates, D original represents the original data set, and unique represents the function that returns all unique records in the data set.
[0020] Preferably, the formula for filling missing values is as follows:
[0021]
[0022] In the formula, D filled (i) represents the dataset after filling missing values, D riginal represents the original dataset, NaN represents the symbol of missing values, Imputation() represents the algorithm for filling missing values, and i represents the digital flag bit.
[0023] Preferably, the formula for correcting the data format is as follows:
[0024] D corrected (j) = format(D original (j))
[0025] D corrected (j) represents the dataset after format correction, D original represents the original dataset, format() represents the processing function for converting data into a specific format, and j represents the position of the data in the dataset.
[0026] Preferably, the formula for normalizing the data is as follows:
[0027]
[0028] In the formula, D standardized (i) represents the data value after normalization, D original (i) represents the original data value, μ represents the mean of the original dataset, and σ represents the standard deviation of the original dataset.
[0029] Preferably, the formula for calculating the data missing ratio is as follows:
[0030]
[0031] In the formula, Missing Rate represents the data missing ratio, Number of Missing Values represents the number of missing values in the dataset, and Total Number of Values represents the total number of values in the dataset.
[0032] Preferably, the calculation formula for the data conversion rate is as follows:
[0033]
[0034] In the formula, Conversion Rate represents the data conversion rate, Number of Conversion represents the number of users who have achieved a specific goal, and Total visito represents the total number of users who have accessed the data.
[0035] Preferably, the calculation formula of the user growth rate is as follows:
[0036]
[0037] In the formula, User Griwth Rate represents the user growth rate, New Users represents the number of new users obtained in the current time period, and Total User represents the total number of users in the previous time period.
[0038] Preferably, the formula for the key trend analysis of the data is as follows:
[0039] Trend Analysis = Time Series(D original )
[0040] In the formula, Trend Analysis represents the result of the key trend analysis, and Time Series() represents the analysis method applied to time series data.
[0041] Compared with the prior art, the present invention provides a data flow intelligent monitoring system based on a data middle platform, which has the following beneficial effects:
[0042] Through the real-time data collection and verification mechanism, the present invention ensures the reliability of data quality, reduces abnormal situations caused by data quality problems. The cleaning and standardization processing of the data preprocessing module improve the consistency and availability of the data, thus providing a solid foundation for subsequent analysis. The data analysis module helps enterprises gain insights into key indicators and timely discover potential risks through systematic analysis and trend mining. The real-time status monitoring and performance index tracking of the monitoring module ensure the efficiency and stability of data flow. The alarm feedback module feeds back information to the operation and maintenance personnel in the first time when an abnormal situation occurs, effectively improving the reaction speed of the response mechanism. It not only improves the transparency and real-time nature of data flow, but also enhances the enterprise's trust in data, greatly improving the data utilization efficiency and helping the enterprise gain a competitive advantage in data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the system flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] In view of the fact that there are still technical bottlenecks in the data flow monitoring of many current enterprises, and there is often a lack of effective monitoring mechanisms in the data collection, preprocessing, and storage links, resulting in low data quality and the inability to timely detect abnormal situations in the data flow. For this reason, a data flow intelligent monitoring system based on a data middle platform is proposed. Please refer to Figure 1 , the system includes a data acquisition module, a data preprocessing module, a data storage module, a data analysis module, a data monitoring module, and an alarm feedback module;
[0046] The data acquisition module is the core component of the system and is responsible for collecting raw data from multiple heterogeneous data sources in real time. These data sources include business platforms (such as e-commerce systems and customer relationship management systems), sensor devices (such as industrial Internet of Things sensors and environmental monitoring sensors), and external market data (such as social media APIs and financial data services). In order to ensure the validity and accuracy of the data, the data acquisition module uses a series of advanced technical means for data verification, including using a rule engine to verify the data format and type, applying anomaly detection algorithms to identify and exclude noisy data, and introducing data flow monitoring technologies (such as Apache Kafka and Apache Flink) to achieve real-time monitoring and processing of the data flow. In addition, the system will perform metadata tagging on each piece of data to record the data source, type, and collection time to improve the subsequent data governance and traceability capabilities. After these comprehensive processes, the acquisition module transmits the valid data to the data preprocessing module, laying a solid foundation for subsequent data processing and analysis;
[0047] The data preprocessing module is a crucial link in the data processing flow. It receives the raw output from the data acquisition module and is responsible for comprehensively cleaning and organizing the data to ensure a more solid foundation for subsequent analysis and decision-making. In this module, first, the data denoising operation is performed, which usually uses filtering techniques (such as wavelet transform or moving average method) to eliminate random noise, thereby improving the signal-to-noise ratio of the data. Then, by applying duplicate data removal algorithms, such as hash detection and counting sort, redundant records are effectively identified and removed to ensure the uniqueness and validity of each piece of data. In addition, for existing missing values, the preprocessing module uses filling methods to handle them to maintain the integrity of the data. The step of correcting the data format is executed through data format conversion tools and rule engines to ensure that all fields follow a unified standard and format, simplifying the complexity of subsequent data processing. At the same time, the preprocessing module also standardizes the data to eliminate the influence caused by different dimensions and data distributions;
[0048] Among them, the formula for data denoising is as follows:
[0049]
[0050] Denoising can significantly reduce the random errors and biases in the data, ensuring that the data relied on for subsequent analysis is accurate and reliable. Noise will obscure the true trends and patterns in the data, so denoising ensures that the analysis is based on clear information. In the formula, Y(t) represents the denoised data value at time point t, X(t) represents the original data value at time point t, N represents the window size used for calculating the moving average, and i represents the subscript index. Removing unnecessary noise and outliers can reduce the amount of data to be processed, thereby reducing the consumption of computing resources, improving the processing efficiency of the system, and accelerating the decision-making speed;
[0051] The formula for removing duplicate data is as follows:
[0052] D unique = unique(D original )
[0053] Duplicate data will occupy additional storage resources. By removing redundant data, the system can save storage costs and optimize the use of storage resources for subsequent data expansion and management. In the formula, D unique represents the unique data set after removing duplicates, D original represents the original data set, and unique represents a function that returns all unique records in the data set. In the process of data analysis and report generation, duplicate data will lead to misleading results and low efficiency. After removing duplicate data, the clarity and accuracy of the data are higher, and analysis can be carried out more concentratedly, improving the accuracy of decision-making;
[0054] The formula for filling missing values is as follows:
[0055]
[0056] The existence of missing values can lead to incomplete information, thus affecting the comprehensiveness of data analysis. Filling in the missing values can enhance the integrity of the dataset, making the analysis results more credible and representative. In the formula, D filled (i) represents the dataset after filling in the missing values, D original represents the original dataset, NaN represents the symbol of the missing value, Imputation() represents the algorithm for filling in the missing values, i represents the digital flag bit. After handling the missing values, the analysis bias caused by the missing values can be reduced. Especially in statistical analysis, filling in the missing values helps to more accurately estimate the data distribution and feature relationships;
[0057] The formula for correcting the data format is as follows:
[0058] D corrected (j) = format(D original (j))
[0059] Data with inconsistent formats will increase the complexity during processing and analysis. Correcting the data format enables the data processing program to run more simply and efficiently, reducing the time and resources required for data conversion. D corrected (j) represents the dataset after format correction, D orginal represents the original dataset, format() represents the processing function used to convert the data into a specific format, j represents the position of the data in the dataset. By correcting the data format, ensuring that all data follows a unified format standard simplifies the data integration and query processes and improves the data integration ability;
[0060] The formula for standardizing the data is as follows:
[0061]
[0062] Different features may have different dimensions (for example, income and age). Standardization processing can eliminate this influence, enabling comparison of each feature on the same basis and avoiding the situation where certain features dominate the model. In the formula, D standardized (i) represents the standardized data value, D original (i) represents the original data value, μ represents the mean of the original dataset, σ represents the standard deviation of the original dataset. The standardized data enhances the effects of clustering and classification algorithms, making similar data points closer in the feature space, thereby improving the accuracy and performance of classification;
[0063] The data storage module plays a crucial role in the entire data processing architecture. It is responsible for uniformly storing the preprocessed cleaned data in a relational or non-relational database, such as MySQL, PostgreSQL, MongoDB, or other big data storage solutions (e.g., Hadoop Distributed File System), to ensure data security and accessibility. This module adopts efficient data storage designs, such as partitioning, indexing, and compression techniques, to improve the efficiency of data query and retrieval, providing strong support for subsequent data analysis. Subsequently, the data analysis module fetches the stored data from the data storage module and systematically analyzes the data using data analysis tools and libraries (such as Pandas, NumPy in Python, or dedicated data analysis platforms like Tableau, PowerBI). This module calculates key metrics, including the data missing ratio, calculates the data conversion rate and user growth rate by analyzing user behavior patterns identified from historical data, and conducts key data trend analysis by combining time series analysis methods. The analysis results will be integrated and detailed data reports will be generated. These reports not only provide key business insights but also are presented through visualization components (such as charts and dashboards), enabling business decision-makers to quickly understand the trends and patterns behind the data, thus facilitating more accurate strategic decision-making and resource allocation. Throughout the process, the data analysis module can also achieve automated scheduling, through regular tasks or event-driven mechanisms, to ensure the continuity and real-time nature of data analysis, further enhancing the agility of business operations;
[0064] The formula for calculating the data missing ratio is as follows:
[0065]
[0066] Calculating the data missing ratio helps monitor data quality. By identifying the sources of missing values, effective data governance measures can be implemented to ensure data integrity. In the formula, Missing Rate represents the data missing ratio, Numberof Missing Values represents the number of missing values in the dataset, and Total Number of Values represents the total number of values in the dataset. Data with a high missing ratio and lacking relevant information may affect the effectiveness of decision-making. Calculating the data missing ratio can provide a necessary basis for risk assessment in decision-making;
[0067] The formula for calculating the data conversion rate is as follows:
[0068]
[0069] By analyzing the conversion rate, enterprises can identify the pain points of users in the conversion process, optimize website design and user experience based on data, reduce user churn, and improve conversion efficiency. In the formula, Conversion Rate represents the data conversion rate, Bynber of Conversion represents the number of users who have completed a specific goal, and Total visito represents the total number of users who have accessed the data. The data conversion rate can effectively evaluate the effectiveness of marketing activities, help enterprises understand the success or failure of marketing strategies, optimize future marketing methods, and increase the return on investment;
[0070] The formula for calculating the user growth rate is as follows:
[0071]
[0072] The user growth rate is a direct indicator of a company's growth and health. Analyzing the user growth rate can intuitively reflect changes in the user base and help enterprises determine whether they are growing steadily. In the formula, User Growth Rate represents the user growth rate, New Users represents the number of new users acquired in the current time period, and Total User represents the total number of users in the previous time period. By monitoring the user growth rate, enterprises can identify trends in user churn and then take measures to improve user retention rates, ensuring that while attracting new users, they do not neglect the value of existing users;
[0073] The formula for data key trend analysis is as follows:
[0074] Trend Analysis=Time Series(D original )
[0075] Trend analysis can reveal the patterns in which data changes over time, provide enterprises with profound business insights, and help enterprises adjust strategies and operational directions in a timely manner to maintain their competitiveness in the market. In the formula, Trend Analysis represents the result of key trend analysis, and Time Series() represents the analysis method applied to time series data;
[0076] The data monitoring module plays a crucial real-time monitoring role in the entire data flow system. Based on the analysis results provided by the data analysis module, it continuously tracks and monitors the status of the data flow to ensure the efficiency and accuracy of data flow. By using a stream processing framework (such as Apache Kafka or Apache Flink), this module can analyze the data flow speed in real time to ensure that the data transmission time in the system meets the preset standards. At the same time, it also monitors the abnormal data volume and identifies abnormal fluctuations or incorrect data in the data flow by setting thresholds and anomaly detection algorithms (such as statistics-based methods or machine learning models) so as to take timely measures. In addition to monitoring the data flow, the data monitoring module is also responsible for real-time tracking of system performance metrics, including resource utilization rates (such as CPU and disk I / O), memory occupancy, and system response time. The monitoring of these performance metrics is achieved by integrating performance monitoring tools (such as Prometheus, Grafana, or ELK Stack) to ensure that the system can still operate stably under high load. When any abnormal situation is detected, the data monitoring module will automatically generate a data monitoring report and send it to the alarm feedback module. The alarm feedback module then conducts real-time detection on the received monitoring report. Once it finds that the data flow is abnormal or the system performance does not meet the standard, it will automatically trigger alarm information and feedback it to the operation and maintenance personnel in a timely manner. The operation and maintenance personnel can quickly locate the problem through these alarm messages, take corresponding measures for troubleshooting and system optimization, thereby ensuring the stability and reliability of the entire data flow system.
[0077] Through the comprehensive application of the above system, not only the transparency and real-time nature of data flow are improved, but also the enterprise's trust in data is enhanced, greatly improving the data utilization efficiency and helping the enterprise gain a more competitive advantage in data-driven decision-making.
[0078] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data flow intelligent monitoring system based on a data middle station, characterized by: It includes data acquisition module, data preprocessing module, data storage module, data analysis module, data monitoring module and alarm feedback module; The data acquisition module is used to collect raw data from multiple data sources in real time, including business platforms, sensor devices and external market data, and send the collected data to the data preprocessing module after verification; The data preprocessing module receives the output of the data acquisition module and performs data cleaning tasks, including data denoising, removing duplicate data, filling missing values, correcting data formats, and standardizing data; The data storage module is used to uniformly store the pre-processed cleaned data in a NoSQL database for subsequent data query and analysis; The data analysis module obtains data from the data storage module, performs a system analysis on the stored data, calculates the data missing ratio, data conversion rate, user growth rate and data key trend analysis, and generates data reports; The data monitoring module monitors the status of the data flow in real time according to the analysis results of the data analysis module, continuously analyzes the data flow speed and the amount of abnormal data, and monitors the system performance indicators, including CPU usage, memory usage and response time, and generates a data monitoring report and sends it to the alarm feedback module; The alarm feedback module detects the data monitoring report and automatically triggers the alarm information to be fed back to the operation and maintenance personnel when abnormal conditions in data flow and system performance are found.
2. According to claim 1, a data flow intelligent monitoring system based on a data middle station is characterized in that: The formula for data denoising is as follows: In the formula, Y(t) represents the denoised data value at time point t, X(t) represents the original data value at time point t, N represents the window size used to calculate the moving average, and i represents the subscript index.
3. According to claim 2, a data flow intelligent monitoring system based on a data middle station is characterized in that: The formula for removing duplicate data is as follows: D unique unique(D original ) In the formula, D unique represents the unique data set after removing duplicates, D original Represents the original data set, and unique represents a function that returns all unique records in the data set.
4. According to claim 3, a data flow intelligent monitoring system based on a data middle station is characterized in that: The formula for filling missing values is as follows: In the formula, D filled (i) represents the data set after filling in missing values, D original Represents the original data set, NaN represents the symbol of missing values, Imputation() represents the algorithm for filling missing values, and i represents the digital flag.
5. According to claim 4, a data flow intelligent monitoring system based on a data middle station is characterized in that: The formula for the correction data format is as follows: D corrected (j)=format(D original (j)) D corrected (j) represents the dataset after format correction, D original Represents the original data set, format() represents the processing function used to convert the data into a specific format, and j represents the position of the data in the data set.
6. According to claim 5, a data flow intelligent monitoring system based on a data middle station is characterized in that: The formula for normalizing the data is as follows: In the formula, D standardized (i) represents the standardized data value, D original (i) represents the original data value, μ represents the mean of the original data set, and σ represents the standard deviation of the original data set.
7. The data flow intelligent monitoring system based on the data middle station according to claim 6 is characterized by: The formula for calculating the missing data ratio is as follows: In the formula, Missing Rate represents the proportion of missing data, Number of Missing Values represents the number of missing values in the data set, and Total Number of Values represents the total number of values in the data set.
8. The data flow intelligent monitoring system based on the data middle station according to claim 7 is characterized by: The calculation formula of the data conversion rate is as follows: In the formula, Conversion Rate represents the data conversion rate, Number of Conversion represents the number of users who complete a specific goal, and Total visito represents the total number of users who visit the data.
9. The data flow intelligent monitoring system based on the data middle station according to claim 8 is characterized by: The calculation formula of the user growth rate is as follows: In the formula, User Growth Rate represents the user growth rate, New Users represents the number of new users obtained in the current time period, and Total User represents the total number of users in the previous time period.
10. The data flow intelligent monitoring system based on the data middle station according to claim 9 is characterized in that: The formula for the key trend analysis of the data is as follows: Trend Analysis=Time Series(D original ) In the formula, Trend Analysis represents the result of key trend analysis, and Time Series() represents the analysis method applied to time series data.
Citation Information
Patent Citations
Enterprise big data analysis system based on artificial intelligence
CN117993737A
Industrial and commercial intelligent management system based on big data
CN118967210A
Automatic intelligent equipment operation and maintenance control system
CN119396037A
An intelligent data quality management system for improved AI / ML model performance across various industries.
DE202024104069U1
Cited By
Multi-level data intelligent statistical analysis system based on machine learning
CN121502159A