Real-time data stream processing method and device, equipment and medium
By constructing data acquisition pipelines and time window technology, real-time data streams are processed and visualized in real time, solving the problems of information overload and the limitations of anomaly detection, improving the intelligence level of data stream processing, and meeting the needs of real-time decision-making.
Patent Information
- Application Number
- CN202510987270.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies suffer from problems such as information overload and inefficient extraction, limitations in anomaly detection, lack of deep insight capabilities, and insufficient real-time performance and scalability when processing real-time data streams. They are unable to quickly capture core information, identify complex anomalies, and provide real-time decision support.
Data streams are collected in real time through a pre-built data acquisition pipeline, data is encapsulated and preprocessed, feature data is determined using time windows, data summaries are generated and anomalies are detected, and anomaly reports are generated in conjunction with a knowledge base to achieve data visualization.
It enables efficient processing of massive real-time data streams, automatically extracts core information, improves the comprehensiveness and accuracy of anomaly identification, reduces false alarm rate, and meets the real-time requirements of scenarios such as industrial monitoring and financial risk control.
Smart Images

Figure CN120849469A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data processing, and in particular to a real-time data stream processing method, apparatus, device, and medium. Background Technology
[0002] Against the backdrop of the rapid development of big data and the Internet of Things (IoT), real-time data streams, such as industrial sensor data, financial transaction streams, web logs, and IoT device data, exhibit characteristics of high speed, massive volume, and heterogeneity. Current processing of real-time data streams faces the following key challenges: 1. Information overload and inefficient extraction: The raw data streams are massive in scale and complex in structure, making it difficult to quickly capture core information. Traditional methods rely on manual analysis or simple statistics, failing to automatically extract key patterns, trends, or anomalies from the data, leading to delayed decision-making. 2. Limitations of anomaly detection: Existing systems often focus on single-dimensional anomaly identification, struggling to handle contextual or collective anomalies in complex scenarios, resulting in high false alarm rates and difficulty adapting to dynamically changing data streams. 3. Lack of deep insight capabilities: Even when anomalies are detected, traditional systems can only output simple alerts, unable to combine domain knowledge to analyze root causes, scope of impact, and response suggestions, leading to low decision-making efficiency. 4. Insufficient real-time performance and scalability: Faced with high-speed data streams of tens of thousands of records per second, traditional batch processing architectures suffer from high latency and are difficult to scale to adapt to diverse data sources, failing to meet the real-time requirements of scenarios such as industrial monitoring and financial risk control. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a real-time data stream processing method, apparatus, device, and medium, which can efficiently process real-time data streams and improve the intelligence level of real-time monitoring and decision-making. The specific solution is as follows:
[0004] Firstly, this application provides a real-time data stream processing method, including:
[0005] Data streams are acquired in real time through a pre-built data acquisition pipeline, and the acquired data streams are encapsulated to obtain the encapsulated initial data stream.
[0006] The initial data stream obtained through a preset unified interface is preprocessed to obtain the target data stream; the preprocessing operations include format standardization, noise filtering, outlier handling, and missing value imputation;
[0007] The target feature data corresponding to the target data stream is determined based on a pre-set time window, a data summary corresponding to the target data stream is generated based on the target feature data, and anomaly detection is performed on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window;
[0008] If the anomaly detection result indicates the presence of an anomaly, a corresponding anomaly report is generated based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence.
[0009] Based on the data summary and the anomaly report, the data stream information corresponding to the target data stream is visualized and displayed to the user in real time.
[0010] Optionally, the step of acquiring data streams in real time through a pre-built data acquisition pipeline and encapsulating the acquired data streams to obtain an encapsulated initial data stream includes:
[0011] The system collects diverse data streams in real time based on a pre-built data acquisition pipeline. These diverse data streams include sensor network time-series data, network logs, financial transaction streams, IoT device status, and user behavior data. The data acquisition pipeline supports different data transmission protocols and interface standards, and uses asynchronous I / O technology and batch processing technology to collect data streams.
[0012] The data acquisition pipeline is used to identify the format of the diverse data streams and encapsulate them into a preset standardized format to obtain the encapsulated initial data stream.
[0013] Optionally, determining the target feature data corresponding to the target data stream based on a pre-set time window, and generating a data summary corresponding to the target data stream based on the target feature data, includes:
[0014] The data stream types corresponding to the target data streams are determined, and the feature extraction methods corresponding to the target data streams are determined according to the data stream types. The target feature data corresponding to the target data streams is determined based on the feature extraction methods and a pre-set time window.
[0015] Based on the target feature data and streaming processing technology, the effective information in the target data stream is compressed and summarized to obtain the data summary corresponding to the target data stream.
[0016] Optionally, the anomaly detection of the target data stream based on the target feature data includes:
[0017] Based on the data stream type corresponding to the target data stream, an anomaly detection algorithm is determined for the target data stream. Anomaly detection is performed on the target data stream based on the anomaly detection algorithm and the target feature data to determine whether there is an anomaly in the target data stream. The types of anomalies include point anomalies, context anomalies, and collective anomalies.
[0018] Among them, the point anomaly is the appearance of a single burst of high-value data or low-value data in the target data stream. The high-value data is data whose value meets the preset high-value judgment condition, and the low-value data is data whose value meets the preset low-value judgment condition. The context anomaly is an anomaly in the target data determined based on the context information corresponding to the target data in the target data stream. The collective anomaly is an anomaly in the combined data after combining the interrelated data in the target data stream.
[0019] Optionally, generating a corresponding anomaly report based on relevant data streams and corresponding feature data before and after the anomaly occurrence time includes:
[0020] The target anomaly information is obtained by analyzing the relevant data streams and corresponding feature data before and after the anomaly occurrence time. The preliminary cause of the anomaly is determined based on the target anomaly information and the historical anomaly information stored in the target knowledge base.
[0021] The preliminary cause of the anomaly is verified using the anomaly association rules stored in the target knowledge base to obtain the target cause of the anomaly, and an anomaly report corresponding to the anomaly is generated based on the target cause of the anomaly; the anomaly report includes anomaly details, confidence level, cause of anomaly, scope of impact of anomaly, and handling measures;
[0022] The content of the anomaly report is fed back to the target knowledge base to update the target knowledge base.
[0023] Optionally, the real-time data stream processing method further includes:
[0024] The system obtains rule modification instructions sent by the user and defines, modifies, enables, and disables target rules in the target knowledge base according to the rule modification instructions. The target knowledge base includes historically accumulated professional knowledge, historical system operation data, historical anomaly information, device dependencies, and target rules. The target rules are used to generate data summaries and anomaly cause inferences.
[0025] Optionally, the step of visually displaying the data stream information corresponding to the target data stream to the user terminal in real time based on the data summary and the anomaly report includes:
[0026] The data summary corresponding to the target data stream is graphically displayed to the user.
[0027] The abnormal situations in the abnormal reports are classified according to a preset classification standard to obtain an abnormal alarm list. The abnormal alarm list is then displayed to the user terminal so that the user terminal can drill down to view the abnormal information associated with each abnormal situation in the abnormal alarm list by clicking.
[0028] Secondly, this application provides a real-time data stream processing apparatus, comprising:
[0029] The data encapsulation module is used to acquire data streams in real time through a pre-built data acquisition pipeline and encapsulate the real-time acquired data streams to obtain the encapsulated initial data stream.
[0030] The data preprocessing module is used to perform preprocessing operations on the initial data stream obtained through a preset unified interface to obtain the target data stream; the preprocessing operations include format standardization, noise filtering, outlier handling, and missing value imputation;
[0031] The summary generation and anomaly detection module is used to determine the target feature data corresponding to the target data stream based on a pre-set time window, generate a data summary corresponding to the target data stream based on the target feature data, and perform anomaly detection on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window;
[0032] The report generation module is used to generate a corresponding anomaly report based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence if the anomaly detection result indicates that an anomaly exists.
[0033] The information display module is used to visualize and display the data stream information corresponding to the target data stream to the user terminal in real time based on the data summary and the anomaly report.
[0034] Thirdly, this application provides an electronic device, comprising:
[0035] Memory, used to store computer programs;
[0036] A processor is used to execute the computer program to implement the aforementioned real-time data stream processing method.
[0037] Fourthly, this application provides a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the aforementioned real-time data stream processing method.
[0038] In this application, a data stream is acquired in real time through a pre-constructed data acquisition pipeline, and the acquired data stream is encapsulated to obtain an encapsulated initial data stream. The initial data stream, obtained through a pre-defined unified interface, undergoes preprocessing operations to obtain a target data stream. These preprocessing operations include format standardization, noise filtering, outlier handling, and missing value imputation. Target feature data corresponding to the target data stream is determined based on a pre-set time window. A data summary corresponding to the target data stream is generated based on the target feature data, and anomaly detection is performed on the target data stream based on the target feature data. The time window includes a sliding time window or a rolling time window. If the anomaly detection result indicates the presence of an anomaly, a corresponding anomaly report is generated based on the relevant data streams and corresponding feature data before and after the anomaly occurrence time. The data stream information corresponding to the target data stream is visualized and displayed to the user in real time based on the data summary and the anomaly report. As can be seen, this application determines target feature data based on a time window and generates a data summary based on the target feature data. By leveraging time windows to extract key features, massive amounts of raw data streams are condensed into concise summary information, replacing traditional manual analysis or simple statistics. This automatically extracts the core information of the data stream, enabling users to quickly grasp the overall data flow and addressing decision-making delays caused by information overload. Anomaly detection based on target feature data covers point anomalies, contextual anomalies, and collective anomalies. Features are updated in real-time via time windows to adapt to dynamic changes in the data stream, reducing false alarms caused by changes in data distribution. This overcomes the limitations of traditional single-dimensional detection, improving the comprehensiveness and accuracy of anomaly identification and lowering the false alarm rate. If an anomaly is detected, a corresponding anomaly report is generated based on the relevant data streams and corresponding feature data before and after the anomaly occurrence time. By integrating contextual information, it provides detailed information far exceeding that of simple alerts, avoiding the output of only superficial information of abnormal alerts. At the same time, it assists in tracing the root cause through correlation data and feature analysis, reducing the cost of manual analysis and improving decision-making efficiency. It uses a pre-built data acquisition pipeline to efficiently access data streams, combined with streaming preprocessing and real-time feature calculation within a time window, ensuring low latency from data acquisition to analysis. This meets the real-time requirements of scenarios such as industrial monitoring and financial risk control. Meanwhile, the initial data stream obtained through a pre-set unified interface can be compatible with diverse data sources, solving the problem of data heterogeneity. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0040] Figure 1 This is a flowchart of a real-time data stream processing method disclosed in this application;
[0041] Figure 2 This is a schematic diagram of the structure of a real-time data stream processing device disclosed in this application;
[0042] Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in this application. Detailed Implementation
[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] Against the backdrop of the rapid development of big data and the Internet of Things, real-time data streams exhibit characteristics of high speed, massive volume, and heterogeneity. Current processing methods for real-time data streams suffer from problems such as information overload and inefficient extraction, limitations in anomaly detection, lack of deep insight capabilities, and insufficient real-time performance and scalability. To address these issues, this application provides a real-time data stream processing method that can efficiently process real-time data streams and improve the intelligence level of real-time monitoring and decision-making.
[0045] See Figure 1 As shown in the figure, this application discloses a real-time data stream processing method, including:
[0046] Step S11: Collect data streams in real time through a pre-built data acquisition pipeline, and encapsulate the real-time collected data streams to obtain the encapsulated initial data stream.
[0047] In this embodiment, diverse data streams can first be acquired in real time based on a pre-built data acquisition pipeline. These diverse data streams include, but are not limited to, sensor network time-series data, network logs, financial transaction streams, IoT device status, and user behavior data. The data acquisition pipeline supports different data transmission protocols and interface standards, and employs asynchronous I / O (Input / Output) and batch processing technologies to acquire the data streams. Then, the data acquisition pipeline can be used to identify the formats of the diverse data streams and encapsulate them into a preset standardized format to obtain the encapsulated initial data stream.
[0048] It should be noted that this embodiment mainly consists of a data acquisition and access module, a data preprocessing and cleaning module, a real-time feature engineering module, an intelligent summary generation module, an anomaly detection and identification module, an anomaly insight and root cause analysis module, a knowledge base and rule management module, and a visualization and interaction module. Among these, real-time data can first be acquired through the data acquisition and access module.
[0049] The data acquisition and access module is primarily used to build a high-efficiency, stable, and high-throughput real-time data stream ingestion pipeline. This pipeline connects to various data sources of different origins and formats, including but not limited to: time-series data streams generated by various sensor networks, such as industrial equipment operating parameters and environmental monitoring data, complex network traffic logs, financial market transaction data streams, various status information sent by IoT devices, and real-time user behavior data in applications. To accommodate diverse data sources, this module supports multiple data transmission protocols and technologies, such as TCP / IP (Transmission Control Protocol / Internet Protocol), UDP (User Datagram Protocol), MQTT (Message Queuing Telemetry Transport), mainstream message queuing systems such as Apache Kafka, RabbitMQ, and Apache Pulsar, the ingestion API (Application Programming Interface) of stream processing platforms, various database connectors, and file listening. During data acquisition, this module employs asynchronous I / O, batch processing, or micro-batch processing techniques to minimize data latency, focusing on data real-time performance and integrity. Meanwhile, to ensure continuous operation even under conditions of fluctuating data sources, network instability, or even the failure of some acquisition nodes, this module incorporates fault tolerance and retry mechanisms to guarantee reliable data transmission. Furthermore, this module is responsible for preliminary format recognition and encapsulation of the raw data, providing a unified data interface for subsequent preprocessing stages. The efficient operation of the data acquisition and access module is a prerequisite for the entire system to handle massive, high-speed, real-time data streams, determining the real-time nature and coverage of the system's analysis. Utilizing the dynamic scalability of the data acquisition and access module adapts to the ever-increasing data volume and data source types, thereby ensuring that the system stably acquires the required data in various application scenarios.
[0050] Step S12: Perform preprocessing operations on the initial data stream obtained through a preset unified interface to obtain the target data stream; the preprocessing operations include format standardization, noise filtering, outlier handling, and missing value filling.
[0051] After the data acquisition and access module obtains the initial data stream, it enters the data preprocessing and cleaning module. Since the initial data stream often contains various defects, such as inconsistent formats, noise, outliers, missing values, and redundant information, these problems can severely affect the accuracy of subsequent analysis. Therefore, the core task of this module is to perform a series of transformations, cleaning, and standardization operations on the initial data stream. The specific process may include: first, parsing the data to unify different formats into a standardized internal format; then, real-time data verification to identify and process data with format errors, invalid values, or data that does not conform to business logic; next, applying denoising techniques, such as moving averages, exponential smoothing, or wavelet analysis, to filter out random noise in the data; and finally, identifying and processing outliers, such as using preliminary real-time outlier detection methods based on statistics, distance, or models to mark, correct, or remove extreme values. For missing values in the data stream, this module can fill them according to preset strategies, such as using the previous valid value, historical average, or using predicted values based on simple time series models (such as EWMA, Exponential Weighted Moving Average).
[0052] It should be noted that all the above preprocessing operations must be performed in a streaming manner to ensure extremely low processing latency and keep up with the speed of real-time data streams. The data preprocessing and cleaning module lays a solid foundation for subsequent accurate operation by providing a clean, consistent, and structured data stream, and is a key link in improving the overall analytical capabilities of the system.
[0053] Step S13: Determine the target feature data corresponding to the target data stream based on a pre-set time window, generate a data summary corresponding to the target data stream based on the target feature data, and perform anomaly detection on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window.
[0054] In this embodiment, after preprocessing the initial data stream to obtain the target data stream, the data stream types corresponding to each target data stream are first determined, and feature extraction methods corresponding to each target data stream are determined based on the data stream type. Target feature data corresponding to the target data stream is then determined based on the feature extraction methods and a pre-set time window. Next, effective information in the target data stream is compressed and summarized based on the target feature data and streaming processing technology to obtain a data summary corresponding to the target data stream. Further, an anomaly detection algorithm corresponding to the target data stream is determined based on the data stream type corresponding to each target data stream. Anomaly detection is performed on the target data stream based on the anomaly detection algorithm and the target feature data to determine whether anomalies exist in the target data stream. The types of anomalies include, but are not limited to, point anomalies, context anomalies, and collective anomalies. Point anomalies are single bursts of high or low value data in the target data stream; high value data is data whose values meet preset high value criteria, and low value data is data whose values meet preset low value criteria. Context anomalies are anomalies in the target data determined based on the context information corresponding to the target data in the target data stream. Collective anomalies are anomalies that occur after combining interrelated data in the target data stream.
[0055] It's important to note that the preprocessed target data stream first enters the real-time feature engineering module. This module aims to extract or construct real-time features from the target data stream that reflect the essence and potential patterns of the data. These features can be used for intelligent summarization and anomaly detection. Unlike offline feature engineering, streaming feature engineering needs to be completed immediately upon data arrival or within a very short time window. Therefore, this module calculates various features based on a predefined time window, such as a sliding window or a flipped window. These features include, but are not limited to, statistics such as the mean, median, maximum, minimum, standard deviation, variance, counts or frequencies (e.g., the number of times a specific event occurs, the frequency of a certain value), trends (e.g., the slope of a linear regression, differences between adjacent data points), time-series characteristics (e.g., autocorrelation coefficient, specific frequency components of Fourier transform results), and aggregation features (e.g., the aggregated values of multiple related data streams within the same time window). This module applies different feature extraction strategies and algorithms for different types of data streams and different analysis objectives. For example, for sensor data, it can extract its volatility and periodicity features; for log data, it can extract the sequence patterns of events, error code frequencies, etc. An efficient feature computation framework, such as state management based on a stream processing engine, is key to this module. It ensures that the corresponding target feature data can be calculated in real time while data is flowing in at high speed, providing high-quality input for subsequent analysis.
[0056] Secondly, the target feature data extracted by the real-time feature engineering module will be sent to the intelligent summary generation module. The core function of this module is to analyze the target feature data using streaming processing and machine learning algorithms to effectively compress and summarize information from the continuous real-time data stream, generating concise and informative summaries to help users quickly grasp the overall state and main trends of the data stream. These summaries can be presented in various forms, such as: real-time updated key indicators, such as real-time aggregated values of average throughput, current processing rate, and resource utilization; window-based data distribution overviews, such as using histograms, quantile sketches, etc., to display the distribution characteristics of data values; real-time topic words or keyword clouds extracted from text or unstructured data streams; and the main patterns or groups in the data stream and their proportion changes identified through streaming clustering or classification algorithms. The algorithms used in this module can handle high-speed, unlimited data streams and have low memory requirements. For example, online clustering algorithms, such as StreamKMeans, can be used to identify the main clusters in the data stream, or lightweight online learning models can be applied to identify the main patterns in the data stream. The intelligent summary generation process is real-time, and the generated summary can dynamically reflect the latest changes in the data stream, providing users or more advanced decision-making systems with a data overview and greatly improving the efficiency of information acquisition.
[0057] Secondly, in parallel or serial with the intelligent summarization generation module, the output target feature data of the real-time feature engineering module also enters the anomaly detection and identification module. This module is the core engine for discovering potential problems. By monitoring the target data stream and the corresponding target feature data, it identifies data points, event sequences, or overall patterns that deviate from normal behavior patterns. This module incorporates a variety of anomaly detection algorithms suitable for streaming data. These algorithms include, but are not limited to: statistical model-based methods, such as detecting values exceeding confidence intervals and EWMA-based bias detection; distance or density-based methods, such as real-time calculation of the Local Outlier Factor (LOF) or density-based clustering of the data stream; time-series analysis-based methods, such as using real-time prediction models to mark points with excessively large differences between actual and predicted values as anomalies; and machine learning-based methods, such as online-trained Isolation Forest and One-Class SVM; or using Recurrent Neural Networks (RNNs) to capture anomalous patterns in sequences. This module has high detection accuracy and low false alarm rate, and can adapt to the non-stationary nature of data streams. It can detect anomalies in real time and accurately, which is crucial for the system's early warning and risk control capabilities.
[0058] Step S14: If the anomaly detection result indicates the presence of an anomaly, then generate a corresponding anomaly report based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence.
[0059] In this embodiment, if the anomaly detection result indicates the presence of an anomaly, the target anomaly information can first be obtained by analyzing the relevant data streams and corresponding feature data before and after the anomaly occurrence time. A preliminary cause of the anomaly is then determined based on the target anomaly information and historical anomaly information stored in the target knowledge base. Next, the preliminary cause of the anomaly can be verified using the anomaly association rules stored in the target knowledge base to obtain the target anomaly cause. An anomaly report corresponding to the anomaly is then generated based on the target anomaly cause. The anomaly report includes, but is not limited to, anomaly details, confidence level, cause of the anomaly, scope of impact, and handling measures. Finally, the content of the anomaly report can be fed back to the target knowledge base for updating.
[0060] It should be noted that the system can obtain rule modification instructions sent by the user in real time, and define, modify, enable and disable target rules in the target knowledge base according to the rule modification instructions; the target knowledge base includes, but is not limited to, historically accumulated professional knowledge, historical system operation data, historical anomaly information, device dependencies and target rules; the target rules can be used to generate data summaries and anomaly cause inference.
[0061] Understandably, when the anomaly detection and identification module issues an anomaly alarm, the system will immediately activate the anomaly insight and root cause analysis module. The goal of this module is to transform the raw anomaly signal into meaningful and understandable insight information and analyze its potential root causes as much as possible. This typically involves in-depth contextual analysis of relevant data streams, real-time feature data, and system status before and after the anomaly occurred. This module also interacts with a knowledge base to query relevant domain knowledge, historical anomaly information, device dependencies, business rules, etc. This module employs various analytical techniques, such as: correlation analysis to identify other data streams or system indicators strongly correlated with the current anomaly; pattern matching to search for similar patterns in historical anomaly information stored in the knowledge base to infer the possible causes of the current anomaly; and a rule-based reasoning engine to apply expert rules stored in the knowledge base or anomaly association rules discovered through machine learning for logical reasoning, such as "If sensor A reading increases abnormally and device B load drops sharply, the possible cause is valve C failure." This module outputs an insight report, i.e., an anomaly report, containing anomaly details, confidence level, cause of the anomaly, scope of impact, and handling measures. This module enhances the system's practical value, transforming low-level "alarms" into advanced "intelligent insights" to help users quickly understand problems and formulate response strategies.
[0062] The core supporting module for the anomaly insight and root cause analysis module is the knowledge base and rule management module. This module supports the accuracy and intelligence of intelligent summarization, anomaly detection, and especially anomaly insight. The knowledge base stores various types of key information, including but not limited to knowledge accumulated by domain experts, such as the normal operating parameter range of specific equipment, the steps and indicators of key processes, the priority and processing flow of different alarms, historical system operation data (e.g., historical normal and abnormal patterns after aggregation or characterization), topological relationships and interdependencies between devices or system components, and target rules obtained through data mining or expert definition, such as "abnormal temperature sensor readings are usually related to cooling fan failures." The rule management function of this module allows system administrators or domain experts to define, modify, enable, or disable various target rules. These rules can be used to assist in intelligent summarization generation, filtering, or aggregating anomaly alarms, such as notifying only once of the same type of alarm that recurs within a short period, and performing reasoning and judgment in the anomaly insight module. By providing rich background information and reasoning capabilities, the knowledge base enables the system to integrate actual business logic and domain experience, significantly improving the accuracy and practicality of anomaly insight.
[0063] Step S15: Based on the data summary and the anomaly report, visualize the data stream information corresponding to the target data stream to the user terminal in real time.
[0064] In this embodiment, after obtaining the data summary and anomaly report, the data summary corresponding to the target data stream can be graphically displayed to the user terminal. Simultaneously, the anomalies in the anomaly report can be categorized based on preset classification criteria to obtain an anomaly alert list. This list is then displayed to the user terminal, allowing them to drill down and view the anomaly information associated with each anomaly in the alert list.
[0065] It should be noted that the analysis results and user interaction needs of the entire system can be realized through the visualization and interaction module. This module provides a real-time updated visualization dashboard to display core information in an intuitive and easy-to-understand way, including but not limited to: graphically displaying key indicators and intelligent summaries of real-time target data streams, such as through line charts, area charts, heatmaps, and dashboard components; prominently displaying a list of real-time anomaly alerts and classifying anomalies according to severity or type; providing detailed information on anomaly events and anomaly reports generated by the anomaly insight and root cause analysis module, and supporting users to drill down to view raw data and contextual information related to the anomaly. In addition, this module provides system configuration, management, and interaction functions. For example, users can configure data source connection parameters, define or modify feature engineering strategies, adjust parameters or thresholds of anomaly detection algorithms, manage rules in the knowledge base, set alert notification methods such as email, SMS, and instant messaging, and view the system's historical operation records and analysis reports. Through a user-friendly interface and powerful visualization capabilities, this module enables users to quickly and comprehensively grasp complex real-time data situations and respond promptly to anomaly events.
[0066] As can be seen from the above, this embodiment constructs an end-to-end, intelligent real-time data stream processing and analysis pipeline through the collaborative work of the above modules. Based on the integrated system of efficient processing of real-time data streams, automatic generation of information summaries, accurate identification of complex anomalies, and in-depth insights combined with domain knowledge, it realizes effective understanding of massive high-speed data streams, extraction of key information, discovery of abnormal behavior, and in-depth analysis of causes, providing timely and accurate intelligent support for decision-making and improving the level of intelligence of real-time monitoring and decision-making.
[0067] The following example, using a bank's real-time risk control for credit card transactions as an example in a real-time anti-fraud monitoring scenario for financial transactions, illustrates the technical solution in this application.
[0068] 1. Data collection and access: Banks process tens of thousands of credit card transactions daily and need to monitor transaction risks in real time.
[0069] The data acquisition and access module connects to the bank's core transaction system via the HTTP / S protocol (HTTP, or Hypertext Transfer Protocol), synchronously accessing the real-time data stream of each transaction, including transaction amount, time, location, merchant type, cardholder's historical behavior, etc. It also accesses auxiliary data streams such as user device fingerprints, recent login logs, and changes in credit blacklists. High-throughput buffering is achieved through a Kafka message queue, and fault tolerance mechanisms are enabled to ensure no data loss.
[0070] 2. Data preprocessing and cleaning: The original transaction data contains disordered format and noise.
[0071] The preprocessing module standardizes timestamp format, standardizes merchant type coding, filters invalid transactions, and fills in missing device information using rules, such as associating commonly used devices based on IP (Internet Protocol) addresses, and performs preliminary deduplication marking on abnormal high-frequency transactions to avoid duplicate analysis.
[0072] 3. Real-time feature engineering: Extract risk features from the cleaned transaction data.
[0073] Basic features: Calculate "number of cross-regional transactions" and "percentage of nighttime transactions" based on a 10-minute sliding window; Behavioral features: Extract "deviation from historical consumption locations", "number of sudden changes in merchant type", and "ratio of single transaction amount to cardholder's average monthly consumption"; Association features: Calculate "matching degree between current IP and frequently used IP" and "consistency between device fingerprint and the last 3 transactions".
[0074] 4. Parallel processing of intelligent summarization and anomaly detection:
[0075] (1) Intelligent summary generation: Aggregates all transaction features to generate a real-time risk overview:
[0076] Statistical analysis of high-risk transactions, i.e., the proportion of transactions with abnormal characteristics;
[0077] A heat map showing the top five regions for nighttime cross-border transactions;
[0078] Streaming clustering is used to identify transaction patterns, displaying the percentage of categories such as "normal consumption," "suspected fraud," and "test transactions."
[0079] (2) Anomaly detection: Multiple algorithms are used to monitor the characteristics of a single transaction.
[0080] For a late-night overseas transaction, such as one involving 50,000 yuan, far exceeding the cardholder's historical single-transaction limit of 20,000 yuan, the One-Class SVM model detected a deviation from the normal distribution of features and marked it as an anomaly. Combined with related features, such as an unfamiliar overseas IP address and a device fingerprint not appearing in historical records, the LSTM (Long Short-Term Memory) time-series model was used for verification, such as the match rate with the cardholder's spending history over the past three months. Confirm the validity of the anomaly.
[0081] 5. Anomaly Detection and Root Cause Analysis: When the above-mentioned overseas transactions trigger anomaly alerts:
[0082] The insights module automatically links the cardholder's historical data, such as no overseas spending records in the past year, a good credit score, and real-time auxiliary data, such as three fraudulent transaction reports associated with this IP address in the past seven days. After querying the knowledge base, it matches the rule "overseas unfamiliar device + transactions exceeding historical amounts". The message "Suspected fraudulent use" was retrieved from historical cases. "In transactions with similar characteristics in 2024, 85% were ultimately confirmed as fraudulent use." An insight report was generated: "Abnormal transaction, 95% confidence level. It may be due to card information leakage leading to fraudulent use. We recommend immediately freezing the card and contacting the cardholder for verification. Historically, similar cases have resulted in an average loss of 32,000 yuan."
[0083] 6. Visualization and Interaction
[0084] The bank's risk control center dashboard displays a real-time "Full Transaction Risk Index," for example, currently at 65 points, with a warning line of 80 points. High-risk transactions are marked with a flashing red indicator. Clicking on an abnormal transaction entry displays a pop-up window showing a complete insight report, including transaction details, anomaly feature comparison charts, and case summaries matched to the knowledge base. Risk control personnel can trigger a "temporary card freeze" operation with a single click through the interface, or add a "manual review" note. The system simultaneously sends an SMS notification to the cardholder, achieving a closed-loop risk management process.
[0085] As shown above, through the fully automated processing described, the system completes the closed loop from transaction collection to risk warning within 100 milliseconds, successfully intercepting suspected fraudulent transactions and preventing losses for cardholders. Simultaneously, intelligent summaries help the risk control team grasp the overall risk situation in real time, and continuous iteration of the knowledge base constantly improves the system's anti-fraud accuracy.
[0086] See Figure 2 As shown in the embodiments, this application also discloses a real-time data stream processing apparatus, including:
[0087] The data encapsulation module 11 is used to collect data streams in real time through a pre-built data acquisition pipeline and encapsulate the real-time collected data streams to obtain the encapsulated initial data streams.
[0088] Data preprocessing module 12 is used to perform preprocessing operations on the initial data stream obtained through a preset unified interface to obtain a target data stream; the preprocessing operations include format standardization, noise filtering, outlier handling, and missing value imputation;
[0089] The summary generation and anomaly detection module 13 is used to determine the target feature data corresponding to the target data stream based on a preset time window, generate a data summary corresponding to the target data stream based on the target feature data, and perform anomaly detection on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window;
[0090] The report generation module 14 is used to generate a corresponding anomaly report based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence if the anomaly detection result indicates that an anomaly exists.
[0091] The information display module 15 is used to visualize and display the data stream information corresponding to the target data stream to the user terminal in real time based on the data summary and the anomaly report.
[0092] As can be seen from the above, this application determines target feature data based on a time window and generates a data summary based on the target feature data. By utilizing a time window to extract key features, massive amounts of raw data streams are condensed into concise summary information, replacing traditional manual analysis or simple statistics. This automatically extracts the core information of the data stream, enabling users to quickly grasp the overall situation of the data stream and solving the problem of decision-making lag caused by information overload. Anomaly detection based on target feature data can cover point anomalies, contextual anomalies, and collective anomalies. Features are updated in real time through the time window to adapt to the dynamic changes in the data stream, reducing false alarms caused by changes in data distribution. This overcomes the limitations of traditional single-dimensional detection, improves the comprehensiveness and accuracy of anomaly identification, and reduces the false alarm rate. If an anomaly is detected, a corresponding anomaly report is generated based on the relevant data streams and corresponding feature data before and after the anomaly occurrence time. By integrating contextual information, it provides detailed information far exceeding that of simple alerts, avoiding the output of only superficial information of abnormal alerts. At the same time, it assists in tracing the root cause through correlation data and feature analysis, reducing the cost of manual analysis and improving decision-making efficiency. It uses a pre-built data acquisition pipeline to efficiently access data streams, combined with streaming preprocessing and real-time feature calculation within a time window, ensuring low latency from data acquisition to analysis. This meets the real-time requirements of scenarios such as industrial monitoring and financial risk control. Meanwhile, the initial data stream obtained through a pre-set unified interface can be compatible with diverse data sources, solving the problem of data heterogeneity.
[0093] In some specific embodiments, the data encapsulation module 11 includes:
[0094] The data acquisition unit is used to acquire diverse data streams in real time based on a pre-built data acquisition pipeline. The diverse data streams include sensor network time-series data, network logs, financial transaction streams, IoT device status, and user behavior data. The data acquisition pipeline supports different data transmission protocols and interface standards, and uses asynchronous I / O technology and batch processing technology to acquire data streams.
[0095] The data encapsulation unit is used to identify the format of the diverse data streams using the data acquisition pipeline, and encapsulate the diverse data streams into a preset standardized format to obtain the encapsulated initial data stream.
[0096] In some specific embodiments, the summary generation and anomaly detection module 13 includes:
[0097] The feature extraction unit is used to determine the data stream type corresponding to the target data stream, determine the feature extraction method corresponding to the target data stream according to the data stream type, and determine the target feature data corresponding to the target data stream based on the feature extraction method and a preset time window;
[0098] The summary generation unit is used to compress and summarize the effective information in the target data stream based on the target feature data and streaming processing technology to obtain a data summary corresponding to the target data stream.
[0099] In some specific embodiments, the summary generation and anomaly detection module 13 includes:
[0100] An anomaly detection unit is used to determine the anomaly detection algorithm corresponding to the target data stream based on the data stream type corresponding to the target data stream, and to perform anomaly detection on the target data stream based on the anomaly detection algorithm and the target feature data to determine whether there is anomaly in the target data stream; the types of anomalies include point anomalies, context anomalies and collective anomalies;
[0101] Among them, the point anomaly is the appearance of a single burst of high-value data or low-value data in the target data stream. The high-value data is data whose value meets the preset high-value judgment condition, and the low-value data is data whose value meets the preset low-value judgment condition. The context anomaly is an anomaly in the target data determined based on the context information corresponding to the target data in the target data stream. The collective anomaly is an anomaly in the combined data after combining the interrelated data in the target data stream.
[0102] In some specific embodiments, the report generation module 14 includes:
[0103] The cause determination unit is used to obtain target anomaly information based on the analysis of relevant data streams and corresponding feature data before and after the anomaly occurrence time, and to determine the preliminary cause of the anomaly based on the target anomaly information and the historical anomaly information stored in the target knowledge base.
[0104] The report generation unit is used to verify the preliminary cause of the anomaly using the anomaly association rules stored in the target knowledge base, obtain the target cause of the anomaly, and generate an anomaly report corresponding to the anomaly based on the target cause of the anomaly; the anomaly report includes anomaly details, confidence level, cause of anomaly, scope of impact of anomaly, and handling measures;
[0105] The knowledge base update unit is used to feed back the content of the anomaly report to the target knowledge base to update the target knowledge base.
[0106] In some specific embodiments, the real-time data stream processing device further includes:
[0107] The rule modification unit is used to obtain rule modification instructions sent by the user terminal, and to define, modify, enable, and disable target rules in the target knowledge base according to the rule modification instructions; the target knowledge base includes historically accumulated professional knowledge, historical system operation data, historical anomaly information, device dependencies, and target rules; the target rules are used to generate data summaries and anomaly cause inferences.
[0108] In some specific embodiments, the information display module 15 includes:
[0109] The information display unit is used to graphically display the data summary corresponding to the target data stream to the user terminal;
[0110] The information viewing unit is used to classify the corresponding abnormal situations in the abnormal report based on a preset classification standard to obtain an abnormal alarm list, and to display the abnormal alarm list to the user terminal so that the user terminal can drill down to view the abnormal information associated with each abnormal situation in the abnormal alarm list by clicking.
[0111] Furthermore, embodiments of this application also disclose an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0112] Figure 3 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the real-time data stream processing method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0113] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0114] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0115] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including computer programs capable of performing the real-time data stream processing method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0116] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned real-time data stream processing method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0117] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0118] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0119] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0120] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0121] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A real-time data stream processing method, characterized in that, include: Data streams are acquired in real time through a pre-built data acquisition pipeline, and the acquired data streams are encapsulated to obtain the encapsulated initial data stream. The initial data stream obtained through a preset unified interface is preprocessed to obtain the target data stream; the preprocessing operations include format standardization, noise filtering, outlier handling, and missing value imputation; The target feature data corresponding to the target data stream is determined based on a pre-set time window, a data summary corresponding to the target data stream is generated based on the target feature data, and anomaly detection is performed on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window; If the anomaly detection result indicates the presence of an anomaly, a corresponding anomaly report is generated based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence. Based on the data summary and the anomaly report, the data stream information corresponding to the target data stream is visualized and displayed to the user in real time.
2. The real-time data stream processing method according to claim 1, characterized in that, The process of acquiring data streams in real time through a pre-built data acquisition pipeline and encapsulating the acquired data streams to obtain an encapsulated initial data stream includes: The system collects diverse data streams in real time based on a pre-built data acquisition pipeline. These diverse data streams include sensor network time-series data, network logs, financial transaction streams, IoT device status, and user behavior data. The data acquisition pipeline supports different data transmission protocols and interface standards, and uses asynchronous I / O technology and batch processing technology to collect data streams. The data acquisition pipeline is used to identify the format of the diverse data streams and encapsulate them into a preset standardized format to obtain the encapsulated initial data stream.
3. The real-time data stream processing method according to claim 1, characterized in that, The step of determining the target feature data corresponding to the target data stream based on a pre-set time window, and generating a data summary corresponding to the target data stream based on the target feature data, includes: The data stream types corresponding to the target data streams are determined, and the feature extraction methods corresponding to the target data streams are determined according to the data stream types. The target feature data corresponding to the target data streams is determined based on the feature extraction methods and a pre-set time window. Based on the target feature data and streaming processing technology, the effective information in the target data stream is compressed and summarized to obtain the data summary corresponding to the target data stream.
4. The real-time data stream processing method according to claim 3, characterized in that, The anomaly detection of the target data stream based on the target feature data includes: Based on the data stream type corresponding to the target data stream, an anomaly detection algorithm is determined for the target data stream. Anomaly detection is performed on the target data stream based on the anomaly detection algorithm and the target feature data to determine whether there is an anomaly in the target data stream. The types of anomalies include point anomalies, context anomalies, and collective anomalies. Among them, the point anomaly is the appearance of a single burst of high-value data or low-value data in the target data stream. The high-value data is data whose value meets the preset high-value judgment condition, and the low-value data is data whose value meets the preset low-value judgment condition. The context anomaly is an anomaly in the target data determined based on the context information corresponding to the target data in the target data stream. The collective anomaly is an anomaly in the combined data after combining the interrelated data in the target data stream.
5. The real-time data stream processing method according to claim 1, characterized in that, The generation of corresponding anomaly reports based on relevant data streams and corresponding feature data before and after the anomaly occurrence time includes: The target anomaly information is obtained by analyzing the relevant data streams and corresponding feature data before and after the anomaly occurrence time. The preliminary cause of the anomaly is determined based on the target anomaly information and the historical anomaly information stored in the target knowledge base. The preliminary cause of the anomaly is verified using the anomaly association rules stored in the target knowledge base to obtain the target cause of the anomaly, and an anomaly report corresponding to the anomaly is generated based on the target cause of the anomaly; the anomaly report includes anomaly details, confidence level, cause of anomaly, scope of impact of anomaly, and handling measures; The content of the anomaly report is fed back to the target knowledge base to update the target knowledge base.
6. The real-time data stream processing method according to claim 1, characterized in that, Also includes: The system obtains rule modification instructions sent by the user and defines, modifies, enables, and disables target rules in the target knowledge base according to the rule modification instructions. The target knowledge base includes historically accumulated professional knowledge, historical system operation data, historical anomaly information, device dependencies, and target rules. The target rules are used to generate data summaries and anomaly cause inferences.
7. The real-time data stream processing method according to claim 1, characterized in that, The step of visually displaying the data stream information corresponding to the target data stream to the user terminal in real time based on the data summary and the anomaly report includes: The data summary corresponding to the target data stream is graphically displayed to the user. The abnormal situations in the abnormal reports are classified according to a preset classification standard to obtain an abnormal alarm list. The abnormal alarm list is then displayed to the user terminal so that the user terminal can drill down to view the abnormal information associated with each abnormal situation in the abnormal alarm list by clicking.
8. A real-time data stream processing device, characterized in that, include: The data encapsulation module is used to acquire data streams in real time through a pre-built data acquisition pipeline and encapsulate the real-time acquired data streams to obtain the encapsulated initial data stream. The data preprocessing module is used to preprocess the initial data stream obtained through a preset unified interface to obtain the target data stream; The preprocessing operations include format normalization, noise filtering, outlier handling, and missing value imputation. The summary generation and anomaly detection module is used to determine the target feature data corresponding to the target data stream based on a pre-set time window, generate a data summary corresponding to the target data stream based on the target feature data, and perform anomaly detection on the target data stream based on the target feature data; the time window includes a sliding time window or a rolling time window; The report generation module is used to generate a corresponding anomaly report based on the relevant data streams and corresponding feature data before and after the time point of the anomaly occurrence if the anomaly detection result indicates that an anomaly exists. The information display module is used to visualize and display the data stream information corresponding to the target data stream to the user terminal in real time based on the data summary and the anomaly report.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the real-time data stream processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the real-time data stream processing method as described in any one of claims 1 to 7.
Citation Information
Cited By
Short message anomaly detection method and device, equipment and storage medium
CN121174156A