Machine learning-based platform big data exception online early warning method

Through the online early warning method of big data anomalies based on machine learning, the problem of insufficient flexibility and adaptability in the existing technology is solved, and more accurate and adaptive abnormality detection effects are achieved.

CN120045429APending Publication Date: 2025-05-27SHANGHAI YOUJIA NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510150067.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology lacks flexibility and adaptability in platform big data abnormality detection, making it difficult to adapt to data changes in time, and requires manual adjustment of parameters or model structure.

Method used

The online early warning method of big data anomalies based on machine learning is adopted, including data preprocessing module, model training module, anomaly detection and early warning module and monitoring and feedback module. The data patterns and features are automatically learned through machine learning algorithms to identify nonlinear relationships and slight changes.

Benefits of technology

It realizes more accurate abnormality detection, improves detection accuracy, and the machine learning model is adaptable, can automatically adjust parameters and structures to adapt to dynamic changes in data without manual adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045429A_ABST
    Figure CN120045429A_ABST
Patent Text Reader

Abstract

The invention discloses a platform big data anomaly online early warning method based on machine learning, and belongs to the technical field of smart home equipment control, the platform big data anomaly online early warning method based on machine learning comprises four subsystems of a data preprocessing module, a model training module, an anomaly detection and early warning module and a monitoring and feedback module, the data preprocessing module is responsible for collecting data of each platform; the model training module is responsible for selecting a model with optimal performance according to the characteristics of the data and the demand of anomaly detection; the anomaly detection and early warning module is responsible for carrying out anomaly detection on the real-time data; and the monitoring and feedback module is responsible for monitoring the running state of the system in real time. According to the method, complex modes and features in the data can be automatically learned through the addition of a machine learning algorithm, a nonlinear relation and small changes in the data are identified, the accuracy of anomaly detection is greatly improved, manual regulation of rules or parameters is not needed, and good anomaly detection performance is always kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of platform big data, and particularly relates to an online anomaly early warning method for platform big data based on machine learning. Background Art

[0002] During the operation of the platform, data is continuously collected in real time from various channels. These data sources are extensive, including user behavior data, such as browsing, purchasing, and commenting operations of users on e-commerce platforms; sensor data, such as environmental data of temperature, humidity, pressure, etc. uploaded by Internet of Things devices; and data generated by business systems, such as transaction records in banking systems and order status information in logistics systems. These data are generated in real time in a high-frequency and large-scale manner, providing basic materials for subsequent analysis and decision-making of the platform. The operation of the big data platform depends on the collaborative work of multiple components such as hardware devices, network environments, and software systems. When a component in the system fails or has performance problems, it may lead to abnormal data processing. If the data generation speed exceeds the processing capacity of the platform, it may lead to data backlog and congestion, affecting the normal operation of the platform. The online anomaly early warning method for platform big data comprehensively uses a variety of technologies and means to monitor and analyze a large amount of data generated during the operation of the platform in real time, so as to timely detect data anomalies and issue early warning signals.

[0003] Existing early warning methods usually need to make assumptions about the distribution of data, such as assuming that the data follows a normal distribution, etc. The anomaly detection method based on a fixed statistical model may not be able to adapt to these changes in time, and requires manual re-adjustment of parameters or model structure, with insufficient flexibility and self-adaptability. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide an online anomaly early warning method for platform big data based on machine learning.

[0005] The technical solution adopted to solve the above technical problem is: an online anomaly early warning method for platform big data based on machine learning, including four subsystems: a data preprocessing module, a model training module, an anomaly detection and early warning module, and a monitoring and feedback module. The data preprocessing module is responsible for collecting data from each platform; the model training module is responsible for selecting the model with the optimal performance according to the characteristics of the data and the requirements of anomaly detection; the anomaly detection and early warning module is responsible for performing anomaly detection on real-time data; the monitoring and feedback module is responsible for monitoring the running state of the system in real time.

[0006] Further, the data preprocessing module includes data collection, data processing, data storage, and visual display;

[0007] For data collection, various sensors or monitoring devices are used to monitor the parameters of the platform operation in real time to obtain the operation data of the platform and the user behavior data. By selecting an appropriate sampling frequency, sufficient data samples are ensured to reflect the online status of the platform's big data personnel;

[0008] For data processing, the collected data is cleaned to remove duplicate data, noise data and missing values, and the data is transformed by standardization and normalization to provide a basis for subsequent feature engineering and model training;

[0009] For data storage, the collected raw data is recorded and stored in a time series manner and saved through a suitable data storage medium;

[0010] For visual display, chart and dashboard visualization tools are used to present the processed data in an intuitive and easy-to-understand manner, facilitating data analysis and decision-making by online management personnel.

[0011] Furthermore, for each column of data that needs to be standardized in the data processing, its mean μ and standard deviation σ are calculated, using the formula:

[0012]

[0013] Each data point X is converted into a standardized Z value, where the Z value represents how many standard deviations the data point is from the mean;

[0014] The minimum value min and the maximum value max of each column that needs to be normalized are found, using the formula,

[0015]

[0016] Each data point X is converted into a normalized value X norm , whose range is between 0 and 1;

[0017] The data processing matches and merges the attributes of the same entity from different data sources, and at the same time converts data of different data types into types suitable for subsequent analysis.

[0018] Furthermore, the model training module selects a suitable machine learning model according to the characteristics of the data and the requirements of anomaly detection. The model training module selects the linear regression algorithm according to the problem type. Linear regression assumes that there is a linear relationship between the input features and the output labels, that is:

[0019] y = β 0 + β 1 x 1 + β 2 x 2 +…+ βn x n +∈

[0020] where y is the predicted value, x i is the input feature, β i is the model parameter, and ∈ is the error term;

[0021] In simple linear regression (with only one feature), the parameters β 0 (intercept) and β 1 (slope) can be initialized to some random values or given initial values based on experience. In multiple linear regression (with multiple features), all β parameters are also initialized;

[0022] The commonly used loss function is the mean squared error, and its formula is:

[0023]

[0024] where m is the number of samples, y i is the true value, is the predicted value, and the loss function is used to measure the degree of difference between the model prediction result and the true result;

[0025] The gradient descent algorithm is used to update the parameters to minimize the loss function. The basic principle of gradient descent is to update the parameters along the direction where the loss function decreases fastest (i.e., the opposite direction of the gradient). For the parameters β j of linear regression, the update formula is:

[0026]

[0027] where α is the learning rate, which controls the step size of parameter update;

[0028] The parameters are updated through multiple iterations until the loss function converges to a smaller value or reaches the preset number of iterations. In each iteration of the model training module, the gradient of the loss function with respect to each parameter is calculated, and then the parameters are updated according to the gradient descent formula.

[0029] Furthermore, the anomaly detection and warning module inputs the data obtained from the data preprocessing module into the trained model. The model performs anomaly detection on the real-time data based on the learned normal data patterns and features to determine whether the data is abnormal data.

[0030] The anomaly detection and warning module uses the Gaussian mixture model by determining the number K of Gaussian components in the Gaussian mixture model and initializing the parameters of the Gaussian mixture model, including the mean vector μ k of each Gaussian component, the covariance matrix ∑k, and the mixing coefficient π k , k = 1, 2,..., K;

[0031] For each data point x i , calculate the posterior probability p(k|x i ) that it belongs to the k-th Gaussian component. According to Bayes' formula:

[0032]

[0033] where N(x i |μ k ,∑k) is the Gaussian probability density function,

[0034] Update the parameters of the Gaussian mixture model according to the posterior probability calculated in the previous step, and update the mixing coefficients,

[0035]

[0036] where N is the total number of training data points;

[0037] Update the mean vector,

[0038]

[0039] Update the covariance matrix,

[0040]

[0041] Repeat the above two steps until the parameters of the model converge, that is, the change in the parameters is less than a preset threshold;

[0042] The anomaly detection and warning module substitutes the preprocessed real-time data into the trained Gaussian mixture model to calculate its probability density:

[0043]

[0044] According to the analysis of the platform system requirements and historical data, set a probability density threshold T. If p(x)<T, then determine that the real-time data x is abnormal data; otherwise, consider x to be normal data. When the abnormal value exceeds the threshold, the warning module will trigger a warning and send a warning message to relevant personnel.

[0045] Furthermore, the monitoring and feedback module monitors the running state of the system in real time, including the running conditions of each module such as data acquisition, data preprocessing, model training, and anomaly detection, discovers problems and faults in the system in a timely manner, evaluates the performance of the system regularly, calculates various indicators such as the accuracy rate, recall rate, and F1 value of the system, and evaluates whether the performance of the system meets the business requirements.

[0046] The beneficial effects of the present invention are as follows: By incorporating machine learning algorithms, the present invention can automatically learn complex patterns and features in data, identify non-linear relationships and subtle changes in data. Compared with traditional online warning methods, it can more accurately discover abnormal patterns hidden in massive data, not limited to simple statistical features or preset rules, greatly improving the accuracy of anomaly detection. Moreover, the machine learning model has self-adaptability and can automatically adjust the model's parameters and structure by continuously learning new data to adapt to the dynamic changes of data, without manual adjustment of rules or parameters, and always maintain good anomaly detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a schematic diagram of the modules of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] As Figure 1 shown, the online big data anomaly warning method based on machine learning in this embodiment includes four subsystems: a data preprocessing module, a model training module, an anomaly detection and warning module, and a monitoring and feedback module. The data preprocessing module is responsible for collecting data from each platform; the model training module is responsible for selecting the model with the best performance according to the characteristics of the data and the requirements of anomaly detection; the anomaly detection and warning module is responsible for detecting anomalies in real-time data; the monitoring and feedback module is responsible for monitoring the running status of the system in real-time.

[0050] The data preprocessing module includes data collection, data processing, data storage, and visualization display;

[0051] Regarding data collection, various sensors or monitoring devices are used to monitor the operating parameters of the platform in real-time to obtain the operating data and user behavior data of the platform. By selecting an appropriate sampling frequency, sufficient data samples are ensured to reflect the online status of big data personnel on the platform;

[0052] Regarding data processing, the collected data is cleaned to remove duplicate data, noise data, and missing values, and the data is transformed by standardization and normalization to provide a basis for subsequent feature engineering and model training;

[0053] Regarding data storage, the collected raw data is recorded and stored in a time series manner and saved through a suitable data storage medium;

[0054] For visualization, chart and dashboard visualization tools are adopted to present the processed data in an intuitive and understandable way, facilitating online managers to analyze and make decisions based on the data.

[0055] For each column of data that needs to be standardized during data processing, calculate its mean μ and standard deviation σ using the formula:

[0056]

[0057] Convert each data point X into a standardized Z - value, where the Z - value represents how many standard deviations the data point is from the mean;

[0058] Find the minimum value min and maximum value max for each column that needs to be normalized using the formula

[0059]

[0060] Convert each data point X into a normalized value X norm , whose range is between 0 and 1;

[0061] During data processing, match and merge the attributes of the same entity from different data sources, and at the same time convert data of different data types into types suitable for subsequent analysis.

[0062] The model training module selects a suitable machine learning model according to the characteristics of the data and the requirements of anomaly detection. The model training module selects the linear regression algorithm according to the problem type. Linear regression assumes a linear relationship between the input features and the output labels, that is:

[0063] y = β 0 + β 1 x 1 + β 2 x 2 + … + β n x n + ∈

[0064] where y is the predicted value, x i is the input feature, β i is the model parameter, and ∈ is the error term;

[0065] In simple linear regression (with only one feature), the parameters β 0 (intercept) and β 1 (slope) can be initialized to some random values or given initial values according to experience. In multiple linear regression (with multiple features), all β parameters are also initialized;

[0066] The commonly used loss function is the mean squared error, and its formula is:

[0067]

[0068] where m is the number of samples, and y i is the true value, is the predicted value, and the loss function is used to measure the degree of difference between the model prediction result and the true result;

[0069] The gradient descent algorithm is used to update the parameters to minimize the loss function. The basic principle of gradient descent is to update the parameters along the direction where the loss function decreases fastest (i.e., the opposite direction of the gradient). For the parameter β of linear regression j , the update formula is:

[0070]

[0071] where α is the learning rate, which controls the step size of parameter update;

[0072] The parameters are updated through multiple iterations until the loss function converges to a smaller value or reaches the preset number of iterations. In each iteration, the model training module calculates the gradient of the loss function with respect to each parameter, and then updates the parameters according to the gradient descent formula.

[0073] The anomaly detection and warning module inputs the data obtained from the data preprocessing module into the trained model. The model performs anomaly detection on the real-time data according to the learned normal data patterns and features to determine whether the data is abnormal.

[0074] The anomaly detection and warning module uses the Gaussian mixture model. By determining the number K of Gaussian components in the Gaussian mixture model, the parameters of the Gaussian mixture model are initialized, including the mean vector μ of each Gaussian component k , the covariance matrix ∑k, and the mixing coefficient π k , k = 1, 2,..., K;

[0075] For each data point x i , calculate its posterior probability p(k|x i ) belonging to the k-th Gaussian component. According to Bayes' formula:

[0076]

[0077] where N(x i |μ k ,∑k) is the Gaussian probability density function,

[0078] According to the posterior probability calculated in the previous step, update the parameters of the Gaussian mixture model and update the mixing coefficient,

[0079]

[0080] where N is the total number of training data points;

[0081] Update the mean vector,

[0082]

[0083] Update the covariance matrix,

[0084]

[0085] Repeat the above two steps until the parameters of the model converge, that is, the change in the parameters is less than a preset threshold;

[0086] The anomaly detection and warning module substitutes the preprocessed real-time data into the trained Gaussian mixture model to calculate its probability density:

[0087]

[0088] According to the requirements of the platform system and the analysis of historical data, a probability density threshold T is set. If p(x) < T, the real-time data x is determined to be abnormal data; otherwise, X is considered normal data. When the abnormal value exceeds the threshold, the warning module will trigger a warning and send a warning message to relevant personnel.

[0089] The monitoring and feedback module monitors the running state of the system in real time, including the running conditions of each module such as data collection, data preprocessing, model training, and anomaly detection, discovers problems and faults in the system in a timely manner, evaluates the performance of the system regularly, calculates various indicators such as the accuracy rate, recall rate, and F1 value of the system, and evaluates whether the performance of the system meets the business requirements.

[0090] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. An online early warning method for platform big data anomalies based on machine learning, characterized by: It includes four subsystems: data preprocessing module, model training module, anomaly detection and warning module, and monitoring and feedback module. The data preprocessing module is responsible for collecting data from each platform; The model training module is responsible for selecting the model with the best performance based on the characteristics of the data and the needs of anomaly detection; the anomaly detection and early warning module is responsible for anomaly detection of real-time data; and the monitoring and feedback module is responsible for real-time monitoring of the system's operating status.

2. The method for online early warning of platform big data anomalies based on machine learning according to claim 1 is characterized in that: The data preprocessing module includes data collection, data processing, data storage and visual display; In terms of data collection, the platform operation parameters are monitored in real time through sensors or monitoring devices to obtain the platform operation data and user behavior data. The sampling frequency is selected to ensure that enough data samples are collected to reflect the online status of the platform's big data personnel; In terms of data processing, the collected data is cleaned to remove duplicate data, noise data and missing values, and the data is standardized and normalized to provide a basis for subsequent feature engineering and model training; In terms of data storage, the collected raw data is recorded and stored in a time series manner and saved through appropriate data storage media; In terms of visual display, chart and dashboard visualization tools are used to present the processed data in an intuitive and easy-to-understand way, making it easier for online managers to analyze and make decisions on the data.

3. The method for online warning of platform big data anomalies based on machine learning according to claim 2 is characterized in that: The data processing calculates the mean μ and standard deviation σ of each column of data that needs to be standardized, using the formula: Convert each data point X into a standardized Z value, where the Z value indicates how many standard deviations the data point is from the mean; Find the minimum value min and maximum value max of each column to be normalized, using the formula, Convert each data point X to a normalized value X norm , which ranges from 0 to 1; The data processing matches and merges the attributes of the same entity in different data sources, and converts data of different data types into a type suitable for subsequent analysis.

4. The method for online warning of platform big data anomalies based on machine learning according to claim 1 is characterized in that: The model training module selects a suitable machine learning model according to the characteristics of the data and the needs of anomaly detection. The model training module selects a linear regression algorithm according to the problem type. Linear regression assumes that there is a linear relationship between input features and output labels, that is: y=β0+β1x1+β2x2+…+β n x n +∈ Where y is the predicted value, x i is the input feature, β i are model parameters, ∈ is the error term; In simple linear regression (only one feature), the parameters β0 (intercept) and β1 (slope) can be initialized to some random values ​​or given initial values ​​based on experience. In multivariate linear regression (multiple features), all β parameters are initialized in the same way. The commonly used loss function is the mean square error, and its formula is: Where m is the number of samples, y i is the true value, is the predicted value, and the loss function is used to measure the difference between the model prediction result and the actual result; The gradient descent algorithm is used to update the parameters to minimize the loss function. The basic principle of gradient descent is to update the parameters along the direction of the fastest decrease of the loss function (that is, the opposite direction of the gradient). For the linear regression parameter β j , the update formula is: Where α is the learning rate, which controls the step size of parameter updates; The parameters are updated through multiple iterations until the loss function converges to a smaller value or reaches a preset number of iterations. In each iteration, the model training module calculates the gradient of the loss function for each parameter and then updates the parameters according to the gradient descent formula.

5. The method for online early warning of platform big data anomalies based on machine learning according to claim 1 is characterized in that: The anomaly detection and early warning module inputs the data obtained from the data preprocessing module into the trained model. The model performs anomaly detection on the real-time data based on the learned normal data patterns and features to determine whether the data is abnormal data.

6. The method for online warning of platform big data anomalies based on machine learning according to claim 5 is characterized in that: The anomaly detection and warning module uses a Gaussian mixture model, determines the number of Gaussian components K in the Gaussian mixture model, and initializes the parameters of the Gaussian mixture model, including the mean vector μ of each Gaussian component. k , covariance matrix ∑k and mixing coefficient π k , k=1,2,…,K; For each data point x i , calculate the posterior probability p(k|x i ), according to the Bayesian formula: Where N(x i |μ k ,∑k) is the Gaussian probability density function, Update the parameters of the Gaussian mixture model and the mixing coefficients according to the posterior probability obtained from the previous step, where N is the total number of training data points; update the mean vector, update the covariance matrix, repeat the above two steps until the parameters of the model converge, that is, the change in the parameters is less than a preset threshold; The anomaly detection and warning module substitutes the preprocessed real-time data into the trained Gaussian mixture model to calculate its probability density: According to the requirements of the platform system and the analysis of historical data, set a probability density threshold T. If p(x) < T, then determine that the real-time data x is abnormal data; otherwise, consider X as normal data. When the abnormal value exceeds the threshold, the warning module will trigger a warning and send a warning message to relevant personnel.

7. The method for online warning of platform big data anomalies based on machine learning according to claim 1 is characterized in that: The monitoring and feedback module monitors the running status of the system in real time, including the running conditions of each module such as data collection, data preprocessing, model training, and anomaly detection, discovers problems and faults in the system in a timely manner, evaluates the performance of the system regularly, calculates various indicators such as the accuracy rate, recall rate, and F1 value of the system, and evaluates whether the performance of the system meets the business requirements.

Citation Information

Cited By

  • Information operation dimension intelligent management and control platform system and method, electronic equipment and storage medium

    CN121151246A