Big data anomaly detection method and system based on cloud computing

By combining sliding time windows, multidimensional statistical features, and graph database semantic features in a cloud computing environment, the problems of window segmentation adaptability and feature stability of high-frequency streaming data are solved, efficient anomaly detection and false alarm suppression are achieved, and the robustness and interpretability of the system are improved.

CN120744706AActive Publication Date: 2025-10-03WUHAN CITY VOCATIONAL COLLEGE +1

Patent Information

Application Number
CN202511257977.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-03
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Under the sudden fluctuations of high-frequency streaming data, the existing technology's window segmentation is not sufficiently adaptable to the business cycle, and it is difficult to quantify the stability between multi-source heterogeneous features, resulting in limited model robustness. The decision tree algorithm's sensitivity to feature fluctuations causes fluctuations in false alarm rates and deviations in root cause location, which poses particular challenges in real-time financial risk control and industrial Internet operation and maintenance scenarios.

Method used

A cloud computing-based big data anomaly detection method is adopted. The data stream is segmented through a sliding time window. Multidimensional statistical features and semantic features of the graph database are combined to calculate the feature fluctuation coefficient. Mutual information dynamic screening and gradient boosting decision tree model are used to generate anomaly probability values ​​and feature heat maps, realizing dynamic feature screening and coordinated optimization of stability.

Benefits of technology

It significantly enhances the representation robustness and noise immunity of high-dimensional feature space, improves the ability to accurately confirm abnormal signals and suppress false alarms, and improves the generalization adaptability and interpretability of the system in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744706A_ABST
    Figure CN120744706A_ABST
Patent Text Reader

Abstract

The invention provides a big data anomaly detection method and system based on cloud computing, and relates to the technical field of big data anomaly detection.According to the big data anomaly detection method and system based on cloud computing, a complete time sequence mode is effectively captured through a service cycle self-adaptive sliding window segmentation mechanism, and the regularity and continuity of data representation are remarkably enhanced; multi-dimensional statistical features and graph database-driven entity semantic features are deeply fused to construct a joint fluctuation quantitative model, so that the representation robustness and noise immunity of a high-dimensional feature space are greatly improved; a mutual information dynamic threshold screening strategy is combined with a feature stability collaborative optimization mechanism of a gradient boosting decision tree to form an anti-interference enhanced decision normal form, so that accurate confirmation of abnormal signals and efficient suppression of false alarm behaviors are realized in a complex dynamic environment; and based on a root cause positioning system of decision path reconstruction and weighted contribution accumulation, the problem of offset of characteristic contribution evaluation by a traditional method is solved, and the identification precision and interpretability of key abnormal driving factors are qualitatively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data anomaly detection, and in particular to a big data anomaly detection method and system based on cloud computing. Background Art

[0002] The current technological evolution in the field of big data anomaly detection is showing a diversified development trend. As the scale of data carried by cloud computing platforms continues to expand, mainstream solutions generally use sliding time windows for time series segmentation, combining statistical feature analysis and entity relationship mining based on graph databases to construct a multidimensional feature space; at the model level, gradient boosting decision trees have become a commonly used classifier due to their high efficiency, and mutual information feature screening has also been widely applied. However, industry observations show that the sudden fluctuations in high-frequency streaming data can easily lead to the problem of insufficient adaptability of window segmentation to business cycles. The difficulty in quantifying the stability of multi-source heterogeneous features leads to limited model robustness. At the same time, the challenges of false alarm rate fluctuations and root cause location deviation caused by the sensitivity of decision tree algorithms to feature fluctuations continue to attract attention in real-time financial risk control and industrial Internet operation and maintenance scenarios. How to coordinate feature engineering and model decision-making to adapt to complex dynamic environments has become a technical difficulty to be overcome in this field.

[0003] Therefore, it is necessary to provide a big data anomaly detection method and system based on cloud computing to solve the above technical problems. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a big data anomaly detection method and system based on cloud computing, which achieves the beneficial effect of accurate and effective anomaly data detection.

[0005] The present invention provides a big data anomaly detection method based on cloud computing, comprising:

[0006] S1: Split the pre-processed raw data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle;

[0007] S2: Extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment;

[0008] S3: combining the statistical features and semantic features of multiple consecutive standard data segments into a statistical feature matrix and a semantic feature matrix in chronological order, and calculating a feature fluctuation coefficient vector based on the statistical feature matrix and the semantic feature matrix;

[0009] S4: Fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor;

[0010] S5: Calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal labels in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen feature channels above the dynamic screening threshold to obtain a high-value feature set;

[0011] S6: Input the high-value feature set into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

[0012] Preferably, in step S1, the length of the sliding time window is greater than or equal to 3 complete business cycles.

[0013] Preferably, the preprocessing of the original data stream in step S1 includes the following steps:

[0014] Performing timestamp alignment on the original data stream to obtain a time-aligned data stream, and segmenting the time-aligned data stream according to a preset temporary segmentation window to obtain multiple temporary data segments;

[0015] Calculate the upper and lower thresholds for each temporary data segment, and construct a dynamic boundary interval for each temporary data segment based on the lower and upper thresholds. The calculation formulas for the upper and lower thresholds are: ;

[0016] in, is the upper threshold, is the third quartile of the temporary data segment, is the first quartile of the temporary data segment, is the empirical scaling factor, is the interquartile range, is the dynamic expansion coefficient, is the standard deviation of the temporary data segment;

[0017] A dynamic boundary interval is constructed based on the lower and upper thresholds, and abnormal data points that exceed the dynamic boundary interval in each temporary data segment are removed. The temporary data segments after the abnormal data points are removed are merged in timestamp order to generate a standardized data stream.

[0018] Preferably, in step S3, the columns in the statistical feature matrix and the semantic feature matrix represent the values ​​of all standard data segments of a feature dimension.

[0019] Preferably, in step S3, the step of generating the characteristic fluctuation coefficient vector includes:

[0020] The coefficient of variation is calculated for each column of the statistical feature matrix, and the standard deviation is calculated for each column of the semantic feature matrix. The coefficient of variation and the standard deviation are concatenated in dimensional order to form a feature fluctuation coefficient vector.

[0021] Preferably, in step S5, the construction of the dynamic screening threshold comprises the following steps:

[0022] The mean and standard deviation of the mutual information values ​​of all feature channels are calculated, and the value of the mean plus three times the standard deviation is calculated to obtain the first candidate threshold;

[0023] Calculate the ratio of the number of abnormal labels in the pre-stored benchmark label set to the total number of labels in the benchmark label set, and take the natural logarithm of the reciprocal of the ratio to obtain the second candidate threshold;

[0024] The larger value between the first candidate threshold and the second candidate threshold is taken as the dynamic screening threshold.

[0025] Preferably, in step S6, inputting the high-value feature set into the pre-trained gradient boosting decision tree model further comprises:

[0026] Multiply the characteristic fluctuation coefficient and the mutual information value to generate the weighted confidence;

[0027] Reconstruct the input order of high-value feature sets in descending order of weighted confidence.

[0028] Preferably, in step S6, a depth attenuation factor is added to the split gain calculation of the pre-trained gradient boosting decision tree model.

[0029] Preferably, in step S6, the step of obtaining an abnormal probability value and a feature heat map by dynamically weighting the feature contribution by the feature fluctuation coefficient vector includes:

[0030] A weight factor is generated through the feature fluctuation coefficient vector. The decision tree node split evaluation value is corrected in real time based on the weight factor to reconstruct the decision path. The abnormal probability value is output based on the terminal node distribution of the reconstructed decision path, and the path split contribution is accumulated to form a feature heat map.

[0031] The present invention also provides a cloud computing-based big data anomaly detection system, which is applied to a cloud computing-based big data anomaly detection method, including:

[0032] The data segmentation module is used to segment the pre-processed raw data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle;

[0033] The feature joint extraction module is used to extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment;

[0034] A fluctuation feature quantification module is used to combine the statistical features and semantic features of multiple consecutive standard data segments into a statistical feature matrix and a semantic feature matrix in chronological order, and calculate a feature fluctuation coefficient vector based on the statistical feature matrix and the semantic feature matrix;

[0035] Tensor fusion module, used to fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor;

[0036] The mutual information dynamic screening module is used to calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal labels in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen feature channels above the dynamic screening threshold to obtain a high-value feature set;

[0037] The anti-disturbance decision parsing module is used to input high-value feature sets into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

[0038] Compared with related technologies, the cloud computing-based big data anomaly detection method and system provided by the present invention has the following beneficial effects:

[0039] The present invention effectively captures complete time series patterns through a business cycle adaptive sliding window segmentation mechanism and significantly enhances the regularity and continuity of data representation; deeply integrates multi-dimensional statistical features with entity semantic features driven by a graph database to construct a joint fluctuation quantification model, which greatly improves the representation robustness and noise immunity of high-dimensional feature space; adopts a mutual information dynamic threshold screening strategy combined with a feature stability collaborative optimization mechanism of a gradient boosting decision tree to form an anti-interference enhanced decision paradigm, thereby achieving accurate confirmation of abnormal signals and efficient suppression of false alarm behaviors in complex dynamic environments; based on a root cause location system based on decision path reconstruction and weighted contribution accumulation, it breakthroughs the bias problem of feature contribution assessment in traditional methods, and achieves a qualitative improvement in the identification accuracy and interpretability of key abnormal driving factors; the entire technical solution integrates the multi-stage process of data preprocessing, feature engineering and model decision-making into an organically coordinated computing framework, while maintaining real-time processing capabilities, greatly enhancing the system's generalized adaptability to business scenario migration and sudden abnormal patterns, and realizing big data anomaly detection in a cloud computing environment in terms of accuracy, robustness, interpretability and scenario adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is a flow chart of a big data anomaly detection method based on cloud computing of the present invention;

[0041] Figure 2This is a module structure diagram of a cloud computing-based big data anomaly detection system of the present invention. DETAILED DESCRIPTION

[0042] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with one another unless there is a conflict.

[0043] It should also be noted that, for ease of description, only portions relevant to the present invention are shown in the accompanying drawings, rather than all of the contents. Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the various operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the various operations can be rearranged. The process can be terminated when its operations are completed, but may also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0044] Example 1

[0045] A big data anomaly detection method based on cloud computing, in the specific implementation process, such as Figure 1 , which shows a flow chart of a big data anomaly detection method based on cloud computing, including:

[0046] Step S1: Segment the pre-processed original data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle.

[0047] Specifically, in step S1, the length of the sliding time window is greater than or equal to three complete business cycles.

[0048] Specifically, the preprocessing of the original data stream in step S1 includes the following steps:

[0049] Performing timestamp alignment on the original data stream to obtain a time-aligned data stream, and segmenting the time-aligned data stream according to a preset temporary segmentation window to obtain multiple temporary data segments;

[0050] Calculate the upper and lower thresholds for each temporary data segment, and construct a dynamic boundary interval for each temporary data segment based on the lower and upper thresholds. The calculation formulas for the upper and lower thresholds are: ;

[0051] in, is the upper threshold, is the third quartile of the temporary data segment, is the first quartile of the temporary data segment, is the empirical scaling factor, is the interquartile range, is the dynamic expansion coefficient, is the standard deviation of the temporary data segment;

[0052] A dynamic boundary interval is constructed based on the lower and upper thresholds, and abnormal data points that exceed the dynamic boundary interval in each temporary data segment are removed. The temporary data segments after the abnormal data points are removed are merged in timestamp order to generate a standardized data stream.

[0053] In the specific implementation process, the original data stream is first aligned with the millisecond timestamp through the network time protocol to eliminate the cross-source timing deviation and generate a time-aligned data stream. Then, based on the preset temporary segmentation window, the length is usually 1 / 10 to 1 / 5 of the business cycle, and the time-aligned data stream is cut into continuous temporary data segments. The dynamic boundary construction operation is performed on each temporary data segment: the first quartile of the data segment is calculated. and the third quartile , and find the interquartile range , synchronously calculate the standard deviation of the data segment Then, substitute into the dynamic expansion coefficient formula , so that the threshold has data adaptability. When the standard deviation is larger, the expansion coefficient increases nonlinearly to accommodate reasonable fluctuations, and then the coefficient is scaled according to experience. , for example, the typical value is 1.5, and the upper threshold is calculated and lower threshold , forming a dynamic boundary interval for the distribution of the package data body, and treating the discrete data points outside the dynamic boundary interval as outliers and eliminating them. Finally, all the valid data of the temporary data segments after denoising are reassembled in the original timestamp order to generate a standardized data stream; for this standardized data stream, according to the business cycle characteristics, for example, such as financial transactions with a 24-hour cycle and industrial equipment with an 8-hour production shift as a cycle, the sliding window length is determined, and the sliding window length is forced to be greater than or equal to 3 complete business cycles. A number of standard data segments are generated by sliding interception with a fixed step size as the input basis for subsequent feature processing. Through the synergistic effect of the dynamic scaling mechanism of the temporary window and the periodic adaptation mechanism of the business window, both instantaneous noise interference is eliminated and the cross-cycle regularity pattern is completely retained. It can effectively retain the periodic characteristics of the data and reduce the proportion of invalid data segments caused by network jitter in the cloud computing environment.

[0054] Step S2: Extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment.

[0055] During the specific implementation process, firstly, a multi-dimensional statistical feature extraction operation is performed on each standard data segment. The multi-dimensional statistical features include but are not limited to multi-dimensional basic statistical indicators such as time domain mean, variance, skewness and frequency domain wavelet energy ratio. At the same time, the data packet header information is parsed to obtain entity identification fields such as but not limited to device code, operator ID, etc.; based on this entity identification field, the pre-built graph database retrieval process is activated. The construction of the pre-built graph database includes collecting entity relationship data of the entire business chain, for example, equipment number, sensor node, process link in industrial scenarios, account ID, counterparty, funds in financial scenarios. Flow direction, converting entities into graph nodes and assigning static attributes such as device model and account level. The interaction relationship between entities is converted into weighted directed edges, such as the frequency of device data flow transmission and the flow of account funds. The Neo4j engine is used to build a distributed graph structure storage cluster. The node attributes include timeliness labels, such as the update timestamp of the remaining life estimation of the device and the embedding of probabilistic transmission quality parameters in the edge attributes to achieve a millisecond response of billions of nodes; after completing the real-time search of the graph database, a dynamic semantic feature vector strongly associated with the entity identifier is output, and finally a hybrid feature containing numerical statistical features and associative semantic features is formed.

[0056] Step S3: The statistical features and semantic features of a plurality of consecutive standard data segments are combined into a statistical feature matrix and a semantic feature matrix in chronological order, and a feature fluctuation coefficient vector is calculated based on the statistical feature matrix and the semantic feature matrix.

[0057] Specifically, in step S3, the columns in the statistical feature matrix and the semantic feature matrix represent the values ​​of all standard data segments of a feature dimension.

[0058] Specifically, in step S3, the step of generating the characteristic fluctuation coefficient vector includes:

[0059] The coefficient of variation is calculated for each column of the statistical feature matrix, and the standard deviation is calculated for each column of the semantic feature matrix. The coefficient of variation and the standard deviation are concatenated in dimensional order to form a feature fluctuation coefficient vector.

[0060] During the specific implementation process, multiple consecutive standard data segments are first sorted in time series, and data structures are constructed for the statistical feature dimension and the semantic feature dimension respectively. The column vector of the statistical feature matrix represents the value sequence of a specific statistical feature on all data segments, which is recorded as a statistical feature column, and the column vector of the semantic feature matrix represents the value sequence of the preset dimension semantic feature on all data segments, which is recorded as a semantic feature column. The key technical constraint here is that the values ​​of all data segments of each feature dimension must be strictly arranged in timestamp order and must not be misplaced; then the feature fluctuation coefficient calculation is performed, and the dimensionless coefficient of variation calculation method is used for the statistical feature column, and the standard deviation of the column is divided by the mean of the column to eliminate the impact of dimensional differences on the evaluation of statistical feature stability, while the standard deviation of the semantic feature is calculated for the semantic feature column, retaining the original dimensional fluctuation information to reflect the degree of semantic deviation; finally, the feature fluctuation coefficient vector is constructed by splicing in the order of feature dimensions, where the splicing rule requires that the first element is the statistical feature variation coefficient arranged in the order of feature declaration, and the second element is the standard deviation of the semantic feature, arranged in the order of graph database dimension definition, providing a calculation basis for subsequent steps.

[0061] Step S4: Fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor.

[0062] During the specific implementation process, the statistical feature matrix and the semantic feature matrix are first dimensionally aligned, and the two matrices are uniformly expanded to the same feature dimension through zero padding to form an equal-dimensional matrix; at the same time, the feature fluctuation coefficient vector is replicated multiple times along the time dimension through the vertical broadcast mechanism to generate a fluctuation coefficient matrix, and the number of replications is strictly equal to the total number of time windows. Subsequently, a three-channel splicing operation is performed along the channel axis: the first channel is filled with the zero-expanded statistical feature matrix, the second channel is filled with the zero-expanded semantic feature matrix, and the third channel is filled with the broadcast-expanded fluctuation coefficient matrix, and finally a three-dimensional feature tensor is generated. Through tensor fusion, the static features and dynamic stability indicators are temporally and spatially associated in a unified high-dimensional space, so that the subsequent mutual information calculation can synchronously capture the eigenvalue distribution law and the stability pattern across time periods.

[0063] Step S5: Calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal label in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen the feature channels above the dynamic screening threshold to obtain a high-value feature set.

[0064] Specifically, in step S5, the construction of the dynamic screening threshold includes the following steps:

[0065] The mean and standard deviation of the mutual information values ​​of all feature channels are calculated, and the value of the mean plus three times the standard deviation is calculated to obtain the first candidate threshold;

[0066] Calculate the ratio of the number of abnormal labels in the pre-stored benchmark label set to the total number of labels in the benchmark label set, and take the natural logarithm of the reciprocal of the ratio to obtain the second candidate threshold;

[0067] The larger value between the first candidate threshold and the second candidate threshold is taken as the dynamic screening threshold.

[0068] During the specific implementation process, the pre-stored benchmark label set is first loaded. The benchmark label set is a multi-dimensional label library that contains clearly marked abnormal events and normal events through the accumulation of historical data. Its time span must cover at least three typical business cycles and the total sample volume must be no less than 100,000 records to ensure that the distribution of abnormal labels can reflect the statistical laws of real scenarios; then, for each feature channel of the three-dimensional feature tensor, that is, the data slice representing a single feature dimension in the tensor, the mutual information value calculation is performed: a histogram distribution is formed by the continuous values ​​of the discretized feature channel, and the joint probability density is calculated in conjunction with the abnormal label distribution in the benchmark label set. Finally, the statistical correlation strength between the feature channel and the abnormal event is quantified based on the Shannon entropy formula; after completing the calculation of the mutual information values ​​of all channels, the dynamic screening threshold construction stage is entered: first, the arithmetic average of all mutual information values ​​is calculated The first candidate threshold is obtained by adding three times the standard deviation to the mean value, which can cover 99.7% of the regular feature distribution interval. At the same time, the ratio of the number of abnormal labels in the benchmark label set to the total number of labels is counted, and the reciprocal natural logarithm of the ratio is taken as the second candidate threshold. Finally, the larger value of the first candidate threshold and the second candidate threshold is selected as the dynamic screening threshold. The dual threshold mechanism not only captures high-discrimination features by counting outliers, but also adapts to low-frequency and high-risk events through the inverse function of abnormal probability. In the screening stage, feature channels with mutual information values ​​exceeding the dynamic screening threshold are included in the high-value feature set, and invalid channels with mutual information values ​​below 0.01 are forcibly eliminated to avoid noise interference. The generated feature set establishes a mapping relationship with the original feature dimension through hash index, providing a noise-resistant input source for the subsequent decision tree model.

[0069] Step S6: Input the high-value feature set into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

[0070] Specifically, in step S6, inputting the high-value feature set into the pre-trained gradient boosting decision tree model further includes:

[0071] Multiply the characteristic fluctuation coefficient and the mutual information value to generate the weighted confidence;

[0072] Reconstruct the input order of high-value feature sets in descending order of weighted confidence.

[0073] Specifically, in step S6, a depth attenuation factor is added to the split gain calculation of the pre-trained gradient boosting decision tree model.

[0074] Specifically, in step S6, the step of obtaining an abnormal probability value and a feature heat map by dynamically weighting the feature contribution by the feature fluctuation coefficient vector includes:

[0075] A weight factor is generated through the feature fluctuation coefficient vector. The decision tree node split evaluation value is corrected in real time based on the weight factor to reconstruct the decision path. The abnormal probability value is output based on the terminal node distribution of the reconstructed decision path, and the path split contribution is accumulated to form a feature heat map.

[0076] In the specific implementation process, the high-value feature set is first pre-processed and strengthened. The stability parameter of each feature is extracted based on the feature fluctuation coefficient vector and multiplied with the mutual information value of the corresponding feature to generate a weighted confidence score. Then, the input sequence of the high-value feature set is reconstructed in descending order of the score, so that features with high stability and strong correlation are processed by the model first; then the reconstructed feature sequence is input into the pre-trained gradient boosting decision tree model. The gradient boosting decision tree model is constructed through historical multi-source data training. The training process includes three stages. In the first stage, the standardized anomaly detection data set of the past six months is selected as the basic sample. In the second stage, the standardized anomaly detection data set of the past six months is selected as the basic sample. The approximate second-order derivative acceleration algorithm of Levler expansion iteratively optimizes the splitting points of the decision tree. In the third stage, the learning rate and tree complexity are automatically adjusted based on the accuracy of the validation set after each round of improvement calculation until the model converges to the steady-state accuracy threshold; in the model execution stage, a dual optimization mechanism is implemented simultaneously during the node splitting process: on the one hand, a depth attenuation factor is introduced in the splitting gain calculation, and the gain weight decreases exponentially with each additional layer of splitting depth to suppress overfitting. On the other hand, the feature fluctuation coefficient vector is read in real time. The stability parameter is indexed according to the current feature identifier, and the weight factor is generated through negative correlation mapping. For example, when the fluctuation coefficient is 0.3, the weight factor is approximately 1 / 0.3. Equal to 3.33, multiply the weight factor by the benchmark split evaluation value. For example, the benchmark split evaluation value is such as the Gini index gain or the information entropy reduction, and a weighted split decision is constructed to replace the original split rule to reconstruct the entire decision path; finally, the abnormal probability value is calculated based on the terminal node distribution of the reconstructed decision path in the output layer, and the stabilized contribution value of each feature is accumulated backtracking along the weighted split path. The shallower the depth, the higher the contribution weight of the node, and a visual feature heat map is generated through normalized mapping. The abnormal probability value output by the entire process can be used as a real-time risk quantification indicator to directly drive the automatic response strategy. The feature heat map uses a coloring mapping mechanism to map the key root features. The physical location of the anomaly in the system topology is highlighted, and the structured anomaly detection report integrates the original data fragments, probability value evolution curves, thermal root cause distribution maps and historical similar case handling suggestions to form a closed-loop decision chain, so that operation and maintenance personnel can not only confirm the abnormal status in real time, but also accurately locate the location of the faulty equipment or the attack entry point, and generate executable operation instructions in a short time with reference to historical policy templates, and finally realize a fully automatic closed loop from anomaly perception, root cause diagnosis to handling decision in a cloud computing environment in seconds. The whole process reduces the decision interference rate of unstable features and improves the thermal significance of key root cause features through the feature stability perception mechanism.

[0077] The working principle of the cloud computing-based big data anomaly detection method provided by the present invention is as follows:

[0078] The present invention uses timestamp alignment and an adaptive anomaly filtering mechanism based on the interquartile range to denoise and standardize the raw data stream to generate continuous standard data segments. On this basis, multi-dimensional statistical features are extracted and combined with the semantic features of the graph database to construct a dual feature matrix. The coefficient of variation is calculated for the statistical feature matrix and the standard deviation is calculated for the semantic feature matrix to form a joint fluctuation coefficient vector to represent feature stability. The feature matrix and the fluctuation coefficient vector are then fused into a three-dimensional spatiotemporal tensor through zero padding and broadcast expansion, and a high-value feature set is locked through mutual information dynamic threshold screening. When inputting the pre-trained gradient boosting decision tree model, the weighted confidence generated by the product of the fluctuation coefficient and the mutual information is used to optimize the feature input sequence. Dual decision intervention is implemented during the model inference process. In the node splitting stage, a deep attenuation factor is simultaneously applied to suppress overfitting, and a dynamic weight factor based on the fluctuation coefficient is used to reconstruct the splitting path. Finally, the anomaly probability is calculated based on the terminal node distribution of the weighted decision path to quantify the risk level. At the same time, the depth-corrected contribution value of each feature in the weighted splitting point is back-traced to generate a root cause heat map. The two jointly drive the automated operation and maintenance strategy, realizing a full-link closed loop from data preprocessing, feature stability quantification, model anti-disturbance decision-making to interpretable result output.

[0079] Example 2

[0080] A cloud computing-based big data anomaly detection system is applied to a cloud computing-based big data anomaly detection method. In the specific implementation process, Figure 2 As shown, it shows a module structure diagram of a big data anomaly detection system based on cloud computing, including:

[0081] The working principle of the cloud computing-based big data anomaly detection system provided by the present invention is as follows:

[0082] The data segmentation module is used to segment the pre-processed raw data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle;

[0083] The feature joint extraction module is used to extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment;

[0084] A fluctuation feature quantification module is used to combine the statistical features and semantic features of multiple consecutive standard data segments into a statistical feature matrix and a semantic feature matrix in chronological order, and calculate a feature fluctuation coefficient vector based on the statistical feature matrix and the semantic feature matrix;

[0085] Tensor fusion module, used to fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor;

[0086] The mutual information dynamic screening module is used to calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal labels in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen feature channels above the dynamic screening threshold to obtain a high-value feature set;

[0087] The anti-disturbance decision parsing module is used to input high-value feature sets into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

[0088] The working principle of the cloud computing-based big data anomaly detection system provided by the present invention is as follows:

[0089] The present invention generates standard data segments with pattern integrity from the pre-processed raw data stream according to the business cycle through the streaming data segmentation module 100, and the linkage feature joint extraction module 200 analyzes the multi-dimensional statistical features and the semantic features of the graph database to form a bimodal spatiotemporal matrix. The fluctuation feature quantization module 300 performs cross-window variation coefficient and standard deviation calculation on the matrix to construct a global feature fluctuation coefficient vector; the tensor fusion module 400 splices the feature matrix with zero filling and the fluctuation vector with broadcast expansion along the channel axis into a three-dimensional feature tensor, and the mutual information dynamic screening module 500 calculates the mutual information of each channel according to the three-dimensional feature tensor. The abnormal correlation degree is used to screen high-value feature sets through a dual-threshold selection mechanism based on statistical distribution and information theory risk construction; the anti-disturbance decision analysis module 600 reconstructs the split path decision rule of the gradient boosting decision tree with the feature fluctuation coefficient as the weight factor, combines deep attenuation to suppress the overfitting interference of deep nodes, and dynamically weights and strengthens the stable feature decision weights. Finally, the abnormal probability risk value is generated according to the weighted path terminal distribution, and the cumulative depth correction contribution is traced back along the split node to form a root cause heat map, realizing a full-link closed-loop processing framework from dynamic feature stability quantification, model anti-disturbance decision optimization to interpretable result output.

[0090] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0091] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program. The program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, magnetic disk storage, or magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0092] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

Claims

1. A big data anomaly detection method based on cloud computing, characterized in that: The anomaly detection method comprises the following steps: S1: Split the pre-processed raw data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle; S2: Extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment; S3: combining the statistical features and semantic features of multiple consecutive standard data segments into a statistical feature matrix and a semantic feature matrix in chronological order, and calculating a feature fluctuation coefficient vector based on the statistical feature matrix and the semantic feature matrix; S4: Fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor; S5: Calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal labels in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen feature channels above the dynamic screening threshold to obtain a high-value feature set; S6: Input the high-value feature set into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

2. The method for detecting anomalies in big data based on cloud computing according to claim 1, wherein: In step S1, the length of the sliding time window is greater than or equal to three complete business cycles.

3. The big data anomaly detection method based on cloud computing according to claim 2, characterized in that: The pre-processing of the original data stream in step S1 includes the following steps: Performing timestamp alignment on the original data stream to obtain a time-aligned data stream, and segmenting the time-aligned data stream according to a preset temporary segmentation window to obtain multiple temporary data segments; Calculate the upper and lower thresholds for each temporary data segment, and construct a dynamic boundary interval for each temporary data segment based on the lower and upper thresholds. The calculation formulas for the upper and lower thresholds are: ; in, is the upper threshold, is the third quartile of the temporary data segment, is the first quartile of the temporary data segment, is the empirical scaling factor, is the interquartile range, is the dynamic expansion coefficient, is the standard deviation of the temporary data segment; A dynamic boundary interval is constructed based on the lower and upper thresholds, and abnormal data points that exceed the dynamic boundary interval in each temporary data segment are removed. The temporary data segments after the abnormal data points are removed are merged in timestamp order to generate a standardized data stream.

4. The method for detecting anomalies in big data based on cloud computing according to claim 3, wherein: In step S3, the columns in the statistical feature matrix and the semantic feature matrix represent the values ​​of all standard data segments of a feature dimension.

5. The method for detecting anomalies in big data based on cloud computing according to claim 4, characterized in that: In step S3, the step of generating the characteristic fluctuation coefficient vector includes: The coefficient of variation is calculated for each column of the statistical feature matrix, and the standard deviation is calculated for each column of the semantic feature matrix. The coefficient of variation and the standard deviation are concatenated in dimensional order to form a feature fluctuation coefficient vector.

6. The method for detecting anomalies in big data based on cloud computing according to claim 5, characterized in that: In step S5, the construction of the dynamic screening threshold comprises the following steps: The mean and standard deviation of the mutual information values ​​of all feature channels are calculated, and the value of the mean plus three times the standard deviation is calculated to obtain the first candidate threshold; Calculate the ratio of the number of abnormal labels in the pre-stored benchmark label set to the total number of labels in the benchmark label set, and take the natural logarithm of the reciprocal of the ratio to obtain the second candidate threshold; The larger value between the first candidate threshold and the second candidate threshold is taken as the dynamic screening threshold.

7. The method for detecting anomalies in big data based on cloud computing according to claim 6, characterized in that: In step S6, inputting the high-value feature set into the pre-trained gradient boosting decision tree model further includes: Multiply the characteristic fluctuation coefficient and the mutual information value to generate the weighted confidence; Reconstruct the input order of high-value feature sets in descending order of weighted confidence.

8. The method for detecting anomalies in big data based on cloud computing according to claim 7, characterized in that: In step S6, a depth attenuation factor is added to the split gain calculation of the pre-trained gradient boosting decision tree model.

9. The method for detecting anomalies in big data based on cloud computing according to claim 8, characterized in that: In step S6, the step of obtaining an abnormal probability value and a feature heat map by dynamically weighting the feature contribution by the feature fluctuation coefficient vector includes: A weight factor is generated through the feature fluctuation coefficient vector. The decision tree node split evaluation value is corrected in real time based on the weight factor to reconstruct the decision path. The abnormal probability value is output based on the terminal node distribution of the reconstructed decision path, and the path split contribution is accumulated to form a feature heat map.

10. A big data anomaly detection system based on cloud computing, characterized in that: The method for detecting anomalies in big data based on cloud computing according to any one of claims 1 to 9, wherein the anomaly detection system comprises: The data segmentation module is used to segment the pre-processed raw data stream into multiple continuous standard data segments according to the sliding time window determined by the business cycle; The feature joint extraction module is used to extract the multidimensional statistical features and entity identification fields of each standard data segment, and retrieve the pre-built graph database based on the entity identification fields to obtain the semantic features of the preset dimensions of each standard data segment; A fluctuation feature quantification module is used to combine the statistical features and semantic features of multiple consecutive standard data segments into a statistical feature matrix and a semantic feature matrix in chronological order, and calculate a feature fluctuation coefficient vector based on the statistical feature matrix and the semantic feature matrix; Tensor fusion module, used to fuse the statistical feature matrix, semantic feature matrix and feature fluctuation coefficient vector to generate a three-dimensional feature tensor; The mutual information dynamic screening module is used to calculate the mutual information value between each feature channel in the three-dimensional feature tensor and the abnormal labels in the pre-stored benchmark label set, construct a dynamic screening threshold based on the statistical characteristics of the mutual information values ​​of all feature channels, and screen feature channels above the dynamic screening threshold to obtain a high-value feature set; The anti-disturbance decision parsing module is used to input high-value feature sets into the pre-trained gradient boosting decision tree model, dynamically weight the feature contribution through the feature fluctuation coefficient vector to obtain the anomaly probability value and feature heat map, and output the anomaly detection report based on the anomaly probability value and feature heat map.

Citation Information

Patent Citations

  • External data extraction method for retrieval enhancement generation system

    CN119271706A

  • Abnormal short message behavior detection method and system based on multi-dimensional feature fusion

    CN120238869A

  • Abnormality detection method and system for intelligent motor

    CN120408475A

  • Auditing data early warning method and system

    CN120410210A

  • Detecting anomalous sensor data

    US20180268264A1

Cited By

  • Heterogeneous network modeling analysis method based on link entropy

    CN121530857A

  • Traffic data exception type determination method, electronic equipment and storage medium

    CN122093190A