Abnormal data identification method and device, equipment and storage medium
By extracting the data distribution characteristics of high-dimensional time-series data streams and adjusting the dimensionality reduction of the random projection matrix, and building a multi-resolution hash table based on local volatility to identify abnormal data points, it solves the problems of high-dimensional data processing in the existing technology and the low accuracy caused by distribution drift, and achieves efficient abnormal data recognition.
Patent Information
- Application Number
- CN202510630936.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art faces the problem of high-dimensional timing data processing in real-time monitoring of semiconductor manufacturing equipment. Traditional dimensionality reduction methods have poor adaptability to dynamic data, resulting in the loss of key fault features. The detection algorithm with fixed threshold cannot cope with the distribution drift caused by process switching, equipment aging, etc., resulting in low accuracy of abnormal data recognition.
The data distribution characteristics within each time window in the high-dimensional time-series data stream are extracted, and the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream are generated to generate low-dimensional data after dimensionality reduction. The local sensitive hash resolution parameters are configured according to the local fluctuations of the low-dimensional data, and a multi-resolution hash table is constructed, and abnormal data points are identified based on the data density in the bucket of the multi-resolution hash table.
By extracting data distribution characteristics and adjusting the dimensionality reduction process, the calculation complexity is reduced while retaining key information, and a multi-resolution hash table is constructed based on the local volatility of the dimensionality reduction data to realize adaptive analysis of different data areas, which can effectively deal with dimensional difficulties and distribution changes, and improve processing efficiency while ensuring detection accuracy.
Smart Images

Figure CN120145286A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data recognition, and particularly to an abnormal data recognition method, device, equipment and storage medium. Background Art
[0002] In the big data management of smart parks, the real-time monitoring of semiconductor manufacturing equipment faces the problem of processing high-dimensional time-series data. Taking the etcher in a wafer fab as an example, when the equipment runs, sensor data streams in hundreds of dimensions such as temperature, air pressure, and radio frequency power are generated. These data not only contain periodic process fluctuations but also transient anomalies caused by equipment failures. Existing anomaly detection technologies have two major defects: one is that traditional dimensionality reduction methods have poor adaptability to dynamic data, resulting in the loss of key fault features; the other is that the detection algorithm with a fixed threshold cannot cope with distribution drifts caused by process switching, equipment aging, etc. These defects lead to low accuracy in identifying abnormal data. Summary of the Invention
[0003] The main purpose of this application is to provide an abnormal data recognition method, device, equipment and storage medium, aiming to solve the technical problem of low accuracy in abnormal recognition caused by high data dimensions and dynamic distribution changes in the prior art.
[0004] To achieve the above object, this application proposes an abnormal data recognition method, and the method includes: Extract the data distribution characteristics within each time window in the high-dimensional time-series data stream; Generate the reduced-dimensional low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, where the adjusted random projection matrix is a matrix jointly optimized by the data distribution characteristics and the initial random projection matrix of the high-dimensional time-series data stream; Configure the local sensitive hashing resolution parameter according to the local volatility of the low-dimensional data, and construct a multi-resolution hash table; Identify abnormal data points according to the data density within the buckets of the multi-resolution hash table.
[0005] In one embodiment, the step of identifying abnormal data points according to the data density within the buckets of the multi-resolution hash table includes: Determine the data density within each hash bucket in the multi-resolution hash table; Adjust the initial density determination threshold according to the resolution parameter of the multi-resolution hash table to obtain the target density determination threshold; Mark abnormal hash buckets by comparing the data density within the buckets with the target density determination threshold; Perform secondary verification on the data points within the abnormal hash buckets to obtain a verification result, and determine abnormal data points in combination with the verification result and the density distribution characteristics of adjacent hash buckets.
[0006] In one embodiment, the step of adjusting an initial density determination threshold according to a resolution parameter of a multi-resolution hash table to obtain a target density determination threshold includes: Determine a base density threshold according to the resolution parameter of the multi-resolution hash table; Obtain the aggregated distribution features of multiple time windows in the high-dimensional time-series data stream, and determine a density correction coefficient based on the aggregated distribution features; Multiply the base density threshold by the density correction coefficient to obtain the target density determination threshold.
[0007] In one embodiment, the step of obtaining the aggregated distribution features of multiple time windows in the high-dimensional time-series data stream and determining a density correction coefficient based on the aggregated distribution features includes: Obtain the aggregated distribution features of multiple time windows in the high-dimensional time-series data stream, where the aggregated distribution features include a statistical characteristic matrix of data in each dimension and a principal component evolution trajectory; Construct a spatio-temporal association graph model based on the aggregated distribution features. The spatio-temporal association graph model is a model that uses the data distribution features of each time window as nodes and represents the distribution similarity between windows through edge weights. Through the spatio-temporal association graph model, identify the stable evolution stage and the mutation transition stage of the high-dimensional time-series data stream, and output a distribution evolution stage identifier. Determine the density correction coefficient according to the distribution evolution stage identifier.
[0008] In one embodiment, the step of extracting the data distribution features in each time window of the high-dimensional time-series data stream includes: Perform multi-scale window segmentation processing on the input high-dimensional time-series data stream to obtain window analysis results, where the window analysis results include at least one of short-period windows with a fixed duration, event windows automatically divided according to data mutation points, and long-period windows containing historical data; Generate short-period statistical features by determining the statistical quantities of data in each dimension within each short-period window; Generate data event features by performing waveform analysis on each event window; Perform smoothing processing and trend decomposition on the data in the long-period window to generate long-term trend features; Based on the short-period statistical features, the data event features, and the long-term trend features, obtain the data distribution features in each time window of the high-dimensional time-series data stream.
[0009] In one embodiment, the step of generating low-dimensional data after dimensionality reduction based on the data distribution features and the adjusted random projection matrix of the high-dimensional time-series data stream includes: Generate an initial random projection matrix based on the high-dimensional time-series data stream, and analyze the variance contribution of each dimension of data in the principal component direction from the data distribution characteristics to generate a feature importance vector; Determine the sparsification weight of the initial random projection matrix based on the correlation between the feature importance vector and the column vectors of the random projection matrix; Perform structured sparsification processing on the initial random projection matrix based on the sparsification weight to obtain an adjusted random projection matrix; Perform a multiplication operation on the high-dimensional time-series data stream and the adjusted random projection matrix to generate reduced-dimensional low-dimensional data.
[0010] In one embodiment, the step of configuring the local sensitive hashing resolution parameter according to the local volatility of the low-dimensional data and constructing a multi-resolution hash table includes: Analyze the local volatility of each data point in the low-dimensional data within a preset spatio-temporal neighborhood to obtain a volatility eigenvalue; Based on a preset partitioning threshold and the volatility eigenvalue, partition the preset spatio-temporal neighborhood to obtain a data space partition; Configure corresponding local sensitive hashing resolution parameters for the data space partition, where the preset partitioning threshold is a value determined by clustering historical data; Construct a multi-resolution hash table according to the data space partition and the local sensitive hashing resolution parameter.
[0011] In addition, to achieve the above object, the present application also proposes an abnormal data identification device, where the abnormal data identification device includes: A feature extraction module for extracting the data distribution characteristics within each time window in the high-dimensional time-series data stream; A data dimensionality reduction module for generating reduced-dimensional low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, where the adjusted random projection matrix is a matrix jointly optimized by the data distribution characteristics and the initial random projection matrix of the high-dimensional time-series data stream; A parameter adjustment module for configuring local sensitive hashing resolution parameters according to the local volatility of the low-dimensional data and constructing a multi-resolution hash table; An anomaly identification module for identifying abnormal data points according to the data density within the buckets of the multi-resolution hash table.
[0012] In addition, to achieve the above object, the present application also proposes an abnormal data identification device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the abnormal data identification method as described above.
[0013] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the abnormal data recognition method described above are implemented.
[0014] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the abnormal data recognition method described above are implemented.
[0015] The technical solution proposed by the present application extracts the data distribution characteristics within each time window of the high-dimensional time-series data stream, generates the dimension-reduced low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, configures the local sensitive hashing resolution parameters according to the local volatility of the low-dimensional data, constructs a multi-resolution hash table, and identifies abnormal data points according to the data density in the buckets of the multi-resolution hash table. The present application adjusts the dimension reduction process by extracting data distribution characteristics, reduces the computational complexity while retaining key information, constructs a multi-resolution hash table based on the local volatility of the dimension-reduced data, realizes the adaptive analysis of different data regions, and identifies abnormal data points according to the data density in the buckets, which can effectively cope with the dimension problem and the distribution change problem, and improves the processing efficiency while ensuring the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the abnormal data recognition method of the present application; Figure 2 It is a schematic flowchart provided for Embodiment 2 of the abnormal data recognition method of the present application; Figure 3 It is a schematic flowchart provided for Embodiment 3 of the abnormal data recognition method of the present application; Figure 4 It is a schematic block diagram of the module structure of the abnormal data recognition device according to the embodiment of the present application; Figure 5It is a schematic diagram of the device structure of the hardware operating environment involved in the abnormal data recognition method in the embodiments of the present application.
[0019] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0021] To better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.
[0022] There are two major defects in the existing anomaly detection technologies: one is that traditional dimensionality reduction methods have poor adaptability to dynamic data, resulting in the loss of key fault features; the other is that the detection algorithms with fixed thresholds cannot cope with the distribution drift caused by process switching, equipment aging, etc. These defects lead to low accuracy in identifying abnormal data.
[0023] Therefore, to overcome the above defects, the present application provides a solution. By extracting the data distribution characteristics to adjust the dimensionality reduction process, while retaining the key information, the computational complexity is reduced. A multi-resolution hash table is constructed based on the local volatility of the dimensionality-reduced data to achieve adaptive analysis of different data regions. Abnormal data points are identified according to the data density in the bucket, which can effectively cope with the dimensionality problem and distribution change problem, and improve the processing efficiency while ensuring the detection accuracy.
[0024] It should be noted that the execution subject of each embodiment of the present application can be a computing service system with data processing, network communication and program running functions, such as an electronic system, an abnormal data recognition system, etc. that can implement the above functions. The following takes the abnormal data recognition system as an example (hereinafter referred to as the "system") to illustrate the following embodiments.
[0025] Based on this, the embodiments of the present application provide an abnormal data recognition method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the abnormal data recognition method of the present application.
[0026] In this embodiment, the abnormal data recognition method includes steps S10 to S40: Step S10, extract the data distribution characteristics in each time window of the high-dimensional time-series data stream.
[0027] In the fields of industrial Internet of Things and intelligent monitoring, anomaly detection of high-dimensional time-series data streams faces three core challenges: First, there is a significant curse of dimensionality in high-dimensional data from multi-source sensors (such as 300+-dimensional process parameters in semiconductor equipment monitoring), and traditional dimensionality reduction methods will cause the loss of key fault features; Second, the data distribution has time-varying characteristics, and factors such as process switching and equipment aging lead to continuous drift of statistical characteristics; Third, anomaly patterns present multi-scale features, including both millisecond-level sudden anomalies and progressive degradations that last for several months.
[0028] The anomaly data recognition method of this application constructs an adaptive dynamic detection system by innovatively integrating the improved Random Projection (RP) and Locality Sensitive Hashing (LSH) technologies. Compared with traditional technologies, this application dynamically adjusts the sparsification weight of the random projection matrix based on the data distribution characteristics, and intelligently assigns the importance of dimensions based on the contribution degree of the principal components, improving the key feature retention rate compared with the fixed sparse mode, and effectively solving the problem of key feature loss caused by the curse of dimensionality; Based on the adaptive multi-resolution hashing mechanism of the local volatility of the dimensionality-reduced data, by quantifying the spatio-temporal neighborhood volatility index and automatically configuring the differential resolution parameters to replace the fixed resolution design of traditional LSH, the F1 value of the sudden anomaly detection (i.e., the harmonic mean of precision and recall) is improved, and it also meets the detection requirements of the multi-scale features of the anomaly pattern; Introducing a spatio-temporal correlation graph model to achieve the dynamic evolution of the density correction coefficient, by analyzing the change law of the data distribution similarity between windows, it overcomes the problem of the decline in detection accuracy caused by the time-varying characteristics of the data distribution.
[0029] It can be understood that this step first extracts features from the high-dimensional time-series data stream that can reflect the data distribution characteristics. High-dimensional time-series data streams usually contain dynamically changing data in multiple dimensions (such as sensor monitoring, financial transactions, etc.), and directly processing high-dimensional data has high computational complexity and is difficult to capture effective information. Before extracting features, the original data can be preprocessed first. For example, the time-series data can be segmented by a sliding window, or the event window can be dynamically divided by using a change point detection algorithm (such as the CumulativeSum (CUSUM) sequential analysis method, Bayesian change point detection method) to adapt to the non-stationarity of the data stream. Among them, the high-dimensional time-series data stream refers to the continuous time-series data composed of multiple monitoring dimensions (such as temperature, pressure, vibration, etc.), and the time window is a technical means to divide the infinite data stream into finite analysis units.
[0030] For each time window, the extraction of data distribution characteristics may include statistics (mean, variance, skewness), frequency-domain characteristics (which can be the main frequency components extracted through Fourier transform or wavelet analysis), or time-series model parameters (such as the coefficients of the Autoregressive Integrated Moving Average Model (ARIMA)). Further, multi-scale analysis can be introduced, combining short-term statistical characteristics (such as extreme values within a sliding window), waveform characteristics of event windows (such as peak duration), and long-term trend characteristics (which can be decomposed through Hodrick-Prescott (HP) filtering), to comprehensively characterize the spatio-temporal characteristics of data distribution. This step solves the problem of non-robust feature expression caused by the curse of dimensionality and noise interference in high-dimensional time-series data.
[0031] Step S20: Generate the low-dimensional data after dimensionality reduction based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, where the adjusted random projection matrix is a matrix jointly optimized by the data distribution characteristics and the initial random projection matrix of the high-dimensional time-series data stream.
[0032] Traditional random projection maps high-dimensional data to a low-dimensional space through a random matrix, but does not consider the distribution characteristics of the data itself, which may lead to the loss of key feature information. The innovation of this step lies in optimizing the initial random projection matrix through data distribution characteristics.
[0033] In specific implementation, an initial random projection matrix (an initial random Gaussian matrix or a sparse matrix) can be first generated, and then the initial random projection matrix is structurally adjusted in combination with the data distribution characteristics. For example, the column vectors of the initial matrix are weighted using a feature importance vector (calculated through variance contribution) to enhance the projection weight of high-contribution dimensions, or redundant dimension interference is eliminated through sparsification processing (such as L1 regularization). The adjusted random projection matrix can retain the main distribution pattern of the data, so as to better reflect the discriminative features of the original data in the low-dimensional data after dimensionality reduction, solve the problem of decreased abnormal sensitivity caused by traditional dimensionality reduction methods ignoring data distribution, and at the same time improve the calculation efficiency through matrix optimization.
[0034] Step S30: Configure the local sensitive hashing resolution parameter according to the local volatility of the low-dimensional data, and construct a multi-resolution hash table.
[0035] Locality-Sensitive Hashing (LSH) usually adopts fixed resolution parameters and is difficult to adapt to the dynamic fluctuation characteristics of time-series data. In this step, by analyzing the local volatility of the data after dimensionality reduction (such as calculating the standard deviation or entropy value within the neighborhood of data points), multi-resolution hashing parameters are dynamically configured. The construction of the multi-resolution hash table can be achieved through a family of parallel hash functions, where each function corresponds to different resolution parameters, thus accommodating data distributions with different granularities in the same hash structure. This step solves the problem of abnormal missed detection or false detection caused by uneven data density in fixed-resolution hashing, and at the same time improves the adaptability of the hash table to dynamic data through multi-resolution design.
[0036] Step S40: Identify abnormal data points according to the data density within the buckets of the multi-resolution hash table.
[0037] It should be understood that this step realizes the accurate positioning of abnormal points through density comparison. First, the density of each hash bucket is statistically calculated (such as based on nearest neighbor counting or kernel density estimation), and the density determination threshold is dynamically adjusted based on the resolution parameters of the hash table. For example, the threshold for high-resolution buckets is lower to capture sparse anomalies. Subsequently, the abnormal buckets are marked by comparing the bucket density with the threshold, and the data points within the buckets are secondarily verified, and then a list of abnormal points with comprehensive verification results is output. The density hierarchical determination through multi-resolution hashing in this step solves the problem of missed detection or false detection caused by the global density threshold in traditional methods, and is especially suitable for the detection of local anomalies in high-dimensional time-series data.
[0038] In this embodiment, the data distribution characteristics within each time window of the high-dimensional time-series data stream are extracted, and based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, the reduced-dimensional low-dimensional data is generated. According to the local volatility of the low-dimensional data, the locality-sensitive hashing resolution parameters are configured, a multi-resolution hash table is constructed, and abnormal data points are identified according to the data density within the buckets of the multi-resolution hash table. This application adjusts the dimensionality reduction process by extracting data distribution characteristics, reduces the computational complexity while retaining key information, constructs a multi-resolution hash table based on the local volatility of the reduced-dimensional data, realizes the adaptive analysis of different data regions, and identifies abnormal data points according to the data density within the buckets, which can effectively cope with the dimensionality problem and distribution change problem, and improves the processing efficiency while ensuring the detection accuracy.
[0039] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 ..., and the said step S40 may include steps S401 to S404: Step S401: Determine the data density within each hash bucket in the multi-resolution hash table.
[0040] It can be understood that the core of this step lies in quantifying the data aggregation degree of each hash bucket to provide a density benchmark for anomaly detection. In specific implementation, a density estimation algorithm based on neighbor statistics (such as Kernel Density Estimation (KDE) or simple density calculation based on counting) can be adopted, and normalization processing is combined with the spatial volume of the hash bucket to avoid biases caused by differences in bucket sizes under different resolutions. For dynamic time series data, a time decay factor (such as exponentially weighted moving average) can also be introduced to adjust the influence weight of historical data on the current density, so as to more accurately reflect the real-time changes in data distribution.
[0041] Step S402: Adjust the initial density determination threshold according to the resolution parameter of the multi-resolution hash table to obtain the target density determination threshold.
[0042] It should be noted that the initial density determination threshold can be obtained through training with historical normal data (such as based on the 3σ principle or percentile statistics), and the adjustment process needs to consider the resolution parameter of the hash bucket (such as bucket width or hash function parameter). In specific implementation, a resolution-threshold mapping model (such as a logarithmic linear relationship or a regression model based on machine learning) can be established, so that the determination threshold in the high-resolution area is automatically reduced to capture subtle anomalies, while the threshold in the low-resolution area is increased to reduce false alarms. A sliding window mechanism can also be introduced to dynamically update the mapping relationship according to the characteristics of recent data streams.
[0043] As an implementation manner, step S402 in this embodiment may include: determining a basic density threshold according to the resolution parameter of the multi-resolution hash table; obtaining the aggregated distribution characteristics of multiple time windows in the high-dimensional time series data stream, and determining a density correction coefficient based on the aggregated distribution characteristics; multiplying the basic density threshold by the density correction coefficient to obtain the target density determination threshold.
[0044] It can be understood that in order to obtain the target density determination threshold, it is first necessary to establish a quantitative relationship between the hash resolution and the density benchmark. In specific implementation, a threshold generation algorithm driven by resolution parameters can be adopted, where the basic density threshold is directly related to the geometric characteristics of the hash bucket - for the high-resolution area (small bucket volume), a lower basic threshold can be derived through a sparse data model based on the Poisson process; while for the low-resolution area (large bucket volume), a Gaussian mixture model is used to fit the normal data distribution to determine a higher threshold. A resolution-threshold response curve (such as an exponential decay function) can be introduced to adaptively calibrate the parameters by fitting the optimal detection thresholds at different resolution levels in historical data. This method breaks through the limitation of traditional fixed thresholds and solves the problem of density incomparability caused by differences in bucket volumes in a multi-resolution environment.
[0045] Furthermore, the scenario adaptability of the threshold is achieved through spatio-temporal feature fusion. In the feature extraction stage, a sliding window mechanism is adopted to aggregate the three-dimensional feature tensors of the recent time window (including statistical moments, principal component load matrices, and time series autocorrelation coefficients), and the spatio-temporal attention network is used to model the feature evolution pattern. The generation of the density correction coefficient can combine two mechanisms: 1) hard correction based on feature drift detection, and when a distribution mutation is detected, the sharp adjustment coefficient is calculated through the Kullback-Leibler Divergence (KLD); 2) soft correction based on neural differential equations, and the coefficient is fine-tuned by modeling the gradual change of the feature through a continuous-time recurrent neural network.
[0046] Finally, the precise calibration of the detection threshold is achieved through composite calculation, that is, multiplying the basic density threshold by the density correction coefficient to obtain the target density determination threshold. In the multiplication fusion process, logarithmic space operations can be introduced to maintain the stability of thresholds of different orders of magnitude, and dynamic boundary constraints (such as restricting the correction coefficient to the preset interval of [0.5, 2.0] through the sigmoid function) are added to prevent over-adjustment. For time series data with both periodicity and trend, the seasonal factors obtained by the Fourier basis function decomposition can be additionally superimposed for periodic modulation.
[0047] As an implementation manner, the steps of obtaining the aggregated distribution features of multiple time windows in the high-dimensional time series data stream and determining the density correction coefficient based on the aggregated distribution features may include: obtaining the aggregated distribution features of multiple time windows in the high-dimensional time series data stream, where the aggregated distribution features include the statistical property matrix and the principal component evolution trajectory of each dimension data; constructing a spatio-temporal correlation graph model based on the aggregated distribution features, where the spatio-temporal correlation graph model is a model that takes the data distribution features of each time window as nodes and represents the distribution similarity between windows through edge weights; identifying the stable evolution stage and the mutation transition stage of the high-dimensional time series data stream through the spatio-temporal correlation graph model, and outputting the distribution evolution stage identifier; and determining the density correction coefficient according to the distribution evolution stage identifier.
[0048] It can be understood that first, the features of continuous time windows need to be aligned and fused: for the statistical property matrix, the tensor splicing method can be used to retain the high-order statistics such as the mean and variance of each window, and the instantaneous fluctuation is eliminated through moving average filtering; for the principal component evolution trajectory, after aligning the principal components analysis (PCA) components of different windows through dynamic time warping (DTW), the time series features such as the trajectory curvature and the direction change rate are extracted.
[0049] When constructing a spatio-temporal correlation graph model based on aggregated distribution features, the data distribution features of each time window are transformed into graph nodes, and the distribution similarity between windows is calculated as the edge weight. In this process, the model can be further enabled to adaptively capture the dependencies between different time windows by combining a time attention mechanism and a graph neural network. Specifically, a graph attention neural network can be used to calculate the dynamic similarity between nodes, and time series embedding (such as Transformer encoding) can be combined to enhance the expressive ability of temporal features. In addition, the edge weight calculation can be optimized to multi-modal similarity fusion. For example, statistical distance and shape similarity can be combined to more comprehensively characterize the spatio-temporal evolution pattern of data distribution.
[0050] Next, the spatio-temporal correlation graph model is used to analyze the evolution pattern of the data stream, identify the stages of stable data evolution and the transitional stages of mutations, and output the corresponding distribution evolution stage identifiers. Finally, the density correction coefficient is dynamically adjusted according to these stage identifiers, where the stable stage corresponds to a higher coefficient to reduce the detection sensitivity, and the mutation stage corresponds to a lower coefficient to improve the abnormal capture ability.
[0051] Step S403: Mark the abnormal hash buckets by comparing the data density in the buckets with the target density determination threshold.
[0052] It can be understood that in the density-threshold comparison stage, a multi-level determination strategy can be adopted: First, a hard threshold comparison is performed to mark the obvious abnormal buckets, and then a soft determination based on statistical tests (such as t-test (Student's t test) or (Mann-Whitney U test)) is performed on the boundary buckets (densities close to the threshold). For the temporal scenario, temporal consistency verification can be superimposed (such as requiring that N consecutive windows be marked to be confirmed as abnormal) to improve the robustness. The marking results can form an abnormal probability heat map, where the abnormal probability of each bucket is determined by the degree of its density deviation from the threshold. The dual filtering mechanism that combines statistical tests and temporal verification effectively distinguishes real anomalies from instantaneous noise.
[0053] Step S404: Perform secondary verification on the data points in the abnormal hash buckets to obtain the verification results, and determine the abnormal data points by combining the verification results and the density distribution features of adjacent hash buckets.
[0054] This step realizes the accurate positioning of abnormal points and the elimination of false alarms. Among them, two types of technologies can be used in parallel for secondary verification: 1) Model-based verification. For example, the reconstruction error is calculated using a pre-trained autoencoder; 2) Density-based local verification. For example, the density ratio of the target point to its nearest neighbor in the adjacent bucket is compared. Then, by integrating the verification results of the secondary verification and the density gradient features of the adjacent buckets (such as modeling the density propagation relationship between buckets through a graph neural network), special attention is paid to those data points with high self-verification scores and located at the density mutation boundary, which can effectively solve the problem of misjudgment of boundary points caused by complex local density changes in high-dimensional data.
[0055] In this embodiment, by automatically adjusting the decision threshold according to the resolution parameter of the hash table, the detection sensitivity is ensured. Through the secondary verification of abnormal hash buckets and by combining the density distribution characteristics of adjacent buckets, the false alarm rate is reduced. Moreover, by combining the basic density threshold based on the resolution parameter and the correction coefficient of the aggregated distribution characteristics, the dynamic optimization of the abnormal decision threshold is realized. By quantifying the distribution similarity between windows through a spatio-temporal correlation graph model and identifying the stable evolution stage and the mutation transition stage, the density correction coefficient is dynamically adjusted, enabling the system to predict the data change trend in the pre-stage and adjust the detection strategy in advance, significantly improving the response speed to sudden anomalies.
[0056] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , step S10 may include steps S101 to S105: Step S101, perform multi-scale window segmentation processing on the input high-dimensional time-series data stream to obtain window analysis results, where the window analysis results include at least one of short-period windows with a fixed duration, event windows automatically divided according to data mutation points, and long-period windows containing historical data.
[0057] It can be understood that this step adopts a three-level window division strategy to achieve multi-scale feature capture. That is, for short-period windows with a fixed duration, an equally spaced sliding window is used to ensure the basic analysis granularity; event windows are dynamically generated through change point detection based on the CUSUM algorithm to ensure that key data change events are completely included; long-period windows are constructed through an exponentially weighted historical data backtracking mechanism to focus on capturing the macro trend. In addition, a window overlap compensation algorithm can be introduced to automatically adjust the segmentation boundary when the event window overlaps with the fixed window to ensure the data integrity of each scale window. For example, in industrial vibration monitoring, this method can simultaneously capture the instantaneous abnormal vibration of the bearing (event window), the working condition fluctuations per minute (short-period window), and the equipment aging trend (long-period window).
[0058] Exemplarily, taking the monitoring of a semiconductor foundry etching equipment as an example, the system intelligently selects a window combination strategy according to actual requirements: in the conventional process monitoring mode, a short-period window with a 1-second interval is used to capture the instantaneous fluctuations of the plasma, an event window triggered by mutations is used to completely record the abnormal discharge process, and a long-period window of 200 process cycles is used to track the electrode loss; in the equipment health assessment mode, only the long-period window is used to analyze the aging trend; and in the process debugging stage, the short-period window and the event window are combined to capture real-time anomalies.
[0059] Step S102, by determining the statistical quantities of each dimension within each short-period window, short-period statistical features are generated.
[0060] For each short-period window, multi-dimensional statistical quantities including the time domain (mean, variance, kurtosis) and the frequency domain (amplitude of the main frequency of FFT, spectral entropy) are calculated. In particular, for the correlation characteristics between dimensions of high-dimensional data, the mutual information matrix of dimension pairs within the window is calculated to supplement the traditional statistical quantities. Taking chemical process monitoring as an example, not only the mean and variance of the readings of each sensor are recorded, but also the change in mutual information of key dimension pairs such as temperature-pressure is calculated to form a 20-dimensional short-period feature vector.
[0061] Step S103, by performing waveform analysis on each event window, data event features are generated.
[0062] The feature extraction of the event window focuses on the waveform morphological characteristics: first, the data within the window is normalized, and then morphological indicators including the number of waveform extreme points, zero-crossing rate, approximate entropy, etc. are extracted, and data event features at different scales are obtained through continuous wavelet transform.
[0063] Step S104, the data of the long-period window is smoothed and trend-decomposed to generate long-term trend features.
[0064] The STL (Seasonal-Trend decomposition using Loess) (a technique for time series decomposition) is used to decompose the data of the long-period window, separating the trend term, the periodic term, and the residual term. The innovation lies in constructing the derivative features of the trend term, and quantifying the trend change rate and acceleration by calculating the first-order / second-order derivatives of each point on the trend line. Taking the temperature monitoring of a server cluster as an example, not only the 24-hour average temperature trend is extracted, but more importantly, the acceleration feature of the trend change is calculated, and a warning signal of an abnormal increase in the trend curvature can be detected 3 days before the failure of the equipment cooling system.
[0065] Step S105, based on the short-period statistical features, the data event features, and the long-term trend features, the data distribution features within each time window in the high-dimensional time series data stream are obtained.
[0066] Multi-scale feature fusion is achieved through feature cascading and attention weighting. Short-period statistical features (such as 20 dimensions), event features, and long-term trend features are concatenated into the data distribution features within each time window of the high-dimensional time series data stream. The finally output distribution features can include original statistics, cross-scale correlation metrics (such as the covariance between short-period variance and long-term trend slope), and attention weight values.
[0067] As an implementation manner, the step S20 may include: generating an initial random projection matrix according to the high-dimensional time series data stream, and parsing the variance contribution degrees of each dimension data in the principal component direction from the data distribution features to generate a feature importance vector; determining the sparsification weight of the initial random projection matrix based on the correlation between the feature importance vector and the column vectors of the random projection matrix; performing structured sparsification processing on the initial random projection matrix based on the sparsification weight to obtain an adjusted random projection matrix; and performing a multiplication operation on the high-dimensional time series data stream and the adjusted random projection matrix to generate the low-dimensional data after dimensionality reduction.
[0068] It can be understood that the initial random projection matrix is generated using a Gaussian random distribution (with dimensions d×k, where d is the original dimension and k is the target dimension), which conforms to the theoretical guarantee of the Johnson-Lindenstrauss lemma. At the same time, the principal component analysis results are extracted from the data distribution features, and the sum of the squared loadings of each original dimension on the Top-k (i.e., selecting the first k most important principal components) principal components is calculated as the variance contribution degree to form a feature importance vector. In addition, incremental PCA can be used to process streaming data and dynamically update the principal component direction. For non-linear feature relationships, kernel PCA or the feature importance in the latent space extracted by an autoencoder can be introduced. The correlation between the feature importance vector and the column vectors of the random projection matrix can be calculating the cosine similarity between the column vectors of the initial matrix and the feature importance vector, and then generating the sparsification weight of the initial random projection matrix. Among them, for high-importance dimensions (such as cosine similarity > 0.7), the complete projection relationship is retained, medium importance (such as cosine similarity between 0.3 and 0.7) is probabilistically sparsified, and low importance (such as cosine similarity < 0.3) is forced to zero.
[0069] The structured sparsification processing and dimensionality reduction calculation process, that is, using an improved approximate Gram-Schmidt orthogonalization process, under the premise of maintaining the approximate orthogonality of the matrix, zeroing the matrix elements according to the sparsification weight, and retaining the cross-projection relationship between important dimensions (such as the joint projection of temperature and pressure sensors), thereby obtaining the adjusted random projection matrix. Then, the high-dimensional time series data stream is multiplied by the adjusted random projection matrix to generate the low-dimensional data after dimensionality reduction.
[0070] As an implementation manner, the step S30 may include: analyzing the local volatility of each data point in the low-dimensional data within a preset spatio-temporal neighborhood to obtain a volatility eigenvalue; based on a preset partitioning threshold and the volatility eigenvalue, partitioning the preset spatio-temporal neighborhood to obtain a data space partition; configuring a corresponding locality-sensitive hashing resolution parameter for the data space partition, where the preset partitioning threshold is a value determined by clustering historical data; and constructing a multi-resolution hash table according to the data space partition and the locality-sensitive hashing resolution parameter.
[0071] It can be understood that the adaptive multi-resolution hash construction is achieved by analyzing the local volatility characteristics of data, and its core lies in establishing an intelligent mapping relationship between the dynamic characteristics of data and the hash resolution. Specifically, first, for each data point in the low-dimensional data, the local volatility eigenvalue is calculated within its spatio-temporal neighborhood (such as the set of points within 3 time points before and after + the Euclidean space radius r). This value is a weighted combination of the standard deviation, approximate entropy, and trend slope, and can quantify the stability degree of the data in this region. Based on the volatility threshold obtained by clustering historical data (such as the volatility threshold obtained by the Density-Based Spatial Clustering of Applications with Noise (DBSCAN)), the data space is divided into three types of partitions: the first partition, the second partition, and the third partition (high volatility area, transition area, and stable area). The first partition corresponds to the area where data mutations occur frequently (such as the period when industrial equipment fails), and the third partition corresponds to the normal operating state.
[0072] Configure the locality-sensitive hashing resolution parameter for each partition: the first partition (high volatility area) uses a fine resolution (small hash bucket width) to capture subtle anomalies, the third partition (stable area) uses a coarse-grained resolution (large bucket width) to improve the calculation efficiency, and the second partition (transition area) takes the intermediate value. The finally constructed multi-resolution hash table adopts a hierarchical storage structure, and different resolution regions are quickly located through a spatial index.
[0073] For ease of understanding, an example is given as follows, but it does not limit the abnormal data recognition method of the present application. In the intelligent park monitoring scenario of semiconductor manufacturing equipment, the abnormal data recognition method of the present application can effectively solve the problem of equipment status monitoring in the chip production process. Taking a wafer etching machine as an example, when the equipment is running, it will generate a sensor data stream containing hundreds of dimensions such as temperature, air pressure, and radio frequency power. These data not only contain periodic fluctuations of process parameters but may also suddenly have transient mutations caused by equipment failures or process anomalies.
[0074] The system first performs multi-scale window segmentation on the original data stream: a 5-second short-period window is used to capture the instantaneous fluctuations of the plasma state, and the process abnormality event window is dynamically divided through Bayesian change point detection. At the same time, a long window containing 200 production cycles is set to analyze the aging trend of equipment.
[0075] In the feature extraction stage, the short-period window calculates the time-frequency domain characteristics of each sensor reading (such as the spectral entropy of the RF power), the event window analyzes the peak duration and rising slope of the abnormal discharge waveform, and the long-period window extracts the equipment performance attenuation curve through STL decomposition. After these features are reduced in dimension by a random projection matrix based on the optimization of the principal component contribution, the 50-dimensional low-dimensional data can still retain most of the key information. By analyzing the local volatility of the plasma state in the etching chamber, the system divides the data space into a stable process area (resolution parameter 0.1), a transition area (resolution parameter 0.05), and an abnormal discharge area (resolution parameter 0.01), and constructs a corresponding multi-resolution hash table.
[0076] In actual detection, when an etching machine experiences electrode aging, the system first captures the slow rise of electrode impedance in the long-period trend characteristics, and the trigger density correction coefficient is lowered to 0.8 to improve the detection sensitivity. Subsequently, abnormal RF power fluctuations are detected in the short-period window, and the multi-resolution hash table marks the abnormal hash bucket at the fine resolution level (0.01). After automatic encoder reconstruction error verification and plasma density gradient analysis of adjacent buckets, the electrode contact failure in the No. 3 reaction chamber is finally located. Compared with traditional fixed threshold detection, this method can advance the warning time of equipment failure and reduce the false alarm rate in actual measurements at a smart park of TSMC.
[0077] This embodiment adopts the method of multi-scale window segmentation and multi-dimensional feature fusion to fully capture the distribution characteristics of high-dimensional time series data. The sparse weight adjustment of the random projection matrix is driven by the data distribution characteristics to achieve adaptive optimization of the dimensionality reduction process. The sparse pattern is dynamically adjusted according to the contribution of each dimension in the direction of the principal component, so that the low-dimensional data after dimensionality reduction can retain key feature information to the greatest extent, while significantly reducing the computational complexity. The multi-resolution hash table construction method based on local volatility balances the detection accuracy and computational efficiency. By quantifying the local volatility characteristics of each data point, the data space is divided into different areas and differentiated resolution parameters are configured, so that the high-volatility area obtains fine detection capabilities and the low-volatility area maintains efficient screening, which improves the real-time and accuracy of the system as a whole.
[0078] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the abnormal data identification method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0079] The present application also provides an abnormal data recognition device. Please refer to Figure 4 , the abnormal data recognition device includes: A feature extraction module 10, configured to extract data distribution features within each time window of a high-dimensional time-series data stream; A data dimensionality reduction module 20, configured to generate reduced-dimensional low-dimensional data based on the data distribution features and an adjusted random projection matrix of the high-dimensional time-series data stream, where the adjusted random projection matrix is a matrix jointly optimized by the data distribution features and an initial random projection matrix of the high-dimensional time-series data stream; A parameter adjustment module 30, configured to configure local sensitive hashing resolution parameters according to the local volatility of the low-dimensional data and construct a multi-resolution hash table; An abnormality recognition module 40, configured to recognize abnormal data points according to the data density within the buckets of the multi-resolution hash table.
[0080] The abnormal data recognition device provided by the present application adopts the abnormal data recognition method in the above embodiment, and can solve the technical problem of low abnormal recognition accuracy caused by high data dimensions and dynamic distribution changes in the prior art. Compared with the prior art, the beneficial effects of the abnormal data recognition device provided by the present application are the same as those of the abnormal data recognition method provided by the above embodiment, and other technical features in the abnormal data recognition device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0081] The present application provides an abnormal data recognition device. The abnormal data recognition device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the abnormal data recognition method in the first embodiment above.
[0082] Next, refer to Figure 5 , which shows a schematic structural diagram of an abnormal data recognition device suitable for implementing the embodiments of the present application. The abnormal data recognition device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Desctions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The abnormal data recognition device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0083] As shown Figure 5 in the figure, the abnormal data recognition device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the abnormal data recognition device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the abnormal data recognition device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an abnormal data recognition device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0084] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0085] The abnormal data recognition device provided by the present application adopts the abnormal data recognition method in the above-mentioned embodiment, and can solve the technical problem of low abnormal recognition accuracy caused by high data dimensions and dynamic changes in distribution in the prior art. Compared with the prior art, the beneficial effects of the abnormal data recognition device provided by the present application are the same as those of the abnormal data recognition method provided by the above-mentioned embodiment, and other technical features in the abnormal data recognition device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0086] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0087] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0088] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the abnormal data recognition method in the above embodiments.
[0089] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0090] The above computer-readable storage medium can be included in the abnormal data recognition device; it can also exist separately without being assembled into the abnormal data recognition device.
[0091] The above computer-readable storage medium carries one or more programs, which when executed by the abnormal data recognition device, cause the abnormal data recognition device to: extract the data distribution characteristics within each time window in the high-dimensional time-series data stream, generate the reduced-dimensional low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time-series data stream, configure the locality-sensitive hashing resolution parameter according to the local volatility of the low-dimensional data, construct a multi-resolution hash table, and identify abnormal data points according to the data density within the buckets of the multi-resolution hash table.
[0092] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, executed entirely on a remote computer or server, or executed on an ARM (Advanced RISC Machines) development board. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0094] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0095] The readable storage medium provided by the present application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned abnormal data recognition method, and can solve the technical problem of low abnormal recognition accuracy caused by high data dimension and dynamic change of distribution in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the abnormal data recognition method provided by the above embodiments, and will not be elaborated here.
[0096] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the abnormal data recognition method as described above are implemented.
[0097] The computer program product provided by the present application can solve the technical problem of low abnormal recognition accuracy caused by high data dimension and dynamic change of distribution in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the abnormal data recognition method provided by the above embodiments, and will not be elaborated here.
[0098] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the description and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A method for identifying abnormal data, characterized in that: The method comprises the following steps: Extract data distribution features within each time window in high-dimensional time series data stream; Based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time series data stream, generating low-dimensional data after dimensionality reduction, wherein the adjusted random projection matrix is a matrix obtained by jointly optimizing the data distribution characteristics and the initial random projection matrix of the high-dimensional time series data stream; According to the local volatility of low-dimensional data, the local sensitive hashing resolution parameters are configured to construct a multi-resolution hash table. Identify abnormal data points based on the data density in the bucket of the multi-resolution hash table.
2. The abnormal data identification method according to claim 1, characterized in that: The step of identifying abnormal data points according to the data density in the bucket of the multi-resolution hash table includes: Determine the data density within each hash bucket of the multi-resolution hash table; Adjusting the initial density determination threshold according to the resolution parameter of the multi-resolution hash table to obtain the target density determination threshold; By comparing the data density in the bucket with the target density determination threshold, marking an abnormal hash bucket; A secondary verification is performed on the data points in the abnormal hash bucket to obtain a verification result, and the abnormal data point is determined in combination with the verification result and the density distribution characteristics of adjacent hash buckets.
3. The abnormal data identification method according to claim 2, characterized in that: The step of adjusting the initial density determination threshold according to the resolution parameter of the multi-resolution hash table to obtain the target density determination threshold comprises: Determining a base density threshold according to a resolution parameter of the multi-resolution hash table; Obtaining aggregate distribution features of multiple time windows in a high-dimensional time series data stream, and determining a density correction coefficient based on the aggregate distribution features; The basic density threshold is multiplied by the density correction coefficient to obtain a target density determination threshold.
4. The abnormal data identification method according to claim 3, characterized in that: The step of obtaining aggregate distribution features of multiple time windows in the high-dimensional time series data stream and determining a density correction coefficient based on the aggregate distribution features includes: Obtaining aggregate distribution features of multiple time windows in a high-dimensional time series data stream, wherein the aggregate distribution features include a statistical characteristic matrix of each dimensional data and a principal component evolution trajectory; A spatiotemporal association graph model is constructed based on the aggregated distribution characteristics. The spatiotemporal association graph model is a model that uses the data distribution characteristics of each time window as nodes and represents the distribution similarity between windows through edge weights. The stable evolution stage and the sudden transition stage of the high-dimensional time series data stream are identified through the spatiotemporal association graph model, and the distribution evolution stage identifier is output; the density correction coefficient is determined according to the distribution evolution stage identifier.
5. The abnormal data identification method according to any one of claims 1 to 4, characterized in that: The step of extracting data distribution features in each time window in the high-dimensional time series data stream includes: Perform multi-scale window segmentation processing on the input high-dimensional time series data stream to obtain a window analysis result, wherein the window analysis result includes at least one of a short-period window with a fixed length, an event window automatically divided according to a data mutation point, and a long-period window containing historical data; By determining the statistics of each dimension of data in each short-period window, short-period statistical features are generated; By performing waveform analysis on each event window, data event features are generated; Smoothing and trend decomposition of data in long-period windows to generate long-term trend characteristics; Based on the short-term statistical characteristics, the data event characteristics and the long-term trend characteristics, the data distribution characteristics in each time window in the high-dimensional time series data stream are obtained.
6. The abnormal data identification method according to any one of claims 1 to 4, characterized in that: The step of generating the reduced-dimensional low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time series data stream comprises: Generate an initial random projection matrix according to the high-dimensional time series data stream, and analyze the variance contribution of each dimensional data in the direction of the principal component from the data distribution characteristics to generate a feature importance vector; Determining a sparsification weight of an initial random projection matrix based on the feature importance vector and a column vector correlation of the random projection matrix; Performing structured sparse processing on the initial random projection matrix based on the sparse weights to obtain an adjusted random projection matrix; The high-dimensional time series data stream and the adjusted random projection matrix are multiplied to generate low-dimensional data after dimensionality reduction.
7. The abnormal data identification method according to any one of claims 1 to 4, characterized in that: The step of configuring the local sensitive hash resolution parameter according to the local volatility of the low-dimensional data and constructing a multi-resolution hash table includes: Analyze the local volatility of each data point in the low-dimensional data within a preset space-time neighborhood to obtain the volatility characteristic value; Based on a preset partition threshold and the fluctuation characteristic value, obtaining data space partitions by partitioning the preset spatiotemporal neighborhood; Configuring a local sensitive hashing resolution parameter corresponding to the data space partition, wherein the preset partition threshold is a value determined by clustering historical data; A multi-resolution hash table is constructed according to the data space partitions and the locality sensitive hash resolution parameters.
8. An abnormal data identification device, characterized in that: The abnormal data identification device comprises: Feature extraction module, used to extract data distribution features in each time window in high-dimensional time series data stream; A data dimension reduction module, used for generating reduced low-dimensional data based on the data distribution characteristics and the adjusted random projection matrix of the high-dimensional time series data stream, wherein the adjusted random projection matrix is a matrix obtained by jointly optimizing the data distribution characteristics and the initial random projection matrix of the high-dimensional time series data stream; A parameter adjustment module is used to configure the local sensitive hashing resolution parameters according to the local volatility of low-dimensional data and construct a multi-resolution hash table; The anomaly identification module is used to identify abnormal data points according to the data density in the bucket of the multi-resolution hash table.
9. An abnormal data identification device, characterized in that: The abnormal data identification device comprises: a memory, a processor, and an abnormal data identification program stored in the memory and executable on the processor, wherein the abnormal data identification program implements the abnormal data identification method according to any one of claims 1 to 7 when executed by the processor.
10. A storage medium, characterized in that: The storage medium stores an abnormal data identification program, and when the abnormal data identification program is executed by the processor, the abnormal data identification method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Network traffic abnormality rapid detection method based on multilayer partial sensitive hash table
CN107070867A
Data flow abnormity identification method based on random projection angle distribution
CN110311879A
Safety detection time series data real-time anomaly discovery method and electronic device
CN111694860A
Machine learning-based data analyses for outlier detection
US11537942B1
Cited By
Network equipment monitoring method, system and equipment based on industrial internet of things, and medium
CN121309401A
Network equipment monitoring method, system, device and medium based on industrial internet of things
CN121309401B