Hardware simulation platform environment anomaly detection method based on dynamic hierarchical clustering
Through the combination of dynamic hierarchical clustering and constraint-driven mechanisms, the adaptability and real-time problems of abnormal detection of hardware simulation platform are solved, and high-precision abnormal recognition and stability improvement are achieved.
Patent Information
- Application Number
- CN202510434005.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing hardware simulation platform anomaly detection methods have limitations in adaptability, real-timeness and detection accuracy. Fixed thresholds and statistical analysis methods are difficult to process high-dimensional dynamic data. Traditional clustering methods cannot adapt to real-time changes in the simulation environment. Deep learning methods are highly dependent on data and have high computational overhead.
Using a method based on dynamic hierarchical clustering, the clustering hierarchy model is constructed and the constraint-driven mechanism is combined with the clustering hierarchy and classification standards are updated in real time, and anomaly judgment is used using dynamic weights and Mahayana distances to make exception score models to distinguish between minor and severe anomalies.
It improves the accuracy and stability of abnormal detection, can adapt to changes in the hardware simulation environment, accurately identify low-frequency abnormalities and complex failure modes, and reduces computing overhead.
Smart Images

Figure CN120277584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware simulation platforms, and particularly to a method for detecting anomalies in the hardware simulation platform environment based on dynamic hierarchical clustering. Background Art
[0002] With the development of modern hardware simulation technology, hardware simulation platforms are widely used in the fields of integrated circuit design verification, electronic system testing, and embedded system development. Hardware simulation platforms can simulate the system operation state in a real environment and provide efficient debugging means for designers. However, with the expansion of the hardware simulation scale and the increase in complexity, the problem of anomaly detection in the simulation environment has gradually become a research hotspot. A large amount of multi-dimensional data, including voltage, current, signal timing, logical state, and power consumption data, will be generated during the operation of the hardware simulation platform. Abnormal changes in these data will reflect system errors, hardware failures, or instability of the operating environment.
[0003] Currently, the anomaly detection techniques for hardware simulation platforms mainly include detection means based on threshold judgment, statistical analysis, and traditional clustering methods. However, there are still many deficiencies in practical applications. The detection method based on a fixed threshold usually relies on expert experience to set the anomaly determination criteria. Although it is simple and easy to implement, in the face of a complex and changing simulation environment, the fixed threshold is difficult to adapt to different types of data fluctuations and is prone to false positives and false negatives. The statistical analysis method performs anomaly detection by calculating the mean, variance, and other statistical characteristics of the data distribution, but it has poor adaptability to high-dimensional data and is difficult to effectively capture the complex correlation relationships between data. In addition, although traditional clustering methods can, to a certain extent, mine the structural characteristics of data, they usually cannot achieve adaptive adjustment, and the clustering model is prone to failure in the face of the continuously changing real-time data in the hardware simulation environment, resulting in unstable detection results.
[0004] In recent years, anomaly detection methods based on machine learning have gradually attracted attention. For example, deep learning, autoencoders, and isolation forest techniques have been introduced into the field of hardware anomaly detection, which can better learn the potential distribution patterns of data and improve the accuracy of anomaly recognition. However, deep learning methods usually require a large amount of labeled anomaly data for training, and the anomaly data in the hardware simulation environment is often difficult to obtain, resulting in insufficient model training. In addition, deep learning methods have a large computational volume and high requirements for hardware resources, making it difficult to meet the requirements of efficient real-time detection.
[0005] In summary, the existing anomaly detection methods for hardware simulation platforms still have significant limitations in terms of adaptability, real-time performance, and detection accuracy. The existing fixed threshold and statistical analysis methods are difficult to handle high-dimensional dynamic data, and the existing clustering methods cannot adapt to the real-time changes in the simulation environment. Deep learning methods are highly dependent on data and have a large computational overhead. Therefore, there is an urgent need for an anomaly detection method that can adapt to the characteristics of the hardware simulation environment, balance real-time performance and high accuracy, in order to improve the stability and fault identification ability of the hardware simulation platform. Summary of the Invention
[0006] An object of the present invention is to propose an anomaly detection method for the hardware simulation platform environment based on dynamic hierarchical clustering. The present invention can effectively distinguish minor anomalies from severe anomalies and improve the accuracy of anomaly detection.
[0007] An anomaly detection method for the hardware simulation platform environment based on dynamic hierarchical clustering according to an embodiment of the present invention includes the following steps:
[0008] S1. Multidimensional data generated during the operation of the hardware simulation platform are collected in real time, and the multidimensional data are denoised, normalized, and feature-extracted to form a standardized feature vector data stream.
[0009] S2. A multidimensional feature vector set reflecting the state of the hardware simulation platform environment is constructed based on the standardized feature vector data stream.
[0010] S3. Using the multidimensional feature vector set, a hierarchical clustering algorithm is employed to establish an initial hierarchical clustering model, generating an initial clustering center and an initial hierarchical division structure.
[0011] S4. During the continuous operation of the hardware simulation platform, a newly generated standardized feature vector data stream is received in real time, and the newly generated standardized feature vector data stream is compared with the initial hierarchical clustering model. An adaptive hierarchical adjustment operation is performed based on the data distribution change and the feature vector difference to update the hierarchical structure, clustering center, and classification criteria of the initial hierarchical clustering model in real time.
[0012] S5. Based on the hierarchically clustering model updated in real time, the deviation degree between each feature vector and its affiliated clustering center is calculated, and an anomaly judgment is made on the feature vectors whose deviation degree exceeds a preset threshold to realize the identification of anomaly data in the hardware simulation platform environment.
[0013] S6. An anomaly alarm signal is generated for the identified anomaly data.
[0014] Optionally, the S1 includes the following steps:
[0015] S11. Collect multi-dimensional data generated during the operation of the hardware simulation platform. The multi-dimensional data includes voltage, current, signal timing, logical state, and power consumption data, and construct a time series multi-dimensional data set:
[0016] D = {d1, d2,..., d N};
[0017] where D is the collected multi-dimensional data set, d i represents the data point at the i-th time step, and N is the number of time steps collected;
[0018] S12. Denoise the multi-dimensional data set D, suppress signal noise, remove high-frequency interference signals, and reduce transient signal fluctuations through moving average filtering to generate a denoised multi-dimensional data set;
[0019] S13. Normalize the denoised multi-dimensional data set, map all data to the [0, 1] interval, and calculate data in different dimensions on the same scale to obtain a normalized multi-dimensional data set;
[0020] S14. Extract features from the normalized multi-dimensional data set, and generate a standardized feature vector data stream F according to the data type by selecting time domain, frequency domain, and statistical features:
[0021] F = {f1, f2,..., f N};
[0022] where f i represents the feature vector at the i-th time step.
[0023] Optionally, S2 includes the following steps:
[0024] S21. Generate a multi-dimensional feature vector set X of the hardware simulation platform environment based on the standardized feature vector data stream F:
[0025] X = {x1, x2,..., x N};
[0026] where X is the multi-dimensional feature vector set, x i represents the i-th feature vector, and N is the total number of feature vectors;
[0027] S22. Calculate the mean vector μ X and covariance matrix Σ X of the feature vector set X according to the distribution characteristics of the feature vectors, which are used to characterize the statistical features of the hardware simulation platform environment state;
[0028] S23. Optimize the data indexing of the feature vector set X and establish a feature index structure I X :
[0029] I X = {I(x1), I(x2),..., I(x N )};
[0030] Among them, I(x i ) represents the storage location of the feature vector x i in the feature index structure to accelerate subsequent clustering calculations and anomaly detection;
[0031] S24. According to the statistical characteristics of the multi-dimensional feature vector set X, construct the feature space S of the hardware simulation platform environment state X , and define:
[0032] S X = (μ X , Σ X , I X );
[0033] Among them, S X represents the feature space calculated based on the feature vector set X, including the mean vector, covariance matrix, and feature index structure.
[0034] Optionally, the S3 includes the following steps:
[0035] S31. Based on the multi-dimensional feature vector set X, use the dynamic weighted Euclidean distance metric to calculate the distance matrix D of the multi-dimensional feature vector set X X :
[0036]
[0037] Among them, x ik represents the value of x i in the k-th dimension, m is the dimension of the feature vector, and ω k (t) is the dynamic weight:
[0038]
[0039] Among them, σ k (t) is the standard deviation of the k-th dimension data within the current time t, and ε is a small positive number to prevent division by zero;
[0040] S32. Introduce a constraint-driven mechanism, combine the known fault modes and normal modes of the hardware simulation platform to construct the must-merge constraint set ML and the prohibited-merge constraint set CL, and define the constraint penalty function Δ(x i , x j ):
[0041]
[0042] Among them, λML With λ CL and λ ML are the penalty coefficients for the must - merge and forbidden - merge constraints respectively, and η CL and η i are the corresponding scaling parameters, ‖x j - x i ‖ represents the Euclidean distance between x j and x
[0043] S33. Combine the dynamic distance matrix D X with the constraint penalty function Δ(x i , x j ) to construct a constraint - driven clustering merge cost function C(i, j):
[0044] C(i, j) = D X (i, j)+Δ(x i , x j )
[0045] where C(i, j) represents the comprehensive cost generated when merging x i and x j ;
[0046] S34. Perform bottom - up hierarchical clustering based on the constraint - driven merge cost function C(i, j). Update the cost between clustering clusters during each merge operation, and use the dynamic threshold optimization method to determine the optimal hierarchical threshold. By selecting the optimal threshold that distinguishes the normal state from the abnormal state, reasonably partition the clustering clusters, construct the hierarchical clustering tree T and form the initial clustering cluster set.
[0047] Optionally, S34 includes the following steps:
[0048] S341. Based on the constraint - driven merge cost function C(i, j) and the dynamic weighted Euclidean distance metric, construct a comprehensive merge cost function that incorporates the within - cluster dispersion information for selecting the optimal merge pair. For any two clustering clusters C i and C j , define their comprehensive merge cost as:
[0049]
[0050] where and represent the variances of the internal feature vectors of clustering clusters C i and C j respectively, reflecting the dispersion degree of the data within each cluster. δ is a small positive number to prevent the denominator from being zero. Subsequently, select the optimal merge cluster pair to satisfy the following formula and make the selected cluster pair optimal in terms of the constraint - driven merge cost, reflecting the dynamic characteristics of the hardware simulation platform environment data:
[0051]
[0052] Among them, M t represents the set of all cluster clusters that are still in an independent state at time step t;
[0053] S342. After the selected cluster pair (C p , C q ) is merged, the center μ r of the new merged cluster C r is calculated using the exponentially decaying weighted cluster center:
[0054]
[0055] Among them, μ p and μ q are the cluster centers of cluster C p and C q respectively, |C p | and |C q | represent the number of feature vectors in each cluster, and β is a positive parameter used to control the influence of variance on the weight;
[0056] The cluster with a more concentrated data distribution has a high influence weight after merging, thereby reflecting the normal operating state and abnormal differences of the hardware simulation platform;
[0057] S343. During the clustering iteration process, an optimal hierarchical threshold θ * is determined using a dynamic threshold optimization method based on the difference factor correction. The difference factor δ ij is defined for the dynamic adaptation of the data distribution characteristics at different merging stages:
[0058]
[0059] Among them, and are the average internal dispersions of cluster C i and C j respectively, and ζ is a constant to prevent the denominator from being zero;
[0060] The optimal threshold is defined by combining the merging cost and the inter-cluster dispersion difference to distinguish the normal and abnormal states of the hardware simulation platform:
[0061]
[0062] Among them, τ0 is a benchmark threshold preset from historical data;
[0063] S344. According to the determined optimal hierarchical threshold θ *Divide the set M of the final clustering clusters * , define the decision function φ(C i , θ * ):
[0064]
[0065] Then the set of the final clustering clusters is expressed as:
[0066] M * ={C i ∈M t |φ(C i , θ * ) = 1};
[0067] Among them, M t is the set of all unmerged or merged clusters in the clustering iteration process, and M * is the set of the final clusters that meet the condition that the inter-cluster merging cost is higher than the optimal threshold θ * , so that each clustering cluster can reflect the internal consistency of the hardware simulation platform environment data;
[0068] S345. Construct a hierarchical clustering tree T that records the evolution of the hierarchical structure of the hardware simulation platform environment data in the dynamic clustering process as the supporting data structure for the entire clustering process, and completely record the clustering iteration process and its evolution trajectory. Let the clustering tree T be expressed as the set of all merging events:
[0069]
[0070] Among them, E l represents the l-th merging event, and are the clustering clusters merged in the l-th merging, is the new clustering cluster formed after the merging, is the cost of this merging, and it satisfies ω 1 ≤ω 2 ≤…≤ω L .
[0071] Optionally, the S5 includes the following steps:
[0072] S51. According to the final hierarchical clustering model M * and the hierarchical clustering tree T, perform clustering mapping on the newly collected standardized feature vector data stream F′, and calculate the membership degree of each feature vector f′ i to its affiliated clustering cluster C k :
[0073]
[0074] Among them, α i,kDenote the feature vector f′ i The membership degree to the cluster C k , μ k is the cluster center of C k , d(f′ i , μ k ) is the Euclidean distance between f′ i and μ k , and λ is the membership degree scaling coefficient;
[0075] S52. Calculate the deviation degree of each feature vector f′ i in its affiliated cluster C k , and determine whether it belongs to an outlier. Define the deviation measure based on Mahalanobis distance:
[0076]
[0077] where D M (f′ i , C k ) represents the Mahalanobis distance of f′ i in C k , and Σ k is the covariance matrix of the cluster C k ;
[0078] S53. Calculate the outlier scores of all newly collected feature vectors f′ i to distinguish normal data from abnormal data. Define the outlier score A(f′ i ) as:
[0079]
[0080] where A(f′ i ) reflects the degree to which the feature vector f′ i deviates from its affiliated cluster C k , and γ is the outlier score scaling coefficient, such that Mahalanobis distances higher than the threshold correspond to higher outlier scores;
[0081] S54. Set the outlier determination threshold θ A to perform outlier identification on all newly collected feature vectors. Define the outlier judgment function:
[0082]
[0083] where φ(f′ i ) = 1 indicates that f′ i is abnormal data, and φ(f′ i ) = 0 indicates that f i is normal data;
[0084] S55. Construct an abnormal data set F according to the determination result abn And store the abnormal data:
[0085] F abn ={f′ i |φ(f′ i ) = 1};
[0086] Among them, F abn represents the detected abnormal data set.
[0087] The beneficial effects of the present invention are:
[0088] (1) The present invention adopts a dynamic hierarchical clustering method. By constructing a hierarchical clustering model and combining a constraint-driven mechanism, the clustering hierarchy structure and classification criteria are updated in real time in the hardware simulation platform environment. Existing clustering methods usually use a fixed clustering structure and cannot adapt to the changing data distribution in the hardware simulation environment, resulting in a decrease in the sensitivity of anomaly detection. Through adaptive hierarchical adjustment operations, combined with the changes in data distribution and the differences in feature vectors, the division of clustering clusters is continuously optimized during operation, enabling the model to dynamically adjust to adapt to the changes in the simulation platform state, thereby improving the accuracy and stability of anomaly detection.
[0089] (2) The present invention introduces a constraint-driven mechanism. During the hierarchical clustering process, a must-merge constraint set and a prohibited-merge constraint set are constructed by combining known fault modes and normal modes, and the merging cost between clustering clusters is adjusted through a constraint penalty function, enabling the effective distinction between normal and abnormal states during the clustering process. The existing fault knowledge can be used to strengthen the ability to distinguish abnormal patterns at the initial stage of clustering, avoiding recognition biases caused by the mixing of abnormal data and normal data. In addition, combined with the dynamic weight adjustment strategy, the feature weights of different dimensions can be adaptively optimized during the clustering process, thereby more accurately capturing the characteristics of abnormal data and improving the recognition ability for low-frequency anomalies and complex fault patterns.
[0090] (3) In the abnormal determination stage of the present invention, a deviation measurement method based on Mahalanobis distance is used to calculate the deviation degree of each feature vector relative to its belonging clustering cluster, and a refined abnormal determination is realized in combination with an abnormal score mechanism. Existing Euclidean distance or simple statistical threshold methods usually have difficulty accurately measuring the distribution of multi-dimensional feature vectors, resulting in insufficient recognition ability for boundary abnormal points. By measuring the distribution deviation degree of feature vectors in a high-dimensional space using Mahalanobis distance and constructing an abnormal score model in combination with an exponential scaling factor, the determination of abnormal data is made more robust, and it can effectively distinguish between mild anomalies and severe anomalies, improving the accuracy of anomaly detection. Description of the Drawings
[0091] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:
[0092] Figure 1 is a flowchart of a method for detecting anomalies in the hardware simulation platform environment based on dynamic hierarchical clustering proposed by the present invention. Specific embodiments
[0093] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0094] Reference Figure 1 , a method for detecting anomalies in the hardware simulation platform environment based on dynamic hierarchical clustering, includes the following steps:
[0095] S1. Real-time collect multi-dimensional data generated during the operation of the hardware simulation platform, and perform denoising, normalization, and feature extraction on the multi-dimensional data to form a standardized feature vector data stream;
[0096] S2. Construct a multi-dimensional feature vector set reflecting the state of the hardware simulation platform environment based on the standardized feature vector data stream;
[0097] S3. Use the multi-dimensional feature vector set and adopt a hierarchical clustering algorithm to establish an initial hierarchical clustering model, generating initial clustering centers and an initial hierarchical partitioning structure;
[0098] S4. During the continuous operation of the hardware simulation platform, real-time receive the newly generated standardized feature vector data stream, compare the new standardized feature vector data stream with the initial hierarchical clustering model, and perform adaptive hierarchical adjustment operations based on the data distribution changes and feature vector differences to update the hierarchical structure, clustering centers, and classification criteria of the initial hierarchical clustering model in real-time;
[0099] S5. Based on the hierarchically clustering model updated in real-time, calculate the deviation degree between each feature vector and its affiliated clustering center, and perform anomaly judgment on the feature vectors whose deviation degree exceeds the preset threshold to realize the identification of abnormal data in the hardware simulation platform environment;
[0100] S6. Generate an anomaly alarm signal for the identified abnormal data.
[0101] In this embodiment, S1 includes the following steps:
[0102] S11. Collect multi-dimensional data generated during the operation of the hardware simulation platform. The multi-dimensional data includes voltage, current, signal timing, logic state, and power consumption data, and construct a time-series multi-dimensional data set:
[0103] D = {d1, d2, ..., d N};
[0104] Wherein, D is the collected multi-dimensional data set, and d i represents the data point at the i-th time step, and N is the number of collected time steps;
[0105] S12. Denoise the multi-dimensional data set D, suppress the signal noise, remove the high-frequency interference signal, and reduce the transient signal fluctuation through moving average filtering to generate a denoised multi-dimensional data set;
[0106] S13. Normalize the denoised multi-dimensional data set, map all data to the interval [0, 1], and calculate the data of different dimensions on the same scale to obtain a normalized multi-dimensional data set;
[0107] S14. Extract features from the normalized multi-dimensional data set, and select time-domain, frequency-domain, and statistical features according to the data type to generate a standardized feature vector data stream F:
[0108] F = {f1, f2, ..., f N};
[0109] Wherein, f i represents the feature vector at the i-th time step.
[0110] In this embodiment, S2 includes the following steps:
[0111] S21. Generate a multi-dimensional feature vector set X of the hardware simulation platform environment according to the standardized feature vector data stream F:
[0112] X = {x1, x2, ..., x N};
[0113] Wherein, X is the multi-dimensional feature vector set, and x i represents the i-th feature vector, and N is the total number of feature vectors;
[0114] S22. Calculate the mean vector μ X and covariance matrix Σ X of the feature vector set X according to the distribution characteristics of the feature vectors, which are used to characterize the statistical features of the hardware simulation platform environment state;
[0115] S23. Optimize the data indexing of the feature vector set X and establish a feature index structure I X :
[0116] I X = {I(x1), I(x2), ..., I(x N)};
[0117] Among them, I(x i ) represents the storage location of the feature vector x i in the feature index structure to accelerate subsequent clustering calculations and anomaly detection;
[0118] S24. According to the statistical characteristics of the multi-dimensional feature vector set X, construct the feature space S of the hardware simulation platform environment state X , and define:
[0119] S X =(μ X , Σ X , I X );
[0120] Among them, S X represents the feature space calculated based on the feature vector set X, including the mean vector, covariance matrix, and feature index structure.
[0121] In this embodiment, S3 includes the following steps:
[0122] S31. Based on the multi-dimensional feature vector set X, adopt the dynamic weighted Euclidean distance metric to calculate the distance matrix D of the multi-dimensional feature vector set X X :
[0123]
[0124] Among them, x ik represents the value of x i in the k-th dimension, m is the dimension of the feature vector, ω k (t) is the dynamic weight:
[0125]
[0126] Among them, σ k (t) is the standard deviation of the k-th dimensional data within the current time t, and ε is a small positive number to prevent division by zero;
[0127] S32. Introduce a constraint-driven mechanism, combine the known fault modes and normal modes of the hardware simulation platform to construct the must-merge constraint set ML and the forbidden-merge constraint set CL, and define the constraint penalty function Δ(x i , x j ):
[0128]
[0129] Among them, λ ML and λ CL are the penalty coefficients for the must-merge and forbidden-merge constraints respectively, and η ML and ηCL is the corresponding scaling parameter, ‖x i - x j ‖ represents the Euclidean distance between x i and x j ;
[0130] S33. Combine the dynamic distance matrix D X with the constraint penalty function Δ(x i , x j ), and construct a constraint-driven clustering merging cost function C(i, j):
[0131] C(i, j) = D X (i, j)+Δ(x i , x j );
[0132] Among them, C(i, j) represents the comprehensive cost generated when merging x i and x j ;
[0133] S34. Perform bottom-up hierarchical clustering based on the constraint-driven merging cost function C(i, j). Update the cost between clustering clusters during each merging operation, and use the dynamic threshold optimization method to determine the optimal hierarchical threshold. By selecting the optimal threshold that distinguishes the normal state from the abnormal state, reasonably divide the clustering clusters, construct the hierarchical clustering tree T, and form the initial clustering cluster set.
[0134] In this embodiment, S34 includes the following steps:
[0135] S341. Based on the constraint-driven merging cost function C(i, j) and the dynamic weighted Euclidean distance metric, construct a comprehensive merging cost function that incorporates the within-cluster dispersion information, which is used to select the optimal merging pair. For any two clustering clusters C i and C j , define their comprehensive merging cost as:
[0136]
[0137] Among them, and respectively represent the variances of the internal feature vectors of clustering clusters C i and C j , reflecting the dispersion degree of the data within each cluster. δ is a small positive number to prevent the denominator from being zero. Subsequently, select the optimal merging cluster pair to satisfy the following formula and make the selected cluster pair optimal in terms of the constraint-driven merging cost, reflecting the dynamic characteristics of the hardware simulation platform environment data:
[0138]
[0139] Among them, Mt Denote the set of all clustering clusters that are still in an independent state at time step t;
[0140] S342. After the selected clustering cluster pair (C p , C q ) is merged, use the clustering center calculation based on exponential decay weighting to calculate the center μ r of the new clustering cluster C r :
[0141]
[0142] where μ p and μ q are the clustering centers of clustering clusters C p and C q respectively, |C p | and |C q | represent the number of feature vectors in each clustering cluster, and β is a positive parameter used to control the influence of variance on the weight;
[0143] Make the clustering clusters with more concentrated data distribution have high influence weights after merging, thereby reflecting the normal operation state and abnormal differences of the hardware simulation platform;
[0144] S343. During the clustering iteration process, use the dynamic threshold optimization method based on the difference factor correction to determine the optimal hierarchical threshold θ * , and define the difference factor δ ij for the dynamic adaptation of the data distribution characteristics in different merging stages:
[0145]
[0146] where, and are the average internal dispersions of clustering clusters C i and C j respectively, and ζ is a constant to prevent the denominator from being zero;
[0147] Define the optimal threshold, combine the merging cost and the difference in inter-cluster dispersion, and use it to distinguish the normal and abnormal states of the hardware simulation platform:
[0148]
[0149] where, τ0 is a benchmark threshold preset from historical data;
[0150] S344. According to the determined optimal hierarchical threshold θ * divide the final clustering cluster set M * , and define the decision function φ(C i , θ * ):
[0151]
[0152] Then the final cluster set is represented as:
[0153] M * ={C i ∈M t |φ(C i ,θ * ) = 1};
[0154] Among them, M t is the set of all unmerged or merged clusters during the clustering iteration process, and M * is the final cluster set that satisfies that the inter-cluster merging cost is higher than the optimal threshold θ * , so that each cluster can reflect the internal consistency of the hardware simulation platform environment data;
[0155] S345. Construct a hierarchical clustering tree T that records the hierarchical structure evolution of the hardware simulation platform environment data during the dynamic clustering process as the supporting data structure for the entire clustering process, and completely record the clustering iteration process and its evolution trajectory. Let the clustering tree T be represented as the set of all merge events:
[0156]
[0157] Among them, E l represents the l-th merge event, and are the clusters to be merged in the l-th merge, is the new cluster formed after the merge, is the cost of this merge, and satisfies ω 1 ≤ω 2 ≤...≤ω L .
[0158] In this embodiment, S5 includes the following steps:
[0159] S51. According to the final hierarchical clustering model M * and the hierarchical clustering tree T, perform clustering mapping on the newly collected standardized feature vector data stream F′, and calculate the membership degree of each feature vector f′ i to its affiliated cluster C k :
[0160]
[0161] Among them, α i,k represents the membership degree of the feature vector f′ i to the cluster C k , and μ k is Ck as the clustering center, d(f′ i , μ k ) is the Euclidean distance between f′ i and μ k , and λ is the membership scaling factor;
[0162] S52. Calculate the degree of deviation of each feature vector f′ i in its belonging clustering cluster C k , and determine whether it belongs to an outlier. Define the deviation measure based on Mahalanobis distance:
[0163]
[0164] where D M (f′ i , C k ) represents the Mahalanobis distance of f′ i in C k , and Σ k is the covariance matrix of the clustering cluster C k ;
[0165] S53. Calculate the outlier scores of all newly collected feature vectors f′ i to distinguish normal data from outlier data. Define the outlier score A(f′ i ) as:
[0166]
[0167] where A(f′ i ) reflects the degree of deviation of the feature vector f′ i from its belonging clustering cluster C k , and γ is the outlier score scaling factor, such that the Mahalanobis distance higher than the threshold corresponds to a higher outlier score;
[0168] S54. Set the outlier determination threshold θ A to perform outlier recognition on all newly collected feature vectors. Define the outlier judgment function:
[0169]
[0170] where φ(f′ i ) = 1 indicates that f′ i is outlier data, and φ(f′ i ) = 0 indicates that f′ i is normal data;
[0171] S55. According to the determination result, construct the outlier data set F abn and store the outlier data:
[0172] Fabn ={f′ i ∣φ(f′ i )=1};
[0173] Among them, F abn Represents a collection of detected anomaly data.
[0174] The present invention adopts a dynamic hierarchical clustering method. By constructing a hierarchical clustering model and combining it with a constraint-driven mechanism, the clustering hierarchy and classification standards are updated in real time in a hardware simulation platform environment. Existing clustering methods usually use a fixed clustering structure that cannot adapt to the ever-changing data distribution in the hardware simulation environment, resulting in a decrease in the sensitivity of anomaly detection. Through adaptive hierarchical adjustment operations, the division of clustering clusters is continuously optimized during operation in combination with data distribution changes and feature vector differences, so that the model can dynamically adjust to adapt to changes in the simulation platform state, thereby improving the accuracy and stability of anomaly detection.
[0175] The present invention introduces a constraint-driven mechanism, which combines known fault modes and normal modes to construct a must-merge constraint set and a prohibited-merge constraint set in the hierarchical clustering process, and adjusts the merging cost between clustering clusters through a constraint penalty function so that the normal state and the abnormal state can be effectively distinguished in the clustering process. The existing fault knowledge can be used in the early stage of clustering to enhance the ability to distinguish abnormal modes and avoid recognition bias caused by the mixing of abnormal data and normal data. In addition, combined with a dynamic weight adjustment strategy, the feature weights of different dimensions can be adaptively optimized in the clustering process, thereby more accurately capturing the characteristics of abnormal data and improving the recognition ability of low-frequency anomalies and complex fault modes.
[0176] In the anomaly determination stage, the present invention adopts a deviation measurement method based on Mahalanobis distance to calculate the degree of deviation of each feature vector relative to its corresponding cluster, and combines it with an anomaly score mechanism to realize refined anomaly determination. The existing Euclidean distance or simple statistical threshold method is usually difficult to accurately measure the distribution of multi-dimensional feature vectors, resulting in insufficient recognition of boundary anomalies. The distribution deviation of feature vectors in high-dimensional space is measured by Mahalanobis distance, and an anomaly score model is constructed in combination with an exponential scaling factor, so that the determination of abnormal data is more robust, and can effectively distinguish between minor anomalies and severe anomalies, thereby improving the accuracy of anomaly detection.
[0177] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for detecting anomalies in the hardware simulation platform environment based on dynamic hierarchical clustering, characterized in that, It includes the following steps: S1. Collect multi-dimensional data generated during the operation in real time from the hardware simulation platform, and perform denoising, normalization, and feature extraction on the multi-dimensional data to form a standardized feature vector data stream; S2. Construct a multi-dimensional feature vector set reflecting the environmental state of the hardware simulation platform based on the standardized feature vector data stream; S3. Use the multi-dimensional feature vector set to establish an initial hierarchical clustering model using the hierarchical clustering algorithm, generating initial cluster centers and an initial hierarchical partitioning structure; S4. During the continuous operation of the hardware simulation platform, receive the newly generated standardized feature vector data stream in real time, compare the new standardized feature vector data stream with the initial hierarchical clustering model, and perform adaptive hierarchical adjustment operations based on changes in data distribution and differences in feature vectors to update the hierarchical structure, cluster centers, and classification criteria of the initial hierarchical clustering model in real time; S5. Based on the hierarchically clustered model updated in real time, calculate the deviation degree between each feature vector and its affiliated cluster center, and perform anomaly judgment on the feature vectors whose deviation degree exceeds the preset threshold to realize the identification of abnormal data in the hardware simulation platform environment; S6. Generate an anomaly alarm signal for the identified abnormal data.
2. The method for detecting hardware simulation platform environment anomalies based on dynamic hierarchical clustering according to claim 1, wherein The S1 includes the following steps: S11. Collect multi-dimensional data generated during the operation of the hardware simulation platform, where the multi-dimensional data includes voltage, current, signal timing, logic state, and power consumption data, and construct a time-series multi-dimensional data set: D = {d1, d2,..., d N}; where D is the multi-dimensional data set collected, and d i represents the data point at the i-th time step, and N is the number of time steps collected; S12. Perform denoising processing on the multi-dimensional data set D, suppress signal noise, remove high-frequency interference signals, and reduce transient signal fluctuations through moving average filtering to generate a denoised multi-dimensional data set; S13. Perform normalization processing on the denoised multi-dimensional data set, map all data to the [0, 1] interval, and calculate data of different dimensions on the same scale to obtain a normalized multi-dimensional data set; S14. Perform feature extraction on the normalized multi-dimensional data set, and generate a standardized feature vector data stream F by selecting time-domain, frequency-domain, and statistical features according to the data type; F = {f1, f2,..., f N}; Among them, f i represents the feature vector at the i-th time step.
3. A method for detecting hardware simulation platform environment anomalies based on dynamic hierarchical clustering according to claim 1, characterized in that, The S2 includes the following steps: S21. Generate a multi-dimensional feature vector set X of the hardware simulation platform environment based on the standardized feature vector data stream F; X = {x1, x2,..., x N}; where X is a set of multi-dimensional feature vectors, and x i represents the i-th feature vector, and N is the total number of feature vectors; S22. Calculate the mean vector μ of the eigenvector set X based on the distribution characteristics of the eigenvectors X and the covariance matrix Σ X , which is used to characterize the statistical characteristics of the hardware simulation platform environment state; S23. Optimize data indexing for the feature vector set X and establish a feature index structure I X : I X = {I(x1), I(x2),..., I(x N )}; Among them, I(x i ) represents the storage location of the feature vector x i in the feature index structure to accelerate subsequent clustering calculations and anomaly detections; S24. Construct a feature space S of the hardware simulation platform environment state according to the statistical characteristics of the multi-dimensional feature vector set X X , and define: S X =(μ X ,Σ X ,I X ); Among them, S X represents the feature space calculated based on the set of feature vectors X, including the mean vector, covariance matrix, and feature index structure.
4. A method for detecting hardware simulation platform environment anomalies based on dynamic hierarchical clustering according to claim 1, characterized in that The S3 includes the following steps: S31. Calculate the distance matrix D of the multi-dimensional feature vector set X by using the dynamic weighted Euclidean distance metric based on the multi-dimensional feature vector set X X : where x ik represents the value of x i in the k-th dimension, m is the dimension of the feature vector, and ω k (t) is the dynamic weight: where, σ k (t) is the standard deviation of the k-th dimensional data within the current time t, and ∈ is a small positive number to prevent division by zero; S32. Introduce a constraint-driven mechanism, combine the known fault modes and normal modes of the hardware simulation platform to construct a must-merge constraint set ML and a prohibited-merge constraint set CL, and define a constraint penalty function Δ(x i ,x j ): Among them, λ ML and λ CL are the penalty coefficients for the must-merge and forbidden-merge constraints respectively, η ML and η CL are the corresponding scaling parameters, ‖x i - x j ‖ represents the Euclidean distance between x i and x j ; S33. Combine with the dynamic distance matrix D X and the constraint penalty function Δ(x i , x j ), to construct the constraint-driven clustering merging cost function C(i, j): C(i,j) = D X (i,j) + Δ(x i , x j ); Among them, C(i,j) represents the comprehensive cost generated when merging x i and x j ; S34. Perform bottom-up hierarchical clustering based on the constraint-driven merging cost function C(i, j), update the cost between clustering clusters during each merge operation, and use the dynamic threshold optimization method to determine the optimal hierarchical threshold. By selecting the optimal threshold for distinguishing normal and abnormal states, perform reasonable partitioning of clustering clusters, construct a hierarchical clustering tree T, and form an initial clustering cluster set.
5. The method for detecting hardware simulation platform environment anomalies based on dynamic hierarchical clustering according to claim 4, wherein, The S34 includes the following steps: S341. Construct a comprehensive merging cost function that incorporates the within-cluster dispersion information by combining the constraint-driven merging cost function C(i, j) and the dynamic weighted Euclidean distance metric to select the optimal merging pair for any two clusters C i and C j Define its comprehensive merging cost as follows: Among them, and respectively represent the variances of the internal feature vectors of the clustering clusters C i and C j reflecting the degree of dispersion of the data within each cluster. δ is a small positive number to prevent the denominator from being zero. Subsequently, the optimal merged cluster pair is selected to satisfy the following formula and make the selected cluster pair optimal in terms of the constraint-driven merging cost, reflecting the dynamic characteristics of the hardware simulation platform environment data: Among them, M t represents the set of all cluster clusters that are still in an independent state at time step t; S342. After merging the selected cluster pair (C p , C q ), the center μ r of the new merged cluster C r is calculated using exponentially decaying weighted cluster centers: where, μ p and μ q are the cluster centers of clusters C p and C q respectively, |C p | and |C q | represent the number of feature vectors in each cluster, and β is a positive parameter used to control the influence of variance on the weight; Make the clustering clusters with more concentrated data distributions have high influence weights after merging, thereby reflecting the normal operation state and abnormal differences of the hardware simulation platform; S343. During the clustering iteration process, an optimal hierarchical threshold θ is determined by using a dynamic threshold optimization method based on the correction of the difference factor * , and the difference factor δ is defined for the dynamic adaptation of the data distribution characteristics in different merging stages ij : Among them, and are the average internal dispersions of the clustering clusters C i and C j respectively, and ζ is a constant to prevent the denominator from being zero; Define the optimal threshold, combine the merging cost and the difference in inter-cluster dispersion to distinguish the normal and abnormal states of the hardware simulation platform: Among them, τ0 is a benchmark threshold preset from historical data; According to the determined optimal stratification threshold θ * Partition the final cluster set M * , define the decision function φ(C i , θ * ): Then the final clustering cluster set is expressed as: M * = {C i ∈ M t | φ(C i , θ * ) = 1}; Among them, M t is the set of all unmerged or merged clusters during the clustering iteration process, and M * is the final cluster set that satisfies the condition that the merging cost between clusters is higher than the optimal threshold θ * , so that each clustering cluster can reflect the inherent consistency of the hardware simulation platform environment data; S345. Construct a hierarchical clustering tree T that records the evolution of the hierarchical structure of the hardware simulation platform environment data during the dynamic clustering process as the supporting data structure for the entire clustering process, and completely record the clustering iteration process and its evolution trajectory. Let the clustering tree T be represented as a set of all merging events: Among them, E l represents the l-th merging event, and is the cluster to be merged in the l-th merging, is the new cluster formed after the merging, is the cost of this merging, and it satisfies ω 1 ≤ ω 2 ≤... ≤ ω L .
6. The method for detecting abnormal conditions in a hardware simulation platform environment based on dynamic hierarchical clustering according to claim 5, wherein The S5 includes the following steps: S51. According to the final hierarchical clustering model M * and the hierarchical clustering tree T, perform clustering mapping on the newly collected standardized feature vector data stream F′, and calculate the membership degree of each feature vector f′ i to its affiliated clustering cluster C k : Among them, α i,k represents the membership degree of the feature vector f′ i to the clustering cluster C k , μ k is the clustering center of C k , d(f′ i , μ k ) is the Euclidean distance between f′ i and μ k , and λ is the membership degree scaling coefficient; S52. Calculate each eigenvector f′ i in its affiliated cluster C k for the degree of deviation, and determine whether it belongs to an outlier. Define the deviation metric based on Mahalanobis distance: Among them, D M (f′ i , C k ) represents the Mahalanobis distance of f′ i within C k , and Σ k is the covariance matrix of the clustering cluster C k ; S53. Calculate the anomaly scores for all newly collected feature vectors f' i to distinguish normal data from abnormal data. Define the anomaly score A(f' i ) as follows: Among them, A(f′ i ) reflects the degree to which the feature vector f′ i deviates from the cluster C k it belongs to. γ is an abnormal score scaling factor, such that the Mahalanobis distance higher than the threshold corresponds to a higher abnormal score; S54. Set the abnormal determination threshold θ A Perform abnormal recognition on all newly collected feature vectors, and define an abnormal judgment function: Among them, φ(f′ i ) = 1 indicates that f′ i is abnormal data, and φ(f′ i ) = 0 indicates that f′ i is normal data; According to the determination result, construct the abnormal data set F abn And store the abnormal data: F abn = {f' i | φ(f' i ) = 1}; Among them, F abn represents the set of detected abnormal data.
Citation Information
Cited By
Circuit simulation device and circuit simulation program product
CN120524895A
Circuit simulation devices and circuit simulation software products
CN120524895B
File co-processing method and system based on cloud computing
CN120561626A