A method and system for screening device monitoring indicators based on time series modeling
Through the timing modeling method of the NCDE model, the problem of missing key abnormal signals in dynamic environments is solved, efficient and low-overhead equipment monitoring indicator screening is achieved, and fault warning capabilities and system stability are improved.
Patent Information
- Application Number
- CN202510457782.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art is difficult to adaptively adjust in a dynamic environment, resulting in the omission of key abnormal signals in a specific time period. Traditional monitoring solutions bring huge burdens to the processing of massive monitoring data, and redundant data masks key changing signals.
The timing modeling method based on the NCDE model is adopted to map candidate monitoring indicators to the hidden space through continuous time differential equations, capture the prediction residuals and hidden state evolution characteristics, combine service response delays and service alarm information, calculate the time-discrete correlation coefficient and importance score, and filter the target monitoring indicators.
It improves the accuracy and adaptability of key indicator screening, reduces data acquisition and storage overhead, enhances the ability to capture key signals in dynamic environments, and improves the fault warning capability and the stability of the monitoring system.
Smart Images

Figure CN120011775B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of server status analysis, and in particular, to a method and system for screening device monitoring indicators based on time series modeling. Background Art
[0002] With the booming development of cloud computing, edge computing, and the Internet of Things, the number of servers and various devices in modern data centers and distributed systems has increased explosively. There are a wide variety of device monitoring indicators in modern data centers and cloud platforms, covering multiple dimensions such as CPU, memory, disk I / O, and network traffic.
[0003] Most existing feature selection algorithms (such as principal component analysis, correlation coefficient method, information gain, etc.) mainly rely on the statistical analysis of static data, assuming that the data distribution and the correlation between indicators are basically stable throughout the monitoring period. However, in a dynamic environment, the operating state of the server is affected by multiple factors such as business load and network conditions, and the correlation between its indicators will change significantly. Static algorithms are difficult to adjust adaptively, resulting in the possible omission of key abnormal signals during a specific time period.
[0004] In addition, traditional monitoring schemes often require real-time collection and storage of all indicators, which not only brings a huge burden to the network and storage systems, but also in actual fault warning and status analysis, redundant data may mask key change signals.
[0005] Therefore, how to screen out key indicators closely related to the server status from a large amount of monitoring data and achieve efficient monitoring with low overhead has become an urgent problem to be solved in the industry. Summary of the Invention
[0006] This application provides a method, system, storage medium, computer program product, and electronic device for screening device monitoring indicators based on time series modeling, so as to at least solve the problem that it is difficult for current related technologies to screen out core key indicators closely related to the server status from a large amount of monitoring data.
[0007] In a first aspect, an embodiment of the present application provides a method for screening device monitoring indicators based on time series modeling, including: obtaining device monitoring time series data, where the device monitoring time series data covers multiple types of candidate monitoring indicators and multiple types of auxiliary parameters, and the auxiliary parameters include any one of the following: service response delay, service alarm information, and network delay; respectively defining the main input and extended input of the NCDE model based on various candidate monitoring indicators and various auxiliary parameters, so that the NCDE model maps continuously changing candidate monitoring indicators to a latent space using a differential equation form of continuous time, and captures the prediction residuals and latent state evolution characteristics of various candidate monitoring indicators through continuous time modeling; calculating the correlation coefficient between the prediction residuals of various candidate monitoring indicators and the system operating state within a continuous time window, and performing weighted summation of the correlation coefficients of each time window to obtain the time-varying correlation coefficients of the corresponding various candidate monitoring indicators; the system operating state is determined based on the service response delay and request error rate within the corresponding time window; fusing the prediction residuals, time-varying correlation coefficients, and latent state evolution characteristics of each candidate monitoring indicator to obtain the corresponding indicator comprehensive characteristics, and evaluating the importance score corresponding to the indicator comprehensive characteristics; screening target monitoring indicators for device monitoring from each candidate monitoring indicator according to the importance score.
[0008] In a second aspect, an embodiment of the present application provides a system for screening device monitoring indicators based on time series modeling, including: a data acquisition unit for obtaining device monitoring time series data, where the device monitoring time series data covers multiple types of candidate monitoring indicators and multiple types of auxiliary parameters, and the auxiliary parameters include any one of the following: service response delay, service alarm information, and network delay; an NCDE analysis unit for respectively defining the main input and extended input of the NCDE model based on various candidate monitoring indicators and various auxiliary parameters, so that the NCDE model maps continuously changing candidate monitoring indicators to a latent space using a differential equation form of continuous time, and captures the prediction residuals and latent state evolution characteristics of various candidate monitoring indicators through continuous time modeling; a time-varying correlation analysis unit for calculating the correlation coefficient between the prediction residuals of various candidate monitoring indicators and the system operating state within a continuous time window, and performing weighted summation of the correlation coefficients of each time window to obtain the time-varying correlation coefficients of the corresponding various candidate monitoring indicators; the system operating state is determined based on the service response delay and request error rate within the corresponding time window; an importance evaluation unit for fusing the prediction residuals, time-varying correlation coefficients, and latent state evolution characteristics of each candidate monitoring indicator to obtain the corresponding indicator comprehensive characteristics, and evaluating the importance score corresponding to the indicator comprehensive characteristics; a target indicator screening unit for screening target monitoring indicators for device monitoring from each candidate monitoring indicator according to the importance score.
[0009] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to perform the steps of the device monitoring metric screening method based on temporal modeling according to any embodiment of the present application.
[0010] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the device monitoring metric screening method based on temporal modeling according to any embodiment of the present application.
[0011] In a fifth aspect, an embodiment of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the steps of the device monitoring metric screening method based on temporal modeling according to any embodiment of the present application.
[0012] Through a device monitoring metric screening method and system based on temporal modeling provided by the present application, at least the following technical effects can be achieved:
[0013] (1) By introducing a temporal modeling method based on the NCDE (Neural Controlled Differential Equations) model and using continuous-time differential equations to model the dynamic changes of device monitoring metrics, the hidden state evolution characteristics and prediction residual characteristics of the metrics changing over time are fully captured, and the importance of the metrics is further evaluated in combination with system state parameters, thus significantly improving the accuracy and adaptability of key metric screening. In addition, by introducing auxiliary parameters such as service response latency, service alarm information, and network latency into the NCDE model as extended inputs, it helps the model to more accurately characterize the correlation between the server system state and monitoring metrics, effectively improving the accuracy of key metric screening, and thus enhancing the ability to early warn of potential failures.
[0014] (2) By analyzing the correlation between the prediction residuals and the system operating state, the change in metric correlation can be detected in a timely manner, enhancing the ability to capture key signals in a dynamic environment. In addition, when calculating the time-varying correlation coefficient, the data of multiple time windows are integrated in a weighted summation manner, which can effectively balance the short-term fluctuations and long-term trends in the temporal data without increasing a large amount of computational overhead, further improving the stability of metric screening.
[0015] (3) By fusing the prediction residuals, time-varying correlation coefficients, and hidden state evolution characteristics of candidate monitoring indicators, a more comprehensive comprehensive feature of indicators was constructed. Based on this, the importance scores of each indicator were evaluated. By fusing a variety of feature information, the influence degree of various candidate indicators on the system state was comprehensively evaluated, enabling the indicator screening to not only identify indicators strongly correlated with the system state but also capture those hidden indicators that have potential indicative effects on the abnormal changes of the system in complex environments, further improving the comprehensiveness and stability of the screening results.
[0016] Through this technical solution, not only the accuracy and timeliness of the key indicators screened from the massive monitoring parameters of the server system are improved, but also the data acquisition and storage overhead are effectively reduced, providing an efficient and low-overhead solution for equipment monitoring in data centers, cloud platforms, and edge computing environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 FIG. shows a flowchart of an example of a method for screening equipment monitoring indicators based on time series modeling according to an embodiment of the present application;
[0019] Figure 2 FIG. shows an operation flowchart of an example of screening target monitoring indicators according to the importance score according to an embodiment of the present application;
[0020] Figure 3 FIG. shows a structural block diagram of an example of a system for screening equipment monitoring indicators based on time series modeling according to an embodiment of the present application;
[0021] Figure 4 FIG. is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0023] In the technical solution of this application, for the processing of collection, storage, use, processing, transmission, provision, and disclosure of users' personal information involved, it complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0024] Figure 1 FIG. 4 shows a flowchart of an example of a method for screening device monitoring indicators based on time series modeling according to an embodiment of the present application.
[0025] Regarding the execution subject of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by a server system management platform or a cloud platform management system. Through the deep combination of time series modeling and the NCDE model, it effectively makes up for the deficiencies of traditional static feature selection methods in a dynamic environment, realizes the precise screening and real-time dynamic adjustment of key monitoring indicators, and provides low-overhead, high-efficiency, and strong self-adaptive technical support for device monitoring in modern data centers and cloud platforms.
[0026] In some examples, it can be integrated and configured in an electronic device or a terminal in a software, hardware, or software-hardware combination manner, and the types of the terminal or the electronic device can be diverse, such as mobile phones, tablet computers, or desktop computers, etc.
[0027] As Figure 1 shown, in step S110, device monitoring time series data is obtained, and the device monitoring time series data covers multiple types of candidate monitoring indicators and multiple types of auxiliary parameters.
[0028] It should be noted that the types of candidate monitoring indicators can be diverse, and can be computing resource indicators (such as CPU usage rate, CPU load, memory usage rate, etc.), storage resource indicators (such as disk usage rate, number of I / Os per second, I / O wait queue length, etc.), network resource indicators (such as network traffic, network bandwidth usage rate, etc.), etc. It can select all service monitoring indicators as the analysis basis, or can also pre-narrow the scope of automatic screening of core indicators according to service business requirements.
[0029] In addition, the auxiliary parameters include any one of the following: service response latency, service alarm information, and network latency. It should be pointed out that the auxiliary parameters can also partially overlap with some of the candidate monitoring indicators. The purpose of selecting the auxiliary parameters is to borrow the parameters in the auxiliary indicators to more prominently show the business environment change information corresponding to the key anomalies of the system device.
[0030] In some embodiments, timestamp-based log data, performance monitoring tools (such as Prometheus, Zabbix), or system API interfaces can be used to obtain metric data to ensure the continuity and integrity of time-series data. Additionally, according to business requirements, preliminary aggregation can be performed at the collection level (such as the average value, maximum value, etc. per second or per minute), which not only reduces storage overhead but also provides smooth input data for subsequent time-series modeling.
[0031] In step S120, based on various candidate monitoring metrics and various auxiliary parameters, the main input and extended input of the NCDE model are respectively defined, enabling the NCDE model to map continuously changing candidate monitoring metrics to the latent space in the form of a differential equation for continuous time and capture the prediction residuals and latent state evolution characteristics of various candidate monitoring metrics through continuous-time modeling.
[0032] In some embodiments, for each candidate monitoring metric (such as CPU utilization, memory occupancy, disk I / O, network traffic, etc.), a continuous-time signal is constructed through data preprocessing (normalization, denoising, time-series alignment, etc.) and used as the main input of the NCDE model. It should be noted that since actual sampling is discrete data, interpolation or piecewise linear methods can be used to construct a continuous function to ensure the continuity of the model input. Additionally, the auxiliary parameters can also be preprocessed to construct a continuous control signal.
[0033] Here, by using the NCDE model to perform continuous-time modeling on candidate monitoring metrics, prediction residuals and latent state evolution characteristics are extracted, and at the same time, auxiliary parameters are fused to enhance the model's expressive ability. Specifically, the auxiliary parameters are fused with the time-series data of the candidate metrics. For example, the information of the auxiliary parameters is embedded into the latent state update in a concatenation manner, enabling the model to perceive the impact of changes in the business environment on the dynamics of the candidate metrics.
[0034] The NCDE model uses the form of a continuous-time differential equation to map the preprocessed candidate monitoring metrics (main input) to the latent space. At the same time, through the modulation of the auxiliary parameters (extended input), it captures the prediction residuals and latent state evolution characteristics of the metrics, effectively handles the problem of irregular sampling, and can also provide rich and accurate time-series dynamics.
[0035] In step S130, the correlation coefficient between the prediction residuals of various candidate monitoring metrics and the system operating state is calculated within a continuous time window, and the correlation coefficients of each time window are weighted and summed to obtain the time-varying correlation coefficients of the corresponding various candidate monitoring metrics.
[0036] In some embodiments, according to business characteristics and monitoring requirements, continuous time-series data is divided into multiple overlapping or non-overlapping time windows to ensure that the data within each window has sufficient representativeness. Furthermore, within each time window, the trained NCDE model is used to predict candidate monitoring indicators, and the residual between the actual value and the predicted value is calculated. Indicators with larger prediction residuals may have a more significant fluctuating impact on the device state.
[0037] On the other hand, the system operating state is determined based on the service response latency and request error rate within the corresponding time window. Determining the system operating state based on the service response latency and request error rate within the time window, for example, through linear weighted fusion, enables the resulting comprehensive indicator of the system operating state to not only measure the system's response speed to requests but also consider the stability of response correctness.
[0038] Within each time window, the correlation coefficient between the prediction residual of the candidate monitoring indicator and the system operating state is calculated. The measurement method of the correlation coefficient can be the Pearson correlation coefficient, Spearman rank correlation coefficient, etc., to capture the dynamic relationship between the candidate monitoring indicator and the system operating state within the corresponding time window. Furthermore, the weighted summation method is used to aggregate the correlation coefficients of each time window, thereby obtaining the time-varying correlation coefficient of each candidate indicator. In addition, the window weighting strategy can also be diversified, for example, flexibly adjusted according to the length of the time window, adaptive weights, or the exponential smoothing weighting method based on a sliding window to smooth out the accidental fluctuations that may exist in a single time window and improve the robustness of the time-varying correlation coefficient of the overall evaluation.
[0039] In step S140, the prediction residuals, time-varying correlation coefficients, and hidden state evolution features of each candidate monitoring indicator are fused to obtain the corresponding comprehensive feature of the indicator, and the importance score corresponding to the comprehensive feature of the indicator is evaluated.
[0040] Here, feature fusion can adopt simple vector concatenation, weighted averaging, or non-linear mapping and feature weighting based on a multi-layer perceptron, enabling each feature to be reasonably reflected in the comprehensive feature. In terms of the importance score, a non-linear model (such as a lightweight neural network or a regression model) can be used to output the final importance score. Through supervised learning or self-supervised mechanisms of historical data, the model parameters are continuously optimized to enable the model to capture the non-linear relationships between the features.
[0041] Exemplarily, to capture more complex non-linear relationships between the features, a lightweight neural network (such as a one- or two-layer fully connected network) can be used to model the comprehensive feature vector of the candidate monitoring indicator
[0042] , formula (1)
[0043] Wherein represents a scoring network are the parameters of the scoring network, trained by supervised learning or self-supervised methods. represents the importance score of the corresponding output, which can automatically learn the interactions and non-linear relationships between features, thereby improving the scoring accuracy.
[0044] In step S150, the target monitoring indicators for device monitoring are screened from each candidate monitoring indicator according to the importance score.
[0045] Here, multiple screening strategies can be set according to business requirements, such as the fixed threshold method, the ranking method, or the adaptive screening method. In the fixed threshold method, only the indicators with importance scores higher than a certain set threshold are selected. In the ranking method, a preset number of indicators with the top scores are selected. In the adaptive screening method, the number of indicators is dynamically adjusted according to the distribution characteristics of the importance scores, and all are within the scope of implementation of the embodiments of the present application.
[0046] Through the screening method based on the importance score, the key indicators that have a significant impact on the device state can be accurately screened out. The screened target monitoring indicators are integrated into the actual monitoring platform, avoiding the ineffective monitoring of secondary or redundant indicators, greatly reducing the occupancy of network, storage, and computing resources, and improving the overall operation efficiency of the system. In addition, based on the high accuracy and real-time nature of the screened target monitoring indicators, the system can quickly detect abnormal states and provide effective data support for fault warning and performance optimization.
[0047] Figure 2 shows an operation flow chart of an example of screening target monitoring indicators according to the importance score according to an embodiment of the present application.
[0048] As Figure 2 shown, in step S210, the importance of each candidate monitoring indicator is screened according to a preset importance threshold to obtain at least one corresponding potential monitoring indicator.
[0049] Here, the importance threshold can be a preset fixed threshold or a dynamic threshold. For example, by setting a fixed value (such as 0.5 or 0.7) as the screening criterion, the indicators with scores higher than the threshold are screened out, and the threshold can also be adaptively set according to the distribution characteristics (such as mean, standard deviation, or quantile) of the importance scores to ensure that the number of screened indicators meets a specific range.
[0050] In addition, through screening based on importance scores, redundant indicators with little impact on the device status can be effectively eliminated, thereby reducing the scale of the indicator set, lowering the computational complexity of subsequent clustering analysis, ensuring that the selected potential monitoring indicators have high relevance and representativeness, and further improving the efficiency of the monitoring system and data processing performance.
[0051] In step S220, each potential monitoring indicator and its corresponding comprehensive indicator feature are input into the dynamic DBSCAN model to determine at least one corresponding clustering cluster, calculate the center of the clustering feature vector corresponding to each clustering cluster, and select the potential monitoring indicator closest to the center of the clustering feature vector as the clustering representative monitoring indicator for the corresponding clustering cluster.
[0052] Here, DBSCAN (Density-Based Spatial Clustering of Applications with Noise, density-based spatial clustering algorithm with noise) can effectively identify clustering clusters with different densities and exclude noise data. However, traditional DBSCAN is relatively fixed in parameter setting. Through the dynamic DBSCAN model, parameters (such as neighborhood radius and minimum sample number) can be adaptively adjusted to improve the adaptability to data distribution and ensure accurate clustering under different data feature conditions.
[0053] Through the dynamic DBSCAN model, potential monitoring indicators are divided into several clustering clusters according to the similarity between comprehensive indicator features. Calculate the center of the clustering feature vector for each clustering cluster (i.e., the average of the feature vectors in the cluster). In each clustering cluster, select the potential monitoring indicator closest to the center of the cluster as the clustering representative monitoring indicator for the cluster. In this way, by selecting the indicator closest to the clustering center as the clustering representative monitoring indicator, it can be ensured that the indicator has strong central features, represents the overall features of the clustering cluster to the greatest extent, reduces the retention of repetitive and redundant indicators, and improves the accuracy of indicator screening.
[0054] In step S230, determine the target monitoring indicators according to the clustering representative monitoring indicators corresponding to each clustering cluster.
[0055] In some embodiments, the selected cluster representative monitoring indicators in each cluster can be directly used as the target monitoring indicators. Compared with directly screening based on the importance scores, through the optimization screening method provided in this embodiment, by selecting the cluster representative monitoring indicators as the final target monitoring indicators, it is ensured that the selected indicators not only have a high importance, but also are representative and stable in terms of feature distribution, further reducing the redundancy between indicators, reducing the data redundancy and storage overhead caused by repeated features, and at the same time ensuring the independence and effectiveness of each monitoring indicator in describing the system state.
[0056] In the following, the relevant details of some example algorithms that may be involved or applied in the embodiments of the present application will be elaborated. It should be understood that the description of these algorithms is only for the convenience of the public to better understand the spirit of the present application, and is not intended to limit the scope of implementation of the present application, and other non-restrictive and suitable algorithms can also be used.
[0057] Regarding the modeling details of the prediction residuals of the candidate monitoring indicators in step S120, in some embodiments, the discrete sampling data is interpolated and converted into a continuous signal, and normalization processing is performed. Then, the candidate monitoring indicators are used as the main input, and the auxiliary parameters are used as the extended input. The NCDE model is used to model using continuous-time differential equations to capture the temporal dynamic characteristics. Finally, the predicted value is obtained through the decoder, and the predicted value is compared with the true value to calculate the prediction residual.
[0058] More specifically, assume that the discrete sampling data of the candidate monitoring indicators and the auxiliary parameters are represented as and , represents the th sampling time point, represents the total number of sampling time points, and respectively represent the data values of the candidate monitoring indicators and the auxiliary parameters sampled at . The continuous time series signals are constructed by interpolation method and normalized to obtain the continuous time series signal of the corresponding candidate monitoring indicators and the continuous time series signal of the auxiliary parameters to meet the input conditions of the NCDE model.
[0059] Here, in order to meet the requirements of NCDE for "continuous time input", interpolation needs to be performed on the original discrete data. For example, linear interpolation, spline interpolation or more advanced interpolation methods (such as locally weighted regression) can be used to smoothly connect adjacent sampling points. In addition, if there are missing data or sampling anomalies, certain outlier detection and preprocessing need to be performed before interpolation, such as using forward filling or mean filling strategies to ensure data continuity.
[0060] The continuous time-series signals of candidate monitoring indicators are mapped to the hidden state space through the NCDE model, and the continuous time-series signals of auxiliary parameters are used to modulate the model.
[0061] , Equation (2)
[0062] In the formula, represents the hidden state at continuous time , represents the initial time, represents the initial hidden state; is the integration variable, representing any time point between and ; represents the control function, parameterized by a neural network with parameters , and this function is used to dynamically calculate the change rate of the hidden state according to the hidden state at time and the extended input ; represents the increment of the main input signal in a tiny time interval, represents the driving effect of the candidate monitoring indicator on the evolution of the hidden state.
[0063] It should be noted that in traditional discrete time-series models such as RNN / LSTM, the state is updated at discrete time steps. However, in the NCDE, the hidden state can be integrally evolved at each moment in the time series, having the natural advantage of handling irregular sampling and complex time-varying correlations.
[0064] The control function adopts an attention mechanism for parameterization.
[0065] Specifically, within the control function, and can first be mapped to the same hidden space, and then the attention weights are obtained through the attention network. If contains multiple auxiliary sub-channels (such as alarm frequency, network RTT, service KPI), they can be concatenated and then dimension-reduced through a multi-layer perceptron (MLP), and then input into the attention network together with for fusion.
[0066] , Equation (3)
[0067] , Equation (4)
[0068] In the formula, Denotes a multi - layer perceptron in the attention network for non - linear feature mapping; Denotes a vector concatenation operation, Denotes the time instant Of the attention weight vector, 、 And Denote the hidden - state weight matrix, the auxiliary - input weight matrix, and the attention bias term respectively.
[0069] Here, Determines the degree of attention to the hidden state and the auxiliary input at time instant . If a certain dimension in the auxiliary input is particularly sensitive to the current state of the device, the attention will automatically increase its weight, and the importance of different auxiliary parameters will be dynamically scheduled through the attention mechanism. Concatenate To And then input them together into the multi - layer perceptron, and use non - linear mapping to calculate the incremental update of the hidden state. In this way, when the device state is close to the critical value or business alarms occur frequently, this update process will be more intense, which helps to capture the subtle signs before the fault.
[0070] It should be noted that although it is continuous time in theory, in practice, the 4th - order Runge - Kutta (RK4) method is used to iteratively solve the Value at discrete sampling points.
[0071] Based on the hidden state Of the decoder at the corresponding sampling time instant , to infer the predicted value Of the candidate monitoring index at the next sampling time point .
[0072] , Equation (5)
[0073] In the formula, Denotes the decoder function, Denotes the learnable parameters of the decoder.
[0074] Here, the decoder is defined as transforming the hidden state at time instant To the device monitoring index space at time instant . Specifically, the decoder can use a simple MLP or a more complex network structure (such as convolution, Transformer decoder) to improve the capture of non - linear relationships.
[0075] Compare the predicted values of the candidate monitoring indexes at each sampling time point with the true observed values to calculate the corresponding prediction residuals.
[0076] , Equation (6)
[0077] In the formula, represents the sampling moment corresponding to the predicted residual, represents at the sampling moment the true value of the candidate monitoring index collected at the location.
[0078] By analyzing the temporal variation of the residual, sudden fluctuations, abnormal trends, etc. can be detected in a timely manner. When the absolute value remains continuously large, it means that the model has obvious deficiencies in grasping the device behavior in this interval, and often corresponds to potential faults or sudden interferences. In addition, the predicted residuals of different types of indicators can be combined or the norm can be calculated to obtain an overall anomaly measure.
[0079] Through the embodiments of the present application, the continuous-time property of the NCDE and the attention mechanism enable the model to better cope with irregular sampling, load mutations, and various service interferences. In addition, the use of predicted residuals can clearly identify key indicators sensitive to faults, reduce the monitoring frequency of redundant indicators with little value, and save storage and computing overhead.
[0080] In some examples of the embodiments of the present application, the hidden state evolution features include the hidden state mean, the hidden state variance, the instantaneous change rate of the hidden state, and the cumulative change amount of the hidden state.
[0081] In the NCDE model, the hidden state will continuously evolve over time. In order to measure the overall level of the hidden state within a time interval, it is necessary to integrate and calculate the mean value thereof, so as to obtain the hidden state mean, which reflects the basic operation level of the system during this time period.
[0082] , Equation (7)
[0083] In the formula, represents the hidden state mean within the time interval , represents any time point from to in between, represents from to any time point in between of the hidden state; is the infinitesimal time increment of the integration variable, representing the integration of the continuous time change process.
[0084] The value of is that it can quickly evaluate the "average level" of the hidden state of the system within a period of time, and can be used to detect whether there is a significant deviation from the historical benchmark.
[0085] Hidden state variance Reflects the system state relative to its mean value within the time interval The larger the variance, the more obvious the fluctuation of the system's operating status, and the more likely there is a potential abnormality or fault sign.
[0086] , Formula (8)
[0087] In the formula, Indicates the time interval The hidden state variance within .
[0088] The value of is that it is complementary to the mean. It only reveals the "central level" of the system as a whole, while the variance It measures the dispersion around the center, and The abnormal increase of often indicates that the system state is unstable.
[0089] In the NCDE model, the input data meets the continuous time series requirements, so that each moment can be regarded as a differentiable state function, so the instantaneous rate of change can be naturally defined. Through the instantaneous rate of change of the hidden state, it is possible to capture whether the system has signs of sudden changes such as acceleration or deceleration in a short period of time, further improving the sensitivity to abnormalities.
[0090] , Formula (9)
[0091] In the formula, Indicates at time The instantaneous rate of change of the hidden state, and Respectively indicate at time and The hidden state vector of is the preset time interval used for discretized derivative computations.
[0092] It is a very small time interval, which can be set according to the sampling resolution or integration step. When the system state suddenly changes, the instantaneous rate of change often has a peak value, which can reflect the signs of fault more promptly than just looking at the mean or variance.
[0093] Integrate the rate of change of the hidden state within the time interval to form a measure of the "total change amplitude" of the system state, which is similar to the "total distance" traveled by the curve in the latent space, and is used to define the cumulative change of the hidden state.
[0094] , Formula (10)
[0095] In the formula, Indicates the time interval The cumulative change of hidden state in represents the Euclidean norm.
[0096] Focus on the dramatic changes at a certain moment, The focus is on whether the latent state has experienced a large migration or fluctuation during the entire time period. The two complement each other and jointly complete the abnormal identification of "chronic degradation" or "sudden failure".
[0097] Through the embodiments of the present application, four types of evolutionary features of the latent state are extracted, namely, the instantaneous rate of change, the cumulative change, the variance and the mean. Sudden events are detected by the instantaneous rate of change, slow evolution is identified by the cumulative change, the overall fluctuation is revealed by the variance, and the drift of the overall level is reflected by the mean. This can comprehensively cover different types of anomalies (suddenness, gradualness, etc.) in the system, and achieve multi-angle characterization and anomaly capture of the system or device status.
[0098] Regarding the details of the calculation of the time-varying correlation coefficient of the candidate monitoring indicator in step S130, in some embodiments, it can be designed through an exponentially weighted time-varying correlation coefficient to adapt to the dynamic changes of the system operation status in real time.
[0099] More specifically, the service response delay and request error rate at each sampling moment are weighted and fused to determine the corresponding system operation status.
[0100] , Formula (11)
[0101] In the formula, Indicates the sampling time The system operation status, Indicates that the system is Service response delay, Indicates the sampling time The request error rate, is the balance weight parameter.
[0102] Here, the overall operating status of the system is abstracted into a comprehensive indicator of service response delay and request error rate. Through linear weighted fusion, not only the response speed of the system to the request is measured, but also the stability of the response correctness is considered. Therefore, through the composite state representation, system operation anomalies can be captured more sensitively, especially when the response is slow or the system is abnormal.
[0103] Assume Time windows Inside Sampling time, record candidate monitoring indicators Get the discrete prediction residual sequence within the time window , and record the system operation status sequence within the time window ; among them, represents the candidate monitoring index within the time window at the -th sampling moment of the prediction residual, represents the system operation status at the -th sampling moment within the time window .
[0104] Within the time window, calculate the Pearson correlation coefficient between the prediction residual sequence and the system status sequence for each candidate index .
[0105] , Equation (12)
[0106] In the formula, represents the Pearson correlation coefficient of the candidate index within the time window , reflecting the linear relationship between the prediction residual and the system status; is the mean value of the prediction residual of the candidate index within the time window , represents the mean value of the system status within the time window .
[0107] Here, calculate the correlation between the prediction residual of the candidate index and the system operation status sequence, so as to judge whether the fluctuation of the prediction residual is closely related to the change of the overall system status. Thus, identify which monitoring indicators' prediction errors more directly reflect the operation fluctuation or abnormal status of the system.
[0108] Adopt an exponential decay weight mechanism to perform weighted averaging on the Pearson correlation coefficients of each time window, so as to obtain the time-varying correlation coefficient of the candidate monitoring index.
[0109] , Equation (13)
[0110] In the formula, represents the weight of the time window , which decays as the distance between the end time of the window and the current time increases; is a positive parameter for controlling the decay rate.
[0111] In Equation (13), as time goes by, a closer time window should have a greater impact on the current prediction, while the impact of more distant data gradually decreases. The principle of timeliness is reflected by the exponential decay mechanism, which enhances the sensitivity of the model to real-time data, enables it to quickly respond to the system state that changes in real time, and is more suitable for the actual dynamic monitoring environment.
[0112] , Equation (14)
[0113] In the formula, represents the time-varying correlation coefficient of the candidate index , and
[0114] is the total number of time windows.
[0115] In Equation (14), the correlation coefficients of multiple windows are weighted in an exponential weight manner to obtain the time-varying correlation coefficient of the candidate index, dynamically reflecting the changing trend of the system sensitivity of the monitoring index, effectively capturing short-term and long-term trend changes, reducing the uncertainty brought by the fluctuation of a single window, and making the anomaly detection more robust and reliable.
[0116] Through the embodiments of the present application, the prediction residual sequence and the system state sequence are directly subjected to synchronous correlation calculation, clearly positioning the key sensitive indicators, effectively avoiding the situation that irrelevant indicators may be misselected in the traditional method, and improving the monitoring efficiency and warning accuracy. In addition, in the final calculation of the time-varying correlation coefficient, by weighted fusion of the correlation coefficients of multiple windows, the interference of accidental fluctuations or abnormal data is weakened, and the stability and robustness of the system anomaly state judgment are improved.
[0117] Regarding the implementation details of step S220, in some embodiments, by adaptively and dynamically determining the DBSCAN algorithm parameters and , it can effectively adapt to the change of the index feature distribution brought by different service loads and monitoring environment changes, so as to achieve adaptive clustering and improve the index clustering effect and stability.
[0118] More specifically, assume that the set of all potential monitoring indicators is represented as , Denote the total number of potential monitoring indicators screened by the importance threshold, and construct an indicator feature matrix by synthesizing the comprehensive features of each potential monitoring indicator. 。
[0119] Here, first, potential monitoring indicators are preliminarily screened by a preset importance threshold, and then a comprehensive feature matrix is constructed for each potential indicator to accurately characterize the dynamic characteristics of each indicator.
[0120] To adapt to the characteristics of the monitoring indicator features changing with time and environment, a dynamic method is used to calculate the neighborhood radius and the minimum number of neighborhood samples.
[0121] Regarding the dynamic calculation method of the neighborhood radius, specifically, a weighted method of the mean and standard deviation of the nearest neighbor distances between indicators is used, which can adaptively adjust parameters and avoid the problem of poor clustering effect caused by fixed parameters in the traditional DBSCAN algorithm.
[0122] Regarding the dynamic calculation method of the minimum number of neighborhood samples, specifically, a method of taking the logarithm of the total number of potential monitoring indicators and multiplying by a regulation coefficient is used, which can be dynamically adjusted according to the change of the number of indicators to ensure the stability of the clustering effect under different data scales and complexities.
[0123] , Equation (15)
[0124] , Equation (16)
[0125] , Equation (17)
[0126] , Equation (18)
[0127] In the formula, Denote the th potential monitoring indicator, and respectively denote the comprehensive features of the indicator and the indicator corresponding to, Denote the set composed of the nearest indicators to the indicator Denote the Euclidean distance between feature vectors, Denote the average distance from the indicator to its nearest neighbor indicators, Denote the standard deviation of the distances from the indicator to its nearest neighbor indicators, Denote the number of nearest neighbor indicators, denotes the floor function; is the neighborhood radius in the DBSCAN algorithm, representing the neighborhood range of an index point in the feature space; is the adjustment coefficient representing the tightness of the neighborhood radius, with a value range of [1.0, 2.0]; is the minimum number of samples in the neighborhood in the DBSCAN algorithm, used to determine the minimum number of indices required in the neighborhood for a point to become a core point; represents the adjustment coefficient used to control the change speed of the minimum number of neighborhood samples with the total number of indices, with a value range of [1.5, 3.0].
[0128] In the above formulas (15)-(18), to solve the problem that the traditional DBSCAN clustering parameters are fixed and difficult to adapt to dynamic changes, a scheme for dynamically adjusting the DBSCAN clustering parameters (neighborhood radius and minimum number of samples) based on the statistical features (mean and standard deviation) of the nearest neighbor distances between indices is proposed. In this way, it is ensured that the clustering algorithm can reflect the true dynamic distribution state of the monitoring index feature space in real time.
[0129] Perform DBSCAN clustering processing on the index feature matrix using the dynamically calculated neighborhood radius and minimum number of neighborhood samples. Specifically, for each potential monitoring index , calculate the number of feature points in its neighborhood. If it exceeds , then mark it as a core point. Through the neighborhood expansion method, expand the clustering cluster starting from the core point until it can no longer be expanded, thus obtaining the corresponding set of clustering clusters , represents the total number of clustering clusters, represents the th clustering cluster.
[0130] For each clustering cluster, calculate the center of the clustering feature vector according to the mean of the comprehensive features of all indices within the clustering cluster. The clustering feature center is represented by the mean of the feature vectors of all indices within the clustering, and can accurately reflect the common features within the clustering.
[0131] , formula (19)
[0132] In the formula, represents the th clustering cluster, represents the number of indices in represents the center of the feature vector of represents each monitoring index within the comprehensive feature of the index.
[0133] For each cluster, select the metric with the minimum distance from the center of the cluster feature vector as the representative monitoring metric for the cluster.
[0134] Specifically, the Euclidean distance is used to select the metric closest to the cluster center as the representative metric. This method can effectively exclude the interference of outliers and marginal metrics, ensuring the typicality and accuracy of the representative metric.
[0135] , Equation (20)
[0136] In the formula, represents the representative monitoring metric selected in , represents the Euclidean distance between the comprehensive feature of the metric and the center of the cluster feature vector of .
[0137] In the embodiment of the present application, after the dynamic DBSCAN clustering is completed, the center of the feature vector of all metrics within each cluster is further accurately calculated, and then the metric closest to the cluster center is selected as the representative monitoring metric for the cluster, ensuring the representativeness and accuracy of the selected representative metric within the cluster.
[0138] Through the embodiment of the present application, the clustering parameters are adaptively adjusted according to the dynamic environmental changes, ensuring that the clustering results are more in line with the actual equipment operating environment and the true distribution state of the feature space, and effectively improving the adaptability and sensitivity to environmental changes. Through the dynamic DBSCAN clustering parameter calculation and clustering process, the characteristic change trend of potential monitoring metrics evolving over time can be captured in real time, so as to adjust the key monitoring metric set in real time, and the response speed of the monitoring system to abnormal events or faults is improved.
[0139] In addition, an accurate selection method for the clustering representative metrics is adopted. By selecting the representative metrics through the clustering feature center, the representativeness and accuracy of the metric selection are ensured, effectively avoiding the problem of inaccurate selection of key metrics caused by metric redundancy or atypism. Through the precise selection of representative metrics, the monitoring overhead of redundant metrics is greatly reduced, the data storage and processing costs are reduced, and the allocation efficiency of monitoring resources is optimized.
[0140] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of combined actions. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0141] Figure 3 FIG. 4 shows a structural block diagram of an example of a device monitoring index screening system based on time series modeling according to an embodiment of the present application.
[0142] As Figure 3 shown, the device monitoring index screening system 300 based on time series modeling includes a data acquisition unit 310, an NCDE analysis unit 320, a time-varying correlation analysis unit 330, an importance evaluation unit 340, and a target index screening unit 350.
[0143] The data acquisition unit 310 is configured to acquire device monitoring time series data, and the device monitoring time series data covers multiple types of candidate monitoring indexes and multiple types of auxiliary parameters, and the auxiliary parameters include any one of the following: service response delay, service alarm information, and network delay.
[0144] The NCDE analysis unit 320 is configured to respectively define the main input and the extended input of the NCDE model based on various candidate monitoring indexes and various auxiliary parameters, so that the NCDE model maps the continuously changing candidate monitoring indexes to the hidden space in the form of a differential equation of continuous time, and captures the prediction residuals and the hidden state evolution characteristics of various candidate monitoring indexes through continuous time modeling.
[0145] The time-varying correlation analysis unit 330 is configured to calculate the correlation coefficient between the prediction residuals of various candidate monitoring indexes and the system operation state within a continuous time window, and perform weighted summation on the correlation coefficients of each time window to obtain the time-varying correlation coefficients of the corresponding various candidate monitoring indexes; the system operation state is determined according to the service response delay and the request error rate within the corresponding time window.
[0146] The importance evaluation unit 340 is configured to fuse the prediction residuals, the time-varying correlation coefficients, and the hidden state evolution characteristics of each candidate monitoring index to obtain the corresponding index comprehensive characteristics, and evaluate the importance score corresponding to the index comprehensive characteristics.
[0147] The target metric screening unit 350 is configured to screen target monitoring metrics for device monitoring from each of the candidate monitoring metrics according to the importance scores.
[0148] In some embodiments, the embodiments of the present application provide a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to perform the steps of any one of the above-mentioned device monitoring metric screening methods based on time series modeling of the present application.
[0149] In some embodiments, the embodiments of the present application further provide a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the steps of any one of the above-mentioned device monitoring metric screening methods based on time series modeling.
[0150] In some embodiments, the embodiments of the present application further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the device monitoring metric screening method based on time series modeling.
[0151] Figure 4 FIG. is a schematic hardware structure diagram of an electronic device for performing the device monitoring metric screening method based on time series modeling provided by another embodiment of the present application. As Figure 4 shown, the device includes:
[0152] One or more processors 410 and a memory 420. Figure 4 Here, one processor 410 is taken as an example.
[0153] The device for performing the device monitoring metric screening method based on time series modeling may further include: an input device 430 and an output device 440.
[0154] The processor 410, the memory 420, the input device 430, and the output device 440 may be connected by a bus or other means. Figure 4 Here, being connected by a bus is taken as an example.
[0155] The memory 420 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the device monitoring metric screening method based on timing modeling in the embodiments of the present application. The processor 410 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 420, that is, implements the device monitoring metric screening method based on timing modeling in the above method embodiments.
[0156] The memory 420 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device. In addition, the memory 420 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 420 may optionally include a memory remotely provided relative to the processor 410, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0157] The input device 430 can receive input digital or character information and generate signals related to the user settings and function control of the electronic device. The output device 440 may include a display device such as a display screen.
[0158] The one or more modules are stored in the memory 420 and, when executed by the one or more processors 410, execute the device monitoring metric screening method based on timing modeling in any of the above method embodiments.
[0159] The above product can execute the method provided in the embodiments of the present application and has the corresponding functional modules and beneficial effects of the executed method. For technical details not described in detail in this embodiment, reference can be made to the method provided in the embodiments of the present application.
[0160] The electronic device in the embodiments of the present application exists in various forms, including but not limited to:
[0161] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0162] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.
[0163] (3) Portable entertainment devices: Such devices can display and play multimedia content. This type of device includes: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.
[0164] (4) Other airborne electronic devices with data interaction functions, such as in-vehicle device installed on vehicles.
[0165] The device embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0166] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0167] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for screening device monitoring indicators based on time series modeling, comprising: Obtaining device monitoring time series data, where the device monitoring time series data covers multiple types of candidate monitoring indicators and multiple types of auxiliary parameters, and the auxiliary parameters include any one of the following: service response latency, service alarm information, and network latency; Defining the main input and extended input of the NCDE model based on various candidate monitoring indicators and various auxiliary parameters respectively, such that the NCDE model maps continuously varying candidate monitoring indicators to a latent space using a differential equation form of continuous time, and captures the prediction residuals and latent state evolution characteristics of various candidate monitoring indicators through continuous time modeling; Among them, comparing the predicted values of candidate monitoring indicators at each sampling time point with the true observed values to calculate the corresponding prediction residuals: , In the formula, represents the sampling moment corresponding to the predicted residual represents the true value of the candidate monitoring index collected at the sampling moment represents the predicted value of the candidate monitoring index at the sampling time point ; Calculating the correlation coefficients between the prediction residuals of various candidate monitoring indicators and the system operating state within a continuous time window, and performing weighted summation of the correlation coefficients of each time window to obtain the time-varying correlation coefficients of various candidate monitoring indicators; the system operating state is determined based on the service response latency and request error rate within the corresponding time window; Among them, the calculation of the time-varying correlation coefficient for the candidate monitoring indicators includes: , In the formula, represents the time-varying correlation coefficient of the candidate monitoring index , is the total number of time windows represents the time window weight; represents the Pearson correlation coefficient of the candidate index within the time window , reflecting the linear relationship between the prediction residual and the system state; Fusing the prediction residuals, time-varying correlation coefficients, and latent state evolution characteristics of each candidate monitoring indicator to obtain the corresponding comprehensive indicator characteristics, and evaluating the importance score corresponding to the comprehensive indicator characteristics; Among them, a lightweight neural network is used to model the comprehensive feature vectors of the candidate monitoring indicators : , In the formula, represents the scoring network, is the parameter of the scoring network, represents the importance score of the corresponding output; Screening target monitoring indicators for device monitoring from each candidate monitoring indicator according to the importance score.
2. The method according to claim 1, wherein, The latent state evolution characteristics include latent state mean, latent state variance, latent state instantaneous change rate, and latent state cumulative change amount; For the modeling of the latent state evolution characteristics, it includes: , wherein, represents the hidden state mean within the time interval , represents any time point between and ; represents any time point between and ; is the hidden state at time point is the infinitesimal time increment of the integration variable, representing the integration of the continuously changing process of time; , In the formula, represents the hidden state variance within the time interval ; , wherein, represents the instantaneous change rate of the hidden state at time , and represent the hidden state vectors at times and respectively, is a preset time interval for discretized derivative calculation; , In the formula, represents the cumulative change in the hidden state within the time interval , and represents the Euclidean norm.
3. The method according to claim 1 or 2, wherein, The screening of target monitoring indicators for device monitoring from each candidate monitoring indicator according to the importance score includes: Performing importance screening on each candidate monitoring indicator according to a preset importance threshold to obtain at least one corresponding potential monitoring indicator; Inputting each potential monitoring indicator and the corresponding comprehensive indicator characteristics into a dynamic DBSCAN model to determine at least one corresponding clustering cluster, calculating the clustering feature vector center corresponding to each clustering cluster, and selecting the potential monitoring indicator closest to the clustering feature vector center as the clustering representative monitoring indicator of the corresponding clustering cluster; Determining the target monitoring indicator according to the clustering representative monitoring indicators corresponding to each clustering cluster.
4. The method according to claim 3, wherein, The inputting each potential monitoring indicator and the corresponding comprehensive indicator characteristics into a dynamic DBSCAN model to determine at least one corresponding clustering cluster, calculating the clustering feature vector center corresponding to each clustering cluster, and selecting the potential monitoring indicator closest to the clustering feature vector center as the clustering representative monitoring indicator of the corresponding clustering cluster includes: The set of all potential monitoring indicators is denoted as , represents the total number of potential monitoring indicators screened by the importance threshold, and forms an index feature matrix by synthesizing the index comprehensive features of each potential monitoring indicator ; Calculating the neighborhood radius and the minimum number of samples in the neighborhood in a dynamic manner: , , , , In the formula, represents the th potential monitoring index, and respectively represent the comprehensive characteristics of the index and the index corresponding thereto, represents the set composed of the nearest indexes to the index represents the Euclidean distance between feature vectors, represents the average distance from the index to its nearest neighbor indexes, represents the standard deviation of the distances from the index to its nearest neighbor indexes, represents the number of nearest neighbor indexes, represents the floor function; is the neighborhood radius in the DBSCAN algorithm, representing the neighborhood range of an index point in the feature space; represents the adjustment coefficient for the tightness of the neighborhood radius, with a value range of [1.0, 2.0]; is the minimum number of samples in the neighborhood in the DBSCAN algorithm, used to determine the minimum number of indexes required in the neighborhood for a point to become a core point; represents the adjustment coefficient used to control the change speed of the minimum number of neighborhood samples with the total number of indexes, with a value range of [1.5, 3.0]; Perform DBSCAN clustering processing on the index feature matrix using the dynamically calculated neighborhood radius and the minimum number of samples in the neighborhood to obtain the corresponding set of clustering clusters , represents the total number of clustering clusters, denotes the th clustering cluster; For each clustering cluster, calculating the clustering feature vector center according to the mean of all comprehensive indicator characteristics within the clustering cluster: , In the formula, represents the th clustering cluster, represents the number of indicators in represents the center of the feature vector of represents each monitoring indicator in the comprehensive feature of the indicator; For each cluster, select the metric with the minimum distance from the center of the cluster feature vector as the representative monitoring metric for the cluster: , In the formula, represents the selected clustering representative monitoring index in , and represents the Euclidean distance between the comprehensive feature of the index in and the center of the clustering feature vector of .
5. A device monitoring metric screening system based on time series modeling for implementing the method according to any one of claims 1-4; the system includes: A data acquisition unit for acquiring device monitoring time series data, where the device monitoring time series data covers multiple types of candidate monitoring metrics and multiple types of auxiliary parameters, and the auxiliary parameters include any one of the following: service response latency, service alarm information, and network latency; An NCDE analysis unit for respectively defining the main input and extended input of the NCDE model based on various types of the candidate monitoring metrics and various types of the auxiliary parameters, such that the NCDE model maps continuously changing candidate monitoring metrics to a hidden space using a differential equation form in continuous time, and captures the prediction residuals and hidden state evolution characteristics of various types of the candidate monitoring metrics through continuous time modeling; A time-varying correlation analysis unit for calculating the correlation coefficient between the prediction residuals of various types of the candidate monitoring metrics and the system operating state within a continuous time window, and obtaining the time-varying correlation coefficients of the corresponding various types of the candidate monitoring metrics by weighted summation of the correlation coefficients of each time window; The system operating state is determined according to the service response latency and request error rate within the corresponding time window; An importance evaluation unit for fusing the prediction residuals, time-varying correlation coefficients, and hidden state evolution characteristics of each candidate monitoring metric to obtain the corresponding comprehensive metric characteristics, and evaluating the importance score corresponding to the comprehensive metric characteristics; A target metric screening unit for screening target monitoring metrics for device monitoring from each of the candidate monitoring metrics according to the importance score.
Citation Information
Patent Citations
Time sequence prediction method and device
CN116432834A
Data processing system for driving measurement calibration
CN116821729A
Power system dominant instability mode identification method based on deep learning
CN119066561A