Software performance real-time monitoring system and method fused with edge computing

By obtaining multi-dimensional data of edge nodes in distributed systems and Internet of Things environments, dynamically correcting and grouping, and training performance fluctuation prediction models, the problem of inaccurate evaluation in the existing technology is solved, real-time and accurate monitoring and abnormal identification of software performance is achieved, and system stability is ensured.

CN120256239AActive Publication Date: 2025-07-04QINGDAO RUBIKS CUBE INTERACTIVE SOFTWARE TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510140009.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-07-04
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

In distributed systems and IoT environments, it is difficult for the existing technology to comprehensively and in real time to reflect the impact of complex and variable workloads and environments on software performance, resulting in inadequate evaluation and inaccurate response to performance fluctuations in time, affecting the real-time and reliability of the system.

Method used

By obtaining multi-dimensional key data of edge nodes, dynamically selecting the benchmark node for correction, using community detection algorithms to group and train performance fluctuation prediction models, predict future performance values, and identify abnormal nodes.

Benefits of technology

It realizes a comprehensive and accurate evaluation of software performance, improves the system's foresight and real-time response capabilities, and ensures the stable operation of the software under various workloads.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256239A_ABST
    Figure CN120256239A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of edge computing in a distributed system and an internet of things environment, and discloses a software performance real-time monitoring system and method fused with edge computing, in particular to a technology for reflecting software performance by monitoring performance of edge nodes, which comprises the following steps: acquiring key data of each edge node, obtaining a preliminary performance evaluation vector of each edge node; dynamically selecting a reference node, and correcting the initial performance evaluation vector of each edge node to obtain a performance evaluation vector; grouping each edge node by using a community detection algorithm to obtain a comprehensive feature vector; using the comprehensive feature vector to train and obtain a performance fluctuation prediction model, and performing prediction based on the performance fluctuation prediction model to obtain a predicted performance value; and obtaining an actual performance monitoring result according to the predicted performance value, and judging to obtain an abnormal node according to the predicted performance value and the actual monitoring result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing in distributed systems and Internet of Things (IoT) environments. More specifically, the present invention relates to a real-time software performance monitoring system and method integrating edge computing. Background Art

[0002] In modern distributed systems and Internet of Things (IoT) environments, monitoring the performance of edge nodes to reflect software performance is crucial for ensuring the stability and efficient operation of applications. Existing performance evaluation methods usually rely on static thresholds or single-metric analysis, making it difficult to comprehensively and real-time reflect the impact of complex and changing workloads and environmental conditions on software performance.

[0003] Most existing methods use fixed thresholds to determine whether the node performance is normal. This method cannot adapt to dynamic workloads and environments, resulting in insensitive and inaccurate evaluation of software performance; traditional performance evaluation often based on a single performance metric (such as CPU utilization), ignoring the mutual influence between other key performance metrics (such as memory usage, network latency, etc.), and unable to comprehensively reflect the true running state of the software; the internal differences of different communities (i.e., a group of nodes with similar characteristics) are not fully considered, resulting in inaccurate performance evaluation and unable to accurately reflect the performance of the software under different workloads; many existing technologies have delays in identifying anomalies and cannot respond to performance fluctuations in a timely manner, affecting the real-time and reliability of the system, and further affecting the rapid diagnosis and repair of software performance. Summary of the Invention

[0004] To overcome the above defects of the prior art and achieve the above object, the present invention provides the following technical solutions: A real-time software performance monitoring system and method integrating edge computing, including:

[0005] Performance evaluation module: Obtain the key data of each edge node and calculate the preliminary performance evaluation vector of each edge node;

[0006] Evaluation correction module: Dynamically select reference nodes and use the reference performance metric vectors of the reference nodes to correct the preliminary performance evaluation vectors of each edge node to obtain performance evaluation vectors;

[0007] Feature construction module: Use community detection algorithms to group each edge node to obtain different communities, and obtain the comprehensive feature vector of each community according to the performance evaluation vector;

[0008] Model construction module: Use the comprehensive feature vector of each community to train and obtain a performance fluctuation prediction model, and predict the predicted performance values of each community in a future period based on the performance fluctuation prediction model;

[0009] Actual Performance Evaluation Module: According to the predicted performance values for a future period, after the actual period arrives, obtain the actual performance monitoring results. Through the predicted performance values and the actual monitoring results, calculate the expected difference to identify abnormal communities, and calculate the deviation values of each edge node within the abnormal community. Adjust the corresponding predicted performance values and actual performance values according to the deviation values to obtain abnormal nodes.

[0010] Furthermore, the method for obtaining the key data of each edge node and calculating the preliminary performance evaluation vector of each edge node includes:

[0011] Take each edge device as an edge node and use a compatible monitoring agent according to the operating system of the edge node;

[0012] Define the time length of a period as sd, and obtain the key data of each edge node in each period through the monitoring agent, including application response time data, CPU utilization data, average memory usage, throughput data, and average error rate;

[0013] Among them, the application response time data includes the average response time and the 95th percentile response time; the CPU utilization data includes the average utilization and the maximum CPU utilization; the throughput data includes the average throughput and the maximum throughput;

[0014] Calculate the representative indicators according to the key data, including the user experience index, resource utilization efficiency, and software stability index;

[0015] Among them, take the mean of the average response time and the 95th percentile response time within the t period as the user experience index;

[0016] Standardize the average utilization, maximum CPU utilization, and average memory usage within the t period, and then use the weighted average method to calculate the resource utilization efficiency;

[0017] Combine the average throughput, maximum throughput, and average error rate within the t period to calculate the software stability index , where JJTL represents the average throughput, MAXTTL represents the maximum throughput, JCWL represents the average error rate, and SSI represents the software stability index;

[0018] Horizontally concatenate the user experience index, resource utilization efficiency, and software stability index of the t period as the performance indicators of the edge node to form the preliminary performance evaluation vector of the edge node in the t period.

[0019] Furthermore, the method for dynamically selecting the reference node includes:

[0020] Collect the preliminary performance evaluation vectors of all edge nodes within bt time periods before the t time period to form a performance index data set;

[0021] Clean the data in the performance index data set and use the difference filling method to handle missing values;

[0022] According to the processed performance index data set, dynamically select reference nodes, including:

[0023] Define the decision variable as set jz i Indicates whether the i-th edge node is selected as a reference node, where jz i =1 indicates that the i-th edge node is selected, and jz i =0 indicates that the i-th edge node is not selected;

[0024] Define the objective function as minimizing the performance difference between all edge nodes and the selected reference nodes , where, Represents the j-th performance index of the i-th edge node, i represents the index of the edge node, I represents the number of edge nodes, j represents the index of the performance index, and J represents the number of performance indices, Represents the average value of the j-th performance index of all reference nodes, Represents the weight of the j-th performance index;

[0025] Based on the entropy concept in information theory, calculate the information entropy of each performance index , where, Represents the information entropy of the j-th performance index;

[0026] According to the information entropy of each performance index, calculate the weight ;

[0027] Define the constraint conditions as quantity limit, uniform geographical distribution, diverse hardware configuration, and time stability;

[0028] Among them, the quantity limit is defined as restricting the number of selected reference nodes to be less than or equal to jds, and the mathematical expression is: , where jds represents the maximum number of reference nodes allowed to be selected;

[0029] The uniform geographical distribution is defined as the reference nodes being evenly distributed in different geographical locations. Obtain the geographical location where each edge node is located. For each geographical location, there is at least one reference node;

[0030] The diverse hardware configuration is defined as the reference nodes having different hardware configurations. Obtain the hardware configuration of each edge node and classify the hardware configurations of all edge nodes. For each type of hardware configuration, ensure that there is at least one reference node;

[0031] The time stability is defined as calculating the standard deviation of the performance metrics of each edge node within bt time periods before the t time period, such that the standard deviation of the performance metrics of the selected reference nodes is less than or equal to the preset performance metric standard deviation threshold;

[0032] Based on the decision variables, objective function, and constraint conditions, a linear programming solver is used to solve for the optimal solution, which serves as the best combination of reference nodes.

[0033] Furthermore, the method for obtaining the performance evaluation vector includes:

[0034] According to all the selected reference nodes within the best combination of reference nodes, the mean of the preliminary performance evaluation vectors of all reference nodes in the t time period is taken as the reference performance metric vector;

[0035] Collect the historical performance data set, including the preliminary performance evaluation vectors of all edge nodes and reference nodes;

[0036] Use statistical analysis techniques to identify and process the historical performance data set, and identify and determine whether there are long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes;

[0037] If there are long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes, use the ARIMA model as the time series model;

[0038] If there are no long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes, use the exponential smoothing method as the time series model;

[0039] According to the selected time series model, use the historical performance data set to train the selected time series model, and output the preliminary performance evaluation prediction vector of edge nodes within the next time period;

[0040] Calculate the ratio of the preliminary performance evaluation prediction vector in the time period after the t time period to the reference performance metric vector in the t time period to obtain the adjustment factor;

[0041] For each time period t, calculate the performance ratio of edge nodes relative to reference nodes, and introduce the adjustment factor to correct the preliminary performance evaluation vector to obtain the performance evaluation vector , where represents the preliminary performance evaluation vector of the i-th edge node in the t time period, represents the performance evaluation vector, represents the reference performance metric vector in the t time period, represents the adjustment factor in the t time period, represents the performance ratio of edge nodes relative to reference nodes.

[0042] Further, the method of grouping each edge node using the community detection algorithm to obtain different communities and obtaining the comprehensive feature vector of each community according to the performance evaluation vector includes:

[0043] Read and collect the network topology information of each edge node, including IP address, subnet, routing path, bandwidth, and latency;

[0044] Construct a network graph through the network topology information of each edge node, calculate the similarity between each edge node according to the network graph, and form a similarity matrix;

[0045] According to the similarity matrix, use the community detection algorithm, and use the edge node performance evaluation vector as an additional load weight to group the edge nodes to obtain different communities, and obtain the comprehensive feature vector according to different communities.

[0046] Further, the method of constructing a network graph through the network topology information of each edge node, calculating the similarity between each edge node according to the network graph, and forming a similarity matrix includes:

[0047] Determine the connection relationship between nodes through the routing path of each edge node. For edge nodes with IP addresses belonging to the same subnet, add an edge with no edge weight between every two edge nodes;

[0048] For edge nodes with IP addresses not belonging to the same subnet, identify the pairs of edge nodes with connection paths according to the routing path, and add an edge with an edge weight to these pairs of edge nodes;

[0049] Define and obtain the edge weight through the bandwidth and latency , where represents the edge weight between the i-th edge node and the r-th edge node, represents the proportionality factor for adjusting the importance of bandwidth and latency, represents the bandwidth between the i-th edge node and the r-th edge node, represents the latency;

[0050] Based on all edge nodes and their corresponding edges, construct a network graph;

[0051] According to the network graph, calculate the similarity between any pair of edge nodes (i,r), and define that if edge node i and edge node r belong to the same subnet, use an unweighted similarity metric, and if i and r do not belong to the same subnet, use a weighted similarity metric;

[0052] ;

[0053] where Indicates the similarity between any pair of edge nodes i and r. Indicates the similarity of edge node pairs within the same subnet of. Indicates the neighbor set of edge node i. Indicates the neighbor set of edge node r. All nodes directly connected to edge node i in the network graph are used as the neighbor set of edge node i. Indicates the similarity of edge node pairs that do not belong to the same subnet. Indicates the edge weight of the edge between edge node i and edge node c. Indicates the edge weight of the edge between edge node r and edge node c;

[0054] Based on the similarity between all pairs of edge nodes, a similarity matrix SSJ is formed. Each position within the similarity matrix Indicates the similarity between edge node i and edge node r.

[0055] Furthermore, the method for obtaining the comprehensive feature vector includes:

[0056] Use the weighted summation formula to calculate all performance metrics within the performance evaluation vector to obtain the load weight;

[0057] Use the spectral clustering algorithm as the community detection algorithm. Introduce the load weight into the modularity function of the community detection algorithm, and define the grouping objective of the community detection algorithm as making nodes with high load more likely to be assigned to the same community;

[0058] The modularity function after introducing the load weight is ;

[0059] Among them, Indicates the modified modularity function, zzq indicates the total weight of all edges in the network graph, and respectively indicate the load weights of edge node i and edge node r, and respectively indicate the degrees of i and r. The number of edges connected to i is used as the degree of i, Indicates the indicator function for judging whether i and r belong to the same community, and respectively indicate the community labels of i and r within the community detection algorithm. If , it is determined that i and r belong to the same community, then , if , it is determined that i and r do not belong to the same community, then , Indicates whether there is an edge between i and r. If there is an edge between i and r, then , if there is no edge between i and r, then ;

[0060] Apply the modified modularity function and the similarity matrix SSJ to the community detection algorithm to group the edge nodes and obtain different communities, where each community contains a group of edge nodes that are closely connected in the network topology and have similar loads;

[0061] For each community, calculate the average value of the corresponding performance metrics in the performance evaluation vectors of each edge node to form the comprehensive feature vector of the community.

[0062] Furthermore, the training method of the performance fluctuation prediction model includes:

[0063] Step A1: Collect a sample set, including the comprehensive feature vector sequences of each community and the corresponding labels, where the label is the true performance value of the next time period corresponding to the comprehensive feature vector;

[0064] According to the comprehensive feature vector of each community, integrate the comprehensive feature vectors of any time period and the previous qc time periods in chronological order to form the comprehensive feature vector sequence of each community at any time period;

[0065] Step A2: Normalize the sample set, and use the interpolation method to handle missing values. Divide the processed sample set into a training set and a test set according to a ratio;

[0066] Step A3: Build a performance fluctuation prediction model, use the comprehensive feature vector sequence of each community as the input, use the predicted performance value of the community in the next time period as the output, and use the predicted performance of the community as the predicted performance value of all edge nodes in the community. The performance fluctuation prediction model is an LSTM model;

[0067] Step A4: Initialize the hyperparameters of the model, and use Bayesian optimization to tune the hyperparameters. Use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations, and select the optimal parameter combination;

[0068] Step A5: Use the optimal parameter combination as the initial parameters of the model, define Adam as the optimizer, and define the mean absolute error as the loss function for evaluating the prediction accuracy of the model , where represents the predicted performance value of the ve-th sample, represents the true performance value of the ve-th sample;

[0069] Step A6: Input the comprehensive feature vector sequence in the training set into the model, perform forward propagation, calculate the predicted performance value, and then use the loss function to calculate the loss between the predicted performance value and the true performance value. Update the model parameters through backpropagation, and repeat the forward propagation and backpropagation iteratively;

[0070] Step A7: For each iteration, use the R 2 score as the evaluation metric and calculate the R 2 score value on the validation set;

[0071] Based on the R 2 score value on the validation set, calculate the difference between the R 2 score value after the current iteration and the R 2 score value of the previous iteration, and denote it as the iteration difference;

[0072] Set an iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved;

[0073] If the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved;

[0074] If the performance of the model does not improve on the validation set in consecutive DC iterations, stop training and obtain the trained performance fluctuation prediction model.

[0075] Furthermore, the method for obtaining the abnormal nodes includes:

[0076] Step B1: According to the predicted performance values of each community in the next period predicted by the performance fluctuation prediction model, after the next period actually arrives, obtain the monitoring results of the actual performance values of each community through the real-time monitoring system;

[0077] Step B2: For each community, calculate the absolute difference between its predicted performance value and the actual performance value as the expected difference of the actual performance;

[0078] Set an expected difference threshold. If the expected difference is less than or equal to the expected difference threshold, it is determined that the actual performance of the community is within the normal range, and use the actual performance value of the community as the final performance evaluation value of each edge node in the community;

[0079] If the expected difference is greater than the expected difference threshold, it is determined that the actual performance value of the community deviates from the normal range, and mark the community as an abnormal community;

[0080] Step B3: For an abnormal community, calculate the standard deviation of each performance metric of the edge nodes in the abnormal community for each performance metric of the edge nodes in the abnormal community;

[0081] Combining the expected difference threshold of the community and the standard deviation, set a node threshold for each edge node in the abnormal community , where represents the node threshold, represents the expected difference threshold of the community, A constant term representing the adjustment threshold looseness, representing the standard deviation;

[0082] Step B4: For each abnormal community, collect the historical performance evaluation vectors of each edge node in the past PCD time periods, integrate the historical performance evaluation vectors of each edge node into an overall vector, and calculate the historical average performance;

[0083] Step B5: Use the Euclidean distance formula to calculate the historical average performance of each edge node and the actual performance value of the abnormal community, and obtain the performance deviation value of the edge node relative to the abnormal community;

[0084] Sum the predicted performance value and the actual performance value of the abnormal community with the performance deviation value respectively to obtain the predicted performance value and the actual performance value of each edge node in the abnormal community;

[0085] Step B6: Calculate the absolute difference between the predicted performance value and the actual performance value of each edge node in the abnormal community respectively, and compare it with the node threshold. If it is greater than the node threshold, it is considered that the performance of the edge node deviates from the normal range and is marked as an abnormal node.

[0086] Furthermore, a real-time software performance monitoring method integrating edge computing is characterized by including:

[0087] Step S1: Obtain the key data of each edge node and calculate the preliminary performance evaluation vector of each edge node;

[0088] Step S2: Dynamically select benchmark nodes, and use the benchmark performance index vectors of the benchmark nodes to correct the preliminary performance evaluation vectors of each edge node to obtain the performance evaluation vectors;

[0089] Step S3: Use the community detection algorithm to group each edge node to obtain different communities, and obtain the comprehensive feature vector of each community according to the performance evaluation vector;

[0090] Step S4: Use the comprehensive feature vector of each community to train and obtain a performance fluctuation prediction model, and predict the predicted performance value of each community in the next time period based on the performance fluctuation prediction model;

[0091] Step S5: According to the predicted performance value in the next time period, after the actual time period arrives, obtain the actual performance monitoring result. By comparing the predicted performance value with the actual monitoring result, calculate the expected difference to identify the abnormal community. For the abnormal community, adjust its predicted and actual performance values by calculating the deviation value of the edge node, and determine the abnormal node based on the set node threshold.

[0092] Technical effects and advantages of a software performance real-time monitoring system and method integrating edge computing according to the present invention:

[0093] The present invention accurately reflects software performance by using performance monitoring of edge nodes: First, collect multi-dimensional key data (such as CPU utilization rate, memory usage, disk I / O, network bandwidth, etc.), calculate the preliminary performance evaluation vector of each edge node to comprehensively reflect the actual performance of the node and avoid the limitation of a single indicator; Second, dynamically select benchmark nodes, and correct the preliminary performance evaluation vectors of each edge node based on their benchmark performance indicator vectors to improve the flexibility and accuracy of evaluation and reduce errors caused by environmental changes; Then, use the community detection algorithm to group edge nodes, construct the comprehensive feature vector of each community, realize refined management and adapt to different workload patterns, and more accurately reflect the performance of the software in different scenarios; Next, train a machine learning model based on the community comprehensive feature vector to predict the performance value in the future period, identify potential problems in advance, optimize the predictability and initiative of the system, and reduce the probability of unexpected failures; Finally, compare the predicted value with the actual monitoring result, calculate the expected difference to identify abnormal communities, and adjust and determine abnormal nodes through the overall deviation value to ensure immediate response and high-precision anomaly detection, and timely solve problems affecting software performance. This method not only improves the comprehensiveness and accuracy of performance evaluation, but also enhances the predictability and real-time response ability of the system, ensuring the stable operation of the software under various workloads. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] Figure 1 Schematic diagram of a software performance real-time monitoring system integrating edge computing according to the present invention;

[0095] Figure 2 Schematic diagram of a software performance real-time monitoring method integrating edge computing according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0096] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0097] Embodiment 1

[0098] Please refer to Figure 1 as shown, the software performance real-time monitoring system integrating edge computing described in this embodiment includes:

[0099] Performance Evaluation Module: Obtain the key data of each edge node and calculate the preliminary performance evaluation vector of each edge node (providing basic data for subsequent calibration and analysis);

[0100] Evaluation and Calibration Module: Dynamically select a reference node, and use the reference performance index vector of the reference node to calibrate the preliminary performance evaluation vectors of each edge node to obtain the performance evaluation vectors (ensuring the accuracy and consistency of performance evaluation and eliminating errors caused by environmental differences);

[0101] Feature Construction Module: Used to group each edge node using the community detection algorithm to obtain different communities, and based on the performance evaluation vectors, obtain the comprehensive feature vectors of each community (capturing the spatial correlation between nodes and enhancing the effectiveness of feature representation);

[0102] Model Construction Module: Train and obtain a performance fluctuation prediction model using the comprehensive feature vectors of each community, and predict the predicted performance values of each community in a future period based on the performance fluctuation prediction model (realizing the prediction of future performance and discovering potential problems in advance);

[0103] Actual Performance Evaluation Module: According to the predicted performance values in a future period, after the actual period arrives, obtain the actual performance monitoring results. By comparing the predicted performance values with the actual monitoring results, calculate the expected difference to identify abnormal communities. For abnormal communities, adjust their predicted and actual performance values by calculating the deviation values of the edge nodes, and determine the abnormal nodes based on the set node thresholds;

[0104] It should be noted that edge nodes, as key components of the software system, software systems in the edge computing environment usually rely on multiple edge nodes to execute tasks. Therefore, the performance of edge nodes directly reflects the operating conditions of the entire software system. If some edge nodes are nodes on the critical path of the software system, their performance changes will have a significant impact on the performance of the entire system. By monitoring the performance of all edge nodes, an overall performance view can be obtained. This comprehensive evaluation can help identify system bottlenecks and optimization points. For distributed applications, monitoring the performance of edge nodes helps understand the health status and quality of service of the entire application;

[0105] Through this method, the performance status of the entire software system can be effectively reflected. Especially in the edge computing environment, by monitoring and evaluating the performance of edge nodes, the comprehensive monitoring of software performance can be indirectly achieved to ensure the stability and efficient operation of the system;

[0106] The methods for obtaining the key data of each edge node and calculating the preliminary performance evaluation vector of each edge node include:

[0107] It should be noted that traditional performance evaluations often rely on a single performance metric (such as CPU utilization), ignoring the mutual influence among other key performance metrics (such as memory usage, network latency, etc.), and are unable to comprehensively reflect the true running state of the software. Based on this, the solution of the present invention is as follows:

[0108] Each edge device is regarded as an edge node, and a compatible monitoring agent is used according to the operating system of the edge node (such as Linux, Windows, RTOS, etc.) (ensuring that the agent supports multiple data collection methods, such as SNMP (Simple Network Management Protocol), WMI (Windows Management Instrumentation), syslog (system log), JMX (Java Management Extensions), etc., to adapt to different operating systems and application environments);

[0109] Define the time length of a period as sd, and obtain the key data of each edge node in each period through the monitoring agent, including application response time data (the time from when the user issues a request to receiving a complete response), CPU utilization data (the percentage of CPU occupied during the software operation), average memory usage (the amount of memory used during the software operation), throughput data (the number of requests processed per unit time), and average error rate (the proportion of failed requests);

[0110] Among them, the application response time data includes the average response time (used to reflect the overall performance) and the 95th percentile response time (indicating that within the t period, the response time of 95% of the requests does not exceed this value, that is, only 5% of the requests will have a response time exceeding this value. This is a statistic used to describe the distribution of software response times. The 95th percentile response time covers 95% of the requests, which means it can well represent the experience of most users. Especially in high-concurrency scenarios, it can better reflect the stability and consistency of the system. The reason for using the 95th percentile response time is that the average response time may be affected by a very small number of extremely slow requests, resulting in its inability to accurately reflect the actual experience of most users, while the 95th percentile response time can better exclude these extreme values and provide a true feedback close to the user experience); the CPU utilization data includes the average utilization (used to reflect the overall computing resource usage within the t period) and the maximum CPU utilization (used to identify high-load moments and help discover possible overload problems); the throughput data includes the average throughput (used to evaluate the processing efficiency within the t period) and the maximum throughput (used to identify peak periods and help evaluate the ultimate processing capacity of the software);

[0111] Representative metrics are calculated based on key data, including the user experience index (used to measure the overall user experience), resource utilization efficiency (used to measure the resource utilization of the system), and software stability index (used to measure the stability and reliability of the software);

[0112] Among them, the mean of the average response time and the 95th percentile response time within the t period is taken as the user experience index;

[0113] The average utilization, maximum CPU utilization, and average memory usage within the t period are standardized, and then the resource utilization efficiency is calculated using the weighted average method (the weighted average after standardization can eliminate the dimension difference, enabling different metrics to be compared on the same scale, making the results more intuitive and easy to interpret to reflect their relative importance);

[0114] Combining the average throughput, maximum throughput, and average error rate within the t period, the software stability index is calculated , where JJTL represents the average throughput, MAXTTL represents the maximum throughput, JCWL represents the average error rate, and SSI represents the software stability index;

[0115] The user experience index, resource utilization efficiency, and software stability index of the t period are horizontally concatenated as the performance metrics of the edge node to form the preliminary performance evaluation vector of the edge node in the t period;

[0116] It should be noted that the key data of each edge node in each period are obtained, including application response time, CPU utilization, average memory usage, throughput, and average error rate. These multi-dimensional data are calculated to obtain the user experience index, resource utilization efficiency, and software stability index, forming the preliminary performance evaluation vector of the edge node. Compared with single-metric analysis, multi-dimensional performance evaluation can more comprehensively reflect the true operating state of the software, avoid ignoring the impact of key performance metrics, and provide a more accurate performance evaluation;

[0117] The methods for dynamically selecting reference nodes include:

[0118] It should be noted that in existing edge computing and distributed system management, most existing methods use fixed thresholds to judge whether the node performance is normal. This method cannot adapt to dynamically changing workloads and environments, resulting in insufficient sensitivity and accuracy in evaluating software performance. Based on this, to solve the static threshold problem, the solution of the present invention is:

[0119] Collect the preliminary performance evaluation vectors of all edge nodes within bt periods (such as the past week or month) before the t period to form a performance metric dataset;

[0120] Clean the performance metric dataset and handle missing values using the difference filling method;

[0121] Based on the processed performance metric dataset, dynamically select reference nodes, including:

[0122] Define the decision variable as set jz i indicating whether the i-th edge node is selected as a reference node, where jz i =1 means the i-th edge node is selected, and jz i =0 means the i-th edge node is not selected;

[0123] Define the objective function as minimizing the performance difference between all edge nodes and the selected reference nodes , where represents the j-th performance metric of the i-th edge node, i represents the index of the edge node, I represents the number of edge nodes, j represents the index of the performance metric, J represents the number of performance metrics, represents the average value of the j-th performance metric of all reference nodes, represents the weight of the j-th performance metric;

[0124] Based on the entropy concept in information theory, calculate the information entropy of each performance metric , where represents the information entropy of the j-th performance metric;

[0125] Based on the information entropy of each performance metric, calculate the weight (By calculating the weights of each index based on the entropy concept in information theory, it can ensure that the selected reference nodes are not only stable in performance but also have sufficient representativeness, thus improving the accuracy of performance evaluation and calibration, avoiding the interference of human factors, reducing the influence of personal experience and preference on the results compared with methods such as human judgment or analytic hierarchy process, making the weights more fair and scientific, being able to handle multiple performance metrics, and automatically adjusting the weight of each metric according to the degree of data variation without the need to preset fixed weight values);

[0126] Define the constraint conditions as quantity limitation, uniform geographical distribution, hardware configuration diversity, and time stability;

[0127] Among them, the quantity limitation is to limit the number of selected reference nodes to be less than or equal to jds (control the number of reference nodes to avoid excessive or insufficient selection and ensure the operability and management feasibility of the results), and the mathematical expression is: , where jds represents the maximum number of reference nodes allowed to be selected;

[0128] Geographical distribution is uniform. To ensure that the reference nodes are distributed in different geographical locations (such as data centers 1, 2, and 3 to cover a wider area), obtain the geographical location where each edge node is located. For each geographical location, ensure that there is at least one reference node (to increase the comprehensiveness of the evaluation, ensure that the performance differences under different hardware configurations can be fully considered, and avoid biases caused by hardware differences), that is , geographical location, and , where represents the set of nodes located at geographical location gd, represents a binary variable indicating whether there is at least one reference node at geographical location gd, = 1 means yes, = 0 means no, and xzl represents the minimum number of geographical locations required to be covered;

[0129] Hardware configuration diversity. To ensure that the reference nodes have different hardware configurations (such as server types 1, 2, and 3 to reflect the performance under different hardware conditions), obtain the hardware configuration of each edge node and classify the hardware configurations of all edge nodes. For each hardware configuration, ensure that there is at least one reference node (to increase the comprehensiveness of the evaluation, ensure that the performance differences under different hardware configurations can be fully considered, and avoid biases caused by hardware differences), that is , hardware configuration, , where represents the set of edge nodes with hardware configuration hv, represents a binary variable indicating whether there is at least one reference node for hardware configuration hv, = 1 means yes, = 0 means no, and zxy represents the minimum number of hardware configurations required to be covered;

[0130] Time stability. To calculate the standard deviation of the performance metrics of each edge node within bt time periods before time t, ensure that the standard deviation of the performance metrics of the selected reference nodes is less than or equal to the preset performance metric standard deviation threshold (the performance metric standard deviation threshold is a small value used to measure the performance stability of edge nodes, ensure that the selected reference nodes have relatively stable performance within the past BT time periods, avoid selecting nodes with large fluctuations, improve the reliability of the reference nodes, and ensure that they can represent the stable state of the system rather than temporary peaks or troughs), that is , , represents the standard deviation of the performance metrics of the i-th edge node, used to measure the performance fluctuation of this node, yz represents the set performance metric standard deviation threshold, and i represents the index of the candidate node, that is, all edge nodes are used as candidate nodes so that all edge nodes have the opportunity to be selected as reference nodes;

[0131] Based on decision variables, objective functions, and constraint conditions, use a linear programming solver (such as GLPK, CPLEX, Gurobi) to solve and obtain the optimal solution as the best combination of reference nodes;

[0132] The method for obtaining the performance evaluation vector includes:

[0133] According to all the selected reference nodes in the best combination of reference nodes, take the mean of the preliminary performance evaluation vectors of all reference nodes at time period t as the reference performance index vector;

[0134] Collect historical performance datasets (such as historical data within the past year), including the preliminary performance evaluation vectors of all edge nodes and reference nodes;

[0135] Use statistical analysis techniques (which are used to systematically collect, organize, analyze, and interpret historical data to identify patterns, trends, and relationships in the data. In this method, statistical analysis techniques are mainly used to process the historical performance data of edge nodes and reference nodes to identify whether there are long-term trends and periodic fluctuations) to identify and process the historical performance datasets, and identify and judge whether there are long-term trends and periodic fluctuations in the performance indicators of edge nodes and reference nodes;

[0136] If there are long-term trends and periodic fluctuations in the performance indicators of edge nodes and reference nodes, use the ARIMA model as the time series model;

[0137] If there are no long-term trends and periodic fluctuations in the performance indicators of edge nodes and reference nodes, use the exponential smoothing method as the time series model;

[0138] According to the selected time series model, use the historical performance datasets to train the selected time series model and output the preliminary performance evaluation prediction vector of edge nodes within the next time period;

[0139] Calculate the ratio of the preliminary performance evaluation prediction vector of the time period after time period t to the reference performance index vector of time period t to obtain the adjustment factor;

[0140] For each time period t, calculate the performance ratio of edge nodes relative to reference nodes, introduce the adjustment factor, and correct the preliminary performance evaluation vector to obtain the performance evaluation vector , where represents the preliminary performance evaluation vector of the i-th edge node at time period t, represents the performance evaluation vector, represents the reference performance index vector at time period t, represents the adjustment factor at time period t;

[0141] It should be noted that by collecting the preliminary performance evaluation vectors of all edge nodes within the bt time periods before the t time period, a performance index data set is formed, and a linear programming solver is used to dynamically select the reference nodes. This ensures that the selection of reference nodes can adapt to changing workloads and environmental conditions. Using the reference performance index vectors of the selected reference nodes to correct the preliminary performance evaluation vectors of each edge node, the final performance evaluation vectors are obtained. This dynamic adjustment mechanism improves the sensitivity and accuracy of the evaluation results; compared with a fixed threshold, the dynamic reference node selection and performance evaluation correction mechanism of the present invention can better reflect the changes in real-time workloads and improve the flexibility and accuracy of performance evaluation;

[0142] Specifically, by collecting the preliminary performance evaluation vectors of all edge nodes within bt time periods before the t time period, a performance index data set is formed, and a linear programming solver is used to dynamically select the reference nodes to ensure that the selection process can adapt to changing workloads and environmental conditions;

[0143] Detailed constraints are defined, including quantity limitations, uniform geographical distribution, diverse hardware configurations, and time stability. These constraints ensure the representativeness of the selected reference nodes and improve the comprehensiveness and accuracy of performance evaluation;

[0144] The objective function is defined as minimizing the performance differences between all edge nodes and the selected reference nodes to ensure that the selection of reference nodes can maximize the reflection of the actual performance situation and improve the accuracy of the evaluation results;

[0145] The linear programming solver is used for solving, and the optimal solution is obtained as the best combination of reference nodes to ensure that the solving process is efficient and the result is optimal;

[0146] The method of using the community detection algorithm to group each edge node to obtain different communities and obtaining the comprehensive feature vector of each community according to the performance evaluation vector includes:

[0147] It should be noted that most current existing methods fail to fully consider the internal differences of different communities (i.e., a group of nodes with similar characteristics), resulting in inaccurate performance evaluation and inability to accurately reflect the performance of software under different workloads; based on this, the solution of the present invention is:

[0148] Read and collect the network topology information of each edge node, including IP address, subnet, routing path, bandwidth, and latency;

[0149] Construct a network graph through the network topology information of each edge node, calculate the similarity between each edge node according to the network graph, and form a similarity matrix;

[0150] According to the similarity matrix, using community detection algorithms (such as the Louvain algorithm, Girvan - Newman algorithm. For example, assume we have multiple edge nodes distributed in different subnets. Using the Louvain algorithm to identify the community structure in the network, first construct a network graph, then run the Louvain algorithm, and finally obtain several communities. Each community contains a group of nodes that are closely connected in the network topology. The community detection algorithm aims to divide nodes into several communities with close internal connections and similar performance by analyzing the connection relationships and similarities between nodes in the network graph, combining the load weights, and using methods such as spectral clustering, so as to achieve refined management and efficient performance evaluation of complex networks. This method not only considers the topological structure of nodes but also incorporates performance metrics, ensuring a high degree of consistency in network and load among the nodes within each community, providing a solid foundation for subsequent detection), and using the edge node performance evaluation vector as an additional load weight to group the edge nodes (so that edge nodes with high load are assigned to the same group), obtaining different communities, and getting the comprehensive feature vector according to different communities;

[0151] It should be noted that using the community detection algorithm to group the edge nodes, obtaining different communities, and getting the comprehensive feature vector of each community according to the performance evaluation vector. The community detection algorithm combines network topology information and performance evaluation vectors to ensure that the nodes within each community are closely connected in the network topology and have similar loads. Compared with the method that does not consider the internal differences of communities, the community detection and grouping strategy of this method improves the accuracy of performance evaluation, can more accurately reflect the performance of the software under different workloads, and realizes refined management;

[0152] The method of constructing a network graph based on the network topology information of each edge node and calculating the similarity between each edge node according to the network graph to form a similarity matrix includes:

[0153] Determine the connection relationship between nodes through the routing path of each edge node. For edge nodes whose IP addresses belong to the same subnet, add an edge without edge weight between every two edge nodes (if the IP addresses of two nodes belong to the same subnet (i.e., the prefix is the same), they can usually communicate directly through a layer - 2 switch without passing through a router, indicating that there is a direct connection relationship between them);

[0154] For edge nodes whose IP addresses do not belong to the same subnet, identify the pairs of edge nodes with connection paths according to the routing path, and add an edge with edge weight to these pairs of edge nodes (if the IP addresses of two nodes do not belong to the same subnet, their communication needs to pass through a router. At this time, the connection path between them can be determined according to the routing table or other information);

[0155] It should be noted that existing methods usually can only identify simple direct connection relationships and cannot accurately reflect complex network topologies. By distinguishing the connection methods of the same subnet and different subnets, the connection relationships between nodes can be more accurately identified, especially for indirect connections across subnets, ensuring the accuracy of the network diagram;

[0156] Define and obtain edge weights through bandwidth and latency , where represents the edge weight between the i-th edge node and the r-th edge node, represents the scaling factor for adjusting the importance of bandwidth and latency (freely adjusted according to actual needs, and 0 < < 1), represents the bandwidth between the i-th edge node and the r-th edge node, represents the latency;

[0157] It should be noted that most existing methods ignore the influence of key performance indicators such as bandwidth and latency, resulting in inaccurate similarity calculations. By introducing bandwidth and latency as components of edge weights and allowing the scaling factor to be adjusted according to actual needs , the similarity calculation is made closer to the actual network performance, improving the accuracy of the evaluation;

[0158] Construct a network diagram based on all edge nodes and corresponding edges;

[0159] It should be noted that when existing methods calculate the similarity matrix, the efficiency is low and errors are prone to occur. By using an efficient algorithm based on the network diagram, quickly calculate the similarity between all node pairs, form an accurate similarity matrix, and improve the calculation efficiency and the reliability of the results;

[0160] According to the network diagram, calculate the similarity between any edge node pair (i, r), and define that if edge node i and edge node r belong to the same subnet, an unweighted similarity metric is used, and if i and r do not belong to the same subnet, a weighted similarity metric is used;

[0161] ;

[0162] where represents the similarity between any edge node pair i and r, represents the similarity of edge node pairs within the same subnet of, represents the neighbor set of edge node i, represents the neighbor set of edge node r, and all nodes directly connected to edge node i in the network diagram are used as the neighbor set of edge node i, represents the similarity of edge node pairs that do not belong to the same subnet, represents the edge weight of the edge between edge node i and edge node c, represents the edge weight of the edge between edge node r and edge node c;

[0163] Based on the similarity between all pairs of edge nodes, a similarity matrix SSJ is formed, and each position in the similarity matrix represents the similarity between edge node i and edge node r;

[0164] It should be noted that traditional methods often use a single similarity measurement method and cannot adapt to different network environments. The segmented similarity measurement method is adopted, and unweighted or weighted similarity measurement is selected according to different subnet situations, ensuring the flexibility and adaptability of similarity calculation;

[0165] When constructing a network diagram, existing methods often ignore some important network characteristics, such as bandwidth and latency. By comprehensively considering the connection relationship and performance metrics between nodes, a more comprehensive and accurate network diagram is constructed, providing a solid foundation for subsequent community detection and performance evaluation;

[0166] This method constructs a network diagram through network topology information, calculates the similarity between each edge node, and forms a similarity matrix, solving the problems in the prior art such as inaccurate recognition of connection relationships, single definition of edge weights, inflexible similarity measurement, incomplete construction of network diagrams, and low efficiency of forming similarity matrices. Specifically:

[0167] Refined connection relationship recognition improves the accuracy of the network diagram. Dynamic edge weight definition makes similarity calculation closer to the actual network performance. Flexible similarity measurement enhances the adaptability and flexibility of the method. Comprehensive construction of the network diagram provides a reliable foundation for subsequent analysis. Efficient formation of the similarity matrix improves the calculation efficiency and reliability of the results;

[0168] The method for obtaining the comprehensive feature vector includes:

[0169] Use the weighted summation formula to calculate all performance metrics in the performance evaluation vector to obtain the comprehensive value of the performance evaluation vector as the load weight (this step ensures that nodes with high load have higher weights in subsequent analysis);

[0170] It should be noted that most existing methods ignore the differences in node loads during community detection, resulting in inaccurate grouping. By introducing load weights and considering the differences in node loads in the modularity function, nodes with high load are more likely to be assigned to the same community, improving the rationality and accuracy of grouping;

[0171] Use the spectral clustering algorithm as the community detection algorithm, introduce the load weight into the modularity function of the community detection algorithm, and define the grouping objective of the community detection algorithm as making the nodes with high load more likely to be assigned to the same community;

[0172] It should be noted that traditional community detection methods are only based on the network topology structure and cannot fully reflect the actual performance of nodes. Combining the performance evaluation vector and the load weight enhances the accuracy of community detection, ensuring that the nodes within each community are not only closely connected in terms of network topology but also have similarity in performance;

[0173] The modularity function after introducing the load weight is ;

[0174] Among them, represents the modified modularity function, zzq represents the total weight of all edges in the network graph, and respectively represent the load weights of edge node i and edge node r, and respectively represent the degrees of edge node i and edge node r. The number of edges connected to edge node i is used as the degree of edge node i, represents the indicator function for judging whether edge node i and edge node r belong to the same community, and respectively represent the community labels of edge node i and edge node r within the community detection algorithm. If , it is determined that edge node i and edge node r belong to the same community, then , if , it is determined that edge node i and edge node r do not belong to the same community, then , represents whether there is an edge between edge node i and edge node r. If there is an edge between edge node i and edge node r, then , if there is no edge between edge node i and edge node r, then ;

[0175] It should be noted that the existing modularity function usually only considers the connection relationship between nodes and ignores the performance differences between nodes. The modified modularity function not only considers the connection relationship between nodes but also introduces the load weight, optimizing the objective function of community division and making the grouping more reasonable and effective;

[0176] Apply the modified modularity function and the similarity matrix SSJ to the community detection algorithm to group the edge nodes and obtain different communities. Each community contains a group of nodes that are closely connected in network topology and have similar loads;

[0177] For each community, calculate the average value of the corresponding performance metrics in the performance evaluation vectors of each edge node to form the comprehensive feature vector of the community (this step ensures that the comprehensive feature vector of each community can represent the overall performance characteristics of the community);

[0178] It should be noted that the traditional method is less efficient in calculating the comprehensive feature vector and is difficult to process large-scale data sets. By efficiently calculating the comprehensive value of the performance evaluation vector and combining with the spectral clustering algorithm, the calculation efficiency is improved, and large-scale data sets can be processed in a relatively short time to generate accurate comprehensive feature vectors. The comprehensive feature vectors generated by the existing methods often cannot comprehensively reflect the characteristics of the community. By calculating the average value of the performance evaluation vectors of the edge nodes within each community, the generated comprehensive feature vector can comprehensively reflect the overall performance characteristics of the community, providing a reliable basis for subsequent analysis;

[0179] The training method of the performance fluctuation prediction model includes:

[0180] A1: Collect a sample set, including the comprehensive feature vector sequence of each community and the corresponding labels, where the label is the true performance value of the next time period corresponding to the comprehensive feature vector;

[0181] According to the comprehensive feature vector of each community, integrate the comprehensive feature vectors of any time period and the previous qc time periods in chronological order to form the comprehensive feature vector sequence of each community at any time period;

[0182] A2: Normalize the sample set and use the interpolation method to process the missing values, and divide the processed sample set into a training set and a test set according to a certain proportion;

[0183] A3: Build a performance fluctuation prediction model, use the comprehensive feature vector sequence of each community as the input, use the predicted performance value of the community in the next time period as the output, and use the predicted performance of the community as the predicted performance value of all edge nodes within the community. The performance fluctuation prediction model is an LSTM model;

[0184] A4: Initialize the hyperparameters of the model, use Bayesian optimization to tune the hyperparameters, use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations, and select the optimal parameter combination;

[0185] A5: Use the optimal parameter combination as the initial parameters of the model, define Adam as the optimizer, and define the mean absolute error as the loss function for evaluating the prediction accuracy of the model , where, represents the predicted performance value of the ve-th sample, represents the true performance value of the ve-th sample;

[0186] A6: Input the sequence of comprehensive feature vectors in the training set into the model for forward propagation to calculate the predicted performance value. Then, use the loss function to calculate the loss between the predicted performance value and the true performance value, and update the model parameters through backpropagation. Repeat the forward propagation and backpropagation iteratively;

[0187] A7: For each iteration, use the 2 R-score as the evaluation metric to calculate the 2 R-score value on the validation set;

[0188] According to the 2 R-score value on the validation set, calculate the difference between the 2 R-score value after the current iteration and the 2 R-score value of the previous iteration, denoted as the iteration difference;

[0189] Set an iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved;

[0190] If the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved;

[0191] If the performance of the model does not improve on the validation set in consecutive DC iterations, stop the training to obtain the trained performance fluctuation prediction model;

[0192] The method for obtaining the abnormal nodes includes:

[0193] B1: According to the predicted performance values of each community in the next period predicted by the performance fluctuation prediction model, after the next period actually arrives, obtain the monitoring results of the actual performance values of each community through the real-time monitoring system;

[0194] B2: For each community, calculate the absolute difference between its predicted performance value and the actual performance value as the expected difference of the actual performance;

[0195] Set an expected difference threshold (set by professionals in the industry based on experience). If the expected difference is less than or equal to the expected difference threshold, it is determined that the actual performance of the community is within the normal range, and use the actual performance value of the community as the final performance evaluation value of each edge node in the community;

[0196] If the expected difference is greater than the expected difference threshold, it is determined that the actual performance value of the community deviates from the normal range, and mark the community as an abnormal community;

[0197] B3: For an abnormal community, calculate the standard deviation of each performance metric of the edge nodes in the abnormal community for all the edge nodes in the abnormal community;

[0198] Set the node threshold for each edge node within the abnormal community by combining the expected difference threshold and the standard deviation of the community , where represents the node threshold, represents the expected difference threshold of the community, represents a constant term for adjusting the threshold looseness (such as 1 or 2), represents the standard deviation;

[0199] B4: For each abnormal community, collect the historical performance evaluation vectors of each edge node in the past PCD time periods, integrate the historical performance evaluation vectors of each edge node into an overall vector, and calculate the historical average performance;

[0200] B5: Use the Euclidean distance formula to calculate the historical average performance of each edge node and the actual performance value of the abnormal community, and obtain the performance deviation value of the edge node relative to the abnormal community;

[0201] Sum the predicted performance value and the actual performance value of the abnormal community with the performance deviation value respectively to obtain the predicted performance value and the actual performance value of each edge node within the abnormal community;

[0202] B6: Calculate the absolute difference between the predicted performance value and the actual performance value of each edge node within the abnormal community respectively, and compare it with the node threshold. If it is greater than the node threshold, it is considered that the performance of the edge node deviates from the normal range and is marked as an abnormal node;

[0203] If there are abnormal nodes, it is determined that the software performance is abnormal.

[0204] In this embodiment, the software performance is accurately reflected by using the performance monitoring of edge nodes. First, multi-dimensional key data (such as CPU utilization, memory usage, disk I / O, network bandwidth, etc.) is collected, and the preliminary performance evaluation vector of each edge node is calculated to comprehensively reflect the actual performance of the node and avoid the limitations of a single indicator. Second, benchmark nodes are dynamically selected, and the preliminary performance evaluation vectors of each edge node are corrected based on the benchmark performance indicator vectors of the benchmark nodes to improve the flexibility and accuracy of the evaluation and reduce the errors caused by environmental changes. Then, the community detection algorithm is used to group the edge nodes, and the comprehensive feature vector of each community is constructed to achieve refined management and adapt to different workload patterns, and more accurately reflect the performance of the software in different scenarios. Next, a machine learning model is trained based on the comprehensive feature vector of the community to predict the performance value in the future period, identify potential problems in advance, optimize the predictability and initiative of the system, and reduce the probability of unexpected failures. Finally, the predicted value is compared with the actual monitoring result, the expected difference is calculated to identify abnormal communities, and the abnormal nodes are adjusted and determined through the overall deviation value to ensure immediate response and high-precision anomaly detection, and timely solve the problems affecting software performance. This method not only improves the comprehensiveness and accuracy of performance evaluation, but also enhances the predictability and real-time response ability of the system, ensuring the stable operation of the software under various workloads.

[0205] Embodiment 2

[0206] Please refer to Figure 2 As shown, the parts not described in detail in this embodiment can be found in the description of Embodiment 1. A real-time software performance monitoring method integrating edge computing is provided, including:

[0207] S1: Obtain the key data of each edge node, and calculate the preliminary performance evaluation vector of each edge node;

[0208] S2: Dynamically select benchmark nodes, and use the benchmark performance indicator vectors of the benchmark nodes to correct the preliminary performance evaluation vectors of each edge node to obtain the performance evaluation vectors;

[0209] S3: Use the community detection algorithm to group each edge node to obtain different communities, and obtain the comprehensive feature vector of each community according to the performance evaluation vector;

[0210] S4: Use the comprehensive feature vector of each community to train and obtain a performance fluctuation prediction model, and predict the predicted performance value of each community in a future period based on the performance fluctuation prediction model;

[0211] S5: According to the predicted performance value for a future period, after the actual period arrives, obtain the actual performance monitoring result. By comparing the predicted performance value with the actual monitoring result, calculate the expected difference to identify abnormal communities. For abnormal communities, adjust their predicted and actual performance values by calculating the deviation value of the edge nodes, and determine the abnormal nodes based on the set node threshold.

[0212] Embodiment III

[0213] This embodiment publicly provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the operation mode of the above-provided real-time software performance monitoring method integrating edge computing.

[0214] Since the electronic device introduced in this embodiment is the electronic device adopted for implementing the real-time software performance monitoring method integrating edge computing in the embodiments of the present application, based on the real-time software performance monitoring method integrating edge computing introduced in the embodiments of the present application, those skilled in the art can understand the specific implementation manners and various variations of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of the present application will not be described in detail here. As long as those skilled in the art implement the electronic device adopted for the real-time software performance monitoring method integrating edge computing in the embodiments of the present application, it falls within the scope of protection of the present application.

[0215] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.

[0216] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for ordinary technical users in the technical field, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A real-time software performance monitoring system integrating edge computing, characterized in that, Including: Performance evaluation module: used to obtain the key data of each edge node and calculate the preliminary performance evaluation vector of each edge node; Evaluation correction module: used to dynamically select reference nodes and use the reference performance metric vectors of the reference nodes to correct the preliminary performance evaluation vectors of each edge node to obtain performance evaluation vectors; Feature construction module: uses community detection algorithms to group each edge node to obtain different communities, and based on the performance evaluation vectors, obtains the comprehensive feature vectors of each community; Model construction module: trains and obtains a performance fluctuation prediction model using the comprehensive feature vectors of each community, and predicts the predicted performance values of each community in a future period based on the performance fluctuation prediction model; Actual performance evaluation module: according to the predicted performance values in a future period, after the actual period arrives, obtains the actual performance monitoring results, calculates the expected difference through the predicted performance values and the actual monitoring results to identify abnormal communities, calculates the deviation values of each edge node in the abnormal communities, and adjusts the corresponding predicted performance values and actual performance values according to the deviation values, and then obtains abnormal nodes.

2. The real-time software performance monitoring system integrating edge computing according to claim 1, characterized in that, The method of obtaining the key data of each edge node and calculating the preliminary performance evaluation vector of each edge node includes: Regarding each edge device as an edge node and using a compatible monitoring agent according to the operating system of the edge node; Defining the time length of a period as sd, and obtaining the key data of each edge node in each period through the monitoring agent, including application response time data, CPU utilization data, average memory usage, throughput data, and average error rate; Among them, the application response time data includes the average response time and the 95th percentile response time; the CPU utilization data includes the average utilization and the maximum CPU utilization; the throughput data includes the average throughput and the maximum throughput; Calculating representative metrics based on the key data, including user experience index, resource utilization efficiency, and software stability index; Among them, taking the mean of the average response time and the 95th percentile response time within the t period as the user experience index; Normalizing the average utilization, maximum CPU utilization, and average memory usage within the t period, and then using the weighted average method to calculate the resource utilization efficiency; The software stability index is calculated by combining the average throughput, maximum throughput, and average error rate during the t period ; Among them, JJTL represents the average throughput, MAXTTL represents the maximum throughput, JCWL represents the average error rate, and SSI represents the software stability index; Taking the user experience index, resource utilization efficiency, and software stability index of the t period as the performance metrics of the edge node for horizontal concatenation to form the preliminary performance evaluation vector of the edge node in the t period.

3. The real-time software performance monitoring system integrating edge computing according to claim 2, wherein The method of dynamically selecting reference nodes includes: Collecting the preliminary performance evaluation vectors of all edge nodes in the bt periods before the t period to form a performance metric dataset; Cleaning the performance metric dataset and using the difference filling method for missing value processing; According to the processed performance metric dataset, dynamically selecting reference nodes, including: Define the decision variable as jz i indicating whether the i-th edge node is selected as the reference node, where jz i = 1 means the i-th edge node is selected, and jz i = 0 means the i-th edge node is not selected; Define the objective function as minimizing the performance difference between all edge nodes and the selected reference nodes , where represents the j-th performance metric of the i-th edge node, i represents the index of the edge node, I represents the number of edge nodes, j represents the index of the performance metric, and J represents the number of performance metrics, represents the average value of the j-th performance metric of all reference nodes, represents the weight of the j-th performance metric; Calculate the information entropy of each performance metric based on the entropy concept in information theory , where represents the information entropy of the j-th performance metric The weights are calculated based on the information entropy of each performance metric ; Defining the constraint conditions as quantity limit, uniform geographical distribution, diverse hardware configurations, and time stability; Among them, the quantity limit is defined as restricting the number of reference nodes selected to be less than or equal to jds, and the mathematical expression is: , where jds represents the maximum number of reference nodes allowed to be selected; Geographical distribution uniformity is defined as that the reference nodes are evenly distributed in different geographical locations. Obtain the geographical location where each edge node is located. For each geographical location, there is at least one reference node; Diverse hardware configurations are defined as that the reference nodes have different hardware configurations. Obtain the hardware configuration of each edge node, and classify the hardware configurations of all edge nodes. For each type of hardware configuration, ensure that there is at least one reference node; Time stability is defined as calculating the standard deviation of the performance metrics of each edge node within bt time periods before the t time period, so that the standard deviation of the performance metrics of the selected reference nodes is less than or equal to the preset performance metric standard deviation threshold; Based on the decision variables, objective function, and constraint conditions, use a linear programming solver to solve and obtain the optimal solution as the best combination of reference nodes.

4. A real-time software performance monitoring system integrating edge computing according to claim 3, characterized in that, The method for obtaining the performance evaluation vector includes: According to all the selected reference nodes in the best combination of reference nodes, take the mean of the preliminary performance evaluation vectors of all reference nodes in the t time period as the reference performance metric vector; Collect the historical performance data set, including the preliminary performance evaluation vectors of all edge nodes and reference nodes; Use statistical analysis techniques to identify and process the historical performance data set, and identify and judge whether there are long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes; If there are long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes, use the ARIMA model as the time series model; If there are no long-term trends and periodic fluctuations in the performance metrics of edge nodes and reference nodes, use the exponential smoothing method as the time series model; According to the selected time series model, use the historical performance data set to train the selected time series model, and output the preliminary performance evaluation prediction vector of edge nodes within the next time period; Calculate the ratio of the preliminary performance evaluation prediction vector in the time period after the t time period to the reference performance metric vector in the t time period to obtain the adjustment factor; For each time period t, calculate the performance ratio of the edge node relative to the reference node, and introduce an adjustment factor to correct the preliminary performance evaluation vector to obtain the performance evaluation vector ; Among them, represents the preliminary performance evaluation vector of the i-th edge node at time t, represents the performance evaluation vector, represents the benchmark performance index vector at time t, represents the adjustment factor at time t, represents the performance ratio of the edge node relative to the benchmark node.

5. The real-time software performance monitoring system integrating edge computing according to claim 4, characterized in that The method for grouping each edge node using the community detection algorithm to obtain different communities and obtaining the comprehensive feature vector of each community according to the performance evaluation vector includes: Read and collect the network topology information of each edge node, including IP address, subnet, routing path, bandwidth, and latency; Construct a network graph through the network topology information of each edge node, calculate the similarity between each edge node according to the network graph, and form a similarity matrix; According to the similarity matrix, use the community detection algorithm, and use the edge node performance evaluation vector as an additional load weight to group the edge nodes to obtain different communities, and obtain the comprehensive feature vector according to different communities.

6. The real-time software performance monitoring system integrating edge computing according to claim 5, characterized in that The method for constructing a network graph through the network topology information of each edge node, calculating the similarity between each edge node according to the network graph, and forming a similarity matrix includes: Determine the connection relationship between nodes through the routing path of each edge node. For edge nodes with IP addresses belonging to the same subnet, add an edge without edge weight between every two edge nodes; For edge nodes whose IP addresses do not belong to the same subnet, identify edge node pairs with connection paths based on the routing path, and add an edge with edge weight to these edge node pairs; Define and obtain edge weights through bandwidth and latency , where represents the edge weight between the i-th edge node and the r-th edge node, represents the scaling factor for adjusting the importance of bandwidth and latency, represents the bandwidth between the i-th edge node and the r-th edge node, represents latency; Based on all edge nodes and the corresponding edges, construct a network graph; According to the network graph, calculate the similarity between any edge node pair (i, r), and define that if edge node i and edge node r belong to the same subnet, use the unweighted similarity metric, and if i and r do not belong to the same subnet, use the weighted similarity metric; ; Among them, represents the similarity between any edge node pair i and r, represents the similarity of edge node pairs within the same subnet , represents the neighbor set of edge node i, represents the neighbor set of edge node r. All nodes directly connected to edge node i in the network graph are used as the neighbor set of edge node i, represents the similarity of edge node pairs that do not belong to the same subnet, represents the edge weight of the edge between edge node i and edge node c, represents the edge weight of the edge between edge node r and edge node c; Based on the similarity between all pairs of edge nodes, a similarity matrix SSJ is formed, and each position within the similarity matrix represents the similarity between edge node i and edge node r.

7. A real-time software performance monitoring system integrating edge computing according to claim 6, characterized in that, The method for obtaining the comprehensive feature vector includes: Use the weighted summation formula to calculate all performance metrics in the performance evaluation vector to obtain the load weight; Use the spectral clustering algorithm as the community detection algorithm, introduce the load weight into the modularity function of the community detection algorithm, and define the grouping goal of the community detection algorithm as making nodes with high load more likely to be assigned to the same community; The modularity function after introducing the load weight is ; Among them, represents the modified modularity function, zzq represents the total weight of all edges in the network graph, and represent the load weights of edge nodes i and r respectively, and represent the degrees of i and r respectively. The number of edges connected to i is taken as the degree of i, represents the indicator function for judging whether i and r belong to the same community, and represent the community labels of i and r within the community detection algorithm respectively. If , it is determined that i and r belong to the same community, then , if , it is determined that i and r do not belong to the same community, then , represents whether there is an edge between i and r. If there is an edge between i and r, then , if there is no edge between i and r, then ; Apply the modified modularity function and the similarity matrix SSJ to the community detection algorithm to group the edge nodes, obtaining different communities, where each community contains a group of edge nodes that are closely connected in the network topology and have similar loads; For each community, calculate the average value of the corresponding performance metrics in the performance evaluation vectors of each edge node to form the comprehensive feature vector of the community.

8. The real-time software performance monitoring system integrating edge computing according to claim 7, characterized in that, The training method of the performance fluctuation prediction model includes: Step A1: Collect a sample set, including the comprehensive feature vector sequence of each community and the corresponding label, and the label is the true performance value of the next time period corresponding to the comprehensive feature vector; According to the comprehensive feature vector of each community, integrate the comprehensive feature vectors of any time period and the previous qc time periods in chronological order into the comprehensive feature vector sequence of each community at any time period; Step A2: Normalize the sample set, and use the interpolation method to process the missing values, and divide the processed sample set into a training set and a test set according to a ratio; Step A3: Construct a performance fluctuation prediction model, use the comprehensive feature vector sequence of each community as the input, use the predicted performance value of the community in the next time period as the output, and use the predicted performance of the community as the predicted performance value of all edge nodes in the community. The performance fluctuation prediction model is an LSTM model; Step A4: Initialize the hyperparameters of the model, and use Bayesian optimization to tune the hyperparameters. Use k-fold cross-validation to evaluate the cross-validation scores of the model under different hyperparameter combinations, and select the optimal parameter combination; Step A5: Use the optimal parameter combination as the initial parameters of the model, define Adam as the optimizer, and define the mean absolute error as the loss function for evaluating the prediction accuracy of the model , where represents the predicted performance value of the ve-th sample, represents the true performance value of the ve-th sample; Step A6: Input the comprehensive feature vector sequence in the training set into the model, perform forward propagation, calculate the predicted performance value, and then use the loss function to calculate the loss between the predicted performance value and the true performance value, and update the model parameters through backpropagation. Repeat the forward propagation and backpropagation iteratively; Step A7: For each iteration, use the R 2 score as the evaluation metric and calculate the R 2 score value on the validation set; According to the R 2 score value on the validation set, calculate the R 2 score value after the current iteration and the R 2 score value of the previous iteration, and denote it as the iteration difference; Set the iteration difference threshold. If the iteration difference is greater than the iteration difference threshold, it is determined that the performance of the model has improved; If the iteration difference is less than or equal to the iteration difference threshold, it is determined that the performance of the model has not improved; If the performance of the model on the validation set has not improved in consecutive DC iterations, stop training to obtain the trained performance fluctuation prediction model.

9. The real-time software performance monitoring system integrating edge computing according to claim 8, characterized in that, The method for obtaining abnormal nodes includes: Step B1: According to the predicted performance value of each community in the next time period predicted by the performance fluctuation prediction model, after the next time period actually arrives, obtain the monitoring results of the actual performance values of each community through the real-time monitoring system; Step B2: For each community, calculate the absolute difference between its predicted performance value and the actual performance value as the expected difference in actual performance; Set the expected difference threshold. If the expected difference is less than or equal to the expected difference threshold, it is determined that the actual performance of the community is within the normal range, and the actual performance value of the community is used as the final performance evaluation value for each edge node within the community; If the expected difference is greater than the expected difference threshold, it is determined that the actual performance value of the community deviates from the normal range, and the community is marked as an abnormal community; Step B3: For an abnormal community, calculate the standard deviation of each performance metric of the edge nodes within the abnormal community for all edge nodes within the abnormal community; Set the node threshold for each edge node in the abnormal community in combination with the expected difference threshold and standard deviation of the community , where represents the node threshold, represents the expected difference threshold of the community, represents the constant term for adjusting the threshold looseness, represents the standard deviation; Step B4: For each abnormal community, collect the historical performance evaluation vectors of each edge node in the past PCD time periods, integrate the historical performance evaluation vectors of each edge node into an overall vector, and calculate the historical average performance; Step B5: Use the Euclidean distance formula to calculate the historical average performance of each edge node and the actual performance value of the abnormal community to obtain the performance deviation value of the edge node relative to the abnormal community; Sum the predicted performance value and the actual performance value of the abnormal community with the performance deviation value respectively to obtain the predicted performance value and the actual performance value of each edge node within the abnormal community; Step B6: Calculate the absolute difference between the predicted performance value and the actual performance value of each edge node within the abnormal community respectively, and compare it with the node threshold. If it is greater than the node threshold, it is determined that the performance of the edge node deviates from the normal range and is marked as an abnormal node.

10. A real-time software performance monitoring method integrating edge computing, which is implemented based on the real-time software performance monitoring system integrating edge computing according to any one of claims 1 to 9, and is characterized in that, Including: Step S1: Obtain the key data of each edge node and calculate the preliminary performance evaluation vector of each edge node; Step S2: Dynamically select a reference node and use the reference performance metric vector of the reference node to correct the preliminary performance evaluation vectors of each edge node to obtain the performance evaluation vector; Step S3: Use the community detection algorithm to group each edge node to obtain different communities, and based on the performance evaluation vector, obtain the comprehensive feature vector of each community; Step S4: Use the comprehensive feature vector of each community to train and obtain a performance fluctuation prediction model, and predict the predicted performance value of each community in the next time period based on the performance fluctuation prediction model; Step S5: According to the predicted performance value in the next time period, after the actual time period arrives, obtain the actual performance monitoring result. By comparing the predicted performance value with the actual monitoring result, calculate the expected difference to identify the abnormal community. For the abnormal community, adjust its predicted and actual performance values by calculating the deviation value of the edge node, and determine the abnormal node based on the set node threshold.

Citation Information

Patent Citations

  • Edge node anomaly positioning method, device, equipment and computer program product

    CN115437858A

  • Metering equipment performance monitoring system

    CN117668774A

  • Community division method and device, electronic equipment and readable storage medium

    CN119273488A

  • System and method for data community detection via data network telemetry

    US11983164B1

Cited By

  • Large model dynamic optimization-based abnormal behavior diagnosis system for power internet of things

    CN120744789A

  • Computer system based on deep learning model

    CN120849403A

  • A computer system based on a deep learning model

    CN120849403B