An Efficient Application Performance Monitoring and Prediction Method Based on Anolis OS
Through the Prometheus and pidstat tool combined with the ConvBiGRU model, efficient performance monitoring and prediction of the Dragon Lizard operating system is achieved, solving the problem of insufficient performance prediction of the Dragon Lizard operating system in high load scenarios, and providing efficient and accurate performance monitoring and early warning functions.
Patent Information
- Application Number
- CN202411885737.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing performance monitoring tools are insufficient to adapt and optimize the domestic operating system Long Lizard operating system, especially in high-load scenarios, which lacks effective performance prediction and fault warning capabilities, which makes it difficult to ensure system stability and performance monitoring.
Prometheus and pidstat tools are used for all-round resource monitoring, combined with CNN and BiGRU's ConvBiGRU model for performance data training and prediction, designed data prediction plug-ins and database plug-ins, and provided visual services through APIs and web applications to achieve efficient performance monitoring and prediction of the Dragon Lizard operating system.
It realizes efficient prediction of key performance indicators such as CPU and memory, ensures prediction accuracy of more than 75% and second-level response speed, provides an easy-to-use visual interface and plug-in architecture, supports real-time early warning and expansion, and is suitable for performance monitoring and optimization needs in different scenarios.
Smart Images

Figure CN119336586B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer performance, and specifically, relates to an efficient application performance monitoring and prediction method based on the Anolis OS. Background Art
[0002] With the rapid development of information technology, modern computing systems have become increasingly complex and large-scale. In such a complex environment, ensuring the stability and performance of the system is a huge challenge. Traditional performance monitoring methods usually rely on frequent user-space to kernel-space switches, which easily lead to additional performance overhead, especially in high-load scenarios. This increases the difficulty of real-time performance monitoring and also raises the risk of system problems.
[0003] The domestic operating system Anolis OS, as one of the operating systems that have received much attention in recent years, is widely used in multiple fields. In high-load scenarios, in order to ensure its stability and performance, a set of efficient performance monitoring and prediction tools are needed to provide timely performance data and trend analysis, helping operation and maintenance personnel quickly locate system bottlenecks and predict potential performance problems, thereby reducing the risk of failures.
[0004] However, most of the current mainstream performance monitoring tools and methods are targeted at general operating systems, and there is relatively little research on adaptation and optimization for domestic operating systems, especially domestic systems such as Anolis OS that are oriented to enterprise-level applications. Anolis OS is designed for large-scale, high-performance computing scenarios. How to ensure its stability and performance in high-load scenarios is one of the important research topics currently. Especially in terms of performance prediction and fault warning under high-load conditions, there is currently no mature solution. Although existing tools are relatively perfect in collecting system operation state data, they lack the ability to deeply analyze and predict data.
[0005] Therefore, developing a performance monitoring and prediction analysis framework suitable for Anolis OS has important theoretical and practical significance. This framework can not only monitor the running state of the system in real time, but also perform performance trend prediction through machine learning models, helping operation and maintenance personnel discover potential problems in advance and make corresponding adjustments. Such a tool that combines performance monitoring and data prediction will greatly improve the reliability and stability of the system and provide strong support for the further development of domestic operating systems. Summary of the Invention
[0006] The present invention proposes an efficient application performance monitoring and prediction method based on Anolis OS to solve problems such as difficult data collection, high system resource consumption, inaccurate performance prediction, and poor scalability, improve the resource utilization rate of the data center, and ensure the service quality of applications.
[0007] The present invention is realized through the following technical solutions:
[0008] An efficient application performance monitoring and prediction method based on the Anolis operating system:
[0009] The method specifically includes the following steps:
[0010] S1. Use Prometheus and pidstat tools for comprehensive resource monitoring and data collection;
[0011] S2. Based on the performance data of virtual machines, train and optimize through the ConvBiGRU model combining CNN and BiGRU;
[0012] S3. Design data prediction plugins, database plugins and plugin managers to adapt to monitoring and prediction requirements in different scenarios;
[0013] S4. Provide visualization services through APIs and web applications to facilitate users to access and use the prediction function and view the performance monitoring results and prediction data of the application.
[0014] Further, in S1, it specifically includes:
[0015] S1.1. Tool selection and data collection:
[0016] In the Anolis operating system environment, deploy Prometheus and pidstat tools according to the operating system type;
[0017] Prometheus comprehensively collects system-level performance metrics, sets its data storage rules, query statements and alarm rules; pidstat collects process-level performance data and determines the range and frequency settings of the collected processes;
[0018] S1.2. The collected data is uniformly stored in the local time series database;
[0019] The database plugin executes insert operations using efficient stored procedures according to the predefined database schema and table structure, and maintains data indexes to ensure fast query access.
[0020] Further, in S1.1,
[0021] Prometheus collects container metrics: For Docker containers, Prometheus regularly obtains the number of CPU cores, total frequency in MHz, used frequency in MHz, usage percentage, total memory capacity in KB, used memory in KB, disk read throughput in KB / s, disk write throughput in KB / s, network receive throughput in KB / s, and network transmit throughput in KB / s of the container according to preset rules. The data format follows the Prometheus specification and is transmitted over the network and stored in a specified database.
[0022] pidstat collects process metrics: pidstat collects the process ID, CPU user usage rate, system usage rate, total usage rate, memory usage in KB, disk read throughput in KB / s, and disk write throughput in KB / s at a set frequency and stores them in the native system format for subsequent processing and conversion.
[0023] Further, in S2.1, dataset preprocessing: Sample with a timestamp of 300ms. First, traverse the dataset to clear all Nan values; then merge the scattered data files; then sample the merged data at 1s intervals, and sample and average every three data points for smoothing; finally, normalize to a unified data distribution to generate sub-datasets of fastStorage_3, rnd_7_3, rnd_8_3, and rnd_9_3.
[0024] S2.2, model architecture design: Select the ConvBiGRU model architecture; configure a convolutional layer at the input end of the time series data to capture local features with specific convolutional kernels and strides; then connect a bidirectional GRU layer to process the time series information according to the gating mechanism and integrate forward and backward dependencies; finally, connect a fully connected layer to map the output dimension to the target performance metric dimension.
[0025] S2.3, model training and optimization:
[0026] During the model training process, the Adam optimizer is used, and a learning rate decay strategy is selected to further optimize the model's convergence speed; a Dropout layer is added to prevent overfitting; combined with an early stopping mechanism, continuously monitor the validation loss and terminate the training early when the validation loss of the model no longer decreases.
[0027] S2.4, develop a model evaluation module after the model training is completed to focus on the results of the model under various evaluation metrics and the comparison between the predicted values and the true values. During the model evaluation process, use metrics including mean squared error (MSE), mean absolute error (MAE), and MAPE for evaluation; through these evaluation metrics, comprehensively understand the prediction performance of the model on different datasets.
[0028] Further, in S3,
[0029] The data prediction plugin includes a Prometheus-based docker container performance monitoring plugin and a local performance monitoring plugin based on pidstat;
[0030] The data prediction plugin is developed by configuring a pre-trained onnx model. Raw data for a recent period is retrieved from the database and divided into multiple normalized batch unit data formats according to data characteristics and model input requirements, and then pushed in batches to the pre-trained ConvBiGRU model for prediction. The prediction results are then de-normalized to obtain the predicted data and transmitted to the database;
[0031] The database plugin is used to connect to the local time series database and implement a plugin for adding, deleting, modifying, and querying data in the database. It inserts the local performance data obtained by the data collection plugin into the local time series database, then receives the parameter for retrieving data from the data prediction plugin, retrieves the raw performance data within a certain time range from the local time series database, and transmits it back to the data prediction plugin.
[0032] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0033] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps of the above method are implemented.
[0034] Advantages of the present invention
[0035] The present invention uses tools such as Prometheus and pidstat for comprehensive resource monitoring; realizes efficient prediction of key performance indicators such as CPU and memory through a ConvBiGRU model combining CNN and BiGRU; adopts a simplified dataset scaling method to ensure a prediction accuracy of more than 75% and a second-level response speed; adopts a plugin architecture and a user-friendly visualization interface, supports real-time warning and extension, and is applicable to performance monitoring and optimization requirements in different scenarios.
[0036] The present invention realizes comprehensive performance monitoring: uses tools such as Prometheus and pidstat to collect key performance indicators such as CPU, memory, disk, and network at the operating system level in real time, ensuring comprehensive monitoring of various resources.
[0037] An efficient prediction model is constructed: based on the collected underlying performance data, a machine learning model capable of accurately predicting application performance is designed and trained, and the model accuracy needs to reach more than 75%.
[0038] Provide an easy-to-use visualization interface: By developing a data visualization web page, users can intuitively view performance monitoring results and prediction data, and a real-time warning function is provided so that users can take corresponding measures in a timely manner.
[0039] Support a plug-in architecture: The system will adopt a plug-in design to facilitate function expansion and customization in different environments, ensuring adaptation to diverse application scenarios.
[0040] Ensure cross-platform compatibility: The model and system need to support Linux systems with different kernel versions, especially to achieve stable performance monitoring and prediction functions on the Anolis operating system. Brief Description of the Drawings
[0041] Figure 1 It is a flowchart of the method of the present invention.
[0042] Figure 2 It is the data extracted after preprocessing to solve the difficulty of dataset collection.
[0043] Figure 3 It is a schematic diagram of the model of the present invention.
[0044] Figure 4 It is a schematic diagram of the prediction result of the model of the present invention and the original data.
[0045] Figure 5 It is the Web program page of the present invention.
[0046] Figure 6 It is the evaluation result of the prediction error of the ConvBiGRU model on different datasets.
[0047] Figure 7 It is the Prometheus monitoring of Docker container metrics and pidstat monitoring of process metrics of the present invention. Detailed Embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0049] The experimental methods used in the following embodiments are all conventional methods unless otherwise specified. The materials, reagents, methods and instruments used, unless otherwise specified, are all conventional materials, reagents, methods and instruments in the art, and those skilled in the art can obtain them through commercial channels.
[0050] The present invention first defines a basic model architecture, pre-trains it using the gwa-t-12-bitbrains dataset, exports the pre-trained model as an onnx model after obtaining it, embeds a data prediction plugin, and is managed by a plugin manager. Additionally, data collection plugins based on pidstat and Prometheus are developed. The real data of local application processes is obtained. On the one hand, it is stored in a time series database for convenient subsequent model training and invocation. On the other hand, after data processing, it is transmitted to the data prediction plugin to evaluate the prediction effect, so as to further improve the prediction model. Finally, a data visualization page is developed, and at the same time, the plugin manager is improved to add a data pipeline function, integrating data collection and data prediction and other plugins. The entire application framework is organized to facilitate the expansion of application functions and adapt to the monitoring and prediction requirements in different scenarios.
[0051] As Figure 1 shown, the present invention proposes an efficient application performance monitoring and prediction method based on the Anolis operating system.
[0052] S1. Use Prometheus and pidstat tools for comprehensive resource monitoring and data collection;
[0053] Prometheus is responsible for comprehensively collecting system-level performance metrics and various performance metrics for Docker containers, thereby realizing performance monitoring at the system and container levels. Pidstat focuses on the collection of process-level performance data and monitors the running status of processes by collecting various relevant metrics of processes. The combination of the two realizes comprehensive performance monitoring at multiple levels from the system, containers to processes.
[0054] S1.1, Tool selection and data collection process:
[0055] In the Anolis operating system environment, Prometheus and pidstat tools are accurately deployed according to the operating system type (Debian / Ubuntu or CentOS / RHEL);
[0056] Prometheus is carefully configured to comprehensively collect system-level performance metrics, and parameters such as its data storage rules, query statements, and alarm rules are set;
[0057] Prometheus collects container metrics: For Docker containers, Prometheus obtains metric data such as the number of CPU cores, total frequency (MHz), usage (MHz), usage percentage, total memory capacity (KB), usage (KB), disk read throughput (KB / s), disk write throughput (KB / s), network receive throughput (KB / s), and network transmit throughput (KB / s) of containers at regular intervals according to preset rules. The data format follows the Prometheus specification and is stored in a specified database via network transmission.
[0058] pidstat focuses on process-level performance data and determines key settings such as the scope and frequency of data collection for processes; ensure that the two work together to achieve comprehensive capture of performance data.
[0059] pidstat collects process metrics: pidstat collects data such as process ID, CPU user usage, system usage, total usage, memory usage (KB), disk read throughput (KB / s), and disk write throughput (KB / s) at a set frequency and stores it in the native system format for subsequent processing and conversion;
[0060] S1.2, The collected data is uniformly stored in a local time series database. The database plugin uses efficient stored procedures to perform insert operations according to predefined database schemas and table structures, and at the same time maintains data indexes to ensure fast query access, ensuring data integrity and consistency, and preparing raw materials for subsequent analysis.
[0061] S2, Based on the performance data of virtual machines, train and optimize through the ConvBiGRU model that combines CNN and BiGRU;
[0062] Based on the collected performance data of virtual machines, train and optimize through the ConvBiGRU model that combines CNN and BiGRU, enabling the model to have the ability to predict performance data.
[0063] Using the gwa-t-12-bitbrains dataset as a blueprint, this dataset contains the performance metrics of 1750 VMs in the Bitbrains distributed data center. (Bitbrains is a service provider that specializes in providing hosting and business computing for enterprises. Customers include many major banks (ING), credit card operators (ICS), insurance companies (Aegon), etc. Bitbrains hosts applications used in the solvency field; examples of application vendors include Towers Watson and Algorithmics. These applications are usually used for financial reporting, mainly at the end of the financial quarter).
[0064] Each file contains the performance metrics of the VMs. These files are organized by trace: fastStorage and Rnd. The first trace, fastStorage, consists of 1,250 VMs that are connected to a fast Storage Area Network (SAN) storage device. The second trace, Rnd, consists of 500 VMs that are either connected to a fast SAN device or to a much slower Network Attached Storage (NAS) device. The fastStorage trace contains a higher proportion of application servers and compute nodes compared to the Rnd trace, due to the higher storage performance of the storage connected to the fastStorage machines. In contrast, for the Rnd trace, a higher proportion of management machines are observed, which only require lower performance and less frequently accessed storage.
[0065] In the Rnd directory, the files are organized into 3 subdirectories by the month that records the metrics. The format of each file is line-based, and each line represents an observation of the performance metrics. Each column of a line is separated by "; ", and the format of each line is as follows:
[0066] (1) Timestamp: The number of milliseconds since 1970-01-01.
[0067] (2) Number of CPU cores: The number of virtual CPU cores pre-allocated.
[0068] (3) Pre-allocated CPU capacity (CPU request): The CPU capacity in MHz, equal to the number of cores x the speed of each core.
[0069] (4) CPU usage: In MHz.
[0070] (5) CPU usage: Expressed as a percentage.
[0071] (6) Provisioned memory (requested memory): The memory capacity of the VM in KB.
[0072] (7) Memory usage: The memory actively used in KB.
[0073] (8) Disk read throughput: In KB / s.
[0074] (9) Disk write throughput: In KB / s.
[0075] (10) Network receive throughput: In KB / s.
[0076] (11)Network transmission throughput: in KB / s.
[0077] S2.1 Dataset preprocessing: The original dataset is the workload performance data of the nodes of each virtual machine, and the goal is to predict application performance data. The most common application performance data is the CPU occupancy and memory occupancy of the application, as well as the local disk read and write occupancy. For web applications, there is also network upload and download occupancy. Similarly, the general metrics of application performance can be abstracted:
[0078] (1)CPU usage rate: expressed as a percentage.
[0079] (2)Memory usage: in KB.
[0080] (3)Disk read throughput: in KB / s.
[0081] (4)Disk write throughput: in KB / s.
[0082] (5)Network receive throughput: in KB / s.
[0083] (6)Network transmission throughput: in KB / s.
[0084] This is consistent with the characteristics of the original dataset. The trajectories of the original dataset are sampled at a timestamp of 300ms. First, traverse the dataset to clear all Nan values; then merge the scattered data files; then sample the merged data at 1s intervals (sample every three data and take the average) for smoothing; finally, normalize to a unified data distribution to generate sub-datasets of fastStorage_3, rnd_7_3, rnd_8_3, rnd_9_3, and use professional algorithms to process the magnitude difference between the system and application data to achieve data standardization preprocessing, as Figure 2 shown.
[0085] S2.2 Model architecture design is as Figure 3 shown,
[0086] Select the ConvBiGRU model architecture. This model is similar to ConvBiLSTM, the difference is that GRU is used instead of the LSTM layer. The convolutional layer helps to extract local features, while GRU processes the overall time series information.
[0087] This structure is more computationally efficient than the LSTM version and also shows good performance in tests. Through experimental comparison, the ConvBiGRU model achieves an optimal balance between prediction accuracy and training speed, especially performing well in processing time-series data of application performance. The ConvBiGRU model shows good convergence during training and testing.
[0088] It configures a convolutional layer at the input end of time-series data to capture local features with specific convolutional kernels and strides; then connects to a bidirectional GRU layer to process time-series information according to the gating mechanism and integrate forward and backward dependencies; finally, connects a fully connected layer to map the output dimension to the dimension of the target performance metric. The parameters of each layer are finely tuned according to the characteristics of the dataset and performance goals to build an accurate prediction model framework.
[0089] Within 50 Epochs, the training loss of the model gradually decreases and tends to be stable, while the test loss remains stable within a lower range, indicating that the model does not suffer from overfitting. This shows that the ConvBiGRU model can fit the training data well and also performs excellently on the test data.
[0090] S2.3, Model training and optimization;
[0091] During the model training process, the Adam optimizer is used, and a learning rate decay strategy is selected to further optimize the convergence speed of the model and balance the convergence speed and accuracy. A Dropout layer is added to the model to randomly mask neurons with a certain probability to prevent overfitting; combined with an early stopping mechanism, continuously monitor the validation loss, and once there is no downward trend for several consecutive rounds, terminate the training to lock the optimal model state and ensure the robustness and generalization ability of the model.
[0092] After training is completed, the model will be used to predict the performance of the application in real time and issue an alarm when potential performance problems are detected.
[0093] S2.4, After the model training is completed, develop a model evaluation module to focus on the results of the model under various evaluation metrics and the curve comparison between the predicted values and the true values, such as Figure 6 As shown, during the model evaluation process, a series of evaluation metrics are adopted, including the mean squared error (MSE), mean absolute error (MAE), and MAPE. Through these evaluation metrics, the prediction performance of the model on different datasets can be comprehensively understood.
[0094] (1) CPU performance evaluation:
[0095] For the CPU cores and CPU usage [%] metrics, the models (rnd_7_3, rnd_8_3, rnd_9_3) on the three different datasets showed similar performance in terms of MAE and RMSE, indicating that these models are stable in predicting CPU performance.
[0096] The model trained with the fastStorage dataset had the lowest MAE and RMSE on CPU cores, indicating higher prediction accuracy on this metric.
[0097] (2) Memory performance evaluation:
[0098] For the Memory usage [KB] and Memory capacity provisioned [KB] metrics, the model trained with the rnd_8_3 dataset had significantly higher RMSE values, indicating larger errors in predicting memory usage for the model on this dataset.
[0099] The MAE and RMSE of the fastStorage dataset model were relatively low, indicating that the model is more accurate in predicting memory usage.
[0100] (3) Disk and network performance evaluation:
[0101] The MAPE values for Disk read throughput [KB / s] and Disk write throughput [KB / s] were very high, especially when using the rnd_8_3 dataset, indicating significant deviations in the model's prediction of disk performance.
[0102] The predictions for Network received throughput [KB / s] and Network transmitted throughput [KB / s] also showed high errors, which may be related to the characteristics of the dataset or limitations in the model structure.
[0103] (4) Summary:
[0104] Overall, the model's prediction performance was relatively stable for CPU cores, Memory usage, and Disk write throughput. Although there were higher errors in some metrics (such as disk and network related metrics), these differences can be reduced through further model optimization.
[0105] The model trained with the fastStorage dataset shows better stability and accuracy on multiple performance metrics, which may be related to the characteristics of the dataset and the adaptability of the model.
[0106] When making inferences on the entire dataset, the ConvBiGRU model demonstrates good prediction ability, especially performing well when dealing with large-scale data. When dealing with complex time-series data in a real environment, the ConvBiGRU model can effectively perform performance prediction and has high practical value.
[0107] S3, design a data prediction plugin, a database plugin, and a plugin manager to adapt to the monitoring and prediction requirements in different scenarios; the data prediction plugin retrieves the raw data of the recent period from the database, processes it, and then pushes it to the pre-trained ConvBiGRU model for prediction. Finally, the predicted data is obtained and transmitted to the database, that is, this step performs the performance prediction operation and processes and stores the prediction results; as Figure 4 shown is a schematic diagram of the prediction result of the model of the present invention and the raw data.
[0108] The data prediction plugin includes a docker container performance monitoring plugin based on Prometheus and a local performance monitoring plugin based on pidstat; as Figure 7 shown, where pidstat mainly focuses on the usage of system resources such as CPU, memory, and I / O. The monitoring function of network traffic is not included in the design. Therefore, its main purpose is to help users analyze system bottlenecks, especially in terms of CPU and memory resources.
[0109] The data prediction plugin is developed and configured by a pre-trained onnx model. It retrieves the raw data of the recent period from the database, divides it into multiple normalized batch unit data formats according to the data characteristics and the model input requirements, and batch-pushes it to the pre-trained ConvBiGRU model for prediction. The prediction result is then denormalized to obtain the predicted data and transmitted to the database.
[0110] The database plugin is a plugin that docks with the local time-series database to implement data addition, deletion, modification, and query operations on the database. It inserts the local performance data obtained by the data collection plugin into the local time-series database, and then receives the parameter of the data prediction plugin to retrieve the raw performance data within a certain time range from the local time-series database and transmits it back to the data prediction plugin.
[0111] The plugin manager allows multiple data collection plugins and multiple data prediction plugins to exist and uniformly manages the configuration.
[0112] S4, provide services through API and Web applications: develop data visualization web pages and backend API interfaces, so that users can easily access and use prediction functions and view the performance monitoring results and prediction data of the application.
[0113] The ConvBiGRU model processes input data according to its built-in architecture and parameters, and generates prediction results through convolution, GRU and fully connected layer operations. The prediction results restore the true value domain through normalized inverse operations and are transmitted in a standardized format through the data pipeline, providing accurate data basis for subsequent visualization and decision-making responses, and seamlessly connecting to system performance monitoring and management processes.
[0114] like Figure 5 As shown in the figure, the data visualization page is mainly divided into three major sections and five panels. The upper left side is the system overall status display panel, which displays the total performance usage of the current system. When a certain performance usage is too high, it will be marked in red to prompt the user; the lower left side is the configuration panel of the background plug-in system, which can edit the necessary configuration of the initialization plug-in; in the middle is the performance data collection display panel of the two data collection plug-ins developed, which displays the performance data list of each container id or pid collected, and a display button is set on the left side of each data collection page. Clicking this button will display the monitoring chart of each indicator of the corresponding container id or pid. Above the performance display chart on the right is the prediction and error calculation panel, which displays the error of each indicator between the predicted value and the actual value in real time.
[0115] The data visualization page presents a panoramic view of system performance with an intuitive layout. The system overall status panel prominently displays global resource usage and real-time risk monitoring; the plug-in configuration panel flexibly customizes plug-in parameters; the data collection panel presents container and process performance data in detail, and interactive charts analyze details. The prediction and error calculation panel accurately compares the predicted value with the actual value, and triggers an immediate warning when the error exceeds the threshold, helping operation and maintenance personnel to respond to performance fluctuations agilely and ensure stable and efficient operation of the system.
[0116] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0117] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0118] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory for the methods described in the present invention is intended to include but not be limited to these and any other suitable types of memory.
[0119] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cable, optical fiber, digital subscriber line (DSL), or wireless means such as infrared, wireless, microwave, etc. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that incorporates one or more available media. The available media can be magnetic media such as floppy disks, hard disks, magnetic tapes, optical media such as high-density digital video discs (DVDs), or semiconductor media such as solid state discs (SSDs), etc.
[0120] In the implementation process, the steps of the above method can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by the hardware processor or executed by the combination of the hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0121] It should be noted that the processor in the embodiments of the present application may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method embodiments can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0122] The above has introduced in detail a method for efficient application performance monitoring and prediction based on the OpenAnolis operating system proposed by the present invention, and has elaborated on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An efficient application performance monitoring and prediction method based on the Anolis operating system, characterized in that: The method specifically includes the following steps: S1. Use Prometheus and pidstat tools for resource monitoring and data collection; S2. Based on the performance data of virtual machines, train and optimize through the ConvBiGRU model combining CNN and BiGRU; S2.
1. Dataset preprocessing: Sample with a timestamp of 300ms. First, traverse the dataset to clear all Nan values; then merge scattered data files; then sample the merged data at 1s intervals, and sample and take the average of every three data for smoothing; finally, normalize it to a unified data distribution to generate sub-datasets of fastStorage_3, rnd_7_3, rnd_8_3, and rnd_9_3; S2.
2. Model architecture design: Select the ConvBiGRU model architecture; configure a convolutional layer at the input end of time series data to capture local features with convolutional kernels and strides; then connect a bidirectional GRU layer to process time series information according to the gating mechanism and integrate forward and backward dependencies; finally, connect a fully connected layer at the end to map the output dimension to the target performance metric dimension; S2.
3. Model training and optimization: During the model training process, use the Adam optimizer and select a learning rate decay strategy to optimize the convergence speed of the model; add a Dropout layer to prevent overfitting; combine an early stopping mechanism to continuously monitor the validation loss and terminate training in advance when the validation loss of the model no longer decreases; S2.
4. After the model training is completed, develop a model evaluation module to focus on the results of the model under various evaluation metrics and the comparison between the predicted values and the true values. During the model evaluation process, use MSE, MAE, and MAPE for evaluation; through these evaluation metrics, comprehensively understand the prediction performance of the model on different datasets; S3. Design data prediction plugins, database plugins, and plugin managers to adapt to monitoring and prediction requirements in different scenarios; S4. Provide visualization services through APIs and web applications to facilitate users to access and use the prediction function and view the performance monitoring results and prediction data of the application.
2. The method according to claim 1, wherein: In S1, it specifically includes: S1.
1. Tool selection and data collection: In the Anolis operating system environment, deploy Prometheus and pidstat tools according to the operating system type; Prometheus comprehensively collects system-level performance metrics, sets its data storage rules, query statements, and alarm rules; pidstat collects process-level performance data and determines the scope and frequency of the collected processes; S1.
2. The collected data is uniformly stored in the local time series database; The database plugin executes insert operations using stored procedures according to the predefined database architecture and table structure, and maintains data indexes to ensure fast query access.
3. The method according to claim 2, wherein: In S1.1, Prometheus collects container metrics: For Docker containers, Prometheus regularly obtains the number of CPU cores, total CPU frequency, CPU usage, CPU usage percentage, total memory capacity, memory usage, disk read throughput, disk write throughput, network receive throughput, and network transmit throughput of the container according to preset rules. The data format follows the Prometheus specification and is stored in the specified database via network transmission; pidstat collects process metrics: pidstat collects the process ID, CPU user usage, CPU system usage, total CPU usage, memory usage, disk read throughput, and disk write throughput at a set frequency and stores them in the native system format for subsequent processing and conversion.
4. The method according to claim 3, wherein: In S3, The data prediction plugin includes a Docker container performance monitoring plugin based on Prometheus and a local performance monitoring plugin based on pidstat; The data prediction plugin is developed by configuring a pre-trained onnx model. It retrieves the raw data of the recent period from the database, divides it into multiple normalized batch data according to the data characteristics and model input requirements, and batch-pushes it to the pre-trained ConvBiGRU model for prediction. The prediction results are then denormalized to obtain the prediction data and transmitted to the database; The database plugin is connected to the local time-series database to implement data addition, deletion, modification, and query of the database. It inserts the local performance data obtained by the data collection plugin into the local time-series database, then receives the parameter for retrieving data from the data prediction plugin, retrieves the raw performance data from the local time-series database, and transmits it back to the data prediction plugin.
5. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
6. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Prediction type elastic scaling method and system of Kubernetes
CN115774605A
1200V SiCMOSFET modeling optimization and performance prediction method based on neural network
CN118627445A