Real-time monitoring method and system for cloud computing machine room

By splitting the server into multiple components, the instantaneous sensitivity value and performance change rate of power fluctuations are calculated, and a comprehensive sensitivity value model is established, which solves the problem of inaccurate impact assessment of power fluctuations in the existing technology, and accurately evaluates and early warnings of server performance, improving the stability and service quality of the server.

CN120492259AActive Publication Date: 2025-08-15HUBEI YIKANGSI TECH CO LTD

Patent Information

Application Number
CN202510501738.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-15
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately quantify the impact of power fluctuations on the performance of various components of the cloud computer room server, and it is impossible to fully understand the coupling effects between components, resulting in inaccurate prediction and early warning.

Method used

By splitting the server into multiple single components, calculating instantaneous sensitivity values ​​and performance change rates, quantifying the performance impact values ​​of power fluctuations, establishing a comprehensive sensitive value model, identifying the mutual influence relationship between components, and dynamically adjusting through monitoring tools and neural network models.

Benefits of technology

It realizes an accurate evaluation of power fluctuations on server performance, improves prediction accuracy and early warning efficiency, optimizes resource configuration, and improves server stability and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492259A_ABST
    Figure CN120492259A_ABST
Patent Text Reader

Abstract

The invention relates to a cloud computing machine room real-time monitoring method and system. The method comprises the steps that data are collected, abnormal behaviors of a user are recognized, a server is divided into a plurality of single assemblies, instantaneous sensitive values of the single assemblies are calculated, and the performance change rate is calculated based on the instantaneous sensitive values; calculating a comprehensive state, namely a comprehensive sensitive value, of the single component under different fluctuation frequencies and time; according to the method, the electric power fluctuation performance influence value is calculated based on the instantaneous sensitive value and the performance change rate, and the electric power fluctuation performance influence value is calculated by quantifying the instantaneous sensitive value and the performance change rate, so that the influence of the electric power fluctuation on the performance of the server can be evaluated more accurately; according to the method, the influence of the power fluctuation on the performance of the whole server can be known more comprehensively, a sensitive value model is established, the performance change of the server under different power fluctuation conditions can be predicted, and a more accurate and comprehensive evaluation method for the influence of the power fluctuation on the performance of the server is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of real-time monitoring technology, and in particular to a real-time monitoring method and system for a cloud computing room. Background Art

[0002] In today's digital age, cloud computing technology has been widely adopted. As the foundation of cloud computing, the stable operation of cloud computing rooms is crucial for ensuring the smooth operation of various cloud services. Cloud computing rooms are home to a large number of servers, which host numerous critical businesses and applications, such as enterprise data storage and processing, and the provision of online services. However, cloud computing rooms face numerous challenges, among which power fluctuations have a particularly significant impact on server performance. Power fluctuations are a common problem in cloud computing rooms and can be caused by a variety of factors, including power grid instability, the startup and shutdown of large equipment, and lightning storms. Power fluctuations can lead to unstable power supply to servers, thus impacting their normal operation. Servers are complex electronic devices composed of multiple components, including processors (CPUs), memory, and hard drives. Different components have varying degrees of sensitivity to power fluctuations. Power fluctuations can cause processor operating frequencies to become unstable, resulting in slow or erroneous computing tasks; memory data transmission errors can affect data storage and retrieval efficiency; and hard drive read and write performance can be affected, leading to data loss or increased access latency.

[0003] For example, Chinese patent application No. 202410911366.8 discloses a real-time monitoring system for a cloud computing server room. The system includes: a network anomaly detection module, a security situation awareness and update module, a power stability monitoring and early warning module, a server performance prediction module, an environmental monitoring and adjustment module, an internal behavior monitoring module, and a dynamic response and security adjustment module. In the present invention, by analyzing time series data, capturing network behavior dependencies, identifying abnormal user behavior, and combining with an external threat database, the monitoring strategy and security protection are optimized. By analyzing power supply parameters and combining a linear regression algorithm, the impact of power on server performance is analyzed, performance trends are predicted, and operational efficiency is improved. By adjusting environmental parameters and internal behavior monitoring in real time, the stability and security of the cloud computing server are optimized. By adjusting security maintenance strategies, the operation and data security of the cloud computing server are optimized, and the system protection capability and efficiency are improved.

[0004] Existing patent documents describe cloud computing room monitoring methods that primarily focus on monitoring the overall operational status of servers, such as server power-on status, CPU and memory usage, but lack in-depth analysis and assessment of the specific impact of power fluctuations on the performance of individual server components. These methods struggle to accurately quantify the extent of the impact of power fluctuations on server performance, nor do they fully understand the performance trends of server components under varying power fluctuation conditions. Furthermore, traditional methods struggle to account for the interplay between server components and are unable to accurately identify coupling effects between components, resulting in inaccurate and in-time predictions of server performance and the development of early warning mechanisms. Summary of the Invention

[0005] In response to the technical problems existing in the prior art, the present invention provides a real-time monitoring method and system for a cloud computing room. By quantifying the instantaneous sensitivity value and the performance change rate to calculate the power fluctuation performance impact value, the impact of power fluctuations on server performance can be more accurately evaluated.

[0006] The present invention solves the above technical problems with the following technical solutions: A real-time monitoring method for a cloud computing room, comprising: S11 collects data and identifies abnormal user behavior, splits the server into multiple single components, calculates the instantaneous sensitivity value of each component, and calculates the performance change rate based on the instantaneous sensitivity value; Calculate the comprehensive state of a single component under different fluctuation frequencies and times, that is, the comprehensive sensitivity value; S12, calculating the power fluctuation performance impact value based on the instantaneous sensitivity value and the performance change rate.

[0007] Preferably, the instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single period of time, reflecting the sensitivity of component performance to power fluctuations.

[0008] Preferably, the formula for calculating the power fluctuation performance impact value is: , where P is the impact value of power fluctuation performance, n is the number of components, S i is the instantaneous sensitivity of the i-th component, ΔV i is the power fluctuation amplitude corresponding to the i-th component, τ is the duration of power fluctuation, W i The weight coefficient of the i-th component.

[0009] Preferably, S201, collect data in the time dimension and calculate the comprehensive sensitivity value; S202, calculating the correlation coefficient between the comprehensive sensitivity values of each single component based on the comprehensive sensitivity value, and identifying the mutual influence relationship based on the correlation coefficient; Calculate the combined sensitivity value based on the comprehensive sensitivity value and correlation coefficient; S203: Establish an early warning mechanism based on the comprehensive sensitivity value.

[0010] Preferably, the comprehensive sensitivity value is calculated according to different fluctuation frequencies and different durations, and the formula is: , where S i is the comprehensive sensitivity value of component i; f i is the fluctuation frequency; t i is the duration of abnormal state; α, β are dynamic weight coefficients, and the initial values of α, β are calculated by the hierarchical analysis method.

[0011] Preferably, the comprehensive sensitivity value refers to the comprehensive state of a single component under different fluctuation frequencies and different durations.

[0012] Preferably, the combined sensitivity value is calculated based on the comprehensive sensitivity value and the correlation coefficient, and the formula is: , where S c is the combined sensitivity value; k is the number of single components; w i is the single weight coefficient of component i; S i and S j is the comprehensive sensitivity value of component i and component j; w ij is the combined weight coefficient of component i and component j; r ij is the correlation coefficient between component i and component j.

[0013] Preferably, the combined sensitivity value is used to quantify the joint sensitivity effect between multiple single components, reflecting the comprehensive impact degree caused by the coupling relationship between components.

[0014] Preferably, S301, using monitoring tools to identify the interaction mode of a single component and the order of processing tasks, and draw a task flow chart; S302, dynamically adjusting component work data and resource allocation based on the task flow chart and historical data, and optimizing the combined sensitivity value based on the algorithm to obtain the optimal combined sensitivity value; S303 : Identify a corresponding relationship between the combined sensitivity value and the power fluctuation performance impact value according to the combined sensitivity value and the power fluctuation performance impact value.

[0015] The present application also provides a real-time monitoring system for a cloud computing room, including an abnormal behavior identification module, a relationship identification module, an early warning mechanism establishment module, an identification task process module and a combination sensitive value and power fluctuation performance impact value association module. The combination sensitive value and power fluctuation performance impact value association module, the relationship identification module and the abnormal behavior identification module are all electrically connected to the early warning mechanism establishment module.

[0016] The beneficial effects of the present invention are as follows: by quantifying the instantaneous sensitivity value and the performance change rate to calculate the power fluctuation performance impact value, the impact of power fluctuations on server performance can be more accurately assessed; by splitting the server into multiple single components for analysis, the impact of power fluctuations on the overall server performance can be more comprehensively understood; and by establishing a sensitivity value model, the performance changes of the server under different power fluctuation conditions can be predicted. This provides a more accurate and comprehensive method for assessing the impact of power fluctuations on server performance, helping system administrators to better optimize server performance and stability. By calculating comprehensive sensitivity values and setting thresholds, we can accurately monitor and provide early warnings for the comprehensive status of a single component at different fluctuation frequencies and durations. Based on the changing trends of comprehensive sensitivity values, we can predict the future status of power stability and server performance, improve the efficiency of power stability monitoring, promptly detect and address unstable factors, improve the accuracy of server performance predictions, optimize resource allocation, and enhance service quality and user experience. By setting thresholds and trigger mechanisms, we can provide early warnings before problems occur, reducing the probability of failures and the scope of impact. The combined sensitivity value comprehensively considers the comprehensive sensitivity value of each single component and the mutual influence relationship between them. It can more comprehensively reflect the overall status of the system and more comprehensively evaluate the comprehensive status of multiple components and their impact on system performance. It provides a basis for system optimization, early warning and maintenance, and improves system reliability and performance. By monitoring the fluctuations of processor, memory and hard disk components in real time and dynamically adjusting task processes and resource allocation, the impact of power fluctuations on system performance can be detected more quickly, thereby improving monitoring efficiency. By combining the correlation between sensitive values and power fluctuation performance impact values, a more accurate performance prediction model can be constructed, thereby improving the prediction accuracy of server performance and the monitoring efficiency of power stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a real-time monitoring method for a cloud computing room according to the present invention is shown; Figure 2 A schematic diagram of a process for calculating a comprehensive sensitivity value according to an embodiment of the present invention; Figure 3 A schematic diagram of a process for optimizing a combined sensitivity value according to an embodiment of the present invention; Figure 4 This is a block diagram of a real-time monitoring system for a cloud computing room according to the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0019] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of this application, "plurality" means two or more, unless otherwise specifically specified.

[0020] In the description of this application, the term "for example" is used to mean "used as an example, illustration or explanation". Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art will recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in this application.

[0021] Example 1: Figure 1 1 is a flow chart of a method for real-time monitoring of a cloud computing room according to embodiment 1 of the present invention, comprising the following steps: S101, collecting data and identifying abnormal behavior of users in the cloud computing server based on the data; Specifically, the collected data includes network traffic data and user behavior data. For network traffic data, the port mirroring function of network devices (such as routers and switches) is used to obtain data packets transmitted between cloud computing servers and external networks. The data packets include source IP addresses, destination IP addresses, and The address, source port, destination port and protocol type are determined according to actual needs and network scale. For high-traffic network environments, a higher collection frequency is used to ensure the real-time and accuracy of the data. For user behavior data, the log collection tool Fluentd is used to obtain user behavior data from the operating system logs, application logs, user authentication systems, etc. of the cloud computing server. The operating system logs can record user login, logout, file operations and other behaviors. The application logs can record user operation records in specific applications. User behavior data includes user ID, login time, operation type and operation results. A time series model is used to identify abnormal user behavior in the cloud computing server. The time series model is a short-term memory network (LSTM). Based on the short-term memory network (LSTM), a deep learning method is used to identify user behavior. The deep learning method is an autoencoder. The autoencoder is used to learn normal user behavior patterns. When the reconstruction error between the new user behavior data and the normal pattern exceeds the set threshold, the behavior is judged to be abnormal behavior.

[0022] S102, splitting the server into multiple single components and calculating the instantaneous sensitivity value of the single components; Furthermore, the instantaneous sensitivity value is the degree of change of a single component in a single period (duration) due to a single fluctuation, reflecting the sensitivity of component performance to power fluctuations. The server is divided into processor, memory, and hard disk components, and the instantaneous sensitivity value of the processor is calculated using the formula: , where S CPU is the instantaneous sensitivity value of the processor (CPU) (dimensionless, range 0-1), U curr is the current CPU usage, that is, the load ratio per unit time, U base is the baseline CPU usage, that is, the baseline value when there is no load or light load (such as the CPU usage when the system is idle). max is the maximum theoretical CPU utilization, which is 100% for single-core and 100% × the number of cores for multi-core. α is the load sensitivity coefficient, which is adjusted according to the task type (e.g., α = 1.2 for compute-intensive tasks and α = 0.8 for I / O-intensive tasks). By comparing the difference between the current load and the baseline load, the impact of instantaneous fluctuations on the CPU is quantified. The instantaneous memory sensitivity value is calculated using the formula: , where S Mem is the instantaneous sensitivity value of memory, B curr is the current memory bandwidth, that is, the amount of data transmitted per unit time, Bbase is the benchmark memory bandwidth, stable bandwidth under light load, B max The theoretical maximum bandwidth of memory is determined by the frequency and number of channels (e.g. DDR5 6400MHz dual-channel is 128GB / s). curr is the current memory latency, the time from request issuance to data return, L base is the baseline memory latency, stable latency under light load, L max is the theoretical minimum memory latency, the lower limit determined by hardware specifications, and β is the bandwidth sensitivity weight. Memory performance needs to consider instantaneous changes in bandwidth and latency. Weight allocation needs to be combined with the application scenario. The instantaneous sensitivity value of the hard disk is calculated using the formula: , where S Disk is the instantaneous sensitivity value of the hard disk, I curr is the current IOPS (input / output operations per second), reflecting random access performance. base is the benchmark IOPS (stable value under light load), I max The theoretical maximum IOPS of the hard disk (such as NVMe SSD can reach 1 million), W curr is the current throughput, sequential read and write performance W base is the baseline throughput (stable value under light load), W max is the theoretical maximum throughput of the hard disk (e.g., approximately 200 MB / s for HDD and approximately 3500 MB / s for NVMe SSD), and γ is the IOPS sensitivity weight.

[0023] S103, based on the instantaneous sensitivity value of each component, use an algorithm to establish a sensitivity value model for each component and calculate the performance change rate; Specifically, based on the instantaneous sensitivity values of the processor, memory, and hard disk components calculated in step S102, a neural network model is used to establish sensitivity models for each component. The resulting instantaneous sensitivity values are divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to adjust the model parameters during training, and the test set is used to ultimately evaluate the model's performance. The selected model is trained using the training set to determine the network structure, such as the number of layers and the number of neurons in each layer. The network weights are adjusted using the training data through a backpropagation algorithm to minimize the error between the model's predicted output and the actual output. The operating principles of the processor, memory, and hard disk are analyzed. For the processor, factors such as its architecture design and operating frequency range are considered; for the hard disk, factors such as its read / write mechanism and cache size are considered. Different components have different sensitivities to power fluctuations, and their performance change thresholds also vary. The performance change trends and limits of the processor, memory, and hard disk components under different power fluctuation conditions are analyzed. Based on these performance trends and expert experience, thresholds are set. The set thresholds are then used to determine whether the components are sensitive to power fluctuations.

[0024] The actual power fluctuation parameters are input into the established sensitive value model, and the sensitive value model calculates the performance change rate of each component under power fluctuation through an internal algorithm.

[0025] S104, calculating the power fluctuation performance impact value based on the instantaneous sensitivity value and the performance change rate; Furthermore, based on the calculated instantaneous sensitivity value and the performance change rate obtained through the sensitivity value model, the power fluctuation performance impact value is calculated, and the formula is: , where P is the impact value of power fluctuation performance, n is the number of components (such as CPU, memory, hard disk, etc.), S i is the instantaneous sensitivity value of the i-th component (such as S CPU 、Memory S Mem , Hard disk S Disk ), ΔV i is the power fluctuation amplitude corresponding to the i-th component (unit: volt, such as voltage drop ΔV=V nominal −V actual ), τ is the duration of power fluctuation (unit: seconds, such as the duration of voltage dip), W i The weight coefficient of the i-th component (assigned based on business importance, such as CPU weight W CPU = 0.5, memory weight W Mem = 0.3, and hard disk weight W Disk = 0.2).

[0026] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by quantifying the instantaneous sensitivity value and the performance change rate to calculate the power fluctuation performance impact value, the impact of power fluctuations on server performance can be evaluated more accurately. By splitting the server into multiple single components for analysis, a more comprehensive understanding of the impact of power fluctuations on the performance of the entire server can be obtained. By establishing a sensitivity value model, the performance changes of the server under different power fluctuation conditions can be predicted, providing a more accurate and comprehensive evaluation method for the impact of power fluctuations on server performance, which helps system administrators better optimize server performance and stability.

[0027] Example 2: Based on the change degree of a single component affected by a single fluctuation in a single period of time in the above Example 1, this example analyzes the comprehensive state of a single component under different fluctuation frequencies and durations, and the changes caused by the mutual influence of multiple single components to improve monitoring efficiency and prediction accuracy, such as Figure 2 shown.

[0028] S201, collect data in the time dimension and calculate the comprehensive sensitivity value; Furthermore, the data in the time dimension includes different fluctuation frequencies and durations. The collected data is preprocessed. The comprehensive sensitivity value refers to the comprehensive state of a single component under different fluctuation frequencies and durations. The comprehensive sensitivity value is calculated based on the collected different fluctuation frequencies and durations. The formula is: , where S i is the comprehensive sensitivity value of component i; f i is the fluctuation frequency, such as the number of current fluctuations per second; t i is the duration of abnormal state; α, β are dynamic weight coefficients, and the initial values of α, β are determined by the hierarchical analysis method.

[0029] S202, calculating the correlation coefficient between the comprehensive sensitivity values of each single component based on the comprehensive sensitivity value, and identifying the mutual influence relationship based on the correlation coefficient; Specifically, the calculated comprehensive sensitivity value data of each single component are organized into a table, and the correlation coefficient between the comprehensive sensitivity values of each single component is calculated using the Pearson correlation coefficient formula. The Pearson correlation coefficient formula is a prior art, and those skilled in the art can easily think of it. The correlation coefficient reflects the strength and direction of the linear relationship between the two variables, and the value range is between -1 and 1. When the correlation coefficient |r|≥0.7, the comprehensive sensitivity values of the two components have a strong influence relationship, that is, the state change of one component will significantly affect the comprehensive sensitivity value of the other component. When 0.3≤|r|<0.7, the comprehensive sensitivity values of the two components have a weak influence relationship, that is, the state change of one component has a certain influence on the comprehensive sensitivity value of the other component, but the influence is not great. When |r|<0.3, there is no significant influence relationship between the comprehensive sensitivity values of the two components.

[0030] S203, establish an early warning mechanism based on the comprehensive sensitivity value; Furthermore, based on the calculated comprehensive sensitivity value, historical data related to the comprehensive sensitivity value is obtained from various channels such as system logs, database records, and monitoring tools, the standard deviation of the historical data is calculated, and the normal fluctuation range of the comprehensive sensitivity value in the historical data is determined. Based on the time series analysis of the historical data, a line graph of the comprehensive sensitivity value changing over time is drawn to identify the time period when the comprehensive sensitivity value deviates from the normal fluctuation range. Based on the relevant data of the time period when the abnormal value deviates from the normal fluctuation range, the conditions that cause the comprehensive sensitivity value to deviate from the fluctuation range are found, which include sudden increases in business volume, network attacks, and power failures. Based on the normal fluctuation range of the comprehensive sensitivity value, the upper limit of the normal fluctuation range of the comprehensive sensitivity value is set as a threshold. When the comprehensive sensitivity value exceeds the set threshold, the early warning mechanism is automatically triggered. When the early warning is triggered, a higher level of monitoring and early warning process is initiated, such as increasing the monitoring frequency, starting the fault diagnosis program, etc.

[0031] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: through the calculation of comprehensive sensitive values and the setting of thresholds, accurate monitoring and early warning of the comprehensive status of a single component at different fluctuation frequencies and different durations can be achieved; based on the changing trend of the comprehensive sensitive values, the future status of power stability and server performance can be predicted, thereby improving the efficiency of power stability monitoring, timely discovering and handling unstable factors, improving the prediction accuracy of server performance, optimizing resource allocation, and improving service quality and user experience; by setting thresholds and trigger mechanisms, early warning can be given before problems occur, reducing the probability of failures and the scope of impact.

[0032] Example 3: Calculate the combined sensitivity value based on the correlation coefficient between the comprehensive sensitivity value calculated in Example 2 and the comprehensive sensitivity values of each single component.

[0033] Specifically, the combined sensitivity value is used to quantify the joint sensitivity effect between multiple components, reflecting the comprehensive impact caused by the coupling relationship between components. The combined sensitivity value is calculated based on the comprehensive sensitivity value of a single component and the correlation coefficient between components. The formula is: , where S c is the combined sensitivity value; k is the number of single components; w i is the single weight coefficient of component i, which reflects the importance of the component in the combination. The weight coefficient is confirmed by the hierarchical analysis method; S i and S j is the comprehensive sensitivity value of component i and component j; w ij is the combined weight coefficient of component i and component j, reflecting the contribution of the mutual influence between single components to the combined sensitivity value; r ij is the correlation coefficient between component i and component j.

[0034] For example, in a large data center, the stable operation of the server system is crucial. The server system is mainly composed of three key components: CPU (central processing unit), memory, and hard disk. Select the CPU, memory, and hard disk as three components, numbered i1 (CPU), i2 (memory), and i3 (hard disk), that is, k = 3. The comprehensive sensitivity value S of the CPU has been calculated through the previous steps: i1 =0.6, which reflects the comprehensive state of the CPU under different fluctuation frequencies (such as changes in task processing frequency) and different durations (such as high load duration); the comprehensive sensitivity value of memory S i2 =0.5, which reflects the impact of memory on system performance, such as memory usage fluctuations; the comprehensive sensitivity value of the hard disk S i3 =0.4, indicating the comprehensive status of hard disk read and write performance, etc., and the correlation coefficient between them is calculated: the correlation coefficient between CPU and memory ri1i2 =0.8, indicating that there is a strong mutual influence between them. For example, when the CPU processes tasks, it needs to frequently interact with the memory for data. The correlation coefficient between the CPU and the hard disk is r i1i3 =0.3, indicating that there is a certain correlation between them, but it is relatively weak. For example, the CPU will have a certain impact when processing tasks involving hard disk reading and writing. The correlation coefficient between memory and hard disk is r i2i3 =0.5, indicating that there is a certain mutual influence between them, such as the relationship between memory data cache and hard disk read and write operations. The initial values obtained by the hierarchical analysis method are w i1 =0.4, w i2 =0.3, w i3 =0.3, similarly, w is obtained by analytic hierarchy process i1i2 =0.4, w i1i3 =0.2, w i2i3 =0.3, =0.4×0.6+0.3×0.5+0.3×0.4+0.096+0.0144+0.03=0.6504.

[0035] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: the combined sensitivity value comprehensively considers the comprehensive sensitivity values of each single component and the mutual influence relationship between them, and can more comprehensively reflect the overall state of the system, more comprehensively evaluate the comprehensive state between multiple components and their impact on system performance, provide a basis for system optimization, early warning and maintenance, and improve the reliability and performance of the system.

[0036] Example 4: Based on the above examples 2 to 3, this example dynamically adjusts the working parameters or resource allocation of the processor, memory, and hard disk components by identifying the task flow and fluctuation changes between components, determines the relationship between the combined sensitivity value and the power fluctuation performance impact value, and improves the monitoring efficiency of power stability and the prediction accuracy of server performance. Figure 3 shown.

[0037] S301, using monitoring tools to identify component interaction methods and the order of processing tasks, and draw a task flow chart; Furthermore, the monitoring tool Prometheus is used to monitor in real time the status changes of the processor, memory, and hard disk components when processing tasks. According to the nature of the task and the execution process, the task is divided into different stages. Each task stage should have a start and end mark, as well as specific functions and goals. For each task stage, the interaction method between the processor, memory, and hard disk components is identified. For example, in the data reading stage, the processor sends a read request to the hard disk, the hard disk reads the data into the memory, and the processor then obtains the data from the memory for processing. The dependency relationship between the processor, memory, and hard disk is determined, that is, the operation of one component must be performed after other components complete specific operations.

[0038] Use Microsoft Visio to draw a task flow chart, using rectangular boxes to represent task stages, labeling the name and main function of the task stages within the rectangular boxes, and using arrows to represent the interaction between components and the flow of data. For example, an arrow pointing from the processor to the memory indicates that the processor writes data to the memory, and an arrow from the hard disk to the processor indicates that the hard disk provides data to the processor. Clearly label the processor, memory, and hard disk components in the flowchart, using different colors or shapes to distinguish them. Label each component's specific work content at different task stages, such as the processor executing a specific algorithm at a certain stage, and the memory being used to store specific types of data.

[0039] S302, dynamically adjusting component work data and resource allocation based on the task flow chart and historical data, and optimizing the combined sensitivity value based on the algorithm to obtain the optimal combined sensitivity value; Specifically, based on the drawn task flow chart, the normal working load range of the processor, memory and hard disk components in different task stages is identified. In the data-intensive task stage, the utilization rate of the processor and memory will be relatively high; while in the data storage stage, the I / O operation of the hard disk will be more frequent. The performance indicator data of the processor, memory and hard disk components when the server was running in the past period of time is extracted from the performance monitoring database. The data should include performance under different time periods and different workloads. Thresholds are set according to the statistical characteristics of historical data. When the processor, memory and hard disk exceed the set threshold, the working data and resource allocation of the processor, memory and hard disk are dynamically adjusted, the memory allocation is increased to improve the data processing capability of the processor, or the task priority is adjusted to reduce the execution of low-priority tasks.

[0040] A random forest model is constructed using a machine learning algorithm based on historical data and the calculated combined sensitivity value. The historical data is divided into a training set and a test set to ensure that the data distribution of the training set and the test set is representative. The model is trained using the training set and the test set. The genetic algorithm in the trained model is used to search for the optimal value of the combined sensitivity value. The calculated combined sensitivity value is compared with the optimal value of the combined sensitivity value, and a deviation range is set. When the calculated combined sensitivity value no longer falls within the set deviation range, the working data and resource allocation of the processor, memory, and hard disk are dynamically adjusted.

[0041] S303, identifying a correspondence between the combined sensitivity value and the power fluctuation performance impact value according to the combined sensitivity value and the power fluctuation performance impact value; Furthermore, based on the combined sensitivity values and the power fluctuation performance impact values, the distribution pattern of the data is observed, and a neural network model is used to construct a correlation model between the combined sensitivity values and the power fluctuation performance impact values. The model is trained using historical combined sensitivity values and power fluctuation performance impact values. The data is divided into a training set and a validation set. The training set is used to train the correlation model, and the validation set is used to evaluate the fitting effect of the model to obtain a trained and evaluated correlation model. The combined sensitivity values in the actual monitoring data are obtained and input into the correlation model. The correlation model outputs a prediction result for the power fluctuation performance impact value.

[0042] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: by real-time monitoring of fluctuations in processor, memory and hard disk components, and dynamically adjusting task processes and resource allocation, the impact of power fluctuations on system performance can be detected more quickly, thereby improving monitoring efficiency; by combining the correlation between sensitive values and power fluctuation performance impact values, a more accurate performance prediction model is constructed, thereby improving the prediction accuracy of server performance and improving the monitoring efficiency of power stability.

[0043] Example 5: Figure 4As shown, this embodiment provides a real-time monitoring system for a cloud computing room, including an abnormal behavior identification module, a relationship identification module, an early warning mechanism establishment module, an identification task flow module, and a combination sensitivity value and power fluctuation performance impact value association module. The abnormal behavior identification module is used to collect network traffic data and user behavior data, and identify abnormal behavior of users in the cloud computing server by analyzing data packets and user behavior logs; the relationship identification module is used to calculate the comprehensive sensitivity value of each component and identify the mutual influence relationship between them; the early warning mechanism establishment module is used to establish an early warning mechanism based on the comprehensive sensitivity value and abnormal behavior identification results; the identification task flow module is used to use monitoring tools to identify the interaction mode and sequence of the processor, memory and hard disk components when processing tasks, draw a task flow chart, and clarify the dependency relationship and data flow between components; the combination sensitivity value and power fluctuation performance impact value association module is used to use a neural network model to construct a correlation model between the combination sensitivity value and the power fluctuation performance impact value, and train the model with historical data to predict the impact of power fluctuation on system performance; the combination sensitivity value and power fluctuation performance impact value association module, the relationship identification module, and the abnormal behavior identification module are all electrically connected to the early warning mechanism establishment module.

[0044] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0045] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0046] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0047] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0049] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0050] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A real-time monitoring method for a cloud computing room, characterized in that: include: S11 collects data and identifies abnormal user behavior, splits the server into multiple single components, calculates the instantaneous sensitivity value of each component, and calculates the performance change rate based on the instantaneous sensitivity value; Calculate the comprehensive state of a single component under different fluctuation frequencies and times, that is, the comprehensive sensitivity value; S12, calculating the power fluctuation performance impact value based on the instantaneous sensitivity value and the performance change rate.

2. A cloud computing room real-time monitoring method according to claim 1, characterized in that: The instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single period of time, reflecting the sensitivity of component performance to power fluctuations.

3. A cloud computing room real-time monitoring method according to claim 1, characterized in that: The formula for calculating the impact value of power fluctuation performance is: , where P is the impact value of power fluctuation performance, n is the number of components, S i is the instantaneous sensitivity of the i-th component, ΔV i is the power fluctuation amplitude corresponding to the i-th component, τ is the duration of power fluctuation, W i The weight coefficient of the i-th component.

4. A cloud computing room real-time monitoring method according to claim 1, characterized in that: S201, collect data in the time dimension and calculate the comprehensive sensitivity value; S202, calculating the correlation coefficient between the comprehensive sensitivity values of each single component based on the comprehensive sensitivity value, and identifying the mutual influence relationship based on the correlation coefficient; Calculate the combined sensitivity value based on the comprehensive sensitivity value and correlation coefficient; S203: Establish an early warning mechanism based on the comprehensive sensitivity value.

5. A cloud computing room real-time monitoring method according to claim 4, characterized in that: The comprehensive sensitivity value is calculated according to different fluctuation frequencies and durations. The formula is: , where S i is the comprehensive sensitivity value of component i; f i is the fluctuation frequency; t i is the duration of abnormal state; α, β are dynamic weight coefficients, and the initial values of α, β are calculated by the hierarchical analysis method.

6. A cloud computing room real-time monitoring method according to claim 5, characterized in that: The comprehensive sensitivity value refers to the comprehensive state of a single component under different fluctuation frequencies and durations.

7. A cloud computing room real-time monitoring method according to claim 4, characterized in that: The combined sensitivity value is calculated based on the comprehensive sensitivity value and correlation coefficient. The formula is: , where S c is the combined sensitivity value; k is the number of single components; w i is the single weight coefficient of component i; S i and S j is the comprehensive sensitivity value of component i and component j; w ij is the combined weight coefficient of component i and component j; r ij is the correlation coefficient between component i and component j.

8. A cloud computing room real-time monitoring method according to claim 7, characterized in that: The combined sensitivity value is used to quantify the joint sensitivity effect between multiple single components, reflecting the comprehensive impact caused by the coupling relationship between components.

9. A cloud computing room real-time monitoring method according to claim 8, characterized in that: S301, using monitoring tools to identify the interaction mode of a single component and the order of processing tasks, and draw a task flow chart; S302, dynamically adjusting component work data and resource allocation based on the task flow chart and historical data, and optimizing the combined sensitivity value based on the algorithm to obtain the optimal combined sensitivity value; S303 : Identify a corresponding relationship between the combined sensitivity value and the power fluctuation performance impact value according to the combined sensitivity value and the power fluctuation performance impact value.

10. A cloud computing room real-time monitoring system, applied to a cloud computing room real-time monitoring method according to any one of claims 1 to 9, characterized in that: It includes an abnormal behavior recognition module, a relationship recognition module, an early warning mechanism establishment module, an identification task process module and a combination sensitivity value and power fluctuation performance impact value association module. The combination sensitivity value and power fluctuation performance impact value association module, the relationship recognition module and the abnormal behavior recognition module are all electrically connected to the early warning mechanism establishment module.

Citation Information

Patent Citations

  • Server health assessment method, system, equipment and medium

    CN113806171A

  • Server fault prediction method, system and device and computer readable storage medium

    CN118051408A

  • Electric power data processing system based on cloud computing

    CN118312908A

  • Real-time monitoring system for cloud computing server room

    CN118509248A

  • Electric energy quality monitoring method and system considering broadband transmission characteristics of mutual inductor

    CN119510929A

Cited By

  • Telecommunication room monitoring method, system, equipment and medium

    CN121098691A

  • Power grid energy storage method

    CN121332623A