A cloud computer room real-time monitoring method and system
By quantifying the instantaneous sensitivity and performance change rate of power fluctuations to various components of cloud computer room servers, a sensitivity model is established to identify the mutual influence between components. This solves the problem of inaccurate assessment of the impact of power fluctuations in existing technologies and improves the monitoring efficiency of server performance and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUBEI YIKANGSI TECH CO LTD
- Filing Date
- 2025-04-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies struggle to accurately quantify the impact of power fluctuations on the performance of various components in cloud computing room servers, and cannot fully understand the coupling effects between components, resulting in inaccurate predictions and early warnings.
By quantifying instantaneous sensitivity values and performance change rates, the impact value of power fluctuation performance is calculated. The server is split into multiple individual components for analysis, a sensitivity value model is established, the mutual influence relationship between components is identified, and a correlation model between combined sensitivity values and the impact value of power fluctuation performance is constructed.
It enables a more accurate assessment of the impact of power fluctuations on server performance, provides a more comprehensive method for assessing the impact of power fluctuations, improves the monitoring efficiency of server performance and stability, optimizes resource allocation, and reduces the probability and scope of failures.
Smart Images

Figure CN120492259B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time monitoring technology, and in particular to a method and system for real-time monitoring of cloud computer rooms. Background Technology
[0002] In today's digital age, cloud computing technology is widely used. As the infrastructure of cloud computing, the stable operation of cloud data centers is crucial for ensuring the normal operation of various cloud services. Cloud data centers deploy a large number of servers, which host numerous critical businesses and applications, such as enterprise data storage and processing, and the provision of online services. However, cloud data centers face many challenges, among which power fluctuations have a particularly prominent impact on server performance. Power fluctuations are a common problem in cloud data centers, and they can be caused by various factors, such as unstable power grids, the start-up and shutdown of large equipment, and lightning storms. Power fluctuations lead to unstable power supply to servers, thus affecting their normal operation. Servers, as complex electronic devices, consist of multiple components, including processors (CPUs), memory, and hard drives. Different components have different sensitivities to power fluctuations. Power fluctuations may cause unstable processor operating frequencies, leading to slow or error-prone execution of computing tasks; memory may experience data transmission errors, affecting data storage and retrieval efficiency; and hard drive read / write performance may also be affected, leading to data loss or increased access latency.
[0003] For example, Chinese Patent Application No. 202410911366.8 discloses a real-time monitoring system for cloud computing server rooms. The system includes: a network anomaly detection module, a security situation awareness and update module, a power stability monitoring and early warning module, a server performance prediction module, an environmental monitoring and adjustment module, an internal behavior monitoring module, and a dynamic response and security adjustment module. In this invention, by analyzing time-series data, network behavior dependencies are captured, abnormal user behavior is identified, and an external threat database is used to optimize monitoring strategies and security protection. By analyzing power supply parameters and using linear regression algorithms, the impact of power on server performance is analyzed, performance trends are predicted, and operating efficiency is improved. By adjusting environmental parameters and monitoring internal behavior in real time, the stability and security of cloud computing servers are optimized. By adjusting security maintenance strategies, the operation and data security of cloud computing servers are optimized, thereby improving system protection capabilities and efficiency.
[0004] Existing patent documents primarily focus on monitoring the overall operational status of servers, such as server startup status, CPU and memory usage, but lack in-depth analysis and evaluation of the specific impact of power fluctuations on the performance of individual server components. These methods struggle to accurately quantify the extent of power fluctuations' impact on server performance and fail to comprehensively understand the performance trends of each server component under different power fluctuation conditions. Furthermore, traditional methods are inadequate in handling the interrelationships between server components, failing to accurately identify coupling effects between components, resulting in inaccurate and untimely predictions of server performance and the establishment of early warning mechanisms. Summary of the Invention
[0005] This invention addresses the technical problems existing in the prior art by providing a real-time monitoring method and system for cloud computer rooms. By quantifying instantaneous sensitivity values and performance change rates to calculate the performance impact value of power fluctuations, the impact of power fluctuations on server performance can be assessed more accurately.
[0006] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A method for real-time monitoring of cloud computer rooms, comprising:
[0007] S11, collect data and identify abnormal user behavior, split the server into multiple single components, calculate the instantaneous sensitivity value of a single component, the instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single time period, reflecting the sensitivity of component performance to power fluctuations, and calculate the performance change rate based on the instantaneous sensitivity value;
[0008] S12, calculate the impact value of power fluctuation performance based on the instantaneous sensitivity values of all components;
[0009] S201, collect data in the time dimension and calculate the comprehensive sensitivity value; where the comprehensive sensitivity value is the overall state of a single component under different fluctuation frequencies and times;
[0010] S202, Establish an early warning mechanism based on comprehensive sensitivity values;
[0011] Based on the comprehensive sensitivity value, the correlation coefficient between the comprehensive sensitivity values of each individual component is calculated, and the mutual influence relationship is identified according to the correlation coefficient; the correlation coefficient is the Pearson correlation coefficient, which reflects the strength and direction of the linear relationship between two variables;
[0012] The combined sensitivity value is calculated based on the comprehensive sensitivity value and the correlation coefficient.
[0013] Preferably, the instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single time period, reflecting the sensitivity of the component's performance to power fluctuations.
[0014] Preferably, the formula for calculating the impact value of power fluctuation performance is: Where P is the impact value of power fluctuation performance, n is the number of components, and S i Let ΔV be the instantaneous sensitivity value of the i-th component. i Let W be the power fluctuation amplitude corresponding to the i-th component, τ be the duration of the power fluctuation, and W be the voltage fluctuation amplitude. i The weight coefficient of the i-th component.
[0015] Preferably, the comprehensive sensitivity value is calculated based on different fluctuation frequencies and different durations, using the following formula: , among which, S i f is the overall sensitivity value of component i; i t is the fluctuation frequency; i The duration of the abnormal state is denoted by α and β, which are dynamic weighting coefficients. The initial values of α and β are calculated using the analytic hierarchy process (AHP).
[0016] Preferably, the combined sensitivity value is calculated based on the comprehensive sensitivity value and the correlation coefficient, using the following formula: , among which, S c For combined sensitivity values; k is the number of individual components; w i S is the single weight coefficient for component i; i and S j w is the combined sensitivity value of component i and component j; ij r is the combined weight coefficient for component i and component j; ij Let be the correlation coefficient between component i and component j.
[0017] Preferably, the combined sensitivity value is used to quantify the joint sensitivity effect among multiple individual components, reflecting the degree of comprehensive influence caused by the coupling relationship between components.
[0018] Preferably,
[0019] S301 uses monitoring tools to identify the interaction methods of individual components and the order of task processing, and draws a task flowchart;
[0020] S302, based on the task flowchart and historical data, dynamically adjust the working data and resource allocation of the components, and optimize the combined sensitivity value according to the algorithm to obtain the optimal combined sensitivity value;
[0021] S303, based on the combined sensitivity value and the power fluctuation performance impact value, identify the correspondence between the combined sensitivity value and the power fluctuation performance impact value.
[0022] This application also provides a real-time monitoring system for cloud computer rooms, including an abnormal behavior identification module, a relationship identification module, an early warning mechanism establishment module, an identification task process module, and a module for associating combined sensitive values with power fluctuation performance impact values. The module for associating combined sensitive values with power fluctuation performance impact values, the relationship identification module, and the abnormal behavior identification module are all electrically connected to the early warning mechanism establishment module.
[0023] The beneficial effects of this invention are: by quantifying instantaneous sensitivity values and performance change rates to calculate the performance impact value of power fluctuations, the impact of power fluctuations on server performance can be assessed more accurately; by breaking down the server into multiple individual components for analysis, the impact of power fluctuations on the overall server performance can be understood more comprehensively; and by establishing a sensitivity value model, the performance changes of the server under different power fluctuation conditions can be predicted. This provides a more accurate and comprehensive method for assessing the impact of power fluctuations on server performance, which helps system administrators to better optimize server performance and stability.
[0024] By calculating comprehensive sensitivity values and setting thresholds, the system can accurately monitor and warn of the overall status of a single component under different fluctuation frequencies and durations. Based on the changing trends of comprehensive sensitivity values, it can predict the future status of power stability and server performance, improve the monitoring efficiency of power stability, promptly identify and handle unstable factors, improve the prediction accuracy of server performance, optimize resource allocation, and enhance service quality and user experience. By setting thresholds and triggering mechanisms, it can provide early warnings before problems occur, reducing the probability of failures and the scope of their impact.
[0025] The combined sensitivity value comprehensively considers the combined sensitivity value of each individual component and the mutual influence between them. It can more comprehensively reflect the overall state of the system, more comprehensively evaluate the combined state of multiple components and their impact on system performance, provide a basis for system optimization, early warning and maintenance, and improve the reliability and performance of the system.
[0026] By monitoring real-time fluctuations in processor, memory, and hard drive components and dynamically adjusting task flows and resource allocation, the impact of power fluctuations on system performance can be detected more quickly, thereby improving monitoring efficiency. By combining the correlation between sensitivity values and the performance impact values of power fluctuations, a more accurate performance prediction model can be built, improving the prediction accuracy of server performance and enhancing the monitoring efficiency of power stability. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating a real-time monitoring method for a cloud computer room according to the present invention.
[0028] Figure 2 This is a schematic diagram of the process for calculating the comprehensive sensitivity value according to an embodiment of the present invention;
[0029] Figure 3 This is a schematic diagram illustrating the process of optimizing the combined sensitivity value according to an embodiment of the present invention;
[0030] Figure 4 This is a block diagram of a cloud computer room real-time monitoring system according to the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0032] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0033] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0034] Example 1: Figure 1 This is a flowchart illustrating a real-time monitoring method for a cloud computer room according to Embodiment 1 of the present invention, comprising the following steps:
[0035] S101, collect data and identify abnormal user behavior on the cloud computing server based on the data;
[0036] Specifically, the collected data includes network traffic data and user behavior data. For network traffic data, the port mirroring function of network devices (such as routers and switches) is used to obtain data packets transmitted between the cloud computing server and the external network. These data packets include the source IP address and destination IP address. The collection frequency is determined based on actual needs and network scale, including address, source port, destination port, and protocol type. For high-traffic network environments, a higher collection frequency is used to ensure data real-time performance and accuracy. For user behavior data, the Fluentd log collection tool is used to obtain user behavior data from operating system logs, application logs, and user authentication systems on the cloud computing server. Operating system logs record user login, logout, and file operations, while application logs record user operations within specific applications. User behavior data includes user ID, login time, operation type, and operation result. A time series model, specifically a Short-Term Memory (LSTM) network, is used to identify abnormal user behavior on the cloud computing server. Deep learning methods, specifically autoencoders, are used to identify user behavior patterns. These autoencoders learn normal user behavior patterns, and when the reconstruction error between new user behavior data and the normal pattern exceeds a set threshold, the behavior is considered abnormal.
[0037] S102, the server is split into multiple individual components, and the instantaneous sensitivity value of each individual component is calculated;
[0038] Furthermore, the transient sensitivity value is the degree of change of a single component under a single time period (duration) due to a single fluctuation, reflecting the sensitivity of component performance to power fluctuations. The server is broken down into processor, memory, and hard drive components, and the transient sensitivity value of the processor is calculated using the following formula: , among which, S CPU U represents the instantaneous sensitivity value of the processor (CPU) (dimensionless, ranging from 0 to 1). curr This refers to the current CPU utilization rate, which is the percentage of load per unit of time. base The baseline CPU utilization is the value under no load or light load (such as CPU usage when the system is idle). max The CPU's maximum theoretical utilization is represented by 100% for single-core and 100% × number of cores for multi-core. α is a load sensitivity coefficient, adjusted according to task type (e.g., α = 1.2 for compute-intensive tasks, α = 0.8 for I / O-intensive tasks). The impact of instantaneous fluctuations on the CPU is quantified by comparing the difference between the current load and the baseline load. The instantaneous memory sensitivity value is calculated using the following formula: , among which, S Mem B is a transient memory-sensitive value. currB represents the current memory bandwidth, i.e., the amount of data transferred per unit time. base B represents the baseline memory bandwidth and the stable bandwidth under light load. max The theoretical maximum bandwidth of memory is determined by the frequency and the number of channels (e.g., DDR5 6400MHz dual-channel is 128GB / s). curr L represents the current memory latency, the time from when the request is sent to when the data is returned. base L represents the baseline memory latency and the stable latency under light load. max β represents the theoretical minimum latency of memory, a lower limit determined by hardware specifications. Memory performance must consider instantaneous changes in bandwidth and latency. Weight allocation needs to be combined with the application scenario, and the instantaneous sensitivity value of the hard drive is calculated using the following formula: , among which, S Disk For hard drive transient sensitivity, I curr The current IOPS (Input / Output Operations Per Second) reflects random access performance. base As a baseline IOPS (stable value under light load), I max W represents the theoretical maximum IOPS of a hard drive (e.g., NVMe SSDs can reach 1 million). curr Given the current throughput, the sequential read / write performance W base W is the baseline throughput (stable value under light load). max γ represents the theoretical maximum throughput of the hard drive (e.g., approximately 200MB / s for HDD and approximately 3500MB / s for NVMe SSD), and γ is the IOPS sensitive weight.
[0039] S103, based on the instantaneous sensitivity values of each component, uses an algorithm to establish a sensitivity value model for each component and calculates the performance change rate;
[0040] Specifically, based on the instantaneous sensitivity values of the processor, memory, and hard disk components calculated in step S102, a sensitivity value model for each component is established using a neural network model. The obtained instantaneous sensitivity values are divided into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used to adjust the model parameters during training, and the test set is used to finally evaluate the model's performance. The selected model is trained using the training set to determine the network structure, such as the number of layers and the number of neurons per layer. The weights of the network are adjusted using the backpropagation algorithm based on the training data to minimize the error between the model's predicted output and the actual output. The working principles of the processor, memory, and hard disk are analyzed. For the processor, factors such as its architecture design and operating frequency range are considered; for the hard disk, factors such as its read / write mechanism and cache size are considered. Different components have different sensitivities to power fluctuations, and their performance change thresholds will also differ. The performance change trends and limits of the processor, memory, and hard disk components under different power fluctuation conditions are analyzed. Based on the change trends and expert experience, thresholds are set, and the sensitivity of components to power fluctuations is determined by the set thresholds.
[0041] The actual power fluctuation parameters are input into the established sensitivity model, which uses an internal algorithm to calculate the performance change rate of each component under power fluctuations.
[0042] S104, Calculate the impact value of power fluctuation performance based on instantaneous sensitivity value;
[0043] Furthermore, the impact of the instantaneous sensitivity value on the power fluctuation performance is calculated using the following formula: Where P is the impact value of power fluctuation performance, n is the number of components (such as CPU, memory, hard drive, etc.), and S i The instantaneous sensitivity value of the i-th component (e.g., Si of the CPU). CPU Memory S Mem S of hard drive Disk ), ΔV i The voltage fluctuation amplitude corresponding to the i-th component (unit: volts, e.g., voltage drop ΔV = V). nominal V actual ), τ is the duration of the power fluctuation (unit: seconds, such as the duration of a voltage drop), W i The weight coefficient of the i-th component (assigned according to business importance, such as CPU weight WCPU = 0.5, memory weight WMem = 0.3, and hard disk weight WDisk = 0.2).
[0044] The technical solutions described in the embodiments of this application have at least the following technical effects or advantages: by quantifying the instantaneous sensitivity value and the performance change rate to calculate the performance impact value of power fluctuations, the impact of power fluctuations on server performance can be assessed more accurately; by breaking down the server into multiple individual components for analysis, the impact of power fluctuations on the overall server performance can be understood more comprehensively; by establishing a sensitivity value model, the performance changes of the server under different power fluctuation conditions can be predicted; and a more accurate and comprehensive method for assessing the impact of power fluctuations on server performance is provided, which helps system administrators to better optimize server performance and stability.
[0045] Example 2: Based on the changes in a single component under a single fluctuation in a single time period as described in Example 1, this example analyzes the comprehensive state of a single component under different fluctuation frequencies and durations. The interactions between multiple single components cause changes, improving monitoring efficiency and prediction accuracy. Figure 2 As shown.
[0046] S201, collect data in the time dimension and calculate the comprehensive sensitivity value;
[0047] Furthermore, the time-dimensional data includes different fluctuation frequencies and durations. The collected data undergoes preprocessing. The comprehensive sensitivity value refers to the overall state of a single component under different fluctuation frequencies and durations. The comprehensive sensitivity value is calculated based on the collected data for different fluctuation frequencies and durations, using the following formula: , among which, S i f is the overall sensitivity value of component i; i The frequency of fluctuation, such as the number of fluctuations per second of current; t i The duration of the abnormal state; α, β: dynamic weighting coefficients, the initial values of α, β are determined by the analytic hierarchy process.
[0048] S202, establish an early warning mechanism based on the comprehensive sensitivity value; calculate the correlation coefficient between the comprehensive sensitivity values of each individual component based on the comprehensive sensitivity value, and identify the mutual influence relationship based on the correlation coefficient;
[0049] Specifically, the calculated comprehensive sensitivity values of each individual component are compiled into a table. The Pearson correlation coefficient formula is used to calculate the correlation coefficient between the comprehensive sensitivity values of each individual component. The Pearson correlation coefficient formula is existing technology and can be easily understood by those skilled in the art. The correlation coefficient reflects the strength and direction of the linear relationship between two variables, and its value ranges from -1 to 1. When the correlation coefficient |r| ≥ 0.7, there is a strong influence relationship between the comprehensive sensitivity values of the two components, that is, the state change of one component will significantly affect the comprehensive sensitivity value of the other component. When 0.3 ≤ |r| < 0.7, there is a weak influence relationship between the comprehensive sensitivity values of the two components, that is, the state change of one component has some influence on the comprehensive sensitivity value of the other component, but the degree of influence is not large. When |r| < 0.3, there is no significant influence relationship between the comprehensive sensitivity values of the two components.
[0050] Furthermore, based on the calculated comprehensive sensitivity value, historical data related to the comprehensive sensitivity value is obtained from various channels such as system logs, database records, and monitoring tools. The standard deviation of the historical data is calculated to determine the normal fluctuation range of the comprehensive sensitivity value in the historical data. Based on the time series analysis of the historical data, a line graph of the comprehensive sensitivity value changing over time is plotted to identify the periods when the comprehensive sensitivity value deviates from the normal fluctuation range. Based on the relevant data of the periods when the abnormal value deviates from the normal fluctuation range, the conditions that cause the comprehensive sensitivity value to deviate from the fluctuation range are identified. These conditions include a sudden increase in business volume, network attacks, and power failures. Based on the normal fluctuation range of the comprehensive sensitivity value, the upper limit of the normal fluctuation range of the comprehensive sensitivity value is set as a threshold. When the comprehensive sensitivity value exceeds the set threshold, an early warning mechanism is automatically triggered. When an early warning is triggered, a higher level of monitoring and early warning process is initiated, such as increasing the monitoring frequency and initiating fault diagnosis procedures.
[0051] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: by calculating the comprehensive sensitivity value and setting the threshold, accurate monitoring and early warning of the comprehensive state of a single component under different fluctuation frequencies and durations can be achieved. Based on the changing trend of the comprehensive sensitivity value, the future state of power stability and server performance can be predicted, thereby improving the monitoring efficiency of power stability, timely detection and handling of unstable factors, improving the prediction accuracy of server performance, optimizing resource allocation, and enhancing service quality and user experience. By setting thresholds and triggering mechanisms, early warnings can be given before problems occur, reducing the probability of failure and the scope of impact.
[0052] Example 3: Calculate the combined sensitivity value based on the correlation coefficient between the comprehensive sensitivity value calculated in Example 2 and the comprehensive sensitivity values of each individual component.
[0053] Specifically, the combined sensitivity value is used to quantify the joint sensitivity effect among multiple components, reflecting the comprehensive impact caused by the coupling relationship between components. The combined sensitivity value is calculated based on the comprehensive sensitivity value of a single component and the correlation coefficient between components, using the following formula: , among which, S c For combined sensitivity values; k is the number of individual components; w i S represents the single weight coefficient of component i, reflecting the importance of this component in the composition. The weight coefficients are confirmed using the analytic hierarchy process (AHP). i and S j w is the combined sensitivity value of component i and component j; ij The combined weight coefficients for components i and j reflect the degree to which the interaction between individual components contributes to the combined sensitivity value; r ij Let be the correlation coefficient between component i and component j.
[0054] A specific example is as follows: In a large data center, the stable operation of the server system is crucial. This server system mainly consists of three key components: CPU (Central Processing Unit), memory, and hard disk. We select these three components, numbered i1 (CPU), i2 (memory), and i3 (hard disk), respectively, i.e., k=3. Through the preceding steps, we have calculated the overall sensitivity value S of the CPU. i1 =0.6, which reflects the overall state of the CPU under different fluctuation frequencies (such as changes in task processing frequency) and different durations (such as high load duration); the overall sensitivity value of memory S i2 =0.5, reflecting the impact of memory on system performance, such as fluctuations in memory usage; the overall sensitivity value S of the hard drive. i3 =0.4, representing the overall state of hard drive read / write performance, etc., and the correlation coefficient between them is calculated: the correlation coefficient r between CPU and memory. i1i2 =0.8, indicating a strong mutual influence between them; for example, the CPU needs to frequently interact with memory when processing tasks. The correlation coefficient r between the CPU and hard drive is... i1i3 =0.3 indicates that there is some correlation between them, but it is relatively weak. For example, the CPU will have some influence when processing tasks involving hard disk read and write; the correlation coefficient r between memory and hard disk is... i2i3 =0.5, indicating that there is some mutual influence between them, such as the relationship between memory data caching and hard disk read / write operations. The initial values obtained through the analytic hierarchy process are w. i1 =0.4, w i2 =0.3, w i3 =0.3, and similarly, w is obtained through the analytic hierarchy process. i1i2 =0.4, wi1i3 =0.2, w i2i3 =0.3, =0.4×0.6+0.3×0.5+0.3×0.4+0.096+0.0144+0.03=0.6504.
[0055] The technical solutions in the above embodiments of this application have at least the following technical effects or advantages: the combined sensitivity value comprehensively considers the comprehensive sensitivity values of each individual component and the mutual influence between them, which can more comprehensively reflect the overall state of the system, more comprehensively evaluate the comprehensive state between multiple components and their impact on system performance, provide a basis for system optimization, early warning and maintenance, and improve the reliability and performance of the system.
[0056] Example 4: Based on Examples 2 and 3 above, this example dynamically adjusts the operating parameters or resource allocation of processor, memory, and hard disk components by identifying task flows and fluctuations between components. It determines the relationship between combined sensitivity values and the impact values of power fluctuation performance, thereby improving the monitoring efficiency of power stability and the prediction accuracy of server performance. Figure 3 As shown.
[0057] S301, using monitoring tools to identify the interaction methods of components and the order of task processing, and drawing task flowcharts;
[0058] Furthermore, the monitoring tool Prometheus is used to monitor the state changes of the processor, memory, and hard disk components in real time during task processing. Based on the nature and execution process of the task, the task is divided into different stages. Each task stage should have start and end markers, as well as specific functions and objectives. For each task stage, the interaction mode between the processor, memory, and hard disk components is identified. For example, in the data reading stage, the processor sends a read request to the hard disk, the hard disk reads the data into memory, and the processor then retrieves the data from memory for processing. The dependencies between the processor, memory, and hard disk are determined, that is, the operation of one component can only be performed after other components have completed specific operations.
[0059] Use Microsoft Visio to draw task flowcharts. Use rectangles to represent task stages, and label the name and main function of each task stage inside the rectangle. Use arrows to indicate the interaction and data flow between components. For example, an arrow from the processor to the memory indicates that the processor writes data to the memory, and an arrow from the hard drive to the processor indicates that the hard drive provides data to the processor. Clearly label the processor, memory, and hard drive components in the flowchart, using different colors or shapes to distinguish them. Next to each component, label its specific work in different task stages, such as the processor executing a specific algorithm in a certain stage, or the memory being used to store specific types of data.
[0060] S302, based on the task flowchart and historical data, dynamically adjust the working data and resource allocation of the components, and optimize the combined sensitivity value according to the algorithm to obtain the optimal combined sensitivity value;
[0061] Specifically, based on the drawn task flowchart, the normal operating load range of processor, memory, and hard disk components in different task stages is identified. In the data-intensive task stage, the utilization rate of processor and memory will be relatively high; while in the data storage stage, hard disk I / O operations will be more frequent. Performance index data of processor, memory, and hard disk components during server operation over a period of time are extracted from the performance monitoring database. The data should include performance under different time periods and different workloads. Thresholds are set based on the statistical characteristics of historical data. When the processor, memory, and hard disk exceed the set thresholds, the working data and resource allocation of the processor, memory, and hard disk are dynamically adjusted. Memory allocation is increased to improve the processor's data processing capability, or task priorities are adjusted to reduce the execution of low-priority tasks.
[0062] A random forest model is constructed using machine learning algorithms based on historical data and calculated combined sensitivity values. The historical data is divided into training and test sets to ensure that the data distribution of the training and test sets is representative. The model is trained using the training and test sets, and the genetic algorithm in the trained model is used to search for the optimal value of the combined sensitivity value. The calculated combined sensitivity value is compared with the optimal value of the combined sensitivity value, and a deviation range is set. When the calculated combined sensitivity value is no longer within the set deviation range, the working data and resource allocation of the processor, memory, and hard disk are dynamically adjusted.
[0063] S303, based on the combined sensitivity value and the power fluctuation performance impact value, identify the correspondence between the combined sensitivity value and the power fluctuation performance impact value;
[0064] Furthermore, based on the combined sensitivity value and the impact value of power fluctuation performance, the distribution pattern of the data is observed. A neural network model is used to construct a correlation model between the combined sensitivity value and the impact value of power fluctuation performance. The model is trained using historical combined sensitivity values and the impact value of power fluctuation performance. The data is divided into training set and validation set. The correlation model is trained using the training set and the model's fitting effect is evaluated using the validation set. The trained and evaluated correlation model is obtained. The combined sensitivity value in the actual monitoring data is obtained and input into the correlation model. The correlation model outputs the prediction result of the impact value of power fluctuation performance.
[0065] The technical solutions described in the above embodiments of this application have at least the following technical effects or advantages: by monitoring the fluctuations of processor, memory and hard disk components in real time and dynamically adjusting task flow and resource allocation, the impact of power fluctuations on system performance can be detected more quickly, thereby improving monitoring efficiency. By combining the correlation between sensitive values and the performance impact values of power fluctuations, a more accurate performance prediction model can be constructed, improving the prediction accuracy of server performance and improving the monitoring efficiency of power stability.
[0066] Example 5: Figure 4 As shown, this embodiment provides a real-time monitoring system for cloud computer rooms, including an abnormal behavior identification module, a relationship identification module, an early warning mechanism establishment module, an identification task flow module, and a module for associating combined sensitivity values with power fluctuation performance impact values. The abnormal behavior identification module is used to collect network traffic data and user behavior data, and identify abnormal user behavior in the cloud computing server by analyzing data packets and user behavior logs. The relationship identification module is used to calculate the comprehensive sensitivity values of each component and identify the mutual influence relationships between them. The early warning mechanism establishment module is used to establish an early warning mechanism based on the comprehensive sensitivity values and abnormal behavior identification results. The identification task flow module is used to use monitoring tools to identify the interaction methods and sequences of processor, memory, and hard disk components when processing tasks, draw task flow diagrams, and clarify the dependencies and data flow between components. The module for associating combined sensitivity values with power fluctuation performance impact values is used to construct an association model between combined sensitivity values and power fluctuation performance impact values using a neural network model, and train the model using historical data to predict the impact of power fluctuations on system performance. The module for associating combined sensitivity values with power fluctuation performance impact values, the relationship identification module, and the abnormal behavior identification module are all electrically connected to the early warning mechanism establishment module.
[0067] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0069] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0072] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0073] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for real-time monitoring of a cloud computer room, characterized in that, include: S11, collect data and identify abnormal user behavior, split the server into multiple single components, calculate the instantaneous sensitivity value of a single component, the instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single time period, reflecting the sensitivity of component performance to power fluctuations, and calculate the performance change rate based on the instantaneous sensitivity value; S12, calculate the impact value of power fluctuation performance based on the instantaneous sensitivity values of all components; S201, collect data in the time dimension and calculate the comprehensive sensitivity value; where the comprehensive sensitivity value is the overall state of a single component under different fluctuation frequencies and times; S202, Establish an early warning mechanism based on comprehensive sensitivity values; Based on the comprehensive sensitivity value, the correlation coefficient between the comprehensive sensitivity values of each individual component is calculated, and the mutual influence relationship is identified according to the correlation coefficient; the correlation coefficient is the Pearson correlation coefficient, which reflects the strength and direction of the linear relationship between two variables; The combined sensitivity value is calculated based on the comprehensive sensitivity value and the correlation coefficient.
2. The method for real-time monitoring of a cloud computer room according to claim 1, characterized in that, The instantaneous sensitivity value is the degree of change of a single component under a single fluctuation in a single time period, reflecting the sensitivity of component performance to power fluctuations.
3. The method for real-time monitoring of a cloud computer room according to claim 1, characterized in that, The formula for calculating the performance impact value of power fluctuations is: Where P is the impact value of power fluctuation performance, n is the number of components, and S i Let ΔV be the instantaneous sensitivity value of the i-th component. i Let W be the power fluctuation amplitude corresponding to the i-th component, τ be the duration of the power fluctuation, and W be the voltage fluctuation amplitude. i The weight coefficient of the i-th component.
4. The method for real-time monitoring of a cloud computer room according to claim 1, characterized in that, The comprehensive sensitivity value is calculated based on different fluctuation frequencies and durations, using the following formula: , among which, S i f is the overall sensitivity value of component i; i t is the fluctuation frequency; i The duration of the abnormal state is denoted by α and β, which are dynamic weighting coefficients. The initial values of α and β are calculated using the analytic hierarchy process (AHP).
5. The method for real-time monitoring of a cloud computer room according to claim 1, characterized in that, The combined sensitivity value is calculated based on the comprehensive sensitivity value and the correlation coefficient, using the following formula: , among which, S c For combined sensitivity values; k is the number of individual components; w i S is the single weight coefficient for component i; i and S j w is the combined sensitivity value of component i and component j; ij r is the combined weight coefficient for component i and component j; ij Let be the correlation coefficient between component i and component j.
6. The method for real-time monitoring of a cloud computer room according to claim 5, characterized in that, The combined sensitivity value is used to quantify the joint sensitivity effect among multiple individual components, reflecting the degree of comprehensive influence caused by the coupling relationship between components.
7. A method for real-time monitoring of a cloud computer room according to claim 6, characterized in that, S301 uses monitoring tools to identify the interaction methods of individual components and the order of task processing, and draws a task flowchart; S302, based on the task flowchart and historical data, dynamically adjust the working data and resource allocation of the components, and optimize the combined sensitivity value according to the algorithm to obtain the optimal combined sensitivity value; S303, based on the combined sensitivity value and the power fluctuation performance impact value, identify the correspondence between the combined sensitivity value and the power fluctuation performance impact value.
8. A cloud computer room real-time monitoring system, applied to the cloud computer room real-time monitoring method as described in any one of claims 1-7, characterized in that, It includes an abnormal behavior identification module, a relationship identification module, an early warning mechanism establishment module, an identification task process module, and a module for associating combined sensitive values with the impact value of power fluctuation performance. The module for associating combined sensitive values with the impact value of power fluctuation performance, the relationship identification module, and the abnormal behavior identification module are all electrically connected to the early warning mechanism establishment module.
Citation Information
Patent Citations
Real-time monitoring system for cloud computing server room
CN118509248A
Server fault prediction method, system and device and computer readable storage medium
CN118051408A
Electric power data processing system based on cloud computing
CN118312908A