CPU frequency monitoring method in cloud computing scene and related device

By using a lightweight module to obtain CPU frequency from the host core in cloud computing scenarios, monitoring the virtual machine CPU in real time and generating down-frequency alarms, the problems of insufficient monitoring interference and accuracy in the existing technology are solved, and efficient and stable monitoring and alarm feedback on the virtual machine CPU frequency is achieved.

CN120371460AActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510519554.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-25
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the scenario of full resource utilization of cloud computing, the existing CPU frequency monitoring methods have performance interference and insufficient monitoring accuracy, and it is impossible to accurately distinguish the virtual machine CPU running on the host side from other non-virtual machine CPU resources, resulting in the impact of the performance stability and reliability of the virtual machine.

Method used

The lightweight module is used to obtain the calibration frequency of the CPU from the host kernel, and the actual frequency of the virtual machine CPU is monitored in real time through the CPUfreq-monitor module in the kernel space, compared with the calibration frequency, generate a down-frequency alarm information, and send it to the smart card for user feedback.

Benefits of technology

It realizes accurate monitoring of the CPU frequency of virtual machines, reduces system resource usage, ensures the stability of virtual machines, provides detailed alarm information so that users can handle CPU exceptions in a timely manner, and improves monitoring efficiency and system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371460A_ABST
    Figure CN120371460A_ABST
Patent Text Reader

Abstract

The invention discloses a CPU frequency monitoring method in a cloud computing scene and a related device, which are applied to a virtual machine. The method comprises the following steps: acquiring a lightweight module for CPU frequency monitoring from a kernel of a host machine; obtaining a calibration frequency of a virtual machine CPU based on the lightweight module; the actual frequency of the virtual machine CPU is compared with the calibration frequency based on a lightweight module, and frequency reduction alarm information of the virtual machine CPU is obtained; and sending the underclocking alarm information to the smart card based on the lightweight module, and sending the underclocking alarm information to a user by the smart card. By means of the mode, the performance influence caused by switching of the user space and the kernel space is reduced, and efficient monitoring of the frequency of the host machine side CPU is achieved while the stability of the virtual machine is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and particularly to a method and related device for monitoring CPU frequency in a cloud computing scenario, which is applied to virtual machines. Background Art

[0002] In the traditional cloud computing scenario, cloud service providers use virtualization technology to divide a single physical server into multiple virtual machines for users to use. However, in this virtualized scenario, not all physical CPU cores on the server can be fully allocated to the user's virtual machines. Because cloud service providers need to reserve a portion of CPU resources to run some network virtualization processes, storage virtualization processes, system monitoring processes, etc. The processes running on these reserved CPU resources ensure the stability, security, and efficiency of cloud services, but at the same time, they also bring overhead of virtualized resources. Now, a relatively common idea in the industry to reduce this overhead of virtualized resources is to offload these processes and some other non-essential processes from the physical CPU to a smart card, and let the smart card undertake tasks such as network virtualization, storage virtualization, and monitoring, and only let the host run some necessary processes to maintain the operation of the virtual machine. This scenario of selling all the CPU resources of the host to virtual machines as much as possible is called the full resource utilization scenario.

[0003] In the scenario of full resource utilization in cloud computing, only necessary processes that can maintain the operation of the host and virtual machines should be run on the physical CPU on the host side, and other processes should be trimmed or offloaded to the smart card as much as possible to reduce the jitter caused to the virtual machine. Therefore, some monitoring processes in the traditional virtualized scenario are not suitable to be placed on the host in this scenario, because these processes cause too much jitter to virtualization, affect the stability of virtual machine operation, and will bring a bad experience to users.

[0004] In the scenario of full resource utilization in cloud computing, there are strict requirements for accurate monitoring of CPU frequency, because it is directly related to the stability and reliability of virtual machine performance. However, in the existing traditional virtualized scenario, the CPU frequency monitoring method has disadvantages such as performance interference and insufficient monitoring accuracy. First, the existing CPU frequency monitoring method is mainly called in the user space. This way causes frequent switching between the user space and the kernel space, seriously affecting the system performance, and thus significantly interfering with the virtual machine performance in this scenario. Second, the traditional method cannot accurately distinguish the CPU running the virtual machine on the host side from other non-virtual machine CPU resources, and will treat all CPU resources equally when performing frequency monitoring, resulting in a large number of ineffective monitoring operations. Third, currently, most monitoring processes only periodically record the running frequency of the CPU, and cannot actively alarm when the CPU frequency is abnormal. Summary of the Invention

[0005] In view of this, embodiments of the present application provide a CPU frequency monitoring method, apparatus, electronic device, computer storage medium, and computer program product in a cloud computing scenario, aiming to solve related problems such as performance interference and insufficient monitoring accuracy of traditional CPU frequency monitoring methods.

[0006] In a first aspect, embodiments of the present application provide a CPU frequency monitoring method in a cloud computing scenario, which is applied to a virtual machine. The method includes:

[0007] Obtain a lightweight module for CPU frequency monitoring from the host kernel; obtain the calibrated frequency of the CPU based on the lightweight module; compare the actual frequency and the calibrated frequency of the CPU based on the lightweight module to obtain the CPU down-frequency warning information; send the down-frequency warning information to a smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user.

[0008] In some embodiments, the obtaining a lightweight module for CPU frequency monitoring from the host kernel includes: searching for and loading the lightweight module from a module storage path specified by the host kernel through a kernel module loading mechanism.

[0009] In some embodiments, the obtaining the calibrated frequency of the CPU based on the lightweight module includes: the lightweight module obtains the calibrated frequency of the CPU by reading the hardware register of the virtual machine CPU.

[0010] In some embodiments, before comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module, the method further includes: the lightweight module collects the actual frequency of the virtual machine CPU at a preset time interval.

[0011] In some embodiments, the comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module includes: calculating the difference between the calibrated frequency and the actual frequency based on the lightweight module, and comparing whether the difference exceeds a preset threshold.

[0012] In some embodiments, the obtaining the CPU down-frequency warning information includes: if the difference is greater than a preset threshold, generating down-frequency warning information; the down-frequency warning information includes the CPU number of the down-frequency, the down-frequency amplitude, and the time when the down-frequency occurs.

[0013] In some embodiments, before obtaining the calibrated frequency of the CPU based on the lightweight module, the method further includes: obtaining a mask vector of the CPU; the lightweight module obtains the CPU to be monitored for frequency based on the mask vector.

[0014] In some embodiments, obtaining the mask vector of the CPU includes: obtaining the number of virtual machines on each CPU; setting the mask of the CPUs with the number of virtual machines greater than zero to 1.

[0015] In some embodiments, the lightweight module is periodically awakened to perform the following operations: first, obtain the actual performance counter value and the maximum performance counter value of the current CPU, then sleep for 1 second, and then obtain the actual performance counter value and the maximum performance counter value of the current CPU again, and calculate the actual frequency of the CPU according to the actual performance counter values and the maximum performance counter values of the two samplings.

[0016] In a second aspect, an embodiment of the present application provides a CPU frequency monitoring device in a cloud computing scenario, which is applied to a virtual machine. The device includes:

[0017] An initialization module that obtains a lightweight module for CPU frequency monitoring from the host kernel;

[0018] A calibrated frequency acquisition module that obtains the calibrated frequency of the CPU based on the lightweight module;

[0019] A frequency comparison module that compares the actual frequency of the CPU with the calibrated frequency based on the lightweight module to obtain the CPU down-frequency warning information;

[0020] A down-frequency warning module that sends the down-frequency warning information to a smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor. Among them, when the processor runs the computer program, it executes the method described in the first aspect above.

[0022] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method described in the first aspect above.

[0023] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by an electronic device, it implements the steps of the method described in the first aspect above.

[0024] An embodiment of the present application provides a CPU frequency monitoring method, device, electronic device, computer storage medium, and computer program product in a cloud computing scenario. The present application solves the following technical problems.

[0025] In the cloud computing full resource utilization scenario, the virtual machine CPU frequency is effectively monitored, and the CPU frequency reduction alarm information is obtained in time and fed back to the user. The lightweight module loading method adopted by the present invention greatly reduces the occupation of system resources. Running in the kernel space enables the monitoring process to interact more closely with the hardware and capture subtle changes in the CPU frequency in real time without causing any interference to the normal operation of the virtual machine. This almost "invisible" monitoring mode provides unprecedented stability guarantees for cloud computing service providers, ensuring that in a multi-tenant environment, each virtual machine can obtain stable computing resources and avoid performance fluctuations caused by monitoring behavior.

[0026] The present invention improves the precision monitoring capability. By accurately identifying the running status of the virtual machine, the present invention can quickly locate the CPU core running the virtual machine and carry out targeted frequency monitoring. This precision-focused monitoring strategy not only improves monitoring efficiency and reduces system overhead, but also enables cloud computing operation and maintenance personnel to clearly understand the CPU performance status of each virtual machine, providing accurate data support for resource allocation and scheduling.

[0027] The active alarm mechanism improves the security of the system. Once the CPU frequency is abnormal, the system will immediately push the alarm information to the user. At the same time, the alarm information also comes with detailed CPU status data, including the normal frequency before the frequency reduction, the frequency reduction range, the time of occurrence, etc., to help users quickly determine the severity of the problem and take targeted solutions, effectively reducing the risk of business interruption caused by CPU failure and ensuring the high availability of cloud computing services.

[0028] With the continuous development of cloud computing technology, the present invention is expected to play a role in a wider range of fields, such as edge computing, big data processing clusters, etc., and continue to provide reliable technical support for CPU frequency monitoring in complex computing environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 A flowchart of a method for monitoring CPU frequency in a cloud computing scenario provided by an embodiment of the present application;

[0030] Figure 2 A general flow chart of a method for monitoring CPU frequency in a cloud computing scenario provided by an embodiment of the present application;

[0031] Figure 3 A host-side flow chart of a CPU frequency monitoring method in a cloud computing scenario provided by an embodiment of the present application;

[0032] Figure 4 A lightweight module flow chart of a CPU frequency monitoring method in a cloud computing scenario provided by an embodiment of the present application;

[0033] Figure 5 Schematic diagram of the CPU frequency monitoring device provided by the embodiment of the present application in a cloud computing scenario;

[0034] Figure 6 Schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners

[0035] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0037] As Figure 1 shown, the embodiment of the present application provides a CPU frequency monitoring method in a cloud computing scenario, which is applied to a virtual machine. The method includes the following steps:

[0038] Step S110: Obtain a lightweight module for CPU frequency monitoring from the host kernel.

[0039] This step is the starting point of the entire monitoring method. The lightweight module is the basis for subsequent implementation of various functions of CPU frequency monitoring. By obtaining this module, it provides the necessary tools for subsequent operations such as obtaining the calibrated frequency and actual frequency of the CPU, as well as frequency comparison and alarm.

[0040] Step S120: Obtain the calibrated frequency of the CPU based on the lightweight module.

[0041] The calibrated frequency is an important reference value for subsequent frequency comparison, and it provides a benchmark for judging whether the CPU is downclocked.

[0042] Step S130: Compare the actual frequency of the CPU with the calibrated frequency based on the lightweight module, and obtain the downclocking alarm information of the CPU.

[0043] Through such a comparison operation, it can be timely discovered whether the CPU is downclocked, and the downclocking alarm information can accurately inform the user of the specific information related to downclocking.

[0044] Step S140: Send the downclocking alarm information to the smart card based on the lightweight module, and the smart card sends the downclocking alarm information to the user.

[0045] This enables the user to timely obtain the downclocking situation of the CPU so as to take corresponding measures.

[0046] In some embodiments, the method includes obtaining a lightweight module for CPU frequency monitoring from the host kernel.

[0047] The lightweight module is the cpufreq - monitor module. The cpufreq - monitor module is a kernel module designed specifically for monitoring the host CPU frequency in the scenario of selling all cloud computing resources. Its function is to implement the frequency monitoring function for the specified CPU. After loading this module, a delayed_work (delayed work) is created for each CPU, and these delayed_work are periodically awakened to perform subsequent frequency monitoring operations.

[0048] In a virtual machine environment, the CPU mask is a mechanism for specifying which physical CPU cores a virtual machine can use. Usually, one flag bit corresponds to one CPU core, and by setting the corresponding bit, it indicates whether the core is used by the virtual machine.

[0049] For example, for a system with 8 CPU cores, the CPU mask can be an 8 - bit data structure. If the mask is 10101010, it means that there are virtual machines on the 2nd, 4th, 6th, and 8th CPU cores, while the 1st, 3rd, 5th, and 7th cores do not. In this way, the access to physical CPU resources can be precisely controlled, achieving reasonable resource allocation and isolation, and improving the performance and stability of the system.

[0050] For example, libvirt is an open - source tool for managing virtualization platforms. When libvirt starts a virtual machine, libvirt senses that there is a new virtual CPU (vCPU) running on a physical CPU, then calculates the cpumask (CPU mask) and writes it to the interface exposed by the cpufreq - monitor module; after receiving the cpumask, the cpufreq - monitor module starts to monitor the frequency of the specified CPU.

[0051] By obtaining this lightweight module, it is possible to minimize the interference to the virtual machine performance while ensuring the monitoring efficiency, and achieve lightweight monitoring of the host CPU frequency.

[0052] This method can effectively monitor the operating frequency of the host - side CPU in the scenario of full - resource utilization of cloud computing. By working in coordination with libvirt, it can accurately monitor the frequency of the CPU running the virtual machine. Once it detects that the CPU has a frequency drop, this method can actively send an alarm to the user to ensure that the user can detect and handle related problems in a timely manner.

[0053] In some embodiments, the method includes obtaining the calibrated frequency of the CPU based on the lightweight module.

[0054] The lightweight module cpufreq - monitor obtains the calibrated frequency of the CPU through a specific mechanism, and this calibrated frequency serves as the benchmark for subsequent comparison of the actual operating frequency of the CPU.

[0055] For example, after the cpufreq - monitor module is loaded and receives the cpumask passed in by libvirt to start monitoring the specified CPU, it will obtain the calibrated frequency of the CPU according to the internally set process.

[0056] Obtaining the calibrated frequency of the CPU provides an accurate reference standard for subsequent judgment of whether the CPU is down - clocked, making the down - clocking judgment more accurate and reliable.

[0057] In some embodiments, the method includes comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module to obtain the down - clocking warning information of the CPU.

[0058] APERF (Actual Performance Frequency) and MPERF (Maximum Performance Frequency) are hardware performance counters provided by Intel and AMD CPUs, which are used to measure the actual operating frequency of the CPU and the number of operating cycles at the maximum frequency. The APERF counter records the number of cycles of the CPU in the actual operating state. It reflects the actual operating frequency of the CPU under the current load and frequency adjustment strategy. The MPERF counter records the number of cycles of the CPU at the maximum frequency (i.e., the nominal main frequency). It represents the number of operating cycles of the CPU in the ideal state. Therefore, the current operating frequency of the CPU can be obtained by calculating through these two counters.

[0059] The cpufreq - monitor module obtains the aperf and mperf values of the current CPU. Through two samplings (first obtain the aperf and mperf values of the current CPU, then sleep for 1 second, and obtain the aperf and mperf values of the current CPU again), calculate the current operating frequency of this CPU according to the aperf and mperf values of the two samplings, and compare this actual operating frequency with the calibrated frequency. When it is detected that the actual operating frequency of a certain CPU is lower than the preset threshold (i.e., the threshold related to the calibrated frequency), an alarm is triggered to obtain the down - clocking warning information.

[0060] For example, during the execution of the delayed_work that wakes up periodically, in accordance with the above - mentioned method of sampling and calculating the frequency, compare with the calibrated frequency. If the frequency is lower than the preset threshold, it is considered that down - clocking occurs, and the module will print an alarm log.

[0061] By comparing the actual frequency and the calibrated frequency of the CPU in this way, it is possible to promptly detect the CPU frequency reduction situation, obtain the frequency reduction warning information, so as to notify the user in a timely manner to take corresponding measures.

[0062] In some embodiments, the method includes sending the frequency reduction warning information to the smart card based on the lightweight module, and the smart card sends the frequency reduction warning information to the user.

[0063] After the cpufreq - monitor module obtains the frequency reduction warning information (prints the warning log), it will forward the warning information to the smart card side. After receiving the warning information, the smart card sends an alarm message to the user according to its set communication method.

[0064] For example, after the cpufreq - monitor module detects the CPU frequency reduction and prints the warning log, it sends the warning log to the smart card through a specific communication interface, and the smart card then sends the frequency reduction warning information to the user by means of text messages, push messages, etc.

[0065] Through this information transmission method, it is possible to promptly feedback the situation of CPU frequency reduction to the user, enabling the user to timely understand the system status and make corresponding decisions.

[0066] As Figure 2 shown, it is the overall flowchart of the CPU frequency monitoring method in the cloud computing scenario. The host, smart card, and user are closely associated through a series of collaborative processes. The host, as the hardware foundation, runs components such as the cpufreq - monitor module in the kernel space and qemu, libvirt, etc. in the user space on it. QEMU (Quick Emulator) is an open - source general - purpose hardware simulation and virtualization management component. The cpufreq - monitor module is responsible for monitoring the frequencies of the physical CPUs (cpu0, cpu1, etc.) in the hardware layer of the host and transmitting the frequency reduction warning information related to the CPU frequencies collected during the monitoring to the smart card. The smart card plays its log filtering function, screening out the frequency reduction warning information from a large amount of log data, and these frequency reduction warning information often reflect abnormal problems such as CPU frequency reduction. Finally, the user is at the receiving end of the entire system, and the abnormal information filtered by the smart card will trigger the alarm mechanism, promptly feedbacking these abnormal situations to the user so that the user can quickly take measures to deal with them and ensure the stable operation of the virtual machines in the cloud computing environment.

[0067] In the virtual machine CPU frequency monitoring system in the cloud computing environment, the host covers three important levels: the user space, the kernel, and the hardware, and they each undertake key responsibilities in the entire monitoring system and cooperate with each other to ensure the system operation.

[0068] In user space, the QEMU module is used for machine emulation and virtualization. It emulates the CPU running environment of the virtual machine through multiple virtual CPUs (vCPUs), providing basic support for the operation of the virtual machine. Libvirt, on the other hand, is used to manage the open-source API, daemon process, and tools of the virtualization platform. By interacting with the kernel space through the CPU mask (cpumask), it can accurately identify the CPUs running the virtual machines, laying the foundation for subsequent precise monitoring work.

[0069] The cpufreq-monitor module at the kernel level is the core for implementing CPU frequency monitoring. It runs in the kernel space, obtains the CPU information to be monitored through cpumask, and then monitors the frequencies of the physical CPUs (such as cpu0, cpu1, cpu2... cpuN) at the host hardware layer, collects CPU frequency-related log information, and passes this information to the smart card for further processing.

[0070] At the hardware level, it is the actual physical CPU of the host machine, which is the target object of the entire monitoring system. The running frequency of the physical CPU directly affects the performance of the virtual machine. By monitoring its frequency, the system can promptly detect and handle possible anomalies such as CPU frequency throttling, ensuring the stable operation of the cloud computing environment.

[0071] As Figure 3 shown, it is a schematic diagram of the host side process of the CPU frequency monitoring method in the cloud computing scenario. This diagram is a refined version of the host side of the overall flowchart, describing the collaboration relationships of modules such as Libvirt, QEMU, and cpufreq-monitor at the user space, kernel, and hardware levels. Figure 2

[0072] libvirt is the core management module. First, two key variables, cpumask and ref, are defined. Among them, cpumask represents the CPUs to be monitored that are finally notified to the cpufreq - monitor module, and ref represents the number of vCPUs running on each CPU. The operations of the libvirt on the virtual machine QEMU process include start (virsh start), destroy (virsh destroy), shutdown (virsh shutdown), live migration, etc. In addition, the situation where QEMU exits abnormally needs to be handled. In the functions of these operations, the processing logic for cpumask and ref is added. For example, when starting a new vCPU on CPU X, check whether the X bit in cpumask is 1. If it is 0, set it to 1; if it is 1, increment the ref count. If cpumask changes, write the new cpumask to the cpufreq - monitor module. When destroying a vCPU on CPU X, decrement the ref count. If the ref count is 0, set the X bit in cpumask to 0 and write the new cpumask to the cpufreq - monitor module. In addition, the situation of the libvirt process restart needs to be considered. After the libvirt restarts, traverse and reconnect each domain to obtain the corresponding cpumask, recalculate cpumask and ref, and then write the new cpumask to the cpufreq - monitor module.

[0073] Among them, cpumask and ref are used to manage the CPU mask reference and interact with the kernel through cpumask. In the qemu - related operation modules, the qemuProcessLaunch module (QEMU process start module) is responsible for starting the qemu instance, which is achieved through the virsh start command or by live migration operation to migrate the qemu instance. The qemuProcessStop module (QEMU process stop module) is used to stop the qemu instance, which is achieved by means of the virsh destroy command. The processMonitorEOFEvent (process end event monitoring) is used to monitor the qemu process end event, which is triggered when the qemu instance is shut down through virsh shutdown and is also associated with this component in case of an exception. The qemuProcessReconnect module (QEMU process reconnect module) is used to handle matters related to the qemu process reconnect.

[0074] The libvirt module connects the user space and the kernel through sysfs. Sysfs is a virtual file system used to transfer information between the kernel space and the user space. Here, it is responsible for transferring cpumask information and connecting the libvirt and the cpufreq - monitor module in the kernel.

[0075] The cpufreq - monitor is a CPU frequency monitoring module in the kernel. It receives cpumask information from libvirt through sysfs to determine the CPUs to be monitored, and then monitors the frequencies of physical CPUs such as cpu0, cpu1, cpuX, etc. at the hardware layer.

[0076] As Figure 4 shown, it is a schematic diagram of the lightweight module of the CPU frequency monitoring method in the cloud computing scenario provided by the embodiment of the present application.

[0077] The lightweight module cpufreq - monitor processes based on the delayed - work mechanism in CPU frequency monitoring.

[0078] Figure 4 shows that different physical CPUs (cpu0, cpu1, ……, cpuN) respectively correspond to their own aperf and mperf metrics (aperf0, mperf0; aperf1, mperf1; ……; aperfN, mperfN), indicating that each CPU has independent performance metric sampling points for targeted monitoring and calculation of their respective frequencies.

[0079] There are multiple identical delayed - work processing flows on different physical CPUs (cpu0, cpu1, ……, cpuN) for the delayed - work processing flow. Each process has the same steps, representing periodic tasks for CPU frequency monitoring.

[0080] The process of the periodic task is as follows.

[0081] 1. Wake up: Start the delayed - work task and enter the execution state from the waiting state.

[0082] 2. Obtain aperf mperf samples: aperf (Architectural Performance) and mperf (Micro - architectural Performance) are metrics related to CPU performance. This step obtains the sampling data of these metrics for subsequent frequency calculation.

[0083] 3. Sleep: According to the sampling needs, let the task enter the sleep state and wait for a period of time before continuing to sample.

[0084] 4. Obtain aperf and mperf samples again: Collect aperf and mperf data again. The CPU operating frequency between two samplings can be obtained through at least two samplings.

[0085] 5. Calculate the CPU frequency: Use the obtained aperf and mperf sampling data to calculate the current operating frequency of the CPU according to a specific algorithm.

[0086] 6. Schedule and wait to be woken up again: After completing one frequency calculation, set the task to the waiting state. The scheduler will wake up the task again at an appropriate time to start the next round of frequency monitoring process.

[0087] Create a delayed_work for each CPU. These delayed_works are woken up periodically to perform the following operations: First, obtain the aperf and mperf values of the current CPU, then sleep for 1 second, obtain the aperf and mperf values of the current CPU again, and calculate the current operating frequency of this CPU based on the aperf and mperf values of the two samplings. If the frequency is lower than the preset threshold, an alarm is triggered; if it meets the expectation, it is scheduled to the next cycle for continuous monitoring.

[0088] Figure 4 The lightweight module shown brings beneficial effects in many aspects. First, accurate monitoring. By using cpumask through libvirt to identify the CPUs running virtual machines, the monitoring target is more targeted, avoiding ineffective monitoring of non-virtual machine related CPU resources, greatly improving the accuracy of monitoring, and helping cloud computing service providers accurately grasp the CPU performance status of virtual machines. Second, high efficiency and stability. The cpufreq - monitor module runs in the kernel space. This method can effectively reduce the impact on the stability of virtual machines compared with traditional user - space monitoring, reduce performance jitter, ensure the stable operation of virtual machines in a multi - tenant environment, and at the same time improve the monitoring efficiency, enabling real - time capture of CPU frequency changes. Through this lightweight kernel module design, the present invention can minimize the interference to the performance of virtual machines while ensuring the monitoring efficiency, and is suitable for the CPU frequency monitoring requirements of the host machine in the cloud computing scenario.

[0089] In some embodiments, the CPU frequency monitoring method in the cloud computing scenario further includes: obtaining the lightweight module for CPU frequency monitoring from the host kernel, including finding and loading the lightweight module from the specified module storage path of the host kernel through the kernel module loading mechanism.

[0090] Obtaining the lightweight module in this specific way can ensure the correct loading and use of the module, guaranteeing the stability and reliability of the system. This additional feature closely cooperates with the previous steps of obtaining the lightweight module, further clarifying the specific way to obtain the module, which helps to improve the efficiency and accuracy of obtaining the module. Obtaining the module in this way can make more efficient use of system resources and reduce unnecessary errors and conflicts.

[0091] In some embodiments, the method further includes: obtaining the calibrated frequency of the CPU based on the lightweight module, including the lightweight module obtaining the calibrated frequency of the CPU by reading the hardware register of the virtual machine CPU.

[0092] This way of obtaining the calibrated frequency is direct and accurate, and the information stored in the hardware register can truly reflect the calibrated frequency of the CPU. This additional feature is closely related to the steps of obtaining the calibrated frequency, clarifying the specific way of obtaining, and improving the accuracy of obtaining the calibrated frequency. By reading the hardware register to obtain the calibrated frequency, the most accurate calibrated frequency value can be obtained, providing a reliable data basis for subsequent frequency comparison.

[0093] In some embodiments, the method further includes: before comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module, the method further includes the lightweight module collecting the actual frequency of the virtual machine CPU at a preset time interval.

[0094] This can timely obtain the change situation of the actual frequency of the CPU, providing the latest data for accurate frequency comparison. This additional feature cooperates with the frequency comparison step, providing real-time and accurate data support for frequency comparison, which helps to more timely detect the CPU downclocking situation. By collecting the actual frequency at a preset time interval, the running state of the CPU can be dynamically tracked, and abnormal downclocking can be detected in time.

[0095] In some embodiments, the method further includes: comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module, including calculating the difference between the calibrated frequency and the actual frequency based on the lightweight module, and comparing whether the difference exceeds a preset threshold.

[0096] Through this quantitative comparison method, it can be more clearly judged whether the CPU is downclocked. This additional feature clarifies the specific calculation and judgment method of frequency comparison, making the frequency comparison more scientific and accurate, which helps to improve the reliability of downclocking judgment.

[0097] In some embodiments, the method further includes: obtaining the downclock warning information of the CPU, including generating downclock warning information if the difference is greater than a preset threshold; the downclock warning information includes the CPU number of the downclock, the downclock amplitude, and the time when the downclock occurs.

[0098] These detailed downclock warning information enables users to comprehensively understand the downclock situation, so as to make more appropriate decisions. This additional feature is closely related to the step of obtaining the downclock warning information, refining the content of the downclock warning information, improving the practicality of the warning information, and helping users better cope with the CPU downclock problem.

[0099] In some embodiments, the method further includes: before obtaining the calibrated frequency of the CPU based on the lightweight module, the method further includes obtaining the mask vector of the CPU; the lightweight module obtains the CPUs to be monitored for frequency based on the mask vector.

[0100] In this way, operations can be performed on the CPUs to be monitored in a targeted manner, improving the monitoring efficiency. This additional feature cooperates with the step of obtaining the calibrated frequency to determine the range of CPUs to be monitored, which helps to concentrate resources for effective monitoring, improving the pertinence and efficiency of monitoring.

[0101] In some embodiments, obtaining the mask vector of the CPU specifically includes the following steps: First, obtain the number of virtual machines on each CPU. This step is implemented through specific system interfaces or monitoring tools, traversing each CPU in the system, querying and recording the number of virtual machines running on it. Then, set the mask of the CPUs with the number of virtual machines greater than zero to 1. This operation is to clearly distinguish which CPUs have running virtual machines in the future. The CPUs with a mask of 1 are the CPUs with running virtual machines. In this way, it is convenient to perform targeted monitoring and management operations on the CPUs with virtual machine loads in the future.

[0102] These two steps are essential technical features for obtaining the CPU mask vector. Cooperating with each other can accurately identify the CPUs with running virtual machines, providing basic data support for subsequent CPU-related operations, enabling the system to allocate resources and monitor performance more precisely for the CPUs running virtual machines, and improving the management efficiency and accuracy of the system.

[0103] In some embodiments, the method further includes the periodic operation of a lightweight module. The lightweight module is periodically awakened to perform the following operations: First, obtain the actual performance counter value and the maximum performance counter value of the current CPU. This is achieved by interacting with the CPU hardware and using the performance counter reading interface provided by the system to obtain the corresponding values, which reflect the current actual operating condition of the CPU and its theoretical maximum operating capacity. Then, sleep for 1 second. This short sleep time is to allow the CPU to run within a relatively stable time period to obtain more accurate performance data. Obtain the actual performance counter value and the maximum performance counter value of the current CPU again. Finally, calculate the actual frequency of the CPU based on the actual performance counter values and the maximum performance counter values of the two samplings. Through such an operation process, the actual operating frequency of the CPU can be obtained dynamically and accurately, so as to timely grasp the operating state of the CPU.

[0104] This operation cooperates with the previous operation of obtaining the CPU mask vector to jointly improve the monitoring system of the CPU operating state. Obtaining the CPU mask vector determines the CPUs that need to be monitored key points, while the lightweight module periodically obtaining the actual CPU frequency further deeply understands the operating performance of these CPUs, which helps to timely detect CPU performance anomalies and improve the stability and reliability of the system.

[0105] As Figure 5 shown, an embodiment of the present application further provides a CPU frequency monitoring device in a cloud computing scenario, which is applied to a virtual machine and is characterized in that the device includes:

[0106] An initialization module 210, which obtains a lightweight module for CPU frequency monitoring from the host kernel;

[0107] A calibrated frequency acquisition module 220, which obtains the calibrated frequency of the CPU based on the lightweight module;

[0108] A frequency comparison module 230, which compares the actual frequency of the CPU with the calibrated frequency based on the lightweight module to obtain the CPU down-frequency warning information;

[0109] A down-frequency warning module 240, which sends the down-frequency warning information to a smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user.

[0110] The device has the following beneficial effects.

[0111] Initialization module 210 obtains a lightweight module for CPU frequency monitoring from the host kernel. The initialization module 210 delves into the host kernel through a specific system call interface to obtain a pre-configured lightweight module for CPU frequency monitoring, providing a basic tool for subsequent monitoring operations. This module is the starting point for the normal operation of the entire device and is an essential technical feature.

[0112] Calibration frequency acquisition module 220 obtains the calibration frequency of the CPU based on the lightweight module. The calibration frequency acquisition module 220 communicates with the CPU hardware by means of the lightweight module to obtain the calibration frequency set at the factory of the CPU or pre-set by the system. This frequency is a reference value for judging whether the actual operation of the CPU is normal. It cooperates with the initialization module. After the initialization module obtains the lightweight module, the calibration frequency acquisition module can obtain the calibration frequency based on this, providing basic data for subsequent frequency comparison, enabling the device to effectively evaluate the operating condition of the CPU.

[0113] Frequency comparison module 230 compares the actual frequency of the CPU with the calibration frequency based on the lightweight module to obtain the CPU down-frequency warning information. After obtaining the actual frequency and calibration frequency of the CPU, the frequency comparison module 230 conducts a comparative analysis through specific algorithms and logics to determine whether the actual frequency is lower than the calibration frequency. If it is lower, it generates down-frequency warning information. This module is the key link for the entire device to implement the warning function. It works in coordination with the previous modules, utilizes the lightweight module obtained by the initialization module and the calibration frequency obtained by the calibration frequency acquisition module, and combines its own comparative analysis of the actual frequency to promptly detect the down-frequency situation of the CPU.

[0114] Down-frequency warning module 240 sends the down-frequency warning information to the smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user. After obtaining the down-frequency warning information, the down-frequency warning module 240 sends the information accurately to the smart card through the communication interface with the smart card by means of the lightweight module, and the smart card then passes the warning information to the user, ensuring that the user can promptly learn about the down-frequency situation of the CPU. This module completes the interaction link between the entire monitoring device and the user, and together with the previous modules, constitutes a complete monitoring and warning system, enabling the user to promptly understand the operating condition of the CPU where the virtual machine is located, so as to take corresponding measures.

[0115] As Figure 6 shown, the embodiment of the present application further provides an electronic device 600. Figure 4 Only the exemplary structure of the electronic device 600 is shown, rather than all structures. According to needs, Figure 4 the shown partial structure or all structures can be implemented. As Figure 4As shown in the figure, the electronic device 600 provided by the embodiment of the present application includes: at least one processor 601, a memory 602, a user interface 603, and at least one network interface 604. Each component in the electronic device 600 is coupled together through a bus system 605. It can be understood that the bus system 605 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 4 all kinds of buses are labeled as the bus system 605.

[0116] Among them, the user interface 603 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a touchpad, or a touch screen, etc.

[0117] The memory 602 in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.

[0118] The intelligent diagnosis guidance method of the electronic device disclosed in the embodiment of the present application can be applied to the processor 601 or implemented by the processor 601. The processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the intelligent diagnosis guidance method of the electronic device can be completed by the integrated logic circuit in the hardware of the processor 601 or the instructions in the form of software. The above-mentioned processor 601 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 601 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiment of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 602. The processor 601 reads the information in the memory 602 and combines its hardware to complete the steps of the intelligent diagnosis guidance method of the electronic device provided by the embodiment of the present application.

[0119] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components for performing the foregoing method.

[0120] It can be understood that the memory 602 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or... The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory). The memory described in the embodiments of the present application is intended to include but not be limited to these and any other suitable types of memory.

[0121] In an exemplary embodiment, the embodiment of the present application further provides a computer storage medium, specifically a computer-readable storage medium, on which a computer program is stored. The computer program can be executed by a processor to complete the steps of the method of the embodiment of the present application. The computer-readable storage medium can be a memory such as ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0122] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program. The computer program can be executed by the processor 601 of an electronic device to complete the steps of the method of the embodiment of the present application.

[0123] It should be noted that "first", "second", etc. are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0124] In addition, the technical solutions described in the embodiments of the present application can be arbitrarily combined without conflict.

[0125] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0126] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A CPU frequency monitoring method in a cloud computing scenario, applied to a virtual machine, characterized in that The method includes: Obtaining a lightweight module for CPU frequency monitoring from the host kernel; Obtaining the calibrated frequency of the CPU based on the lightweight module; Comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module to obtain the CPU down-frequency warning information; Sending the down-frequency warning information to the smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user.

2. The method according to claim 1, wherein The obtaining a lightweight module for CPU frequency monitoring from the host kernel includes: Searching for and loading the lightweight module from the module storage path specified by the host kernel through the kernel module loading mechanism.

3. The method according to claim 1, characterized in that The obtaining the calibrated frequency of the CPU based on the lightweight module includes: The lightweight module obtains the calibrated frequency of the CPU by reading the hardware registers of the virtual machine CPU.

4. The method according to claim 1, wherein Before comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module, the method further includes: The lightweight module collects the actual frequency of the virtual machine CPU at a preset time interval.

5. The method according to claim 1, wherein The comparing the actual frequency and the calibrated frequency of the CPU based on the lightweight module includes: Calculating the difference between the calibrated frequency and the actual frequency based on the lightweight module, and comparing whether the difference exceeds a preset threshold.

6. The method according to claim 5, wherein The obtaining the CPU down-frequency warning information includes: If the difference is greater than the preset threshold, generating down-frequency warning information; the down-frequency warning information includes the CPU number with down-frequency, the down-frequency amplitude, and the time when the down-frequency occurs.

7. The method according to claim 1, wherein Before obtaining the calibrated frequency of the CPU based on the lightweight module, the method further includes: Obtaining a mask vector of the CPU; The lightweight module obtains the CPUs to be monitored for frequency based on the mask vector.

8. The method according to claim 7, wherein The obtaining a mask vector of the CPU includes: Obtaining the number of virtual machines on each CPU; Setting the mask of the CPU with the number of virtual machines greater than zero to 1.

9. The method according to claim 1, wherein The lightweight module is periodically awakened to perform the following operations: First, obtaining the current actual performance counter value and the maximum performance counter value of the CPU, then sleeping for 1 second, and then obtaining the current actual performance counter value and the maximum performance counter value of the CPU again, and calculating the actual frequency of the CPU according to the actual performance counter values and the maximum performance counter values of the two samplings.

10. A CPU frequency monitoring device in a cloud computing scenario, applied to a virtual machine, characterized in that The device includes: An initialization module that obtains a lightweight module for CPU frequency monitoring from the host kernel; A calibrated frequency obtaining module that obtains the calibrated frequency of the CPU based on the lightweight module; A frequency comparison module that compares the actual frequency and the calibrated frequency of the CPU based on the lightweight module to obtain the CPU down-frequency warning information; A down-frequency warning module that sends the down-frequency warning information to the smart card based on the lightweight module, and the smart card sends the down-frequency warning information to the user.

11. An electronic device, characterized in that, Includes: A processor and a memory for storing a computer program that can run on the processor, wherein the processor, when running the computer program, executes the steps of the method according to any one of claims 1 to 9.

12. A computer storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by an electronic device, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Hardware counter virtualization-based performance analysis method for multiple virtual machines

    CN102073535A

  • Method and interrupt controller for treating i / o operation interrupt requests in a virtual machine system

    EP0419723A1