A data processing method and device

By predicting the number of times the virtual machine operates across cache lines on the host CPU, and turning off the exception throwing function and activate the detection thread when the preset threshold is above the preset threshold, the performance degradation caused by the virtual machine's frequent cross-cache lines is solved, and the detection and performance improvement of the actual number of executions is achieved.

CN114595037BActive Publication Date: 2025-05-09ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210267806.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-05-09
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

In a virtualized environment, the virtual machine frequently operates across cache lines on the host's CPU, causing the host's CPU memory access bus to be locked and an exception is thrown, thereby reducing the overall performance of the host and virtual machine.

Method used

By predicting the number of times the VCPU allocated by the virtual machine is expected to perform cross-cache line operations on the host's CPU within the first time period after the current time, if the estimated number of times is greater than or equal to the preset threshold, the function of the CPU throwing an exception due to the locking of the memory bus, and the detection thread state is switched from the silent state to the activated state, and the operation data recorded by the performance monitor of the CPU is polled to obtain the actual number of executions.

Benefits of technology

It effectively avoids the host CPU throwing exceptions when the memory bus is locked, reduces the performance loss of the host and virtual machine, and realizes the detection of the number of times the virtual machine actually performs cross-cache line operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595037B_ABST
    Figure CN114595037B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method and device. In the present application, if the first estimated number of executions of the cross-cache line operation on the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform in the first time period thereafter is greater than a preset threshold, the function of the CPU of the host machine throwing an exception due to the locking of the CPU's memory access bus is first turned off, and then the state of the detection thread is switched from a silent state to an activated state; so that the detection thread polls the CPU operation data recorded in the PMU corresponding to the CPU in the host machine, and obtains the actual number of executions according to the polled CPU operation data. Through the present application, in the case where the VCPU allocated to the virtual machine frequently executes the cross-cache line operation on the CPU of the host machine, in the scenario of detecting the actual number of executions of the cross-cache line operation on the CPU of the host machine that the VCPU allocated to the virtual machine actually executes, the overall performance of the host machine and the performance of the virtual machine can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data processing method and device. Background Art

[0002] With the rapid development of technology, the application scope of virtualization technology is becoming wider and wider. Virtualization is the key technology of cloud computing. Virtualization technology can virtualize a physical machine (host machine) into one or more virtual machines. Each virtual machine has its own virtual hardware, such as VCPU (Virtual Central Processing Unit), virtual memory, and virtual I / O devices, thus forming an independent virtual machine execution environment. Virtualization technology is widely used in cloud computing and high-performance computing due to its high fault tolerance and high resource utilization.

[0003] In a virtualized environment, VMM (Virtual Machine Management) is a software management layer located between the host machine's hardware and the virtual machine. It is mainly responsible for managing the host machine's hardware, such as the host machine's CPU (Central Processing Unit), memory, and I / O devices, and abstracting the host machine's hardware into corresponding virtual device interfaces for use by virtual machines. Summary of the invention

[0004] The present application shows a data processing method and device.

[0005] In the first aspect, the present application shows a data processing method, which is applied to a host machine, in which at least a virtual machine and a detection thread are running; the method includes: predicting a first expected number of executions of a cross-cache line operation on a central processing unit (CPU) of the host machine that a VCPU allocated to the virtual machine is expected to perform within a first time period after a current moment; when the first expected number of executions is greater than or equal to a preset threshold, shutting down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked; so that the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked; and switching the state of the detection thread from a silent state to an active state; so that the detection thread polls the CPU's operating data recorded in a performance monitor (PMU) corresponding to the CPU in the host machine, and obtaining the actual number of executions of the cross-cache line operation on the host machine's CPU that the VCPU allocated to the virtual machine actually performs within the first time period based on the polled CPU's operating data.

[0006] In the second aspect, the present application shows a data processing device, which is applied to a host machine, in which at least a virtual machine and a detection thread are running; the device includes: a first prediction module, which is used to predict the first expected number of executions of the cross-cache line cache line operation on the central processing unit CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform within a first time period after the current moment; a shutdown module, which is used to shut down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked when the first expected execution number is greater than or equal to a preset threshold; so that the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked; and a first switching module, which is used to switch the state of the detection thread from a silent state to an active state; a polling module, which is used to poll the CPU's operating data recorded in the performance monitor PMU corresponding to the CPU in the host machine, and an acquisition module, which is used to obtain the actual number of executions of the cross-cache line operation on the host machine's CPU that the VCPU allocated to the virtual machine actually performs within the first time period according to the polled CPU's operating data.

[0007] In a third aspect, the present application shows an electronic device, the electronic device comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the method shown in any of the aforementioned aspects.

[0008] In a fourth aspect, the present application shows a non-temporary computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform a method as shown in any of the aforementioned aspects.

[0009] In a fifth aspect, the present application illustrates a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform a method as shown in any of the aforementioned aspects.

[0010] Compared with the prior art, this application has the following advantages:

[0011] In the present application, the first estimated number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform in the first time period after the current moment. When the first estimated number of executions is greater than or equal to the preset threshold, the function of the CPU of the host machine throwing an exception due to the memory access bus of the CPU being locked is turned off. So that the CPU of the host machine does not throw an exception when the memory access bus of the CPU of the host machine is locked. And, the state of the detection thread is switched from a silent state to an active state. So that the detection thread polls the operation data of the CPU recorded in the PMU corresponding to the CPU in the host machine, and obtains the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs in the first time period according to the operation data of the polled CPU. Through the present application, in the case where the VCPU allocated to the virtual machine performs the cross-cache line operation on the CPU of the host machine very frequently (for example, tens of thousands or hundreds of thousands of times per second, etc.), in the scenario of detecting the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs, the overall performance of the host machine and the performance of the virtual machine can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 It is a step flow chart of a data processing method of the present application.

[0013] Figure 2 It is a step flow chart of a data processing method of the present application.

[0014] Figure 3 It is a step flow chart of a data processing method of the present application.

[0015] Figure 4 It is a structural block diagram of a data processing device of the present application.

[0016] Figure 5 It is a structural block diagram of a device of the present application. DETAILED DESCRIPTION

[0017] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0018] Sometimes, the CPU in the host machine allows unaligned memory access. In the scenario of unaligned memory access, the operand of the atomic operation (due to address misalignment) will span two cache lines of the host machine's CPU. That is, an operation to access the CPU's cache spans two cache lines, which will trigger a split lock event.

[0019] For example, in one example, in the CPU cache, a cache line consists of 64 bytes, a member in structcounter occupies 8 bytes, and buf fills 62 bytes. Therefore, once this member is accessed, it involves the concatenation of the contents of two cache lines, so executing an atomic operation will trigger a split lock event.

[0020] However, in general, cache uniformity protocols can only guarantee consistency at the cache line level. Accessing two cache lines simultaneously cannot guarantee consistency at the cache line level. To ensure the atomicity of split lock, special logic (such as cold path) can be used to handle the situation where the accessed operands span two cache lines: for example, locking the host CPU's memory access bus (BUSLOCK).

[0021] When the memory bus of the host machine's CPU is locked, access to the memory bus by other threads in the host machine / other cores in the host machine's CPU will be blocked, resulting in interruption of data processing processes of other threads in the host machine / other cores in the host machine's CPU.

[0022] Due to the interruption of data processing processes of other threads in the host machine / other cores in the CPU of the host machine, the average memory access latency of the CPU of the host machine will increase significantly, and since the action of "intercepting access to the memory bus by other threads in the host machine / other cores in the CPU of the host machine" will also consume some computing resources of the CPU of the host machine, the overall performance of the host machine will be reduced.

[0023] Thus, a demand is put forward to improve the overall performance of the host machine.

[0024] In order to improve the overall performance of the host machine, in one possible way, the average memory access latency of the host machine's CPU can be reduced as much as possible, and the number of times the action of "intercepting access to the memory bus by other threads in the host machine / other cores in the CPU in the host machine" can be reduced as much as possible.

[0025] In order to achieve the purpose of “reducing the average memory access latency of the CPU of the host machine”, in one approach, the number of interruptions of data processing processes of other threads in the host machine / other cores in the CPU of the host machine can be reduced.

[0026] In order to achieve the purpose of "reducing the number of interruptions to data processing processes of other threads in the host machine / other cores in the CPU of the host machine", in one method, the number of times access to the memory bus by other threads in the host machine / other cores in the CPU of the host machine is intercepted can be reduced.

[0027] Among them, reducing the number of times that access to the memory bus by other threads in the host machine / other cores in the CPU of the host machine is intercepted can save computing resources required for the interception action.

[0028] In order to achieve the purpose of "reducing the number of times that access to the memory bus by other threads in the host machine / other cores in the CPU of the host machine is intercepted", in one method, the number of times the memory bus of the CPU of the host machine is locked can be reduced.

[0029] In order to achieve the purpose of “reducing the number of times the memory access bus of the host machine's CPU is locked”, in one method, the actual number of times the VCPU allocated to the virtual machine actually executes cross-cache line operations on the host machine's CPU can be detected.

[0030] In the case that the VCPU allocated to the virtual machine actually executes a high number of cross-cache line operations on the host machine's CPU, some countermeasures can be taken to reduce the actual number of cross-cache line operations actually executed by the VCPU allocated to the virtual machine on the host machine's CPU, thereby reducing the number of times the host machine's CPU's memory bus is locked.

[0031] This can reduce the number of times that access to the memory bus by other threads in the host machine / other cores in the CPU of the host machine is intercepted, which can reduce the number of times that the data processing process of other threads in the host machine / other cores in the CPU of the host machine is interrupted, and can reduce the average memory access latency of the CPU of the host machine, thereby achieving the goal of improving the overall performance of the host machine.

[0032] The countermeasures may include: reducing the CPU utilization of the host machine by the virtual machine on the host machine or the access frequency of the virtual machine to the CPU of the host machine, etc.

[0033] Through the above analysis, it is known that it is more important to detect the actual number of times the VCPU allocated to the virtual machine actually executes the cross-cacheline operation on the CPU of the host machine.

[0034] Thus, there is a need to "detect the actual number of times the VCPU allocated to the virtual machine actually executes the cross-cache line operation on the CPU of the host machine."

[0035] In order to achieve the purpose of "detecting the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine", the inventors found that: when the VCPU allocated to the virtual machine actually executes cross-cache line operations on the host machine's CPU, the memory access bus of the host machine's CPU will be locked. When the memory access bus of the host machine's CPU is locked, the host machine's CPU will throw an exception, such as Bus Lock Exception or #DBException.

[0036] The kernel-state VMM on the host machine is subsequently required to handle the exception, so the virtual machine will be exited to the kernel-state VMM (for example, the virtual machine is suspended and the kernel-state VMM is resumed). After exiting to the kernel-state VMM (that is, after resuming the kernel-state VMM), the kernel-state VMM will attempt to handle the exception.

[0037] In one possible case, the kernel-mode VMM may determine that the kernel-mode VMM cannot handle the exception or that the exception should be handled by the user-mode VMM. In this case, the kernel-mode VMM may notify the user-mode VMM to handle the exception. After receiving the notification, the user-mode VMM will obtain the exception or try to handle the exception.

[0038] Among them, after the user-state VMM obtains the exception, it can also obtain relevant information of the exception (for example, the cause of the exception can be recorded in the shared area of ​​the user-state VMM and the kernel-state VMM, from which it can be analyzed whether the cause of the exception is due to the triggering of a split lock event, etc.), and based on the relevant information of the exception, it can be determined whether the VCPU allocated to the virtual machine has executed a cache line operation across the CPU of the host machine (whether the VCPU allocated to the virtual machine has triggered a split lock event), and the actual number of executions of the cross-cache line operation of the CPU of the host machine by the VCPU allocated to the virtual machine can be counted by counting.

[0039] However, the inventors have discovered that in the above method, each time the VCPU allocated to the virtual machine executes a cross-cache line operation on the host machine's CPU, the host machine's CPU will throw an exception (Bus Lock Exception or #DBException, etc.), and then exit from the virtual machine to the kernel-state VMM (for example, through the Vmexit function, etc.) (for example, suspend the virtual machine and resume the operation of the kernel-state VMM).

[0040] When the VCPU assigned to the virtual machine frequently performs cross-cache line operations on the host machine's CPU (e.g., tens of thousands or hundreds of thousands of times per second), on the one hand, the VMM in the host machine will frequently enter the exception handling process, reducing the overall performance of the host machine, and further indirectly causing the host machine to enter a state similar to being attacked by DOS (Denial of Service Attack). On the other hand, the performance of the virtual machine will be reduced due to the VMM exiting from the virtual machine to the kernel state multiple times.

[0041] In view of this, the inventors believe that, when the VCPU assigned to the virtual machine executes cross-cache line operations on the host machine's CPU very frequently (for example, tens of thousands or hundreds of thousands of times per second), it is inappropriate to detect the actual number of times the VCPU assigned to the virtual machine actually executes cross-cache line operations on the host machine's CPU using the above method.

[0042] In this way, the requirement is proposed: "When the VCPU assigned to the virtual machine executes cross-cache line operations on the host machine's CPU very frequently (for example, tens of thousands or hundreds of thousands of times per second, etc.), in the scenario where it is necessary to detect the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU assigned to the virtual machine, the overall performance of the host machine and the performance of the virtual machine are improved."

[0043] In order to achieve the purpose of "improving the overall performance of the host machine and the performance of the virtual machine in a scenario where it is necessary to detect the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the CPU of the host machine when the VCPU allocated to the virtual machine performs cross-cache line operations on the CPU of the host machine frequently (for example, tens of thousands or hundreds of thousands of times per second), the inventors conducted a statistical analysis on the above method and found that:

[0044] The function of causing the host CPU to throw an exception when the host CPU's memory access bus is locked has a function switch, which can be turned on or off.

[0045] According to the actual situation, if it is required that "when the memory access bus of the host machine's CPU is locked, the host machine's CPU throws an exception", the function switch can be turned on, that is, the function of causing the host machine's CPU to throw an exception when the memory access bus of the host machine's CPU is locked is started by turning on the function switch.

[0046] If it is required that "the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked", the function switch can be turned off, that is, the function of the host machine's CPU throwing an exception when the host machine's CPU's memory access bus is locked can be turned off by turning off the function switch.

[0047] In this way, the inventors discovered that the function switch can be turned off. In this way, when the VCPU allocated to the virtual machine performs cross-cache line operations on the host machine's CPU very frequently (for example, tens of thousands or hundreds of thousands of times per second), the host machine's CPU will not throw an exception every time the memory access bus of the host machine's CPU is locked, thereby avoiding reducing the overall performance of the host machine and avoiding reducing the performance of the virtual machine.

[0048] However, although the purpose of "avoiding reducing the overall performance of the host machine and avoiding reducing the performance of the virtual machine" is achieved, the host machine's CPU does not throw an exception when the host machine's CPU memory bus is locked, which will lead to the inability to detect the actual number of executions of the cross-cache line operations of the VCPU assigned to the virtual machine on the host machine's CPU. As a result, it is impossible to achieve the purpose of "improving the overall performance of the host machine and improving the performance of the virtual machine in the scenario where it is necessary to detect the actual number of executions of the cross-cache line operations of the VCPU assigned to the virtual machine on the host machine's CPU when the VCPU assigned to the virtual machine executes cross-cache line operations on the host machine's CPU frequently (for example, tens of thousands or hundreds of thousands of times per second).

[0049] In view of this, the inventors further abandoned the detection method of "after the user-mode VMM obtains the exception, it can also obtain the relevant information of the exception (for example, the cause of the exception can be recorded in the shared area of ​​the user-mode VMM and the kernel-mode VMM, from which the cause of the exception can be analyzed to be due to the triggering of a split lock event, etc.), and based on the relevant information of the exception, it can be determined whether the VCPU assigned to the virtual machine has executed a cache line operation across the host machine's CPU (whether the VCPU assigned to the virtual machine has triggered a split lock event), and the actual number of executions of the cross-cache line operation on the host machine's CPU by the VCPU assigned to the virtual machine is counted by counting" to detect the actual number of executions of the cross-cache line operation on the host machine's CPU by the VCPU assigned to the virtual machine, and instead thought of creating a detection thread on the host machine, which is used to detect the actual number of executions of the cross-cache line operation on the host machine's CPU by the VCPU assigned to the virtual machine.

[0050] Specifically, see Figure 1 , shows a data processing method of the present application, which is applied to a host machine, and is applied to a host machine, in which at least a virtual machine and a detection thread are running. The method includes:

[0051] In step S101, a first estimated number of executions of a cross-cache line operation on a CPU of a host machine that a VCPU allocated to a virtual machine is expected to perform within a first time period after a current moment is predicted.

[0052] In the present application, performing a cross-cache line operation on the host CPU will trigger a split lock event. When the split lock event is triggered, the memory access bus of the host CPU will be locked, which will reduce the overall performance of the host.

[0053] In the present application, in order to improve the overall performance of the host machine, the actual number of cross-cache line operations actually performed by the VCPU assigned to the virtual machine (which may be assigned to the virtual machine by the kernel-state VMM) on the host machine's CPU can be detected. And when the actual number of cross-cache line operations actually performed by the VCPU assigned to the virtual machine on the host machine's CPU is high, some countermeasures (such as reducing the CPU utilization of the virtual machine on the host machine, or the access frequency of the virtual machine to the host machine's CPU, etc.) are used to reduce the number of cross-cache line operations performed by the VCPU assigned to the virtual machine on the host machine's CPU, thereby reducing the number of split lock events triggered later, and thereby reducing the number of times the host machine's CPU access bus is locked later, thereby improving the overall performance of the host machine.

[0054] In one embodiment of the present application, in order to detect the actual number of times a VCPU allocated to a virtual machine actually executes cross-cache line operations on a CPU of a host machine, time can be divided into a plurality of adjacent time periods in sequence, the duration of each time period can be the same, and adjacent time periods can be connected end to end. For example, in adjacent time periods, the end time of an earlier time period can be the same as the start time of a later time period, etc.

[0055] The actual number of times the VCPU allocated to the virtual machine actually executes cross-cache line operations on the host machine's CPU can be detected using the time period as a reference time unit. For example, the actual number of times the VCPU allocated to the virtual machine actually executes cross-cache line operations on the host machine's CPU in each time period can be detected.

[0056] Of course, in a possible case, the actual number of executions of the cross-cache line operations performed by the VCPU assigned to the virtual machine on the host machine's CPU can also be detected based on a benchmark that is less than the time period, so as to improve the real-time performance of detecting the actual number of executions of the cross-cache line operations performed by the VCPU assigned to the virtual machine on the host machine's CPU.

[0057] In view of this, in the present application, the actual number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine can be detected in at least two ways.

[0058] In order to determine a suitable detection method, in one embodiment, if it is necessary to detect the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine within a time period, before the time period, the estimated number of executions of cross-cache line operations on the host machine's CPU by the virtual VCPU allocated to the virtual machine within the time period can be predicted.

[0059] Then, referring to the predicted number of cross-cache line operations that the virtual VCPU allocated to the virtual machine is expected to perform on the host machine's CPU within the time period, one of the methods is selected to detect the actual number of cross-cache line operations that the VCPU allocated to the virtual machine actually performs on the host machine's CPU within the time period.

[0060] For example, a preset threshold may be set according to actual conditions.

[0061] In the case where the predicted number of cross-cache line operations on the host machine's CPU that the virtual VCPU allocated to the virtual machine is expected to perform during the time period is greater than or equal to a preset threshold, one of the methods can be used to detect the actual number of cross-cache line operations that the VCPU allocated to the virtual machine actually performs during the time period on the host machine's CPU. Alternatively, in the case where the predicted number of cross-cache line operations that the virtual VCPU allocated to the virtual machine is expected to perform during the time period is less than a preset threshold, another method can be used to detect the actual number of cross-cache line operations that the VCPU allocated to the virtual machine actually performs during the time period on the host machine's CPU. For details, please refer to the description of the subsequent steps, which will not be described in detail here.

[0062] For example, in one embodiment, there is a first time period after the current moment. Before detecting the actual number of executions of the cross-cache line operation on the host machine's CPU performed by the VCPU allocated to the virtual machine within the first time period, a first estimated number of executions (the predicted number of executions, not the actual number of executions) of the cross-cache line operation on the host machine's CPU performed by the virtual VCPU allocated to the virtual machine within the first time period after the current moment can be predicted, and then step S102 is executed.

[0063] Herein, historical data may be used to predict a first estimated number of times that a virtual VCPU allocated to a virtual machine is expected to perform a cross-cache line operation on a CPU of a host machine within a first time period after a current moment.

[0064] The historical data may include the historical number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine in at least one historical time period before the current moment, and then obtain the first estimated number of times the virtual VCPU allocated to the virtual machine is expected to perform cross-cache line operations on the CPU of the host machine in a first time period after the current moment based on the historical data. For example, the historical data may be analyzed to analyze the regularity of the historical number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine in the historical process, and obtain the first estimated number of times the virtual VCPU allocated to the virtual machine is expected to perform cross-cache line operations on the CPU of the host machine in the first time period after the current moment based on the regularity.

[0065] In one embodiment of the present application, the historical number of times a VCPU allocated to a virtual machine actually performs cross-cache line operations on a host machine's CPU in at least one historical time period before a current moment can be obtained; and then a first estimated number of times a virtual VCPU allocated to the virtual machine is expected to perform cross-cache line operations on a host machine's CPU in a first time period after the current moment is obtained based on the historical number of executions.

[0066] In step S102, when the first expected number of executions is greater than or equal to a preset threshold, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked is turned off. So that the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked. And, the state of the detection thread is switched from a silent state to an active state. So that the detection thread polls the CPU's operating data recorded in the PMU (Performance Monitoring Unit) corresponding to the CPU in the host machine, and obtains the actual number of executions of the cross-cache line operation on the host machine's CPU actually performed by the VCPU allocated to the virtual machine within the first time period based on the polled CPU's operating data.

[0067] In one embodiment of the present application, shutting down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked and switching the state of the detection thread from a silent state to an active state can be executed in parallel.

[0068] Alternatively, in another embodiment of the present application, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked can be disabled first, and then the state of the detection thread can be switched from a silent state to an active state.

[0069] Alternatively, in another embodiment of the present application, the state of the detection thread can be switched from a silent state to an active state, and then the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked can be disabled.

[0070] Since the function of the host machine's CPU to throw an exception due to the CPU's memory access bus being locked is disabled, the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked. However, since no exception is thrown, as analyzed above, it is impossible to detect whether the VCPU allocated to the virtual machine actually performs cross-cache line operations on the host machine's CPU, and it is also impossible to detect the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the host machine's CPU during the first time period.

[0071] Therefore, in order to avoid reducing the overall performance of the host machine and the performance of the virtual machine and also detect the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the host machine's CPU within the first time period when the function of throwing an exception by the host machine's CPU due to the CPU's memory access bus being locked is turned off, in the present application, a detection thread can be created in the host machine in advance, and the detection thread can be used to detect the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the host machine's CPU within the first time period.

[0072] The detection thread has multiple states, for example, a silent state and an activated state.

[0073] The detection thread in the silent state is not working and may be low power or low resource consuming, for example, it may not occupy the CPU overhead (computing resources) of the host machine.

[0074] The detection thread in the activated state is operational.

[0075] In one embodiment of the present application, if the state of the detection thread is activated at this time, the state of the detection thread may not be switched, and the detection thread will automatically poll the CPU operation data recorded in the PMU corresponding to the CPU in the host machine, and obtain the actual number of executions of the cross-cache line operation of the host machine's CPU by the VCPU allocated to the virtual machine within the first time period based on the polled CPU operation data.

[0076] Afterwards, the detection thread can store "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" in the log, so that when "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" is needed later, the detection thread can retrieve "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" from the log.

[0077] Alternatively, in another embodiment of the present application, if the state of the detection thread is a silent state at this time, since the detection thread in the silent state is not working, it is impossible to detect the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the CPU of the host machine in the first time period based on the detection thread in the silent state, so the state of the detection thread can be switched from the silent state to the active state. In this way, the detection thread will automatically poll the CPU operation data recorded in the PMU corresponding to the CPU in the host machine, and obtain the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the CPU of the host machine in the first time period according to the polled CPU operation data.

[0078] Afterwards, the detection thread can store "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" in the log, so that when "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" is needed later, the detection thread can retrieve "the actual number of executions of cross-cache line operations on the host machine's CPU by the VCPU allocated to the virtual machine during the first time period" from the log.

[0079] In the present application, the first estimated number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform in the first time period after the current moment. When the first estimated number of executions is greater than or equal to the preset threshold, the function of the CPU of the host machine throwing an exception due to the memory access bus of the CPU being locked is turned off. So that the CPU of the host machine does not throw an exception when the memory access bus of the CPU of the host machine is locked. And, the state of the detection thread is switched from a silent state to an active state. So that the detection thread polls the operation data of the CPU recorded in the PMU corresponding to the CPU in the host machine, and obtains the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs in the first time period according to the operation data of the polled CPU. Through the present application, in the case where the VCPU allocated to the virtual machine performs the cross-cache line operation on the CPU of the host machine very frequently (for example, tens of thousands or hundreds of thousands of times per second, etc.), in the scenario of detecting the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs, the overall performance of the host machine and the performance of the virtual machine can be improved.

[0080] In one embodiment of the present application, a kernel-mode VMM and a user-mode VMM are also running in the host machine.

[0081] In this way, when predicting the first estimated number of executions of cross-cache line operations on the host machine's CPU that the VCPU allocated to the virtual machine is expected to perform within the first time period after the current moment, the user-state VMM can predict the first estimated number of executions of cross-cache line operations on the host machine's CPU that the VCPU allocated to the virtual machine is expected to perform within the first time period after the current moment.

[0082] Accordingly, when shutting down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked, the user-state VMM may send a shutdown request to the kernel-state VMM, and the shutdown request is used to shut down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked.

[0083] When the first expected number of executions is greater than or equal to the preset threshold, as analyzed above, the VMM in the host machine frequently enters the exception handling process, which will reduce the overall performance of the host machine, and will cause the virtual machine to exit the VMM to kernel state multiple times, thereby reducing the performance of the virtual machine.

[0084] Therefore, in order to avoid reducing the overall performance of the host machine and the overall performance of the virtual machine, as analyzed above, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked can be disabled.

[0085] However, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked is under the control of the kernel-state VMM, and the user-state VMM may not control the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked. Therefore, if the user-state VMM determines that it is necessary to shut down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked when the first expected execution number is greater than or equal to the preset threshold, the kernel-state VMM may be requested to shut down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked.

[0086] For example, the VMM in user mode sends a shutdown request to the VMM in kernel mode, where the shutdown request is used to shut down a function of the CPU of the host machine that throws an exception because the memory access bus of the CPU is locked.

[0087] In one embodiment of the present application, the user-mode VMM may place the shutdown request into a shared area (eg, a shared memory page, etc.) between the user-mode VMM and the kernel-mode VMM, and then notify the kernel-mode VMM.

[0088] Then, the VMM in the kernel state receives the shutdown request, and according to the shutdown request, shuts down the function of the CPU of the host machine that throws an exception due to the locking of the memory access bus of the CPU.

[0089] In one embodiment of the present application, after receiving the notification, the kernel-mode VMM may read the shutdown request from a shared area between the user-mode VMM and the kernel-mode VMM.

[0090] There is a function switch corresponding to the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked. If the function switch is turned off, the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked can be turned off. If the function switch is turned on, the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked can be started.

[0091] Thus, when the kernel-mode VMM turns off the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked according to the shutdown request, it can turn off the function switch corresponding to the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked.

[0092] The VMM in kernel mode can turn off the function switch through an externally exposed API (Application Programming Interface) of the function switch.

[0093] Among them, the detection thread can be a kernel-state detection thread. Since the detection thread can be based on the kernel state, the user-state VMM can switch the state of the detection thread from a silent state to an activated state via the kernel-state VMM. For example, the detection thread exposes an API to the outside, and the user-state VMM can send an activation request carrying the API to the kernel-state VMM, so that the kernel-state VMM switches the state of the detection thread from a silent state to an activated state through the API.

[0094] For example, when switching the state of the detection thread from the silent state to the active state, the user state VMM may send an activation request to the kernel state VMM, and the activation request is used to request switching the state of the detection thread from the silent state to the active state. Then the kernel state VMM receives the activation request and switches the state of the detection thread from the silent state to the active state according to the activation request.

[0095] In another embodiment of the present application, after the kernel-state VMM turns off the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked, the kernel-state VMM can send a shutdown response to the user-state VMM, and the shutdown response is used to notify that the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked has been turned off. In one embodiment of the present application, the kernel-state VMM can put the shutdown response (processing result) into a shared area (such as a shared memory page, etc.) between the user-state VMM and the kernel-state VMM, and then notify the kernel-state VMM.

[0096] Then, the user-mode VMM receives the shutdown response, and sends an activation request to the kernel-mode VMM according to the shutdown response. In one embodiment of the present application, after receiving the notification, the user-mode VMM can read the shutdown response from the shared area of ​​the user-mode VMM and the kernel-mode VMM, and learn from the shutdown response that the function of the CPU of the host machine throwing an exception due to the CPU's memory access bus being locked has been shut down.

[0097] In one embodiment of the present application, the detection thread polls the CPU operation data recorded in the PMU corresponding to the CPU in the host machine, that is, it periodically obtains the CPU operation data recorded in the PMU corresponding to the CPU in the host machine.

[0098] Each time the CPU operation data recorded in the PMU corresponding to the CPU is obtained, it includes: the number of times the VCPU allocated to the virtual machine has executed the cross-cache line operation on the CPU of the host machine when the operation data is obtained.

[0099] In this way, the actual number of cross-cache line operations actually executed by the VCPU assigned to the virtual machine on the host machine's CPU during the first time period can be obtained based on the number of cross-cache line operations executed by the VCPU assigned to the virtual machine on the host machine's CPU obtained at the start time of the first time period and the number of cross-cache line operations executed by the VCPU assigned to the virtual machine on the host machine's CPU obtained at the end time of the first time period.

[0100] For example, in one example, the operation data of the polled CPU includes: a first number of executions of cross-cache line operations on the host machine's CPU by the VCPU assigned to the virtual machine at the start time of the first time period and a second number of executions of cross-cache line operations on the host machine's CPU by the VCPU assigned to the virtual machine at the end time of the first time period.

[0101] In this way, when obtaining the actual number of executions based on the polled CPU running data, the difference between the second number of executions and the first number of executions can be calculated, and then the actual number of executions of the cross-cache line operations on the host machine's CPU actually performed by the VCPU assigned to the virtual machine within the first time period can be obtained based on the difference. For example, the difference can be determined as the actual number of executions of the cross-cache line operations on the host machine's CPU actually performed by the VCPU assigned to the virtual machine within the first time period.

[0102] Through the present application, the actual number of times a VCPU allocated to a virtual machine actually performs cross-cache line operations on a CPU of a host machine within a first time period can be obtained by detecting a thread.

[0103] Then, it is necessary to enter a second time period after the first time period, that is, it is necessary to detect the actual number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine in the second time period.

[0104] See also Figure 2 , the specific process may include:

[0105] In step S201, a second estimated number of executions of a cross-cache line operation on a CPU of a host machine that a VCPU allocated to a virtual machine is expected to perform in a second time period after the first time period is predicted.

[0106] Before detecting the actual number of executions of the cross-cacheline operation on the host machine's CPU by the VCPU allocated to the virtual machine within the second time period, a second estimated number of executions (the predicted number of executions, not the actual number of executions) of the cross-cache line operation on the host machine's CPU by the virtual VCPU allocated to the virtual machine within the second time period can be predicted, and then step S202 is executed.

[0107] The historical data may be used to predict a second estimated number of times that the virtual VCPU allocated to the virtual machine is expected to perform a cross-cache line operation on the CPU of the host machine within a second time period.

[0108] The historical data may include the historical number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine in at least one historical time period before the current moment, and then obtain the second estimated number of times the virtual VCPU allocated to the virtual machine is expected to perform cross-cache line operations on the CPU of the host machine in a second time period after the current moment based on the historical data. For example, the historical data may be analyzed to analyze the regularity of the historical number of times the VCPU allocated to the virtual machine actually performs cross-cache line operations on the CPU of the host machine in the historical process, and obtain the second estimated number of times the virtual VCPU allocated to the virtual machine is expected to perform cross-cache line operations on the CPU of the host machine in the second time period after the current moment based on the regularity.

[0109] Among them, in the present application, the historical number of executions of the cross-cache line operations on the host machine's CPU actually performed by the VCPU allocated to the virtual machine in at least one historical time period before the current moment can be obtained; and then, based on the historical number of executions, the second estimated number of executions of the cross-cache line operations on the host machine's CPU expected to be performed by the virtual VCPU allocated to the virtual machine in a second time period after the current moment is obtained.

[0110] In step S202, when the second expected execution number is less than the preset threshold, a function of causing the host machine's CPU to throw an exception due to the CPU's memory access bus being locked is started, so that the host machine's CPU throws an exception when the host machine's CPU's memory access bus is locked.

[0111] In the present application, the specific method of starting the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked can be found in the subsequent description and will not be described in detail here.

[0112] In step S203, when the relevant information of the exception is obtained, it is determined according to the relevant information of the exception whether the exception is thrown by the CPU of the host machine because the memory access bus of the CPU is locked.

[0113] In step S204, when the exception is that the CPU of the host machine is thrown because the memory access bus of the CPU is locked, it is determined that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine within the second time period.

[0114] Furthermore, the actual number of times that the VCPU allocated to the virtual machine actually performs the cross-cachline operation on the CPU of the host machine within the second time period may be counted.

[0115] In one embodiment of the present application, a kernel-mode VMM and a user-mode VMM are also running in the host machine.

[0116] In this way, when predicting the second expected number of executions of cross-cache line operations on the host machine's CPU that the VCPU allocated to the virtual machine is expected to perform in the second time period after the first time period, the user-state VMM can predict the second expected number of executions of cross-cache line operations on the host machine's CPU that the VCPU allocated to the virtual machine is expected to perform in the second time period after the first time period.

[0117] Accordingly, when starting the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked, the user-state VMM may send a start request to the kernel-state VMM, where the start request is used to start the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked.

[0118] In the case where the second expected number of executions is less than the preset threshold, as analyzed above, although the VMM in the host machine will enter the exception handling process, it will not enter the exception handling process frequently, so the overall performance of the host machine is not greatly reduced, and although the virtual machine will exit to the VMM in kernel state, it will not exit to the VMM in kernel state frequently, so the performance of the virtual machine is not greatly reduced. In this way, the overall performance reduction of the host machine and the performance reduction of the virtual machine caused by "the VMM will not frequently enter the exception handling process and the virtual machine will not frequently exit to the VMM in kernel state" are often tolerable.

[0119] On the other hand, the detection thread polls the CPU operation data recorded in the PMU corresponding to the CPU in the host machine, that is, it periodically obtains the CPU operation data recorded in the PMU corresponding to the CPU in the host machine. Affected by the "periodic" cycle (the time interval between two adjacent pollings), it often takes "one cycle" to obtain the actual number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the CPU of the host machine within one cycle, which will affect the timeliness of obtaining the actual number of cross-cache line operations performed by the VCPU allocated to the virtual machine on the CPU of the host machine to a certain extent (for example, the division of the time period described above cannot be freely divided according to the actual situation, and the division of the time period can only be divided according to the "periodic" cycle, and the minimum length of the time period can only be the "periodic" cycle, which cannot be smaller, so the timeliness will be reduced).

[0120] Therefore, since the degradation of the overall performance of the host machine and the degradation of the performance of the virtual machine caused by "the VMM does not frequently enter the exception handling process and the virtual machine does not frequently exit to the kernel-state VMM" when the second expected number of executions is less than the preset threshold, it is often tolerable. In this way, in order to improve the timeliness of the actual number of executions of the cross-cache line operation of the host machine's CPU actually performed by the VCPU allocated to the virtual machine, it is acceptable to use "the user-state VMM to obtain the exception, and then obtain the relevant information of the exception (for example, the cause of the exception can be recorded in the shared area of ​​the user-state VMM and the kernel-state VMM, from which it can be analyzed whether the cause of the exception is due to the triggering of a split lock event, etc.), and according to the relevant information of the exception, it can be determined whether the VCPU allocated to the virtual machine has executed the cache line operation of the host machine's CPU (whether the VCPU allocated to the virtual machine has triggered a split lock event), and the number of cross-cache operations actually performed by the VCPU allocated to the virtual machine on the host machine's CPU can be counted by counting. The actual number of executions of the cross-cache line operation on the CPU of the host machine is detected by using the detection method of "actual execution times of the cross-cache line operation" to detect the actual execution times of the cross-cache line operation actually executed by the VCPU allocated to the virtual machine on the CPU of the host machine.

[0121] In this way, it is necessary to enable the host machine's CPU to throw an exception due to the CPU's memory access bus being locked.

[0122] However, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked is under the control of the kernel-state VMM, and the user-state VMM may not control the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked. Therefore, if the user-state VMM determines that it is necessary to start the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked when the second expected execution number is less than the preset threshold, the kernel-state VMM may be requested to start the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked.

[0123] For example, the VMM in user mode sends a start request to the VMM in kernel mode, where the start request is used to start a function in which the CPU of the host machine throws an exception because the memory access bus of the CPU is locked.

[0124] In one embodiment of the present application, the user-mode VMM may place the start request into a shared area (eg, a shared memory page, etc.) between the user-mode VMM and the kernel-mode VMM, and then notify the kernel-mode VMM.

[0125] Then, the VMM in the kernel state receives the start request and starts, according to the start request, the function of the CPU of the host machine throwing an exception due to the locking of the memory access bus of the CPU.

[0126] In one embodiment of the present application, after receiving the notification, the kernel-mode VMM may read the start request from a shared area between the user-mode VMM and the kernel-mode VMM.

[0127] There is a function switch corresponding to the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked. If the function switch is turned off, the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked can be turned off. If the function switch is turned on, the function in which the host machine's CPU throws an exception due to the CPU's memory access bus being locked can be started.

[0128] Thus, when the kernel-mode VMM starts the function of causing the host machine's CPU to throw an exception due to the CPU's memory access bus being locked according to the start request, it can start the function switch corresponding to the function of causing the host machine's CPU to throw an exception due to the CPU's memory access bus being locked.

[0129] The VMM in kernel mode can start the function switch through the externally exposed API of the function switch.

[0130] Accordingly, when determining whether the exception is thrown by the host machine's CPU because the CPU's memory access bus is locked based on the exception-related information, the user-mode VMM may determine whether the exception is thrown by the host machine's CPU because the CPU's memory access bus is locked based on the exception-related information.

[0131] Accordingly, when it is determined that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the host machine's CPU within the second time period, the user-state VMM may determine that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the host machine's CPU within the second time period.

[0132] Furthermore, since the second estimated execution number is less than the preset threshold, the detection thread may not be used to detect the actual execution number of cross-cache line operations performed by the VCPU allocated to the virtual machine on the host machine's CPU within the second time period. In this way, the state of the detection thread can be switched from an active state to a silent state.

[0133] For example, a kernel-mode VMM and a user-mode VMM are also running in the host machine.

[0134] Thus, when the state of the detection thread is switched from the active state to the silent state, in one embodiment, the VMM in the user state may send a silent request to the VMM in the kernel state, and the silent request is used to request that the state of the detection thread be switched from the active state to the silent state. The VMM in the kernel state may receive the silent request and switch the state of the detection thread from the active state to the silent state according to the silent request.

[0135] Alternatively, in another embodiment, the detection thread may directly request the VMM in kernel mode to switch the state of the detection thread from an active state to a silent state. The present application does not limit the specific switching method.

[0136] See also Figure 3 , an embodiment is used to illustrate the present solution, but it is not intended to limit the protection scope of the present solution.

[0137] Specifically, this embodiment includes the following process:

[0138] In step S301, the VMM in user mode predicts a first estimated number of executions of a cross-cache line operation on a CPU of a host machine that a VCPU allocated to a virtual machine is expected to perform within a first time period after a current moment.

[0139] In step S302, when the first expected number of executions is greater than or equal to a preset threshold, the VMM in user mode sends a shutdown request to the VMM in kernel mode, where the shutdown request is used to shut down the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked; so that the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked.

[0140] In step S303, the VMM in kernel mode receives a shutdown request, and according to the shutdown request, shuts down the function of the CPU of the host machine that throws an exception due to the memory access bus of the CPU being locked.

[0141] In step S304, the kernel-mode VMM sends a shutdown response to the user-mode VMM, where the shutdown response is used to notify the shutdown host machine's CPU of the function of throwing an exception due to the CPU's memory access bus being locked.

[0142] In step S305, the user-mode VMM receives the shutdown response, and sends an activation request to the kernel-mode VMM according to the shutdown response, where the activation request is used to request to switch the state of the detection thread from the silent state to the active state.

[0143] In step S306, the VMM in kernel state receives the activation request, and switches the state of the detection thread from the silent state to the active state according to the activation request.

[0144] In step S307 , when the state of the detection thread is switched to the activated state, the detection thread polls the CPU operation data recorded in the PMU corresponding to the CPU in the host machine.

[0145] In step S308, the detection thread obtains the actual number of executions of the cross-cache line operation on the CPU of the host machine actually performed by the VCPU allocated to the virtual machine within the first time period according to the polled CPU operation data.

[0146] In the present application, the VMM in user mode predicts the first estimated number of executions of the cross-cachline operation on the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform within a first time period after the current moment. When the first estimated number of executions is greater than or equal to a preset threshold, the VMM in user mode sends a shutdown request to the VMM in kernel mode, and the shutdown request is used to shut down the function of the CPU of the host machine throwing an exception due to the CPU's memory access bus being locked; so that the CPU of the host machine does not throw an exception when the CPU's memory access bus of the host machine is locked. The VMM in kernel mode receives the shutdown request, and shuts down the function of the CPU of the host machine throwing an exception due to the CPU's memory access bus being locked according to the shutdown request. The VMM in kernel mode sends a shutdown response to the VMM in user mode, and the shutdown response is used to notify that the function of the CPU of the host machine throwing an exception due to the CPU's memory access bus being locked has been shut down. The VMM in user mode receives the shutdown response, and sends an activation request to the VMM in kernel mode according to the shutdown response, and the activation request is used to request to switch the state of the detection thread from the silent state to the activated state. The VMM in kernel mode receives the activation request, and switches the state of the detection thread from the silent state to the activated state according to the activation request. When the state of the detection thread switches to the activated state, the detection thread polls the CPU operation data recorded in the PMU corresponding to the CPU in the host machine. The detection thread obtains the actual number of executions of the cross-cachline operation on the CPU of the host machine actually performed by the VCPU allocated to the virtual machine within the first time period based on the polled CPU operation data. Through the present application, in the case where the VCPU allocated to the virtual machine performs cross-cach line operations on the CPU of the host machine very frequently (for example, tens of thousands or hundreds of thousands of times per second, etc.), in the scenario of detecting the actual number of executions of the cross-cach line operation on the CPU of the host machine actually performed by the VCPU allocated to the virtual machine, the overall performance of the host machine and the performance of the virtual machine can be improved.

[0147] It should be noted that, for the method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by the present application.

[0148] Reference Figure 4 , shows a structural block diagram of a data processing device of the present application, which is applied to a host machine, in which at least a virtual machine and a detection thread are running; the device includes:

[0149] A first prediction module 11 is used to predict a first estimated number of executions of a cross-cache line operation on a central processing unit (CPU) of a host machine that a VCPU allocated to a virtual machine is expected to perform within a first time period after a current moment; a shutdown module 12 is used to shut down a function of the host machine's CPU throwing an exception due to a locked memory access bus of the CPU when the first estimated number of executions is greater than or equal to a preset threshold, so that the host machine's CPU does not throw an exception when the memory access bus of the host machine's CPU is locked; and a first switching module 13 is used to switch the state of a detection thread from a silent state to an active state; a polling module 14 is used to poll the CPU's operating data recorded in a performance monitor PMU corresponding to the CPU in the host machine; an acquisition module 15 is used to acquire, based on the polled CPU's operating data, the actual number of executions of the cross-cache line operation on the host machine's CPU that the VCPU allocated to the virtual machine actually performs within the first time period.

[0150] In an optional implementation, a kernel-state VMM and a user-state VMM are also running in the host machine; the first prediction module includes: a first prediction unit of the user-state VMM; the first prediction unit is used to predict a first expected number of executions of a cross-cache line cache line operation on the host machine's central processing unit CPU by the VCPU allocated to the virtual machine within a first time period after the current moment; accordingly, the shutdown module includes: a first sending unit included in the user-state VMM, and also includes a first receiving unit and a shutdown unit included in the kernel-state VMM; the first sending unit is used to send a shutdown request to the first receiving unit, the shutdown request is used to shut down the function; the first receiving unit is used to receive the shutdown request, and the shutdown unit is used to shut down the function according to the shutdown request.

[0151] In an optional implementation, the first switching module includes: a second sending unit included in the user-state VMM, and also includes a second receiving unit and a first switching unit included in the kernel-state VMM; the second sending unit is used to send an activation request to the second receiving unit, and the activation request is used to request to switch the state of the detection thread from a silent state to an activated state; the second receiving unit is used to receive the activation request, and the first switching unit is used to switch the state of the detection thread from a silent state to an activated state according to the activation request.

[0152] In an optional implementation, the first switching module also includes: a third sending unit included in the user-mode VMM, and a third receiving unit included in the kernel-mode VMM; the third sending unit is used to send a shutdown response to the third receiving unit, and the shutdown response is used to notify that the function has been shut down; the third receiving unit is used to receive the shutdown response, and the second sending unit is also used to send an activation request to the second receiving unit based on the shutdown response.

[0153] In an optional implementation, the device also includes: a second prediction module, used to predict a second expected number of executions of the cross-cache line operation on the host machine's CPU that the VCPU allocated to the virtual machine is expected to perform in a second time period after the first time period; a startup module, used to start the function when the second expected number of executions is less than a preset threshold; so that the host machine's CPU throws an exception when the host machine's CPU's memory access bus is locked; a first determination module, used to determine whether the exception is thrown by the host machine's CPU due to the CPU's memory access bus being locked based on the relevant information of the exception when the relevant information of the exception is obtained; and a second determination module, used to determine that the VCPU allocated to the virtual machine actually performed the cross-cache line operation on the host machine's CPU in the second time period when the exception is thrown by the host machine's CPU due to the CPU's memory access bus being locked.

[0154] In an optional implementation, a kernel-state VMM and a user-state VMM are also running in the host machine; the second prediction module includes: a second prediction unit of the user-state VMM; the second prediction unit is used to predict a second expected number of executions of a cross-cache line operation on the host machine's CPU by the VCPU allocated to the virtual machine in a second time period after the first time period; accordingly, the startup module includes: a fourth sending unit included in the user-state VMM, and also includes a fourth receiving unit and a startup unit included in the kernel-state VMM; the fourth sending unit is used to send a startup request to the fourth receiving unit, the startup request is used to start the function; the fourth receiving unit is used to receive the startup request, and the startup unit is used to start the function according to the startup request.

[0155] In an optional implementation, the first determination module also includes a first determination unit included in the user-state VMM; the first determination unit is used to determine whether the exception is thrown by the CPU of the host machine because the memory access bus of the CPU is locked based on the relevant information of the exception; accordingly, the second determination module includes a second determination unit included in the user-state VMM; the second determination unit is used to determine whether the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine within the second time period.

[0156] In an optional implementation, the device further includes: a second switching module, configured to switch the state of the detection thread from an active state to a silent state when the second expected execution number is less than a preset threshold.

[0157] In an optional implementation, a kernel-state VMM and a user-state VMM are also running in the host machine; the second switching module includes: a fifth sending unit included in the user-state VMM, a fifth receiving unit included in the kernel-state VMM, and a second switching unit; including: a fifth sending unit, used to send a silence request to the fifth receiving unit, the silence request is used to request that the state of the detection thread be switched from an active state to a silent state; a fifth receiving unit, used to receive the silence request, and a second switching unit, used to switch the state of the detection thread from an active state to a silent state according to the silence request.

[0158] In an optional implementation, the first prediction module includes: a first acquisition unit, used to obtain the historical execution times of the VCPU allocated to the virtual machine actually performing cross-cacheline operations on the host machine's CPU in at least one historical time period before the current moment; and a second acquisition unit, used to obtain the first estimated execution times based on the historical execution times.

[0159] In an optional implementation, the operation data of the polled CPU includes: a first number of executions of cross-cache line operations on the CPU of the host machine performed by the VCPU assigned to the virtual machine at the start time of the first time period and a second number of executions of cross-cache line operations on the CPU of the host machine performed by the VCPU assigned to the virtual machine at the end time of the first time period; the acquisition module includes: a calculation unit, used to calculate the difference between the second execution number and the first execution number; and a third acquisition unit, used to obtain the actual execution number according to the difference.

[0160] In the present application, the first estimated number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform in the first time period after the current moment. When the first estimated number of executions is greater than or equal to the preset threshold, the function of the CPU of the host machine throwing an exception due to the memory access bus of the CPU being locked is turned off. So that the CPU of the host machine does not throw an exception when the memory access bus of the CPU of the host machine is locked. And, the state of the detection thread is switched from a silent state to an active state. So that the detection thread polls the operation data of the CPU recorded in the PMU corresponding to the CPU in the host machine, and obtains the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs in the first time period according to the operation data of the polled CPU. Through the present application, in the case where the VCPU allocated to the virtual machine performs the cross-cache line operation on the CPU of the host machine very frequently (for example, tens of thousands or hundreds of thousands of times per second, etc.), in the scenario of detecting the actual number of executions of the cross-cache line operation of the CPU of the host machine that the VCPU allocated to the virtual machine actually performs, the overall performance of the host machine and the performance of the virtual machine can be improved.

[0161] The embodiment of the present application also provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0162] The present application embodiment provides one or more machine-readable media on which instructions are stored, and when executed by one or more processors, the electronic device executes one or more methods in the above embodiments. In the present application embodiment, the electronic device includes a server, a gateway, a sub-device, etc., and the sub-device is an Internet of Things device or other device.

[0163] The embodiments of the present disclosure may be implemented as an apparatus configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the apparatus may include electronic devices such as servers (clusters), terminal devices such as IoT devices, and the like.

[0164] Figure 5 An exemplary device 1300 that can be used to implement various embodiments of the present application is schematically shown.

[0165] For one embodiment, Figure 5An exemplary apparatus 1300 is shown having one or more processors 1302, a control module (chip set) 1304 coupled to at least one of the (one or more) processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.

[0166] The processor 1302 may include one or more single-core or multi-core processors, and the processor 1302 may include any combination of general-purpose processors or special-purpose processors (such as graphics processors, application processors, baseband processors, etc.). In some embodiments, the device 1300 can be used as a server device such as a gateway in the embodiments of the present application.

[0167] In some embodiments, the device 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 configured to execute the instructions 1314 in combination with the one or more computer-readable media to implement a module to perform the actions of the present disclosure.

[0168] For one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1302 and / or any suitable device or component in communication with the control module 1304 .

[0169] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0170] The memory 1306 may be used, for example, to load and store data and / or instructions 1314 for the device 1300. For one embodiment, the memory 1306 may include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 1306 may include a double data rate quad synchronous dynamic random access memory (DDR4 SDRAM).

[0171] For one embodiment, control module 1304 may include one or more input / output controllers to provide an interface to NVM / storage device 1308 and input / output device(s) 1310 .

[0172] For example, NVM / storage 1308 may be used to store data and / or instructions 1314. NVM / storage 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0173] NVM / storage device 1308 may include storage resources that are physically part of the device on which apparatus 1300 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 1308 may be accessed via input / output device(s) 1310 over a network.

[0174] (One or more) input / output devices 1310 may provide an interface for the apparatus 1300 to communicate with any other appropriate device, and the input / output device 1310 may include a communication component, a phonetic component, a sensor component, etc. The network interface 1312 may provide an interface for the apparatus 1300 to communicate through one or more networks, and the apparatus 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0175] For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers (e.g., a memory controller module) of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be packaged together with the logic of one or more controllers of the control module 1304 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304. For one embodiment, at least one of the processor(s) 1302 may be integrated on the same die with the logic of one or more controllers of the control module 1304 to form a system-on-chip (SoC).

[0176] In various embodiments, the device 1300 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the device 1300 may have more or fewer components and / or a different architecture. For example, in some embodiments, the device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touch screen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0177] An embodiment of the present application provides an electronic device, comprising: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, enable the electronic device to perform one or more methods in the present application.

[0178] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0179] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0180] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of the processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable information processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable information processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0181] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable information processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0182] These computer program instructions can also be loaded onto a computer or other programmable information processing terminal device, so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0183] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0184] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0185] The data processing method and device provided by the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A data processing method, characterized in that: Applied to a host machine, the host machine at least runs a virtual machine and a detection thread; the method includes: Predicting a first estimated number of executions of a cross-cache line cache line operation on a central processing unit (CPU) of a host machine that a virtual central processing unit (VCPU) allocated to the virtual machine is expected to perform within a first time period after a current moment; When the first expected number of executions is greater than or equal to a preset threshold, the function of the host machine's CPU throwing an exception due to the CPU's memory access bus being locked is turned off; so that the host machine's CPU does not throw an exception when the host machine's CPU's memory access bus is locked; and the state of the detection thread is switched from a silent state to an active state; so that the detection thread polls the CPU's operating data recorded in the performance monitor PMU corresponding to the CPU in the host machine, and obtains the actual number of executions of the cross-cache line operation of the VCPU allocated to the virtual machine on the host machine's CPU in the first time period based on the polled CPU's operating data.

2. The method according to claim 1, characterized in that The host machine also runs a kernel-mode virtual machine monitor VMM and a user-mode VMM; Predicting a first estimated number of executions of a cross-cache line cache line operation on a central processing unit (CPU) of a host machine that a VCPU allocated to the virtual machine is expected to perform within a first time period after a current moment includes: The VMM in the user state predicts a first estimated number of executions of a cross-cache line operation on a central processing unit (CPU) of the host machine that a VCPU allocated to the virtual machine is expected to perform within a first time period after the current moment; Accordingly, the function of shutting down the CPU of the host machine and throwing an exception due to the memory access bus of the CPU being locked includes: The VMM in user mode sends a shutdown request to the VMM in kernel mode, where the shutdown request is used to shut down the function; The VMM in kernel mode receives the shutdown request and shuts down the function according to the shutdown request.

3. The method according to claim 2, characterized in that The step of switching the state of the detection thread from a silent state to an active state includes: The user-mode VMM sends an activation request to the kernel-mode VMM, where the activation request is used to request that the state of the detection thread be switched from a silent state to an activated state; The VMM in the kernel state receives the activation request, and switches the state of the detection thread from the silent state to the active state according to the activation request.

4. The method according to claim 3, characterized in that The method further comprises: The kernel-mode VMM sends a shutdown response to the user-mode VMM, where the shutdown response is used to notify that the function has been shut down; The user-mode VMM receives the shutdown response, and executes the step of sending an activation request to the kernel-mode VMM according to the shutdown response.

5. The method according to claim 1, characterized in that The method further comprises: Predicting a second estimated number of executions of a cross-cache line operation on a CPU of a host machine that is expected to be performed by the VCPU allocated to the virtual machine in a second time period after the first time period; When the second expected execution number is less than a preset threshold, the function is started, so that the CPU of the host machine throws an exception when the memory access bus of the CPU of the host machine is locked; When the relevant information of the exception is obtained, determining whether the exception is thrown by the CPU of the host machine because the memory access bus of the CPU is locked according to the relevant information of the exception; In the case where the exception is thrown by the CPU of the host machine because the memory access bus of the CPU is locked, it is determined that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine within the second time period.

6. The method according to claim 5, characterized in that The host machine also runs a kernel-mode VMM and a user-mode VMM; The prediction of a second estimated number of times that the VCPU allocated to the virtual machine is expected to perform a cross-cache line operation on the CPU of the host machine in a second time period after the first time period includes: The VMM in the user state predicts a second estimated number of executions of the cross-cache line operation on the CPU of the host machine that the VCPU allocated to the virtual machine is expected to perform in a second time period after the first time period; Accordingly, starting the function includes: The VMM in user mode sends a start request to the VMM in kernel mode, where the start request is used to start the function; The VMM in kernel mode receives the start request and starts the function according to the start request.

7. The method according to claim 6, characterized in that The determining, according to the relevant information of the exception, whether the exception is thrown by the CPU of the host machine because the memory access bus of the CPU is locked includes: The user-mode VMM determines, based on the relevant information of the exception, whether the exception is thrown by the CPU of the host machine because the CPU memory access bus is locked; Accordingly, determining that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine in the second time period includes: The user-mode VMM determines that the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine within the second time period.

8. The method according to claim 5, characterized in that The method further comprises: When the second estimated execution number is less than a preset threshold, the state of the detection thread is switched from an active state to a silent state.

9. The method according to claim 8, characterized in that The host machine also runs a kernel-mode VMM and a user-mode VMM; The step of switching the state of the detection thread from the active state to the silent state includes: The VMM in user mode sends a quiet request to the VMM in kernel mode, where the quiet request is used to request to switch the state of the detection thread from an active state to a quiet state; The VMM in the kernel state receives the quiet request, and switches the state of the detection thread from the active state to the quiet state according to the quiet request.

10. The method according to claim 1, characterized in that Predicting a first estimated number of executions of a cross-cache line cache line operation on a central processing unit (CPU) of a host machine that a VCPU allocated to the virtual machine is expected to perform within a first time period after a current moment includes: Obtain the historical number of cross-cache line operations actually performed by the VCPU allocated to the virtual machine on the CPU of the host machine in at least one historical time period before the current moment; The first estimated execution number is obtained according to the historical execution number.

11. The method according to claim 1, characterized in that: The polled CPU operation data includes: a first number of times the VCPU assigned to the virtual machine has executed a cross-cache line operation on the CPU of the host machine at the start time of the first time period and a second number of times the VCPU assigned to the virtual machine has executed a cross-cache line operation on the CPU of the host machine at the end time of the first time period; The actual number of times the VCPU allocated to the virtual machine actually performs a cross-cache line operation on the CPU of the host machine in the first time period is obtained according to the polled CPU operation data, including: Calculating the difference between the second number of executions and the first number of executions; The actual number of executions is obtained according to the difference.

12. A data processing device, characterized in that: Applied to a host machine, in which at least a virtual machine and a detection thread are running; the device comprises: A first prediction module, used for predicting a first estimated number of executions of a cross-cache line cache line operation on a central processing unit (CPU) of a host machine that a virtual central processing unit (VCPU) allocated to the virtual machine is expected to perform within a first time period after a current moment; A first prediction module is used to predict a first estimated number of executions of a cross-cache line operation on a central processing unit (CPU) of a host machine that a VCPU allocated to a virtual machine is expected to perform within a first time period after a current moment; a shutdown module is used to shut down a function of the host machine's CPU throwing an exception due to a locked memory access bus of the CPU when the first estimated number of executions is greater than or equal to a preset threshold, so that the host machine's CPU does not throw an exception when the memory access bus of the host machine's CPU is locked; and a first switching module is used to switch the state of a detection thread from a silent state to an active state; a polling module is used to poll the CPU's operating data recorded in a performance monitor PMU corresponding to the CPU in the host machine; an acquisition module is used to acquire, based on the polled CPU's operating data, the actual number of executions of the cross-cache line operation on the host machine's CPU that the VCPU allocated to the virtual machine actually performs within the first time period.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium, characterized in that: A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Virtual machine coexisting scheduling method based on processor performance monitoring

    CN103955398A

  • Data processing method and device

    CN111124947A