Memory access conflict detection method based on cloud service, medium, equipment and program product
By collecting and analyzing the performance monitoring counter values of the physical central processing unit cores in cloud services, and combining them with virtual machine binding relationships, accurate detection and suppression of memory access conflicts are achieved, solving the memory access conflict problem between virtual machines and improving the stability of cloud services and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-03-31
AI Technical Summary
In cloud service scenarios, there is a lack of effective identification mechanisms for memory access conflicts between virtual machines, which leads to performance fluctuations and unpredictable business impacts. Existing technologies cannot handle these conflicts in a timely manner.
By collecting the performance monitoring counter values of each physical CPU core on the host machine and combining them with the binding relationship between the virtual machine and the physical CPU core, the number of memory access conflicts of the virtual machine is determined, and the first count value is detected by the performance monitoring counter, so as to achieve non-intrusive memory access conflict detection and trend display.
It achieves accurate detection of memory contention events from hardware signals to the virtual machine level, provides non-intrusive detection and suppression mechanisms, reduces the impact of performance jitter on business, and improves the reliability of cloud services and user experience.
Smart Images

Figure CN121764597A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a method, medium, device, and program product for detecting memory access conflicts based on cloud services. Background Technology
[0002] In cloud service scenarios, service providers use virtualization technology to run multiple virtual machines on the same physical machine to achieve efficient resource utilization. To ensure performance isolation between virtual machines, various isolation measures are typically employed, such as binding virtual CPUs to physical CPUs on a one-to-one basis, thereby avoiding resource contention between different virtual machines.
[0003] However, some hardware resources on the physical machine are still shared, especially the mechanism by which the CPU accesses the memory bus. For example, when operands in an atomic operation executed by the virtual CPU of a virtual machine span two cache lines, a bus lock operation is triggered, leading to performance jitter and unpredictable business impact. Related technologies lack effective mechanisms to quickly identify these memory access conflict events, resulting in the inability to handle them in a timely manner. Summary of the Invention
[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] Firstly, this disclosure provides a cloud service-based memory access conflict detection method, including: Collect the first count value of the performance monitoring counter of each physical central processing unit core on the host machine, wherein the performance monitoring counter records the number of memory access contention events that occur in the physical central processing unit core; Based on the binding relationship between the virtual CPU and the physical CPU core of the virtual machine and the first count value of the performance monitoring counter of each physical CPU core, a second count value corresponding to each virtual machine is determined. The second count value represents the number of times the virtual machine has encountered memory access contention events.
[0006] Secondly, this disclosure provides a cloud service-based memory access conflict detection device, comprising: The acquisition module is used to acquire the first count value of the performance monitoring counter of each physical central processing unit core on the host machine, wherein the performance monitoring counter records the number of times the physical central processing unit core has a memory access contention event; The determination module is used to determine a second count value for each virtual machine based on the binding relationship between the virtual central processing unit and the physical central processing unit core of the virtual machine and the first count value of the performance monitoring counter of each physical central processing unit core. The second count value represents the number of times the virtual machine has encountered a memory access contention event.
[0007] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0008] Fourthly, this disclosure provides an electronic device, comprising: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0009] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] Based on the above technical solution, by collecting the first count value of the performance monitoring counters of each physical CPU core on the host machine, and determining the second count value corresponding to each virtual machine based on the binding relationship between the virtual CPU and the physical CPU cores and the first count value of the performance monitoring counters of each physical CPU core, the first count value corresponding to global memory access contention events can be collected by reading the count value of the performance monitoring counters. Furthermore, through the binding relationship, physically occurring memory access contention events can be precisely associated with specific virtual machines, achieving memory access contention event detection from hardware signal perception to the virtual machine level. This not only enables non-intrusive memory access contention event detection but also provides accurate data support for subsequent suppression of memory access contention events and / or displaying the trend of memory access contention events occurring in virtual machines. Moreover, detecting the first count value through performance monitoring counters is also compatible with different types of CPUs.
[0011] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1This is a flowchart illustrating a cloud service-based memory access conflict detection method according to some embodiments.
[0013] Figure 2 This is a schematic diagram illustrating the principle of suppressing memory contention events according to some embodiments.
[0014] Figure 3 This is a logical diagram illustrating the suppression of memory contention events according to some embodiments.
[0015] Figure 4 This is a schematic diagram of the architecture of a cloud service-based memory access conflict detection method, illustrated according to some embodiments.
[0016] Figure 5 This is a schematic diagram of the module connections of a cloud service-based memory access conflict detection device according to an exemplary embodiment.
[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an exemplary embodiment. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0022] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0023] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0024] Figure 1 This is a flowchart illustrating a cloud service-based memory access conflict detection method according to some embodiments. For example... Figure 1 As shown, this disclosure provides a cloud service-based memory access conflict detection method, which can be executed by a cloud service-based memory access conflict detection device, which can be implemented in software and / or hardware. Figure 1 As shown, the method may include the following steps.
[0025] In step 110, the first count value of the performance monitoring counter of each physical central processing unit core on the host machine is collected, wherein the performance monitoring counter records the number of times memory access contention events occur in the physical central processing unit core.
[0026] Here, the host machine refers to the physical computer or server that directly runs virtual machines (VMs) and provides them with hardware resources (such as CPU (Central Processing Unit), memory, and storage). The host machine is the foundation of the virtualization environment, responsible for managing and allocating resources to support the operation of one or more virtual machines (also known as guest machines). A virtual machine is a computing environment created through software simulation, enabling a single physical computer to run multiple independent operating systems and applications. Virtual machines can provide isolated computing resources on a single hardware platform, simulating multiple independent computer systems, thereby improving hardware resource utilization and flexibility.
[0027] The host machine's physical central processing unit (CPU) refers to the central processing unit in the actual hardware of the host machine, the core component responsible for performing computational tasks. A host machine can have multiple physical CPUs, each containing multiple cores, and each core can handle independent computational tasks. Of course, each core can support multithreading, achieving parallel processing through hyper-threading or multi-core technology.
[0028] In a virtualized environment, a single physical computer (host machine) can run multiple independent virtual machines simultaneously, and each virtual machine can have one or more vCPUs (virtual central processing units). A virtual central processing unit is a logical processing unit based on the physical CPU, partitioned using virtualization technology. It abstracts physical computing resources into dynamically allocable virtual resources, supporting multiple virtual machines sharing hardware. Logically, virtual central processing units are independent of each other; resource isolation between virtual central processing units and physical central processing units is achieved through scheduling and resource management at the virtualization layer.
[0029] The Performance Monitor Unit (PMU) and Performance Monitor Counter (PMC) in a CPU are embedded hardware components used to monitor and measure the CPU's operating performance. The PMU is an integrated hardware unit containing multiple PMCs. Each PMC is the specific counting hardware within the PMU, and each PMC can be programmed to monitor specific performance events. As the management unit for the PMCs, the PMU is responsible for coordinating the operation of all performance monitor counters.
[0030] Memory contention refers to a phenomenon in a computer system where multiple instructions or operations compete for access to the same memory resource simultaneously, leading to execution delays or anomalies. Virtual machine-generated memory contention refers to high-cost memory operation events caused by the virtual machine's workload, resulting in resource contention or performance interference on the underlying physical memory subsystem. Memory contention significantly impacts the performance of the host machine and other virtual machines. For example, in the x86 CPU microarchitecture, unaligned memory access is allowed. When the operands in an atomic operation (such as an atomic read, modify, or write operation) executed by the virtual machine's vCPU span two cache lines, bus locking is triggered. Bus locking significantly increases the latency of CPU memory access. For instance, under the split lock mechanism, when an atomic operation with a lock prefix accesses a memory address that spans two cache line boundaries, the CPU locks the memory bus to ensure the atomicity of the operation. Furthermore, since the memory bus is a shared resource, bus locking severely impacts the performance of other virtual machines on the same physical machine, leading to performance fluctuations and unpredictable business impacts. For example, under high load scenarios, contention for memory access may lead to longer response times for cloud servers, reduced service quality, or even service interruptions.
[0031] In this embodiment of the disclosure, a performance monitoring counter is deployed on the core of each physical central processing unit (CPU) of the host machine. This performance monitoring counter increments its value when a memory access contention event occurs on the corresponding core, thereby counting the number of memory access contention events occurring on the core. For example, when a physical CPU core performs a cross-cache-line atomic operation that generates a memory access contention event, the performance monitoring counter corresponding to that core will automatically increment.
[0032] For example, CPU registers can be configured using register control tools, writing configuration information to the CPU registers. This configuration information indicates the types of events that need to be monitored. The CPU hardware responds to the configuration information by allocating performance monitoring counters for monitoring memory access contention events.
[0033] It should be understood that the PMU has multiple performance monitoring counters, and one can be exclusively used to track memory access contention events. Once a memory access contention event occurs on the CPU core, the performance monitoring counter automatically increments by 1, requiring no software intervention and incurring minimal overhead.
[0034] In this embodiment of the disclosure, the first count value of the performance monitoring counter on all physical CPU cores can be periodically polled and collected. Each time the first count value of the performance monitoring counter is read, it represents the number of memory access contention events that occurred on the corresponding physical CPU core.
[0035] It's worth noting that the system can also record a sequence of associated data, including the acquisition timestamp, the physical CPU core number, and the first count value of the physical CPU core. The acquisition timestamp refers to the system time at which the first count value was acquired. The physical CPU core number can be a unique identifier corresponding to the physical CPU core. When acquiring the first count value of the performance monitoring counter, the acquisition timestamp and the core's corresponding number can be recorded simultaneously, forming an associated data sequence of "core number - timestamp - first count value". Of course, the recorded associated data sequence can be stored in a database.
[0036] In step 120, based on the binding relationship between the virtual CPU and the physical CPU core of the virtual machine and the first count value of the performance monitoring counter of each physical CPU core, the second count value corresponding to each virtual machine is determined. The second count value represents the number of times the virtual machine has encountered memory access contention events.
[0037] Here, the binding relationship between a virtual machine's virtual CPU and its physical CPU cores refers to the record of which physical CPU cores each virtual CPU is scheduled to run on at a given moment or within a given time slice. It should be understood that this binding relationship can be dynamic, meaning that each time a virtual CPU is scheduled, it can be assigned to any available physical CPU core. For example, assuming a host machine has 32 cores and runs 10 virtual machines, each virtual machine configured with 4 virtual CPUs, then a total of 40 vCPUs will compete for the 32 cores.
[0038] For example, the binding relationship between a virtual machine's virtual CPU and physical CPU cores can be queried in real time through the management interfaces provided by Libvirt or KVM (Kernel-based Virtual Machine). Libvirt is a virtualization management toolset and interface framework used for unified management and control of the lifecycle of virtual machines under various virtualization technologies. Libvirt allows querying the binding relationship between a virtual machine's virtual CPU and physical CPU cores. KVM is a full virtualization solution used to efficiently run multiple isolated virtual machines on a host machine. KVM allows querying the binding relationship between a virtual machine's virtual CPU and physical CPU cores.
[0039] By leveraging the binding relationship between the virtual central processing unit (CPU) and the physical CPU core of the virtual machine, the first count value collected on the core is associated with the corresponding virtual machine to obtain the second count value corresponding to the memory access contention events that occur in each virtual machine.
[0040] It should be noted that, through the binding relationship, the first count value of memory access contention events occurring on the core of the physical central processing unit is actually assigned to the specific virtual machine, resulting in a second count value of memory access contention events occurring on each virtual machine. This second count value can be understood as a count value at the virtual machine level.
[0041] In this embodiment of the disclosure, after obtaining the second count value corresponding to the memory access contention event that occurred in the virtual machine, the second count value corresponding to each virtual machine can be stored in the database. Of course, if the collected data is a sequence of associated data of "core number-timestamp-first count value", then after associating it with a specific virtual machine through the binding relationship, the obtained associated data sequence should be "virtual machine identifier-timestamp-second count value", and correspondingly, the associated data sequence of "virtual machine identifier-timestamp-second count value" can be stored in the database.
[0042] Following the above implementation method, when recording the associated data sequence including the collection timestamp, the physical central processing unit core number, and the first count value of the physical central processing unit core, the second count value corresponding to each virtual machine can be determined according to the binding relationship and the associated data sequence.
[0043] The recorded second count value is used to visualize the trend of memory access contention events occurring in the virtual machine and / or to suppress memory access contention events occurring in the virtual machine based on the second count value.
[0044] Therefore, by collecting the first count value of the performance monitoring counters of each physical CPU core on the host machine, and based on the binding relationship between the virtual CPU and the physical CPU cores of the virtual machine, as well as the first count value of the performance monitoring counters of each physical CPU core, the second count value corresponding to each virtual machine is determined. The first count value corresponding to global memory access contention events can be collected by reading the performance monitoring counter values. Furthermore, through the binding relationship, physically occurring memory access contention events can be precisely associated with specific virtual machines, achieving memory access contention event detection from hardware signal perception to the virtual machine level. This not only enables non-intrusive memory access contention event detection but also provides accurate data support for subsequent suppression of memory access contention events and / or displaying trends of memory access contention events occurring in virtual machines. Moreover, detecting the first count value through performance monitoring counters is also compatible with different types of CPUs.
[0045] In some feasible implementations, the occurrence rate of each virtual machine can be determined based on the second count value of each virtual machine and the actual runtime of each virtual machine. Target virtual machines with occurrence rates greater than the alarm threshold can be filtered out, and actions can be performed to suppress the target virtual machines from continuing to generate memory access contention events.
[0046] Here, within the time slice allocated to the virtual machine, the occurrence rate of the virtual machine can be determined based on the virtual machine's second count value and the virtual machine's actual runtime. The occurrence rate represents the frequency at which the virtual machine experiences memory contention events.
[0047] Time slicing refers to the time allocated by the central processing unit (CPU) to each virtual machine (VM). Each VM is allocated a time period, called its time slice, which represents the allowed runtime for the VM. For example, the duration of a time slice allocated to a VM could be as short as 1 second. Of course, the duration of a time slice can be set according to specific requirements.
[0048] The actual runtime of a virtual machine refers to the actual time the virtual machine has run within a corresponding time slice. For example, if a time slice is 1 second, the actual runtime of the virtual machine will be less than or equal to 1 second.
[0049] Within the duration of each time slice, the rate at which memory contention events occur during the actual runtime of the virtual machine can be calculated based on the quotient between the second count value recorded in that time slice and the actual runtime of the virtual machine. For example, assuming that in a 1-second time slice, the actual runtime of the virtual machine is 0.2 seconds, and 300 memory contention events are detected within these 0.2 seconds, the occurrence rate is 300 times / 0.2 seconds = 1500 times / second.
[0050] It should be understood that at the end of each time slice, the performance monitoring counter can be reset, that is, the count value of the performance monitoring counter can be reset to 0. In other words, the performance monitoring counter is used to record the number of memory access contention events that occur within the duration of each time slice.
[0051] An alarm threshold refers to the maximum number of memory access contention events allowed to occur per unit of time. For example, an alarm threshold could be the maximum number of memory access contention events allowed per second, such as 1000 times / second.
[0052] When the occurrence rate exceeds the alarm threshold, an anomaly is indicated, requiring suppression of memory contention events. It's important to note that in this embodiment, the need for suppression is not based on whether the absolute second count value exceeds the threshold, but rather on whether the occurrence rate per unit time exceeds the alarm threshold. Even if the virtual machine actually runs for 0.1 seconds, if the occurrence rate corresponding to 0.1 seconds exceeds the alarm threshold, memory contention events still need to be suppressed. By determining whether the occurrence rate exceeds the alarm threshold, short-term bursts of abuse can be prevented and misjudgments of long-term low-frequency tasks can be avoided. For example, a malicious or vulnerable program might initiate a large number of memory contention events in a very short time. Although the virtual machine may not have completed a full time slice, it has already severely interfered with other cores, and suppression will be triggered in this case.
[0053] In this embodiment of the disclosure, if the occurrence rate of a virtual machine is greater than the alarm threshold, then the virtual machine is the target virtual machine whose occurrence rate is greater than the alarm threshold, and actions to suppress the target virtual machine from continuing to generate memory access contention events can be performed.
[0054] In some embodiments, actions to suppress the target virtual machine from continuing to generate memory contention events may be performed in response to the actual runtime being less than the duration corresponding to the time slice and the occurrence rate being greater than the alarm threshold.
[0055] In this case, if the actual runtime is less than the duration corresponding to the time slice, it means that the time slice allocated to the virtual machine has not yet ended. If the occurrence rate exceeds the alarm threshold while the time slice allocated to the virtual machine has not ended, it can be used to suppress the target virtual machine from continuing to generate memory access contention events. If the actual runtime is equal to the duration corresponding to the time slice, it means that the time slice allocated to the virtual machine has ended. In this case, even if the occurrence rate exceeds the alarm threshold, it is not necessary to use actions to suppress the target virtual machine from continuing to generate memory access contention events.
[0056] It's worth noting that controlling the virtual machine sleep rate corresponding to the occurrence rate is actually an action used to suppress memory access contention events. By controlling the virtual machine sleep rate corresponding to the occurrence rate, it can be forced that the frequency of memory access contention events generated by the vCPU corresponding to the virtual machine within a complete time slice will not exceed the alarm threshold.
[0057] As examples, the target virtual machine can be controlled to enter a sleep state. In other words, an action used to suppress the target virtual machine from continuing to generate memory access contention events can be to control the target virtual machine to enter a sleep state.
[0058] As another example, the operating frequency of the virtual CPU corresponding to the target virtual machine can be reduced. In other words, reducing the operating frequency of the virtual CPU corresponding to the target virtual machine can be an action used to suppress further memory access contention events in the target virtual machine.
[0059] In some embodiments, the sleep duration of the target virtual machine can be set, and the sleep state of the target virtual machine can be controlled according to the sleep duration. The sleep duration is the difference between the preset time slice duration and the actual runtime of the virtual machine.
[0060] The sleep duration of the target virtual machine is the difference between the preset time shard duration and the actual runtime of the target virtual machine. For example, if the preset time shard duration of the target virtual machine is 1 second and the actual runtime is 0.2 seconds, then the sleep duration is 1 - 0.2 = 0.8 seconds.
[0061] It is worth noting that by controlling the virtual machine sleep preset duration corresponding to the occurrence rate, it is possible to force the rate at which the vCPU corresponding to the virtual machine generates memory access contention events within a complete time slice to not exceed the alarm threshold.
[0062] Therefore, through the above implementation method, when the occurrence rate exceeds the alarm threshold, actions can be taken to suppress the target virtual machine from continuing to generate memory access contention events. For example, the corresponding target virtual machine can be controlled to force the rate at which the vCPU corresponding to the virtual machine generates memory access contention events within a complete time slice not to exceed the alarm threshold. Moreover, by suppressing the virtual machine from generating memory access contention events, the performance degradation of the host machine caused by memory access contention events can be avoided, greatly improving the reliability of cloud services and user experience. In addition, in the embodiments of this disclosure, the entire process from detecting the second count value to suppressing the virtual machine from generating memory access contention events is completed automatically, which can reduce the response time from minutes to seconds, greatly reducing the impact of performance jitter on business, and also reducing the reliance on manual intervention, thus reducing the operational complexity and cost of cloud service providers.
[0063] In some feasible implementations, alarm thresholds can be generated based on the host machine's system performance degradation coefficient, the time required for the host machine's physical central processing unit to generate a memory access contention event, and the virtual machine's application performance degradation coefficient.
[0064] Here, alarm thresholds can be generated based on the host machine's system performance degradation coefficient, the time required for the host machine's physical central processing unit to generate a memory access contention event, and the virtual machine's application performance degradation coefficient, combined with a preset performance quantification model.
[0065] The alarm threshold is essentially a rate threshold, meaning it can be understood as an alarm rate, which indicates the maximum rate at which the host machine can accept events that generate memory access contention.
[0066] The performance quantification model represents the system performance degradation coefficient as the product of the alarm threshold, the time required to generate a memory access contention event, and the application performance degradation coefficient.
[0067] For example, the performance quantification model can be represented as Where P is the system performance degradation coefficient, N is the alarm threshold, T is the time required to generate a memory access contention event, and K is the application performance degradation coefficient.
[0068] It should be noted that the performance quantification model can quantify the impact of memory contention events on physical machine performance, providing a unified standard for calculating alarm thresholds.
[0069] The system performance degradation factor can refer to the percentage decrease in various performance indicators of the host machine, such as latency, bandwidth, and latency. The host machine's system performance degradation factor can be set by the user according to their needs. Alternatively, it can refer to the maximum acceptable system performance degradation factor for the host machine.
[0070] In some embodiments, a test environment consistent with the host machine's hardware environment can be constructed, and then different types of workloads can be deployed in the test environment. The workloads can be controlled to generate memory access contention events at different rates, and the system performance degradation coefficients corresponding to the memory access contention events at different rates in the test environment can be collected. Then, based on the collected system performance degradation coefficients, the system performance degradation coefficient of the host machine can be determined.
[0071] The test environment is essentially an isolated simulation environment with hardware identical to the host machine. For example, the physical CPU used in the test environment is the same as that used in the host machine. After building the test environment, different rates of memory access contention events can be injected into it using testing tools. Then, the system performance degradation coefficients corresponding to these different rates of memory access contention events are collected.
[0072] It should be noted that workloads can be deployed to the test environment, and then testing tools can be used to control the workloads to generate memory contention events at different rates, thus simulating different rates of memory contention events in the test environment. The workloads can be memory-intensive applications. For example, workloads could be MySQL (a relational database management system), Nginx (a high-performance HTTP (Hypertext Transfer Protocol) server and reverse proxy server), or Redis (an in-memory key-value database).
[0073] During the process of controlling the workload to generate memory access contention events at different rates, performance metrics of the test environment can be collected using testing tools to obtain the system performance degradation coefficient corresponding to different memory access contention events. These performance metrics can include CPU utilization, memory bandwidth utilization, instruction throughput, transaction processing latency, and other indicators.
[0074] Next, the target system performance degradation coefficient is determined based on the collected system performance degradation coefficients. For example, by analyzing the collected system performance degradation coefficients, the inflection point where the performance of the test environment drops significantly can be determined, and the system performance degradation coefficient corresponding to this inflection point is the host machine's system performance degradation coefficient.
[0075] It should be understood that the host machine's system performance degradation coefficient determined based on the collected system performance degradation coefficient can essentially be understood as the maximum acceptable system performance degradation coefficient of the host machine, which can be represented as P_max.
[0076] The time required for a physical central processing unit (CPU) to generate a memory access contention event refers to the duration of such an event. Different types of CPUs require different amounts of time to generate a memory access contention event.
[0077] It should be understood that the time required for different types of physical central processing units to generate a memory access contention event can be tested in advance, and then the time corresponding to that physical central processing unit can be obtained by querying the type of physical central processing unit included in the host machine.
[0078] For example, a continuous memory access conflict scenario can be constructed, and the total number of memory access contention events occurring within a time slice duration of different types of physical CPUs can be counted. Then, based on the quotient between the time slice duration and the total number of events, the time required for different types of physical CPUs to generate one memory access contention event can be obtained.
[0079] It should be noted that different types of central processing units from different manufacturers may take different amounts of time to generate a memory access contention event.
[0080] The application performance degradation factor of a virtual machine can refer to the performance degradation factor corresponding to the workload deployed on the virtual machine. The application performance degradation factor is used to characterize the degree of relevance of the workload deployed on the virtual machine to memory access operations.
[0081] Different types of workloads may correspond to different application performance degradation coefficients. For example, the application performance degradation coefficient ranges from 0 to 1. For instance, for compute-intensive workloads (such as sysbench, a multi-threaded performance testing tool), the corresponding application performance degradation coefficient can approach 0. For memory-intensive applications such as MySQL, Nginx, and Redis, the corresponding application performance degradation coefficient can approach 1. Of course, in addition to compute-intensive workloads and memory-intensive applications, other types of workloads can also be included, and their corresponding application performance degradation coefficients can also be different. This disclosure does not exhaustively exemplify all possible scenarios. For example, for I / O (Input / Output) intensive workloads (such as Kafka, a high-throughput distributed message middleware), the corresponding application performance degradation coefficient can be between 0.5 and 0.8.
[0082] It should be understood that different types of workloads have different application performance degradation coefficients, mainly due to the different dependency patterns and sensitivities of the workloads to underlying hardware resources. Compute-intensive workloads do not rely on memory bandwidth, caching, etc., so their corresponding application performance degradation coefficients are relatively small. Memory-intensive workloads, on the other hand, rely on memory and are extremely sensitive to the latency of each memory access, so their corresponding application performance degradation coefficients are relatively large.
[0083] In some embodiments, a test environment consistent with the host machine's hardware environment can be constructed. Then, different types of workloads can be deployed in the test environment, and the workloads can be controlled to generate memory access contention events at different rates. The system performance degradation coefficient corresponding to the memory access contention events at different rates in the test environment can be collected. Then, based on the rate, the system performance degradation coefficient, and the time required for the physical central processing unit to generate one memory access contention event, the application performance degradation coefficient can be determined in combination with the performance quantification model.
[0084] For details regarding the test environment and the injection of memory contention events at different rates into the test environment, please refer to the relevant descriptions in the above implementation methods, which will not be repeated here.
[0085] After obtaining the system performance degradation coefficients corresponding to memory access contention events at different rates in the test environment, where the rate and system performance degradation coefficients are known quantities, the time required for the corresponding physical CPU to generate one memory access contention event can then be obtained. Next, through... The performance quantification model can yield K= The relationship, through K= By combining the time (T) required for the physical central processing unit to generate a memory access contention event obtained from the test, and the system performance degradation coefficient (P) corresponding to memory access contention events at different rates (N) obtained from the test, the application performance degradation coefficient corresponding to different types of workloads can be calculated.
[0086] It should be understood that the host machine will have different system performance degradation coefficients under different memory access contention events. By testing the system performance degradation coefficients of the host machine under different memory access contention events, and combining them with the time required for the host machine's physical central processing unit to generate one memory access contention event, the application performance degradation coefficients of different virtual machines can be calculated.
[0087] exist In the performance quantification model, the time consumption (T) and application performance degradation coefficient (K) have been obtained through experimental testing. Then, based on the system performance degradation coefficient set by the host machine, the alarm threshold can be calculated.
[0088] It is worth noting that the calculated alarm threshold can be preset in the alarm rules. The alarm engine collects the occurrence rate in real time, and then compares the occurrence rate with the alarm threshold. When the occurrence rate is greater than the alarm threshold, the corresponding virtual machine can be identified as the target virtual machine, and an alarm event can be generated. In response to the alarm event, actions are triggered to suppress the target virtual machine from continuing to generate memory access contention events.
[0089] Therefore, by using the performance quantification model provided by the above implementation method, combined with standardized testing methods, and by injecting controllable memory contention events at different rates, the alarm threshold for triggering and suppressing memory contention events can be accurately calculated.
[0090] In some feasible implementations, the alarm threshold can also be updated periodically.
[0091] It is worth noting that for each host machine, the calculated alarm threshold for that host machine can also be optimized. After controlling the virtual machine on the host machine to sleep to suppress memory contention events, the accuracy of controlling the virtual machine on the host machine to sleep can be checked periodically. If it is inaccurate, the alarm threshold for the host machine can be recalculated using the above method, thereby updating the alarm threshold.
[0092] Figure 2 This is a schematic diagram illustrating the principle of suppressing memory access contention events, based on some embodiments. For example... Figure 2 As shown, in the test environment, workloads and testing tools can be deployed, and memory contention events at different rates can be injected. The system performance degradation coefficient under these different rates is then collected, and an alarm threshold is determined based on this coefficient. The alarm threshold determined in the test environment is applied to the corresponding host machine. On the host machine, a first count value is collected in real time. A second count value is obtained through the first count value and its binding relationship. The occurrence rate is then calculated from the second count value, and it is determined whether the occurrence rate is greater than the alarm threshold. If the occurrence rate is greater than the alarm threshold, an alarm event is generated; otherwise, the first count value is collected again. During the phase of suppressing memory contention events and updating the alarm threshold, the virtual machine can be put to sleep in response to the alarm event. The accuracy of the alarm threshold is then determined. If the alarm threshold is inaccurate, it can be optimized in the test environment, thereby periodically updating the alarm threshold.
[0093] In some feasible implementations, in step 110, the virtual CPU may time out, control the virtual CPU to exit the virtual machine, and then, in response to the virtual CPU exiting the virtual machine, collect the first count value of the performance monitoring counter of the physical CPU core corresponding to the virtual CPU.
[0094] Here, when the interval between the host machine's system time and the start time of the virtual central processing unit entering the virtual machine for execution reaches a preset time interval, the virtual central processing unit execution timeout is determined.
[0095] Here, system time refers to the real-time time of the host system. The start time of virtual CPU execution within the virtual machine refers to the time at which the virtual CPU is loaded into the virtual machine. The preset time interval can refer to the duration of a time slice allocated to the virtual machine. For example, the preset time interval can be 1 second.
[0096] When the interval between the system time and the start time reaches a preset time interval, the virtual CPU is controlled to exit the virtual machine. It should be understood that controlling the virtual CPU to exit the virtual machine can actually be understood as a VM Exit event. The virtual CPU exiting the virtual machine means controlling the virtual CPU to exit guest mode and return the virtual CPU to the host machine kernel.
[0097] For example, when the virtual central processing unit enters the virtual machine for execution, a scheduling timer can be set with a preset time interval. In response to the scheduling timer timeout, the virtual central processing unit is controlled to exit the virtual machine.
[0098] After the virtual central processing unit exits the virtual machine, the suppression module, which is used to suppress memory access contention events, gains execution rights. At this time, the suppression module can perform the action of collecting the first count value of the performance monitoring counter of the physical central processing unit core corresponding to the virtual central processing unit.
[0099] It is worth noting that by periodically controlling the virtual CPU to exit the virtual machine, the ongoing malicious memory access conflict loop of the virtual machine can be interrupted. When the virtual machine is in a malicious memory access conflict loop, no VM Exit event is generated, causing the suppression module to be unable to obtain execution rights and thus unable to read the first count value recorded by the performance monitoring counter. Therefore, in this embodiment, at preset time intervals, the virtual CPU is forcibly controlled to exit the virtual machine, generating a VM Exit event, so that the suppression module can obtain execution rights, thereby supporting the suppression of subsequent memory access contention events.
[0100] Figure 3This is a logical diagram illustrating the suppression of memory contention events according to some embodiments. For example... Figure 3 As shown, during `vm enter`, the vCPU enters the virtual machine and begins running its workload. Upon detecting a `vm exit` event, the current timestamp `now` is determined using the function `get_time()`, which is used to obtain the system time; that is, `now = get_time()`. If the end time `end_time` corresponding to the time slice allocated to the virtual machine is less than the current timestamp `now`, then the start time `begin_time` corresponding to the next time slice is set to `now`, `end_time` is set to `now + 1s`, and `split_lock_count` is set to 0. Here, `split_lock_count = 0` resets the performance monitoring counters.
[0101] It's important to note that if `end_time < now`, it means the previous time slice has expired, and a new time window needs to be set. This involves setting the start time `begin_time` for the next time slice to `now`, the end time `end_time` to `now + 1s`, and `split_lock_count` to 0. Under normal circumstances, VM exit events occur frequently per second. By checking if `end_time` is less than the current timestamp, we can determine if the VMEXIT event triggering is within the time range corresponding to the time slice. When `end_time < now`, it means a memory contention event suppression has already occurred, so the performance monitoring counter needs to be reset to monitor the first count value of memory contention events occurring in the next second. If `end_time ≥ now`, it means the memory contention event is still within the valid time frame and no memory contention event suppression has occurred, so the first count value of the performance monitoring counter needs to be read (`split_lock_count += read_pmc()`).
[0102] Then, the occurrence rate is calculated using the first count value read, and it is determined whether the occurrence rate is less than the alarm threshold, i.e., whether split_lock_count < N, where split_lock_count is the occurrence rate and N is the alarm threshold. If split_lock_count < N, then end_time = begin_time + 1s, and the virtual machine is put to sleep for a preset sleep duration, which is the difference between the end time and the current time, i.e., sleep(end_time - now). If split_lock_count ≥ N, then at a preset time interval, the virtual CPU is controlled to exit the virtual machine, triggering the vm exit event. It should be understood that by controlling the virtual CPU to exit the virtual machine at preset time intervals, the virtual CPU can be prevented from not exiting voluntarily.
[0103] Therefore, through the above implementation method, when the interval between the host machine's system time and the start time of the virtual central processing unit entering the virtual machine reaches a preset time interval, controlling the virtual central processing unit to exit the virtual machine can interrupt the malicious memory access conflict loop that the virtual machine is currently in progress, so that the memory access contention event can be suppressed.
[0104] In some feasible implementations, the second count value is used to generate a visualization view, which includes the second count value corresponding to each virtual machine at each point in time, and the visualization view is used to display it on the client.
[0105] Here, each collected second count value can be stored in the database. A visualization view can be generated using the historical data stored in the database. This visualization view includes the second count value for each virtual machine at each point in time. The generated visualization view can then be displayed through the client's graphical user interface. Following the above implementation, if the database stores a sequence of associated data: "virtual machine identifier - timestamp - second count value," a visualization view can be generated based on this associated data series.
[0106] Therefore, by generating a visualization view using the second count value, the second count value corresponding to the memory access contention events generated by the virtual machine at various points in time can be displayed to the user, so that the user can intuitively understand the second count value of the memory access contention events generated by each virtual machine and the long-term trend of change.
[0107] The following is in conjunction with the appendix Figure 4 The above-described embodiments will be described in detail.
[0108] Figure 4 This is a schematic diagram illustrating the architecture of a cloud service-based memory access conflict detection method, based on some embodiments. For example... Figure 4 As shown, a workload deployed in a virtual machine on the host machine triggers an unaligned memory access atomic operation. This operation is executed by a vCPU thread, which runs in the physical central processing unit located at the hardware layer. The performance monitoring unit at the hardware layer records the first count value of the memory access contention event. It is important to note that the performance monitoring unit actually records the first count value through a performance monitoring counter.
[0109] When a detection agent located in the host machine's user space triggers a memory contention event, the agent invokes the Hardware Event Abstraction Interface (HITE) located in the kernel space to read a first count value. Furthermore, the agent obtains the binding relationships from the hypervisor and, based on the first count value and the binding relationships, obtains a second count value for each virtual machine. Then, the second count value is stored in a database located at the application layer. It should be noted that the hypervisor is an application used to manage virtual machines; following the above implementation, the hypervisor can be KVM.
[0110] The data stored in the database can be used to generate and display visualizations, showing users the secondary counts and long-term trends of memory contention events occurring in various virtual machines. These secondary counts can also be used to help users suppress memory contention events.
[0111] Figure 5 This is a schematic diagram of the module connections of a cloud service-based memory access conflict detection device according to an exemplary embodiment. Figure 5 As shown, this disclosure provides a cloud service-based memory access conflict detection device 500, which may include: The acquisition module 501 is used to acquire the first count value of the performance monitoring counter of each physical central processing unit core on the host machine, wherein the performance monitoring counter records the number of times the physical central processing unit core has a memory access contention event. The determining module 502 is used to determine a second count value for each virtual machine based on the binding relationship between the virtual central processing unit and the physical central processing unit core of the virtual machine and the first count value of the performance monitoring counter of each physical central processing unit core. The second count value represents the number of times the virtual machine has encountered a memory access contention event.
[0112] Optionally, the cloud service-based memory access conflict detection device 500 may further include: The rate determination unit is used to determine the first rate corresponding to each virtual machine based on the second count value of each virtual machine and the actual runtime of each virtual machine. The first rate represents the frequency of memory access contention events occurring in the virtual machine. The execution unit is used to filter target virtual machines whose first rate is greater than the alarm threshold and to perform actions to suppress the target virtual machines from continuing to generate memory access contention events.
[0113] Optionally, the execution unit is specifically used for: Control the target virtual machine to enter a sleep state.
[0114] Optionally, the execution unit is further configured to: Set the sleep duration of the target virtual machine, and control the sleep state of the target virtual machine according to the sleep duration. The sleep duration is the difference between the preset time slice duration and the actual runtime of the virtual machine.
[0115] Optionally, the cloud service-based memory access conflict detection device 500 may further include: The threshold determination module is used to generate the alarm threshold based on the system performance degradation coefficient of the host machine, the time required for the physical central processing unit of the host machine to generate a memory access contention event, and the application performance degradation coefficient of the virtual machine.
[0116] Optionally, the cloud service-based memory access conflict detection device 500 may further include: An update module is used to periodically update the alarm threshold.
[0117] Optionally, the cloud service-based memory access conflict detection device 500 may further include: The recording module is used to record a data sequence including the acquisition timestamp, the physical central processing unit core number, and the first count value of the physical central processing unit core; The determining module 502 is used to: determine the second count value corresponding to each virtual machine based on the binding relationship and the associated data sequence.
[0118] Optionally, the cloud service-based memory access conflict detection device 500 may further include: The storage module is used to store the second count value corresponding to each virtual machine in the database.
[0119] Optionally, the acquisition module 501 is specifically used for: In response to a timeout in the virtual central processing unit (CPU), control the CPU to exit the virtual machine. In response to the virtual central processing unit exiting to the virtual machine, the first count value of the performance monitoring counter of the physical central processing unit core corresponding to the virtual central processing unit is collected.
[0120] Optionally, the second count value is used to generate a visualization view, which includes the second count value corresponding to each virtual machine at each point in time, and the visualization view is used to display it in the client.
[0121] Regarding the cloud service-based memory access conflict detection device 500 in the above embodiments, the method logic executed by each functional module has been described in detail in the section on methods, and will not be repeated here.
[0122] The following is for reference. Figure 6 It shows a schematic diagram of the structure of an electronic device (e.g., a server) 600 suitable for implementing embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0123] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0124] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0125] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0126] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0127] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0128] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0129] The aforementioned computer-readable medium carries one or more programs. When the aforementioned one or more programs are executed by the electronic device, the electronic device causes the following: to collect a first count value from the performance monitoring counters of each physical central processing unit (CPU) core on the host machine, wherein the performance monitoring counters record the number of memory access contention events that occur in the physical CPU cores; and to determine a second count value corresponding to each virtual machine based on the binding relationship between the virtual CPU and the physical CPU cores and the first count value from the performance monitoring counters of each physical CPU core, wherein the second count value represents the number of memory access contention events that occur in the virtual machine.
[0130] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0132] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0133] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0136] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0137] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A cloud service-based memory access conflict detection method, characterized in that, The method comprises: collecting first count values of performance monitoring counters of each physical central processing unit core on a host machine, wherein the performance monitoring counters record the number of memory access contention events occurring in the physical central processing unit core; determining second count values corresponding to each virtual machine according to the binding relationship between the virtual central processing unit of the virtual machine and the physical central processing unit core and the first count values of the performance monitoring counters of each physical central processing unit core, wherein the second count values represent the number of memory access contention events occurring in the virtual machine.
2. The method of claim 1, wherein, The method further comprises: determining occurrence rates corresponding to each virtual machine according to the second count values of each virtual machine and the actual running duration of each virtual machine, wherein the occurrence rates represent the frequency of memory access contention events occurring in the virtual machine; screening target virtual machines with occurrence rates greater than an alarm threshold, and performing an action for suppressing the target virtual machines from continuously generating memory access contention events.
3. The method of claim 2, wherein, The action for suppressing the target virtual machines from continuously generating memory access contention events comprises: controlling the target virtual machines to enter a sleep state.
4. The method of claim 3, wherein, The method further comprises: setting a sleep duration of the target virtual machines, and controlling the sleep state of the target virtual machines according to the sleep duration, wherein the sleep duration is the difference between a preset time slicing duration and the actual running duration of the target virtual machines.
5. The method of claim 2, wherein, The method further comprises: generating the alarm threshold according to a system performance attenuation coefficient of the host machine, a time consumption required by the physical central processing unit of the host machine for generating a memory access contention event, and an application performance attenuation coefficient of the virtual machine.
6. The method of claim 5, wherein, The method further comprises: periodically updating the alarm threshold.
7. The method of claim 1, wherein, The method further comprises: recording a sequence of associated data comprising a collection timestamp, a number of the physical central processing unit core, and the first count value of the physical central processing unit core; determining the second count values corresponding to each virtual machine according to the binding relationship and the sequence of associated data. The method further comprises:
8. The method according to any one of claims 1-7, characterized in that, storing the second count values corresponding to each virtual machine in a database. The collection of the first count values of the performance monitoring counters of each physical central processing unit core on the host machine comprises:
9. The method according to any one of claims 1-7, characterized in that, in response to the virtual central processing unit executing a timeout, controlling the virtual central processing unit to exit the virtual machine; in response to the virtual central processing unit exiting the virtual machine, collecting the first count value of the performance monitoring counter of the physical central processing unit core corresponding to the virtual central processing unit. The second count values are used to generate a visual view, wherein the visual view comprises the second count values corresponding to each virtual machine at each time point, and the visual view is used to be displayed in a client.
10. The method according to any one of claims 1-7, characterized in that, The method comprises: 11.A cloud service based memory access conflict detection apparatus, characterized in that, a collection module configured to collect first count values of performance monitoring counters of each physical central processing unit core on a host machine, wherein the performance monitoring counters record the number of memory access contention events occurring in the physical central processing unit core; The determining module is configured to determine a second count value corresponding to each virtual machine according to a binding relationship between a virtual central processing unit of the virtual machine and a physical central processing unit core and a first count value of a performance monitoring counter of each physical central processing unit core, wherein the second count value represents a number of times of memory access contention events of the virtual machine.
12. A computer readable medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processing device, implements the steps of the method of any one of claims 1-10.
13. An electronic device, comprising: comprising: a storage device having stored thereon a computer program; a processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of claims 1-10.
14. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processing device, implements the steps of the method of any one of claims 1-10. the computer program, when executed by the processing device, implements the steps of the method of any one of claims 1-10.