A method and corresponding device for handling page fault exceptions

By saving the coroutine context of page-failed exceptions to the shared memory in the computer system, and obtaining the context from the shared memory to trigger the page swap process when switching between kernel state and user state, the thread blocking problem caused by page-failed exception handling in the prior art is solved, and lower latency and higher business throughput are achieved.

CN115599510BActive Publication Date: 2025-06-13HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110774711.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-08
Publication Date
2025-06-13
Estimated Expiration
2041-07-08

AI Technical Summary

Technical Problem

The prior art causes thread blockage when handling page-missing exceptions, resulting in decreased service throughput and long-tail delay.

Method used

Avoid thread blocking by saving the context of the coroutine that triggers the page-failed exception to be triggered to share memory and when switching between kernel state and user state, obtaining the context from shared memory to trigger the page swap process.

Benefits of technology

It shortens the delay of page missing exception handling, reduces the long-tail delay of threads, and improves business throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599510B_ABST
    Figure CN115599510B_ABST
Patent Text Reader

Abstract

The present application discloses a method for handling page fault exceptions, which is applied to a computer system. The method includes: saving the context of a first coroutine that triggers a page fault exception to shared memory, where the first coroutine belongs to a first thread, and the shared memory is memory that can be accessed by the first thread in both kernel mode and user mode; switching from the context of the first coroutine to the context of the first thread, where the context of the first thread is configured to the shared memory during the initialization of the first thread; switching from kernel mode to user mode; and triggering a page-in process by running the first thread to obtain the context of the first coroutine from the shared memory. The solution provided by the present application can reduce the processing latency of page fault exceptions, thereby reducing the IO latency of the first thread and improving the service throughput.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method and corresponding apparatus for handling page faults. Background Art

[0002] The scheduling of mainstream storage mainly uses lightweight threads (LWTs), which are also called coroutines. When the memory accessed by an LWT is swapped out, a page fault (PF) will be triggered, and the page needs to be swapped in. Since each processor is bound to only one thread and multiple LWT tasks are executed on each thread, when a page fault is triggered by a certain LWT task, the page fault event needs to be notified to the user space through the page fault handling mechanism, and the LWT task cannot continue to execute until the page is swapped in by the user space.

[0003] The current way of handling page faults will cause the LWT task with the page fault to block the entire thread, resulting in a decrease in business throughput and a long tail latency of the thread. Summary of the Invention

[0004] An embodiment of this application provides a method for handling page faults, which is used to reduce the latency of page fault handling and improve business throughput. Embodiments of this application also provide corresponding apparatuses, computer devices, computer-readable storage media, computer program products, etc.

[0005] A first aspect of this application provides a method for handling page faults. The method is applied to a computer system, which may be a server, a terminal device, a virtual machine (VM), a container, or the like. The method includes: saving the context of a first coroutine that triggers a page fault into shared memory, where the first coroutine belongs to a first thread, and the shared memory is memory that can be accessed by the first thread in both the kernel space and the user space; switching from the context of the first coroutine to the context of the first thread, where the context of the first thread is configured into the shared memory when the first thread is initialized; switching from the kernel space to the user space; and triggering a page swap-in process by running the first thread to obtain the context of the first coroutine from the shared memory.

[0006] In this application, a page fault (PF) can also be referred to as a page miss, which usually occurs in the kernel mode of an operating system (OS). After a page fault occurs, it is necessary to handle the page fault, and this handling process involves the kernel mode and the user mode. The kernel mode and the user mode are two modes or two states of the OS. The kernel mode is usually also called the privileged state, and the user mode is usually also called the non-privileged state. A thread is the smallest unit of OS scheduling (processor scheduling). A coroutine is a lightweight thread. A thread can include multiple coroutines, and each coroutine can correspond to a task. Sometimes, a coroutine is also referred to as a coroutine task.

[0007] In this application, the "first" in the first thread has no substantial meaning. It is just a thread that has a page fault during runtime. This first thread can also be referred to as a business thread or an application thread.

[0008] In this application, the context of the first coroutine includes the data in the registers of the processor when the first coroutine is running. The context of the first thread includes the data read from the shared memory and written into the registers. Switching from the context of the first coroutine to the context of the first thread means writing the context of the first thread into the registers of the processor. The above registers can include any one or more of general-purpose registers, program counter (PC), program state register (PS), etc.

[0009] In this first aspect, in the kernel mode, the context of the first coroutine is saved to the shared memory. After returning from the kernel mode to the user mode, by running the first thread, the context of the first coroutine can be obtained from the shared memory, and then the page-in process can be executed based on the context of the first coroutine. Compared with the prior art, when a certain coroutine of a thread triggers a page fault, it is necessary to notify the monitor thread in the kernel mode, and then the thread enters the sleep state until the monitor thread completes the page-in through the swap-in thread and then sends a notification message to the kernel mode to wake up the thread and then continue to execute the coroutine. The page fault handling process of this application can shorten the latency of page fault handling, thereby reducing the long-tail latency of the first thread. Shortening the latency correspondingly also improves the business throughput.

[0010] In this application, long-tail latency refers to the following: in a computer system, during the process of running a thread, there will always be a small number of latencies of the responses corresponding to the operations of this thread that are higher than the average latency of the computer system. The latencies of these small number of responses are called long-tail latencies. For example, there are 100 responses in a computer system, and the average latency of these 100 responses is 10 microseconds. Among them, the latency of one response is 50 milliseconds. Then, the latency of this response is the long-tail latency. In addition, there is a commonly used P99 standard for latency in business. The definition of long-tail latency in this P99 standard is that the latency of 99% of the responses in the computer system should be controlled within a certain time-consuming, and only 1% of the responses are allowed to have a latency exceeding this certain time-consuming. The latency of the responses exceeding this certain time-consuming is called long-tail latency.

[0011] In this application, the long-tail latency of a thread can be understood as the long-tail latency when the thread performs input / output (IO) operations. If there is no page fault exception during the running process of the thread, it may take 10 microseconds to complete an IO operation. If a page fault exception occurs, according to the existing technology solutions, it takes hundreds of microseconds to handle the page fault exception, which causes the long-tail latency of this thread to execute this IO. If the page fault exception is handled according to the solution provided in this application, it usually only takes a few microseconds to handle the page fault exception. In this way, the long-tail latency of this thread is greatly reduced.

[0012] In a possible implementation manner of the first aspect, the method further includes: when executing the page-in process, running a second coroutine belonging to the first thread to execute the task corresponding to the second coroutine.

[0013] It should be understood that running the second coroutine when executing the page-in process can be understood as running the second coroutine during the process of executing the page-in process, that is, there is a time overlap between the execution of the page-in process and the running of the second coroutine. However, the start time point when the second coroutine starts running is not limited. The second coroutine can start at the same time as the page-in process, or can start after the page-in process starts.

[0014] In this possible implementation manner, when executing the page-in process, the second coroutine can also be run asynchronously, which can further improve the business throughput.

[0015] In a possible implementation manner of the first aspect, the above step: switching from the context of the first coroutine to the context of the first thread includes: writing the context of the first thread into the register of the computer system through a hook function to replace the context of the first coroutine in the register.

[0016] In this possible implementation, the operating system can perform context switching through a hook function, writing the context of the first thread into the registers of the computer system, thereby overwriting the context of the first coroutine originally stored in the registers.

[0017] In a possible implementation of the first aspect, the above steps: obtaining the context of the first coroutine from the shared memory by running the first thread to trigger the page-in process, include: obtaining the context of the first coroutine from the shared memory by running the first thread, and obtaining the destination address from the context of the first coroutine, where the destination address is the address of the physical page to be accessed when the first coroutine triggers a page fault exception; according to the destination address, execute the page-in process of the corresponding physical page.

[0018] In this possible implementation, the context of the first coroutine contains the address of the physical page to be accessed when the first coroutine triggers a page fault exception, that is, the destination address. In this way, the computer system can swap in the physical page corresponding to the destination address from the disk. This way of directly swapping in the physical page through the destination address can improve the swapping-in speed of the physical page, thereby further reducing the latency of page fault exception handling.

[0019] In a possible implementation of the first aspect, the method further includes: when the physical page is swapped into the memory, adding the first coroutine to the coroutine waiting queue, and the coroutines in the coroutine waiting queue are in a pending scheduling state.

[0020] In this possible implementation, after the physical page is swapped in, the first coroutine can be executed again. The execution order can be to put the first coroutine into the coroutine waiting queue to wait for scheduling. One or more coroutines are placed in the coroutine waiting queue in sequence, and the computer system will schedule and execute the coroutines in the coroutine waiting queue in sequence.

[0021] In a possible implementation of the first aspect, the shared memory is configured for the first thread when the first thread is initialized.

[0022] In a possible implementation of the first aspect, the page fault exception is triggered when the first coroutine is run to access the physical page swapped out in the memory.

[0023] In a possible implementation of the first aspect, the shared memory is configured through an extended Berkeley packet filter (ebpf) of the kernel virtual machine. Of course, this application is not limited to being configured through ebpf, and the shared memory can also be configured through other means.

[0024] In this application, eBPF is a new design introduced in kernel 3.15, which develops the original BPF into a "kernel virtual machine" with a more complex instruction set and a wider application range.

[0025] The second aspect of this application provides a page fault exception handling device, which has the function of implementing the method of the first aspect or any possible implementation manner of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. For example: a first processing unit, a second processing unit, a third processing unit, and a fourth processing unit. These four processing units can be implemented by one processing unit or multiple processing units.

[0026] The third aspect of this application provides a computer device, which includes at least one processor, a memory, an input / output (I / O) interface, and computer execution instructions stored in the memory and executable on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation manner of the first aspect.

[0027] The fourth aspect of this application provides a computer-readable storage medium storing one or more computer execution instructions. When the computer execution instructions are executed by the processor, one or more processors execute the method of the first aspect or any possible implementation manner of the first aspect.

[0028] The fifth aspect of this application provides a computer program product storing one or more computer execution instructions. When the computer execution instructions are executed by one or more processors, one or more processors execute the method of the first aspect or any possible implementation manner of the first aspect.

[0029] The sixth aspect of this application provides a chip system, which includes at least one processor. The at least one processor is used to support the page fault exception handling device to implement the functions involved in the first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory for storing the necessary program instructions and data of the page fault exception handling device. The chip system may be composed of chips or may include chips and other discrete devices.

[0030] In the embodiment of the present application, after a page fault is triggered by the first coroutine, the context of the first coroutine is saved in the shared memory in the kernel mode. After the OS returns from the kernel mode to the user mode, the first thread can obtain the context of the first coroutine from the shared memory by running, and then execute the page-in process according to the context of the first coroutine. Compared with the prior art, in which it is necessary to notify the monitor thread in the kernel mode, then the first thread enters the sleep state until the monitor thread completes the page-in through the swap-in thread and then sends a notification message to the kernel to wake up the first thread and then continue to execute the page fault handling process of the first thread, the latency of page fault handling can be shortened, and the corresponding reduction in latency also improves the service throughput. Description of the Drawings

[0031] Figure 1 is a schematic diagram of an embodiment of a computer system provided by an embodiment of the present application;

[0032] Figure 2 is a schematic diagram of a page fault handling architecture provided by an embodiment of the present application;

[0033] Figure 3 is a schematic diagram of an embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0034] Figure 4 is a schematic diagram of another embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0035] Figure 5 is a schematic diagram of another embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0036] Figure 6 is a schematic diagram of another embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0037] Figure 7 is a schematic diagram of another embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0038] Figure 8 is a schematic diagram of another embodiment of a method for handling a page fault provided by an embodiment of the present application;

[0039] Figure 9 is a schematic diagram of an embodiment of a device for handling a page fault provided by an embodiment of the present application;

[0040] Figure 10 is a schematic diagram of a structure of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0041] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Those of ordinary skill in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0042] The terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0043] The embodiments of the present application provide a method for handling page fault exceptions, which is used to reduce the latency of page fault exception handling and improve service throughput. The embodiments of the present application also provide corresponding devices, computer devices, computer-readable storage media, computer program products, etc. These will be described in detail below.

[0044] The method for handling page fault exceptions provided by the embodiments of the present application is applied to a computer system, which can be a server, a terminal device, or a virtual machine (VM).

[0045] A terminal device (which can also be referred to as a user equipment (UE)) is a device with wireless transceiver functions. It can be deployed on land, including indoor or outdoor, handheld or vehicle-mounted; it can also be deployed on water (such as ships, etc.); it can also be deployed in the air (such as airplanes, balloons, satellites, etc.). The terminal can be a mobile phone, a tablet (pad), a computer with wireless transceiver functions, a virtual reality (VR) terminal, an augmented reality (AR) terminal, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc.

[0046] The architecture of the computer system can be referred to Figure 1 for understanding. Figure 1 It is a schematic diagram of the architecture of a computer system.

[0047] As Figure 1 shown, the architecture of the computer system includes a user layer 10, a kernel 20, and a hardware layer 30.

[0048] The user layer 10 includes multiple applications, and each application corresponds to a thread. A thread is the smallest unit scheduled by the operating system (OS). A thread can include multiple coroutines, and a coroutine is a lightweight thread. Each coroutine can correspond to a task, and sometimes a coroutine is also referred to as a coroutine task. In this application, a thread can also be referred to as a business thread or an application thread, etc.

[0049] The kernel 20 is responsible for managing key resources by the OS and provides an OS call entry for threads in the user state to provide services in the kernel, such as: page fault (PF) handling, page table management, and interrupt control services. In addition, the kernel 20 also processes page faults (PF) that occur in the OS. A page fault (PF) can also be referred to as a page fault, which usually occurs in the kernel state of the operating system. After a page fault occurs, it is necessary to handle the page fault, and this handling process involves the kernel state and the user state. The kernel state and the user state are two modes or two states of the OS. The kernel state is usually also referred to as the privileged state, and the user state is usually also referred to as the non-privileged state.

[0050] The hardware layer 30 includes the hardware resources on which the kernel 20 depends for operation, such as: processors, memory (the memory includes shared memory configured for threads, and in this application, shared memory refers to memory that can be accessed by threads in both the kernel mode and the user mode), a memory management unit (MMU), and input / output (I / O) devices and disks, etc. The processor may include a register set, and the register set may include various types of registers, such as: stack frame registers, general-purpose registers, and non-volatile (callee-saved) registers, etc. The registers are used to store the context of a thread or the context of the coroutine of the thread.

[0051] When a page fault exception occurs during a thread's memory access, the corresponding physical page can be swapped in from the disk through the page fault exception handling mechanism, thereby solving the page fault exception problem.

[0052] The MMU is a computer hardware that is responsible for processing the memory access requests of the central processing unit (CPU). Its functions include virtual address to physical address conversion, memory protection, control of the CPU cache, etc.

[0053] In a computer system, usually one application is bound to one thread, and one thread includes multiple lightweight threads (LWTs), which are also called coroutines. This one thread will execute the tasks corresponding to the multiple coroutines it includes, and the tasks corresponding to the coroutines can also be called coroutine tasks.

[0054] The thread will execute the coroutine tasks one by one. When executing any coroutine task, if a page fault exception occurs, the current page fault exception handling scheme will notify the monitor thread in the kernel mode, and then the thread will enter the sleep state until the monitor thread completes the page swap-in through the swap-in thread and then sends a notification message to the kernel mode to wake up the thread and continue to execute the coroutine task. In the current page fault exception handling scheme, after a page fault exception occurs, the coroutine task that triggers the page fault exception will block the entire thread, thereby resulting in a decrease in business throughput and causing a long tail latency of the thread.

[0055] In this application, long-tail latency refers to the following situation in a computer system: during the process of running a thread, there will always be a small number of response latencies of the operations corresponding to this thread that are higher than the average latency of the computer system. These small numbers of response latencies are called long-tail latencies. For example, there are 100 responses in a computer system, and the average latency of these 100 responses is 10 microseconds. Among them, the latency of one response is 50 milliseconds. Then, the latency of this response is the long-tail latency. In addition, there is a commonly used P99 standard for latency in the service. The definition of long-tail latency in this P99 standard is that the latency of 99% of the responses in the computer system should be controlled within a certain time-consuming, and only 1% of the responses are allowed to have a latency exceeding this certain time-consuming. The latency of the responses exceeding this certain time-consuming is called long-tail latency.

[0056] In this application, the long-tail latency of a thread can be understood as the long-tail latency when the thread performs input / output (IO) operations. If there is no page fault exception during the running process of the thread, it may take 10 microseconds to complete an IO operation. If a page fault exception occurs, according to the existing technology solution, it takes several hundred microseconds to handle the page fault exception, which causes the long-tail latency of this thread to execute this IO. If the page fault exception is handled according to the solution provided in this application, it usually only takes a few microseconds to handle the page fault exception. In this way, the long-tail latency of this thread is greatly reduced.

[0057] To speed up the processing speed of page fault exceptions, an embodiment of this application provides a page fault exception handling architecture as Figure 2 shown. As Figure 2 shown, this page fault exception handling architecture includes:

[0058] Multiple threads, such as thread 1 to thread N. Each thread can include multiple coroutines. For example, thread 1 includes coroutine 1, coroutine 2 to coroutine M, and thread N includes coroutine 1, coroutine 2 to coroutine P. Among them, N, M, and P are all positive integers, and they can be equal or not equal. Among them, each coroutine corresponds to a task. A page fault exception response task can be configured in each thread. This page fault exception response task is used to implement the saving of the context of the coroutine that generates the page fault exception and the scheduling of the page swapping-in thread.

[0059] The kernel-mode memory page fault handling mechanism is used to trigger the kernel page fault notification mechanism when a page fault exception occurs in the kernel mode.

[0060] The kernel page fault notification mechanism is used to quickly switch to the context of the user-mode thread where the coroutine that generates the page fault is located during the kernel-mode page fault exception handling process.

[0061] The page swapping-in thread is used to swap in the corresponding physical page from the disk to the memory based on the page fault exception response task.

[0062] Based on the above Figure 1 computer system and Figure 2 the page fault exception handling architecture shown in

[0063] As Figure 3 shown, an embodiment of the page fault exception handling method provided by an embodiment of the present application includes:

[0064] 101. The computer system saves the context of the first coroutine that triggers the page fault exception into the shared memory. The first coroutine belongs to the first thread, and the shared memory is the memory that can be accessed by the first thread in both the kernel mode and the user mode.

[0065] The context of the first coroutine refers to the data in the registers of the processor when the first coroutine is running.

[0066] The relationship between the first coroutine and the first thread can be understood by referring to Figure 2 the relationship between thread 1 and coroutines 1, 2... M in

[0067] Each thread can have a dedicated shared memory.

[0068] Optionally, the shared memory can be configured for the first thread when the first thread is initialized.

[0069] Optionally, the page fault exception can be triggered when running the first coroutine to access the physical page swapped out in the memory.

[0070] 102. The computer system switches from the context of the first coroutine to the context of the first thread. The context of the first thread is configured into the shared memory when the first thread is initialized.

[0071] The context of the first thread includes the data read from the shared memory and then written into the registers. Switching from the context of the first coroutine to the context of the first thread means writing the context of the first thread into the registers of the processor. The above registers can include any one or more of general-purpose registers, program counter (PC), program state register (PS), etc.

[0072] 103. The computer system switches from the kernel mode to the user mode.

[0073] 104. The computer system triggers the page-in process by running the first thread to obtain the context of the first coroutine from the shared memory.

[0074] Since the first coroutine triggered a page fault, after obtaining the context of the first coroutine from the shared memory by running the first thread in user mode, the subsequent page-in process can be executed according to the context of the first coroutine.

[0075] In the solution provided by the embodiments of the present application, the context of the first coroutine is saved to the shared memory in kernel mode. After returning from kernel mode to user mode, the context of the first coroutine can be obtained from the shared memory by running the first thread, and then the page-in process is executed according to the context of the first coroutine. Compared with the prior art, when a certain coroutine triggers a page fault, it is necessary to notify the monitor thread in kernel mode, and then the first thread enters the sleep state until the monitor thread completes the page-in through the swap-in thread and then sends a notification message to kernel mode to wake up the first thread and then continue to execute the coroutine. The page fault handling process of the present application can shorten the latency of page fault handling, thereby reducing the input / output (IO) long-tail latency of the first thread, and shortening the latency correspondingly improves the service throughput.

[0076] Optionally, the method for handling page faults provided by the embodiments of the present application may further include: when executing the page-in process, running a second coroutine belonging to the first thread to execute the task corresponding to the second coroutine.

[0077] It should be understood that running the second coroutine when executing the page-in process can be understood as running the second coroutine during the execution of the page-in process, that is, there is a time overlap between the execution of the page-in process and the running of the second coroutine, but the start time point when the second coroutine starts running is not limited. The second coroutine can start at the same time as the page-in process or start after the page-in process starts.

[0078] The solution provided by the embodiments of the present application to asynchronously run the second coroutine when executing the page-in process can further improve the service throughput.

[0079] Generally speaking, the method for handling page faults provided by the embodiments of the present application can be understood by referring to Figure 4 for understanding.

[0080] As Figure 4 shown, after the first thread is initialized, it will execute the task corresponding to the first coroutine. The first coroutine triggers a page fault during operation, and then it will execute the page fault handling process, and will also execute the task of the second coroutine when handling the page fault.

[0081] Figure 4The described content can include three stages, namely: 1. Initialization stage; 2. Kernel mode handling of page faults; 3. User mode handling of page faults. The following will be introduced in combination with the accompanying drawings respectively.

[0082] 1. Initialization stage.

[0083] As Figure 5 shown, run the main function of the first thread and execute the following steps:

[0084] 201. Initialize the shared memory, that is, allocate shared memory for this first thread.

[0085] 202. Obtain the context of this first thread during initialization through the getcontext function.

[0086] 203. Set the context of the first thread during initialization into the shared memory.

[0087] 2. Kernel mode handling of page faults.

[0088] As Figure 6 shown, this process includes the following steps:

[0089] 301. In kernel mode, running the first coroutine to access a non-existent physical page in memory triggers a page fault.

[0090] 302. Through the hook function, save the context of the first coroutine into the shared memory.

[0091] 303. Perform context switching through the hook function, and write the context of the first thread in the shared memory into the registers of the computer system.

[0092] That is: write the context of the first thread into the registers of the computer system through the hook function to replace the context of the first coroutine in the registers.

[0093] 304. Return from kernel mode to user mode.

[0094] 3. User mode handling of page faults.

[0095] As Figure 7 shown, this process includes the following steps:

[0096] From the above Figure 6 process description, it can be seen that after the kernel mode page fault handling ends, it will return to user mode, so as to perform page fault handling in user mode.

[0097] 401. In user mode, obtain the context of the first coroutine saved in it from the shared memory by running the first thread.

[0098] 402. Save the context of the first coroutine on the first thread.

[0099] 403. Trigger a page-in process according to the context of the first coroutine.

[0100] This process can be: obtain the destination address from the context of the first coroutine, where the destination address is the address of the physical page to be accessed when the first coroutine triggers a page fault exception; execute the page-in process for the corresponding physical page according to the destination address.

[0101] That is to say: the context of the first coroutine contains the address of the physical page to be accessed when the first coroutine triggers a page fault exception, which is the destination address. In this way, the computer system can swap in the physical page corresponding to the destination address from the disk. This way of directly swapping in the physical page through the destination address can improve the swapping-in speed of the physical page, thereby further reducing the latency of page fault exception handling.

[0102] In addition, when executing the page-in process, the second coroutine of the first thread can also be scheduled and the task corresponding to the second coroutine can be executed, which can further improve the business throughput.

[0103] After the page-in process ends, that is, when the physical page is swapped into the memory, add the first coroutine to the coroutine waiting queue, and the coroutines in the coroutine waiting queue are in a state of waiting to be scheduled.

[0104] That is to say, after the physical page is swapped in, the first coroutine can be executed again. The execution order can be to put the first coroutine into the coroutine waiting queue to wait for scheduling. One or more coroutines are placed in the coroutine waiting queue in order, and the computer system will schedule and execute the coroutines in the coroutine waiting queue in sequence.

[0105] The page fault exception handling process provided by the embodiments of the present application can be implemented through the extended berkeley packet filter (ebpf) mechanism of the kernel virtual machine. The shared memory created through the ebpf mechanism can be called an ebpf map.

[0106] In the present application, ebpf is a new design introduced in kernel 3.15, which develops the original BPF into a "kernel virtual machine" with a more complex instruction set and a wider application range.

[0107] When implemented through the ebpf mechanism, the page fault exception handling process can refer to Figure 8 for understanding.

[0108] Such as Figure 8 shown, this process includes the following steps:

[0109] 501. Inject the eBPF execution function into the first thread and create an eBPF map.

[0110] This eBPF map includes a map for storing the context of the first thread and a map for storing the context of the coroutine that triggers the page fault exception.

[0111] 502. Obtain the context of the first thread.

[0112] 503. Save the context of the first thread to the map for storing the context of the first thread.

[0113] 504. The first thread triggers a page fault exception during the execution in the kernel mode.

[0114] 505. In the kernel's Page fault handling process, the eBPF execution function injected in step 501 will be executed, the context of the coroutine that triggers the page fault exception will be saved to the map for storing the context of the coroutine that triggers the page fault exception, and the context in the eBPF execution function will be modified to the context of the first thread saved in step 503.

[0115] 506. After the kernel mode page fault exception handling is completed, return to the user mode, and the program jumps to the page fault exception handling function for execution.

[0116] 507. The user mode receives the kernel page fault exception notification, obtains the context of the coroutine that triggers the page fault exception from the map for storing the context of the coroutine that triggers the page fault exception, executes the page-in process, and schedules other coroutines for execution.

[0117] After the page-in process is completed, the coroutine that triggers the page fault exception is re-queued and waits for scheduling.

[0118] The page fault exception handling method provided by the embodiments of this application is particularly effective for multiple scenarios where page fault exceptions occur concurrently. Even if hundreds of cores trigger page fault exceptions simultaneously, the page fault exception handling can be completed within a few microseconds (us). Compared with the current scenario where it takes hundreds of microseconds to complete the page fault exception handling process for multiple cores with concurrent page fault exceptions, the processing speed of the solution of this application has been greatly improved, significantly reducing the latency and increasing the throughput, thereby also improving the performance of the computer system.

[0119] To facilitate the description of the effect of this application, taking the scenario where 144 cores concurrently generate page fault exceptions as an example, the following introduces the latency in page fault exception handling and thread blocking using the existing page fault exception handling mechanism and the page fault exception handling mechanism provided by this application through Table 1.

[0120] Table 1: Latency comparison

[0121]

[0122] From the comparison between the second and third columns of Table 1, it can be seen that for the solution provided by this application, the latency in page fault exception handling and thread blocking is much shorter than that of the prior art. In a large-scale high-concurrency environment, in the existing Userfaultfd, the latency for notifying the user space of a page fault exception has exceeded 600 microseconds, which is simply unacceptable for the business. Through analysis, it can be known that in a high-concurrency scenario, the competition for file handles during the latency of notifying the user space of a page fault exception is extremely fierce, and as the number of cores increases, the competition becomes even more intense. The synchronous swap-in feature of Userfaultfd also makes its basic latency not less than 210+ us (i.e., the latency for swapping in physical pages by the SSD medium). By using this application, it is still possible to achieve a notification latency in the microsecond level in a scenario where hundreds of cores concurrently generate page fault exceptions. As the number of host cores increases, the benefits become more obvious.

[0123] The above introduced the method for handling page fault exceptions. Next, the page fault exception handling device provided by the embodiments of this application will be introduced with reference to the accompanying drawings.

[0124] As Figure 9 shown, an embodiment of the page fault exception handling device 60 provided by the embodiments of this application includes:

[0125] A first processing unit 601, configured to save the context of the first coroutine that triggers a page fault exception into the shared memory. The first coroutine belongs to the first thread, and the shared memory is the memory that the first thread can access both in the kernel space and the user space; this first processing unit 601 can execute step 101 in the above method embodiment.

[0126] A second processing unit 602, configured to switch from the context of the first coroutine to the context of the first thread after the first processing unit 601 saves the context of the first coroutine into the shared memory. The context of the first thread is configured into the shared memory when the first thread is initialized; this second processing unit 602 can execute step 102 in the above method embodiment.

[0127] A third processing unit 603, configured to switch from the kernel space to the user space after the second processing unit 602 switches the context; this third processing unit 603 can execute step 103 in the above method embodiment.

[0128] A fourth processing unit 604, configured to, after the third processing unit 603 switches from the kernel space to the user space, trigger a page swap-in process by running the first thread to obtain the context of the first coroutine from the shared memory. This fourth processing unit 604 can execute step 104 in the above method embodiment.

[0129] In the solution provided by the embodiment of the present application, the context of the first coroutine is saved in the shared memory in the kernel mode. After returning from the kernel mode to the user mode, by running the first thread, the context of the first coroutine can be obtained from the shared memory, and then the page-in process can be executed according to the context of the first coroutine. Compared with the prior art, when a certain coroutine triggers a page fault exception, it is necessary to notify the monitor thread in the kernel mode, and then the first thread enters the sleep state until the monitor thread completes the page-in through the swap-in thread, and then sends a notification message to the kernel mode to wake up the first thread and continue to execute the coroutine. The page fault exception handling process of the present application can shorten the latency of page fault exception handling, thereby reducing the input / output (IO) long-tail latency of the first thread, and shortening the latency also improves the service throughput accordingly.

[0130] Optionally, the fourth processing unit 604 is further configured to, when executing the page-in process, run the second coroutine belonging to the first thread to execute the task corresponding to the second coroutine.

[0131] Optionally, the second processing unit 602 is configured to write the context of the first thread into the register of the computer system through a hook function to replace the context of the first coroutine in the register.

[0132] Optionally, the fourth processing unit 604 is configured to obtain the context of the first coroutine from the shared memory by running the first thread, and obtain the destination address from the context of the first coroutine, where the destination address is the address of the physical page to be accessed when the first coroutine triggers a page fault exception; according to the destination address, execute the page-in process of the corresponding physical page.

[0133] Optionally, the fourth processing unit 604 is further configured to, when the physical page is swapped into the memory, add the first coroutine to the coroutine waiting queue, and the coroutines in the coroutine waiting queue are in a pending scheduling state.

[0134] Optionally, the shared memory is configured for the first thread when the first thread is initialized.

[0135] Optionally, the page fault exception is triggered when the first coroutine is run to access the physical page swapped out in the memory.

[0136] Optionally, the shared memory is configured through the kernel virtual machine ebpf.

[0137] Above, the relevant content of the page fault exception handling device 60 provided by the embodiment of the present application can be understood by referring to the corresponding content in the foregoing method embodiment part, and will not be repeated here.

[0138] Figure 10As shown in the figure, it is a possible schematic diagram of the logical structure of the computer device 70 provided by the embodiment of the present application. The computer device 70 includes: a processor 701, a communication interface 702, a memory 703, a disk 704, and a bus 705. The processor 701, the communication interface 702, the memory 703, and the disk 704 are interconnected through the bus 705. In the embodiment of the present application, the processor 701 is used to control and manage the actions of the computer device 70. For example, the processor 701 is used to execute Figures 3 to 8 the steps in the method embodiment. The communication interface 702 is used to support the computer device 70 to communicate. The memory 703 is used to store the program code and data of the computer device 70, and to provide memory space for threads. The memory also includes a shared memory, and the function of the shared memory can be understood by referring to the shared memory in the foregoing method embodiment part. The disk is used to store the physical pages swapped out from the memory.

[0139] Among them, the processor 701 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 701 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 705 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0140] In another embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores computer-executable instructions. When the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figures 3 to 8 steps executed by the processor.

[0141] In another embodiment of the present application, a computer program product is further provided. The computer program product includes computer-executable instructions, and the computer-executable instructions are stored in a computer-readable storage medium; when the processor of the device executes the computer-executable instructions, the device executes the above-mentioned Figures 3 to 8 steps executed by the processor.

[0142] In another embodiment of the present application, a chip system is further provided. The chip system includes a processor, which is used to support the processing device for page fault exception to implement the above Figures 3 to 8 steps executed by the processor in the above. In a possible design, the chip system may further include a memory, which is used to store the necessary program instructions and data for the device that writes data. The chip system may be composed of chips or may include chips and other discrete devices.

[0143] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0144] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0145] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] In addition, the functional units in each embodiment of the embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0148] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

Claims

1. A method for handling page fault exceptions, which is applied to a computer system, characterized in that, the method includes: Saving the context of the first coroutine that triggers the page fault exception to the shared memory, where the first coroutine belongs to the first thread, and the shared memory is the memory that the first thread can access in both kernel mode and user mode; Switching from the context of the first coroutine to the context of the first thread, where the context of the first thread is configured to the shared memory when the first thread is initialized; Switching from the kernel mode to the user mode; Triggering the page-in process by running the first thread to obtain the context of the first coroutine from the shared memory.

2. The handling method according to claim 1, characterized in that, the method further includes: When executing the page-in process, running the second coroutine belonging to the first thread to execute the task corresponding to the second coroutine.

3. The handling method according to claim 1 or 2, characterized in that, the switching from the context of the first coroutine to the context of the first thread includes: Writing the context of the first thread into the register of the computer system through a hook function to replace the context of the first coroutine in the register.

4. The handling method according to claim 1 or 2, characterized in that, the triggering of the page-in process by running the first thread to obtain the context of the first coroutine from the shared memory includes: Running the first thread to obtain the context of the first coroutine from the shared memory, and obtaining the destination address from the context of the first coroutine, where the destination address is the address of the physical page that the first coroutine wants to access when triggering the page fault exception; Executing the page-in process for the corresponding physical page according to the destination address.

5. The handling method according to claim 4, characterized in that, the method further includes: After the physical page is swapped into the memory, adding the first coroutine to the coroutine waiting queue, and the coroutines in the coroutine waiting queue are in a state of waiting to be scheduled.

6. The handling method according to claim 1 or 2, characterized in that, the shared memory is configured for the first thread by the first thread when it is initialized.

7. The handling method according to claim 1 or 2, characterized in that, the page fault exception is triggered by running the first coroutine to access the physical page swapped out in the memory.

8. The handling method according to claim 1 or 2, characterized in that, the shared memory is configured by the kernel virtual machine ebpf.

9. A device for handling page fault exceptions, characterized in that, it includes: A first processing unit for saving the context of the first coroutine that triggers the page fault exception to the shared memory, where the first coroutine belongs to the first thread, and the shared memory is the memory that the first thread can access in both kernel mode and user mode; A second processing unit, configured to switch from the context of the first coroutine to the context of the first thread after the first processing unit saves the context of the first coroutine to the shared memory, where the context of the first thread is configured to the shared memory during the initialization of the first thread; A third processing unit, configured to switch from the kernel mode to the user mode after the second processing unit switches the context; A fourth processing unit, configured to, after the third processing unit switches from the kernel mode to the user mode, obtain the context of the first coroutine from the shared memory by running the first thread, so as to trigger a page-in process.

10. The processing device according to claim 9, wherein, the fourth processing unit is further configured to run a second coroutine belonging to the first thread to execute a task corresponding to the second coroutine when executing the page-in process.

11. The processing device according to claim 9 or 10, wherein, the second processing unit is configured to write the context of the first thread into a register of the computer system through a hook function to replace the context of the first coroutine in the register.

12. The processing device according to claim 9 or 10, wherein, the fourth processing unit is configured to obtain the context of the first coroutine from the shared memory by running the first thread, and obtain a destination address from the context of the first coroutine, where the destination address is the address of the physical page to be accessed when the first coroutine triggers a page fault exception; and execute a page-in process for a corresponding physical page according to the destination address.

13. The processing device according to claim 12, wherein, the fourth processing unit is further configured to add the first coroutine to a coroutine waiting queue after the physical page is swapped into the memory, and the coroutines in the coroutine waiting queue are in a state of waiting to be scheduled.

14. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by one or more processors, the method according to any one of claims 1-8 is implemented.

15. A computing device, wherein, comprises one or more processors and a computer-readable storage medium storing a computer program; when the computer program is executed by the one or more processors, the method according to any one of claims 1-8 is implemented.

16. A chip system, wherein, comprises one or more processors, and the one or more processors are called to execute the method according to any one of claims 1-8.

17. A computer program product, wherein, comprises a computer program, and when the computer program is executed by one or more processors, it is used to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Zero copy message reception method and system

    CN102402487A

  • Shared memory structure used for communications between kernel mode and user mode and application thereof

    CN107577539A