A kernel service transfer method and system for accelerating user space execution
By building a process kernel service transfer management structure and a dual-kernel stack switching mechanism in the Linux operating system, the separation of kernel services and user logic is achieved, the performance degradation caused by kernel services in the existing technology is solved, and the execution efficiency and system performance of user space programs are improved.
Patent Information
- Application Number
- CN202510609457.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The lack of efficient kernel service acceleration mechanism in the existing Linux operating systems has led to a decline in user space program execution performance. Especially when the kernel service switching and resource competition are frequent, the system performance overhead is high, and it cannot meet the transparent and convenient acceleration needs without modifying the program source code.
By building a process kernel service transfer management structure, monitoring the process status in real time, the separation of kernel service logic and user logic is achieved, and the user space is handed over to the user logic process for execution, and the kernel service is handed over to the kernel service logic process for completion, reducing system call overhead.
It improves the execution efficiency of user space programs, reduces the delay caused by resource competition and frequent context switching, and improves system performance and response speed.
Smart Images

Figure CN120123187B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer operating system performance optimization, and in particular to a kernel service transfer method and system for accelerating user space execution. Background Art
[0002] In the Linux operating system, to meet the execution requirements of user-space programs and ensure kernel stability, the system sometimes needs to switch to kernel space to execute related kernel code logic (also known as kernel services) during operation. Kernel services can negatively impact the performance of user-space programs, such as sharing limited hardware resources with user program logic, increasing logic execution latency, and impacting the overall speed and responsiveness of applications. Furthermore, the transition from user space to kernel space (i.e., context switching) during kernel services increases overall system overhead. Frequent context switches significantly increase latency, reducing system throughput and responsiveness. Therefore, the indirect costs of kernel services (such as cache misses caused by resource contention) and the direct overhead (such as the additional logic execution time due to context switches) pose a significant challenge to optimizing system performance.
[0003] In order to reduce the performance overhead associated with kernel service processing, one method is to batch system calls, which reduces the processing cost of each system call by submitting multiple system call requests at one time. For example, the io_uring mechanism integrated into the Linux kernel allows applications to package multiple I / O operation requests and submit them to dedicated threads in the kernel for processing at one time. These batched system calls are completed collaboratively by kernel threads, thereby reducing frequent context switches between user space and kernel space and improving the concurrent processing capability of I / O operations. However, io_uring is only suitable for scenarios that require high-concurrency I / O operations. Developers need to modify the existing application source code and call specific io_uring APIs. Therefore, its programming model is more complex than traditional synchronous system calls.
[0004] To address the indirect performance losses caused by TLB (Translation Lookaside Buffer) and cache misses due to hardware contention, another approach is to intercept system calls made by user programs, record the system call parameters, and pass them along. A proxy thread in kernel space then retrieves the passed parameters and executes the system call on behalf of the user program, thus separating kernel services from user programs. However, current approaches only target system calls and ignore other kernel services (such as interrupts and instruction aborts).
[0005] To sum up, in the current Linux operating system, there is still a lack of an efficient kernel service acceleration mechanism that can optimize various kernel services that occur when user space programs are executed without modifying the program source code, and provide a transparent, convenient and efficient kernel service acceleration mechanism for the execution of user space programs. Summary of the Invention
[0006] The technical problem to be solved by the present invention is: In response to the above-mentioned problems in the prior art, an efficient kernel service transfer method and system for accelerating user space execution are provided to reduce the overall overhead of kernel services and improve the performance of user space programs.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0008] A kernel service transfer method for accelerating user space execution, comprising:
[0009] Step S1: When a program needs to process a kernel service, it is determined whether the kernel has enabled the kernel service transfer function. If so, a process kernel service transfer management structure creation operation is performed to create a process kernel service transfer management structure. The process kernel service transfer management structure is used to create and manage an acceleration process capable of executing the kernel service transfer function.
[0010] Step S2, initializing the process kernel service transfer management structure, obtaining a kernel service logic process for executing the kernel service logic in the acceleration process, a user logic process for executing the user logic in the acceleration process, and an idle execution logic management structure for assisting the kernel service transfer;
[0011] Step S3: During the execution of the acceleration process, the process leaving the kernel is continuously monitored. If the leaving process is detected to be a kernel service logic process, a kernel stack switching operation is performed, so that the kernel service logic process obtains and switches to the kernel stack of the idle execution logic management structure, and the user logic process obtains and switches to the original kernel stack of the kernel service logic process and returns to the user space to execute the user logic;
[0012] Step S4: During the execution of the acceleration process, the process entering the kernel is continuously monitored. When it is detected that the process entering the kernel is a user logic process, the kernel service transfer operation is executed, so that the kernel service logic process obtains the kernel processing related data in the user logic process and executes the kernel service logic transferred from the user logic process in the kernel.
[0013] Furthermore, step S2 includes:
[0014] Create and initialize multiple process management structures in the kernel;
[0015] Use one of the process management structures as the kernel logic management structure and create a corresponding process as the kernel service logic process, use the other process management structure as the user logic management structure and create a daemon process as the user logic process;
[0016] The third process management structure is used as the idle execution logic management structure, an unused kernel stack is allocated, and the position where the process starts executing and the used kernel stack are set to the loop function and the created stack of the stage of the acceleration process for monitoring kernel service transfer respectively.
[0017] Furthermore, after step S2 and before step S3, the following steps are included:
[0018] When the initialization of multiple process management structures monitoring the acceleration process is completed, the kernel service logic process enters the loop function, and the user logic process executes the page table switching logic to switch the user logic mapping page table to the page table of the kernel service logic process; and after the switching is completed, the kernel service logic process leaves the loop function and enters the switching path to leave the kernel space and prepare to enter the user space.
[0019] Furthermore, step S3 includes:
[0020] When it is detected that the process leaving the kernel is a kernel service logic process and the kernel service logic process is in a state of releasing the logic of its own kernel stack and about to leave the kernel, writing the system register into the local variable, thereby storing the first context information of the currently executing process in the kernel stack of the kernel service logic process;
[0021] Execute kernel stack switching, write the current CPU context information into the kernel logic management structure, write the context information in the initialized idle execution logic management structure into the CPU, and make the kernel service logic process enter the loop function, so that the kernel stack of the kernel service logic process is switched to the kernel stack of the idle execution logic management structure;
[0022] When it is monitored that the kernel service logic process has switched the kernel stack and entered the loop function, the context information of the CPU where the user logic process is running is written into the user logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored first context information is written into the user logic process, so that the user logic process obtains the original kernel stack of the kernel service logic process and returns to the user space to execute the user logic.
[0023] Furthermore, step S4 includes:
[0024] When it is detected that the process entering the kernel is a user logic process and a kernel service is required, it is determined whether the kernel service logic process is in a state where the kernel stack has been switched and the loop function has been entered. If so, a kernel service transfer is performed, and the system register is written into a local variable, thereby storing the second context information of the currently executed process including the kernel service category;
[0025] Write the current CPU context information into the kernel logic management structure, and write the user logic management structure context information into the CPU, so that the user logic process can obtain the original kernel stack to continue to execute the monitoring acceleration process transfer state value;
[0026] When it is detected that the kernel stack of the kernel logic management structure is currently in an idle state, the context information of the current CPU is written into the idle execution logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored second context information is written into the kernel service logic process, thereby completing the kernel service transfer and enabling the kernel service logic process to execute the kernel service passed from the user logic process.
[0027] Furthermore, if it is determined that the kernel service logic process is not in a state where the kernel stack has been switched and the loop function has not been entered, it indicates that an error has occurred in switching the accelerated process transfer state, and the kernel service transfer operation is refused.
[0028] Furthermore, step S3 and step S4 are executed alternately, so that after the kernel service logic process completes the execution of the transferred kernel service, the user logic is handed over to the user logic process through step S3 for continued execution, and after the user logic process performs the kernel service, the kernel service transfer is completed through step S4. Step S3 and step S4 are executed alternately until the acceleration process is completed or dies.
[0029] The present invention further provides a kernel service transfer system for accelerating user space execution, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute a kernel service transfer method for accelerating user space execution.
[0030] The present invention further provides a computer-readable storage medium having a computer program / instruction stored therein, wherein the computer program / instruction is programmed or configured to execute, through a processor, a kernel service transfer method for accelerating user space execution.
[0031] The present invention further provides a computer program product comprising a computer program / instruction programmed or configured to execute, by a processor, a kernel service transfer method for accelerating user space execution.
[0032] Compared with the prior art, the advantages of the present invention are:
[0033] The present invention manages the acceleration process of kernel service transfer by constructing a process kernel service transfer management structure, making the process of processing kernel services more efficient and more scalable; by real-time monitoring of the process status entering or leaving the kernel, it ensures that the kernel service transfer logic is executed at the right time, thereby achieving efficient, low-latency, safe and stable kernel service processing; through the dual kernel stack switching formed by the kernel service logic process kernel stack and the idle execution logic management structure kernel stack, the user space part of the acceleration process is handed over to the user logic process for execution, while the kernel service of the acceleration process is handed over to the kernel service logic process for completion, and when it is monitored that the user logic process has a kernel service, the kernel service is handed over to the kernel service logic process for execution, thereby achieving separation of kernel service logic and user logic, reducing resource competition caused by mixed operation, avoiding frequent context switching between kernel and user space, thereby reducing the overhead of system calls and improving program performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is the execution flow chart after the kernel service is triggered in the traditional Linux kernel.
[0035] Figure 2 This is the execution flow of io_uring in the Linux kernel.
[0036] Figure 3 Workflow diagram for kernel threads to process system calls.
[0037] Figure 4 Schematic diagram of the flow of the kernel service transfer method for accelerating user space execution according to the present invention.
[0038] Figure 5 This is a schematic diagram of member variables of a process kernel service transfer management structure in a specific application embodiment of the present invention.
[0039] Figure 6 The figure is a schematic diagram of a specific execution flow of a kernel service transfer method for accelerating user space execution in a specific application embodiment of the present invention.
[0040] Figure 7 This is a schematic diagram of the kernel service acceleration process creation process in a specific application embodiment of the present invention.
[0041] Figure 8 This is a flow chart of handing over user logic to a user logic process for processing through a kernel stack switching operation in a specific application embodiment of the present invention.
[0042] Figure 9This is a flow chart of handing over kernel service logic to a kernel service logic process for processing through a kernel service transfer operation in a specific application embodiment of the present invention. DETAILED DESCRIPTION
[0043] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0044] In the Linux operating system, the system switches to the kernel space to execute related kernel code logic during operation. This usually occurs when the user program explicitly requests certain services, such as system calls, or transparent requests triggered by events such as data aborts, instruction aborts, and interrupts. These kernel code logics that enter the kernel space for execution are collectively referred to as kernel services. The execution of kernel services is an important part of the operating system, involving the management and allocation of system resources, as well as the control of hardware devices. Through kernel services, user programs can request the operating system to provide various services, such as file operations, process management, memory management, etc. The execution of these services needs to be performed in kernel mode to ensure the security and stability of the operating system. In the traditional Linux kernel, the process of kernel service processing is as follows: Figure 1 As shown in Figure 2. Applications can actively trigger switching by executing user instructions, or passively trigger it through the interrupt controller. System calls are the primary way user programs interact with the kernel, allowing them to request specific kernel operations. Interrupt and exception handling mechanisms respond to external events, such as hardware interrupts, software interrupts, and various exceptions. They enable the system to respond to and handle these events promptly, ensuring normal operation.
[0045] However, kernel services can negatively impact the performance of user-space programs during their execution. First, kernel service logic shares limited hardware resources with user program logic, such as the CPU and cache system. This resource contention often leads to increased cache miss rates and TLB (Translation Lookaside Buffer) miss rates. When cache miss rates increase, the processor spends more time fetching data from main memory, increasing data access latency. Similarly, an increased TLB miss rate means that virtual-to-physical address translation takes longer, increasing logic execution latency. This further reduces the application's instructions per cycle (IPC). This reduced IPC directly impacts the overall speed and responsiveness of the application. Furthermore, during kernel services, a transition from user space to kernel space (i.e., a context switch) is unavoidable. This transition involves saving the current user program state, switching to kernel mode, executing the necessary kernel service logic, and then switching back to user mode. This process not only requires multiple logic executions but also causes cache and TLB content to be flushed, increasing overall system overhead. Frequent context switching, especially in highly concurrent or real-time systems, significantly increases latency and reduces system throughput and responsiveness. Therefore, the indirect costs of kernel services (such as cache and TLB misses and branch mispredictions) and the direct overhead (such as the additional logic execution time caused by context switching) together constitute a major challenge in optimizing system performance.
[0046] In order to reduce the performance overhead associated with kernel service processing, the indirect performance loss caused by TLB misses and cache misses can be minimized, or the direct overhead generated when switching from user space to kernel space can be minimized. For example, in order to reduce the high overhead of a single system call, some practitioners have proposed a batch system call method. This method reduces the processing cost of each system call by submitting multiple system call requests at one time. Figure 2As shown, the Linux kernel has integrated a mechanism called io_uring. io_uring allows applications to bundle multiple I / O operation requests and submit them to dedicated kernel threads for processing. These batched system calls are coordinated by kernel threads, reducing frequent context switches between user and kernel space and improving the concurrency of I / O operations. io_uring is specifically optimized for asynchronous I / O system calls, using a ring buffer to efficiently manage large numbers of I / O requests. Applications can place multiple I / O requests in a submission queue, which is then batched and processed by a kernel thread. Upon completion, the kernel thread places the results in a completion queue, from which the application can quickly obtain the results through polling or notification mechanisms. This batching significantly reduces the number of context switches required for each system call and the associated resource overhead, thereby improving overall I / O performance. However, io_uring still has some shortcomings and limitations in practical applications: it is only suitable for scenarios that require high-concurrency I / O operations, and developers need to modify existing application source code and call specific io_uring APIs. Therefore, its programming model is more complex than traditional synchronous system calls.
[0047] For another example, in order to solve the indirect performance loss caused by TLB misses and cache misses caused by hardware contention, some practitioners have proposed the method of transfer execution. The core idea of transfer execution is to separate two parts of code that are not highly correlated and execute them on different CPU cores to improve the locality of the program. For a process, the correlation between its kernel services and user logic is very low, and separate execution can improve the locality of the kernel part and the user part respectively. Figure 3 As shown, one approach to transfer execution is to intercept a system call from a user program, record the system call parameters, and pass them along. A proxy thread in kernel space then retrieves the passed parameters and executes the system call on its behalf, thus separating kernel services from user programs. However, current approaches only target system calls and ignore other kernel services (such as interrupts and instruction aborts), making the transfer of kernel services less applicable.
[0048] In summary, the existing kernel service switching efficiency is low, the overhead is high, and the applicable kernel service types are relatively few, which cannot meet the requirements of an efficient kernel service acceleration mechanism for user programs. Therefore, in order to solve the problems such as performance degradation caused by kernel services when the existing Linux operating system executes user programs, the present invention creates a process kernel service transfer management structure and a corresponding acceleration process to monitor the kernel service logic process leaving the kernel and the user logic process entering the kernel, and performs a kernel stack switching operation or a kernel service transfer operation when the corresponding process is monitored, thereby handing over the user space part of the acceleration process to the user logic process for execution, and handing over the kernel service of the acceleration process to the kernel service logic process for completion, thereby realizing the separation of kernel service logic and user logic, reducing the overhead of system calls, and improving program performance.
[0049] like Figure 4 and Figure 6 As shown, the kernel service transfer method for accelerating user space execution according to an embodiment of the present invention includes:
[0050] Step S1: When a program needs to process a kernel service, it is determined whether the kernel has enabled the kernel service transfer function. If so, a process kernel service transfer management structure creation operation is performed to create a process kernel service transfer management structure. The process kernel service transfer management structure is used to create and manage an acceleration process that can execute the kernel service transfer function.
[0051] Step S2, initializing the process kernel service transfer management structure, obtaining the kernel service logic process for executing the kernel service logic in the acceleration process, the user logic process for executing the user logic in the acceleration process, and the idle execution logic management structure for assisting the kernel service transfer;
[0052] Step S3: During the execution of the accelerated process, the process leaving the kernel is continuously monitored. If the leaving process is detected to be a kernel service logic process, a kernel stack switching operation is performed, causing the kernel service logic process to obtain and switch to the kernel stack of the idle execution logic management structure, and causing the user logic process to obtain and switch to the original kernel stack of the kernel service logic process and return to the user space to execute the user logic;
[0053] Step S4: During the execution of the accelerated process, the process entering the kernel is continuously monitored. When it is detected that the process entering the kernel is a user logic process, the kernel service transfer operation is executed, so that the kernel service logic process obtains the kernel processing related data in the user logic process and executes the kernel service logic transferred from the user logic process in the kernel.
[0054] In a specific application embodiment, a designed user interface is used to pre-enable the kernel's kernel service transfer processing mechanism. This allows the kernel to perform corresponding detection functions when the kernel service transfer program is executed. The designed user interface then passes the kernel the flags for the kernel service transfer processing processes so that the kernel can correctly handle these kernel service transfer processes when executing kernel services. The user interface is implemented by reading and writing to a special file, / proc / trans. Writing "start" tells the kernel to enable the kernel service transfer mechanism, and writing the process name tells the kernel to execute the kernel service transfer operation for that process (member variable "flag").
[0055] Process kernel service transfer management structure such as Figure 5 As shown, its design purpose is to manage special processes (i.e., acceleration processes) that need to complete kernel service transfer. The kernel service transfer management structure of this process (hereinafter referred to as the management structure) contains the following important member variables:
[0056] First, a member variable "flag" is defined to indicate whether the kernel service transfer feature is enabled. Its initial value is False, indicating that kernel service transfer is not enabled. If kernel service transfer is enabled, the transfer operation will be executed at a specific kernel hook point when the acceleration process is executed.
[0057] Second, three Linux kernel process management structures (task_struct—the core structure for kernel process management) were created: the kernel logic management structure "k_task" for the acceleration process (used to execute the kernel service logic of the acceleration process); the user logic management structure "u_task" for the acceleration process (used to execute the user logic of the acceleration process); and the idle execution logic management structure "dual_task" for the acceleration process (used to assist in completing kernel service transfers). In the following text, the kernel logic management structure k_task and the user logic management structure u_task, unless otherwise specified, represent their corresponding kernel service logic processes and user logic processes, and are described using variable names.
[0058] Third, accelerate the transfer state of the process. The kernel service transfer is designed to be divided into multiple stages, which are executed in sequence. Therefore, the state is used to mark the stage of service transfer of the process performing the kernel service transfer.
[0059] Preferably, "state" is aligned to 64 bytes, as it may be read and written multiple times during kernel service transfers. This ensures it meets the kernel's cache line requirements and fully utilizes the cache. The three task_struct structures are also aligned to 64 bytes, as they are not modified after initialization and will be read multiple times during kernel service transfers.
[0060] In this embodiment, step S2 includes:
[0061] Create and initialize multiple process management structures in the kernel;
[0062] Use one of the process management structures as the kernel logic management structure and create a corresponding process as the kernel service logic process, use the other process management structure as the user logic management structure and create a daemon process as the user logic process;
[0063] The third process management structure is used as the idle execution logic management structure, an unused kernel stack is allocated, and the position where the process starts executing and the used kernel stack are set to the loop function and the created stack of the stage of the acceleration process for monitoring kernel service transfer respectively.
[0064] In a specific application embodiment, since the traditional Linux kernel's existing process management structure task_struct is difficult to complete the task of kernel service transfer, when the kernel service transfer program is executed, the kernel will execute the process kernel service transfer management structure creation method to create an acceleration process that can perform kernel service transfer processing, so that when the acceleration process is executed, each kernel service initiated by the special process is efficiently transferred through the kernel service transfer method, thereby improving the performance of the acceleration process. Figure 7 As shown (state is initialized to 0), the initialization steps of the process kernel service transfer management structure and process creation are as follows (corresponding to step S2):
[0065] Step A2.1: Insert a detection point at the end of the kernel process creation point. This detection point continuously detects the created process. If the kernel has enabled kernel service transfer and detects that the process is an accelerated process, the following kernel service accelerated process creation method is executed:
[0066] In step A2.2, a new process (the accelerated process) is created in the kernel and begins execution. After the process is created, the kernel allocates and initializes its corresponding task_struct, setting the value of k_task directly to the value of this task_struct. (The variable k_task is the original process management data structure of the accelerated process, and its function is to execute the kernel logic of the accelerated process. Since this task_struct is created in the kernel, k_task can conveniently and naturally participate in the kernel services of the accelerated process.) A daemon process is created, which is dedicated to executing the user logic of the accelerated process. After u_task is created, it begins execution. (For the variable u_task, since it executes user logic, the daemon process only needs to obtain the context of the accelerated process when it leaves kernel space and enters user space to correctly execute the user logic of the accelerated process. The user logic is the user code, and the context includes the register values and the stack used when the process executes a certain code.)
[0067] Step A2.3: Create a dual_task structure for the acceleration process to assist in completing the kernel service transfer operation. (When u_task obtains the appropriate context of k_task in the acceleration process, it can begin to execute user logic. This will cause k_task to have no executable context, and then it will be an invalid process structure in the kernel. Therefore, in order for k_task to correctly respond and execute the required logic when u_task, which is executing user logic, issues a kernel request, a dual_task structure is created.) The steps for creating and initializing dual_task are as follows:
[0068] Step A2.3.1: Create the necessary stack for dual_task to execute the logic.
[0069] Step A2.3.2, clear the stack contents of dual_task, indicating that this stack has not been used;
[0070] Step A2.3.3: Clear the contents of the thread->cpu_context member variable of the dual_task structure (the contents of this member variable represent information such as the location where the process represented by the structure starts execution. If it is correctly set, it means that it can participate in process scheduling);
[0071] Step A2.3.4, set the thread->cpu_context.pc and thread->cpu_context.sp values in the dual_task structure. (These two values represent the execution start location and the kernel stack to be used, respectively. The pc value is set to a special loop function, specifically the location of trans_logic, which will directly call the loop function dual_logic, and the sp value is set to the stack created in step A2.3.1.)
[0072] In step A2.4, state is set to -1 (the default initial value of state after kernel startup is 0), indicating that the important management structures used by the acceleration process have been initialized, but the creation process has not yet been fully completed.
[0073] In this embodiment, after step S2 and before step S3, the following steps are included:
[0074] When the initialization of multiple process management structures monitoring the acceleration process is completed, the kernel service logic process enters the loop function, and the user logic process executes the page table switching logic to switch the user logic mapping page table to the page table of the kernel service logic process; and after the switching is completed, the kernel service logic process leaves the loop function and enters the switching path to leave the kernel space and prepare to enter the user space.
[0075] In a specific application embodiment, after all member variables involved in the management structure are correctly initialized, the acceleration process needs to begin execution. For the user process, the most important thing is that the user logic is correctly executed, so the acceleration process is created by returning to the user space and starting to execute the user logic. Therefore, after initializing the management structure, it is also necessary to make the acceleration process return to the user space and start executing the user logic. The implementation steps are completed by intercepting the execution logic after the k_task is created, as follows:
[0076] In step A3.1, the process represented by k_task enters an infinite while loop, continuously monitoring the value of state until it is not -1 and then exits the infinite loop. (Since state was set to -1 in step A2.4, k_task will be trapped in this infinite loop.)
[0077] In step A3.2, when the running u_task detects that the state has changed to -1, it executes to obtain the context of the acceleration process (i.e., k_task), as follows:
[0078] Step A3.2.1: To properly execute the user logic of the acceleration process, u_task executes the page table switching logic to switch the user logic's mapping page table to k_task's user page table (this is accomplished by calling the Linux kernel's swith_mm_irqs_off function, whose main logic modifies the current CPU's user space page table register TTBR0_EL1). This function then obtains the acceleration process's page table (representing the correct mapping of the acceleration process's user logic) and the values of k_task's related system registers.
[0079] Step A3.2.2: Call some auxiliary switching functions of the Linux kernel (such as tls_thread_switch and contextidr_thread_switch) to switch some auxiliary information and system register values of the current CPU to the value of k_task;
[0080] Step A3.2.3, set state to 0;
[0081] Step A3.2.4: When k_task detects that the state value has become 0, k_task exits the infinite loop, ending the execution flow of the kernel service acceleration process creation method.
[0082] In step A3.3, k_task enters the switching path from kernel space to user space, enters kernel_exit to leave kernel space and prepares to enter user space to execute user logic.
[0083] After completing steps A2 and A3, the acceleration process obtains a fully initialized management structure, including auxiliary data structures such as dual_task and the state transition flag state, along with two currently executing processes (at this point, both u_task and k_task have completed the preparations for leaving the kernel after the acceleration process is created): u_task, which has already retrieved k_task's page table and other contents; and k_task, the acceleration process itself, which, after intercepting its execution logic after creation and completing step A3, enters the kernel-to-user space switching path. It should be noted that the acceleration process's k_task begins execution upon creation, while u_task begins execution when the kernel enables the kernel service transfer function (i.e., the flag is set to true). Its execution involves continuously monitoring the state value of state, and u_task and k_task are set to run on different cores.
[0084] In this embodiment, step S3 includes:
[0085] When it is detected that the process leaving the kernel is a kernel service logic process and the kernel service logic process is in a state of releasing the logic of its own kernel stack and about to leave the kernel, writing the system register into the local variable, thereby storing the first context information of the currently executing process in the kernel stack of the kernel service logic process;
[0086] Execute kernel stack switching, write the current CPU context information into the kernel logic management structure, write the context information in the initialized idle execution logic management structure into the CPU, and make the kernel service logic process enter the loop function, so that the kernel stack of the kernel service logic process is switched to the kernel stack of the idle execution logic management structure;
[0087] When it is monitored that the kernel service logic process has switched the kernel stack and entered the loop function, the context information of the CPU where the user logic process is running is written into the user logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored first context information is written into the user logic process, so that the user logic process obtains the original kernel stack of the kernel service logic process and returns to the user space to execute the user logic.
[0088] In a specific application embodiment, after step A3 is completed, k_task enters the ret_from_fork position of the Linux kernel, detects whether scheduling occurs, and then enters the ret_to_user position of the kernel (ret_from_fork and ret_to_user are fixed execution routes for processes to leave the kernel space in Linux), completes possible kernel events such as signal processing in ret_to_user, and finally enters the kernel's kernel_exit (kernel_exit is the last piece of logic to leave the Linux kernel and return to the normal execution of the user program). At this time, transparent and fast kernel service transfer processing is further achieved through the relevant member variables in the process kernel service transfer management structure, thereby handing over the user space part of the accelerated process to u_task for execution, and leaving the kernel service of the accelerated process to k_task itself to complete, so as to achieve the purpose of separating user logic and kernel logic, expanding hardware resources, and reducing cache and TLB misses. Figure 8 As shown ( Figure 8 The initial state of the process is 0, and the green arrows indicate the main execution steps of this stage. Specifically, the dual kernel stack mechanism formed by the kernel stack of dual_task and the native kernel stack of k_task is used to achieve the purpose of fast kernel service transfer, so that u_task can correctly execute the user logic of the acceleration process. The implementation steps are as follows (corresponding to step S3):
[0089] In step A4.1, in kernel_exit, a hook point is set in advance. When the flag is turned on, each process leaving the kernel is detected at this point. If a normal process is detected leaving the kernel, its execution path is consistent with the traditional Linux kernel. If the process leaving the kernel is detected as a k_task, special processing is performed as follows:
[0090] Step A4.2, mark state as 1, which indicates the start of k_task's logic of releasing its own kernel stack;
[0091] In step A4.3, k_task writes some system registers into local variables (such as TPIDR_EL0. The values in local variables will be saved in the kernel stack of k_task).
[0092] Step A4.4, call the barrier instruction to ensure that the above logic has been completed;
[0093] Step A4.5: Call the kernel stack switching logic. The specific steps are as follows:
[0094] Step A4.5.1: Write the current CPU's SP and X19 to X30 registers into the current k_task's cpu_context (including information such as the process start location (pc) and the memory stack used (sp). The SP (stack pointer) points to the current stack top, controls stack allocation and deallocation, and is used to store local variables, function call return addresses, context, and other data. X19 to X29 (general registers) are used to store function call parameters, local variables, and calculation results. X30 (LR, link register) is used to store the function return address and is used for function call and return operations to ensure the correctness of control flow.
[0095] Step A4.5.2: Take the corresponding values from dual_task's cpu_context and put them into SP, X19 to X30 registers, so that the values of the CPU's SP, X19 to X30 registers are switched to the information stored in dual_task (the values of these registers represent the context information of the process corresponding to the task_struct structure);
[0096] In step A4.5.3, the ret instruction is called. The ret instruction returns to the address pointed to by the LR (X30) register. The LR register (used to store the return address of the function call) has been set in step A4.5.2. According to step A2.3, the address of LR is set to an assembly address. After entering this assembly address, a function jump is immediately executed to enter the dual_logic (infinite loop) function. At this time, k_task has changed the SP stack register from pointing to its own native kernel stack to pointing to the dual_task kernel stack, thus abandoning the use of its own native kernel stack and starting to execute the dual_task logic. It returns to the set trans_logic location and calls the dual_logic function. The execution logic of the dual_logic function is as follows:
[0097] Dual_logic is an infinite loop and continuously checks the value of state. If the value of state is 1, it means that k_task has given up using its own kernel stack and wants to leave the kernel. At this time, state is set to 2 to notify u_task that k_task has given up using its own kernel stack and entered dual_logic. u_task can safely use the context information in k_task's kernel stack. Then, hardware interrupts (CPU interrupts) and other flags are turned on (because before entering dual_logic, k_task is on the path of switching from kernel space to user space. This path will turn off kernel interrupts, so that dual_logic will not occupy the CPU for a long time, improving system resource utilization).
[0098] In step A4.6, u_task running on another core continuously checks the state. If u_task detects that state is 2, it performs the following operations:
[0099] Step A4.6.1, write the SP, x19 to x30 registers of the current CPU into the cpu_context of the current u_task;
[0100] Step A4.6.2: The corresponding values are taken from k_task's cpu_context and placed into the SP, X19 to X30 registers. This causes the values of the CPU's SP, X19 to X30 registers to be switched to the information stored in k_task (because the existence of state ensures the continuity of state switching, it can be ensured that the kernel stack of k_task is no longer in use). At this point, the CPU registers have been switched, and X30 (also called LR) is the register that stores the return address, so the return address of u_task has been changed to the address in kernel_exit.
[0101] In step A4.6.3, u_task retrieves the system register values written to the local variables in step A4.3. These system register values are related to the correct execution of the process;
[0102] In step A4.6.4, u_task obtains the register value of the kernel stack that was set when the k_task process was normally created, so that k_task successfully abandons the use of its own kernel stack. At the same time, u_task successfully obtains k_task's kernel stack and returns to user space to execute user logic (from step A4.6.2, it can be seen that u_task's SP stack register points to k_task's kernel stack).
[0103] It can be understood that when a user-space program needs to process kernel services, the dual-core stack processing mechanism is designed to efficiently complete the kernel service transfer process, thereby transferring the kernel services to other cores for execution. This can avoid competition for hardware resources (such as cache and TLB) caused by executing kernel services, separate the kernel service part of the user program from the user-space part, expand hardware resources, and thus improve the overall performance of the system.
[0104] In this embodiment, step S4 includes:
[0105] When it is detected that the process entering the kernel is a user logic process and a kernel service is required, it is determined whether the kernel service logic process is in a state where the kernel stack has been switched and the loop function has been entered. If so, a kernel service transfer is performed, and the system register is written into a local variable, thereby storing the second context information of the currently executed process including the kernel service category;
[0106] Write the current CPU context information into the kernel logic management structure, and write the user logic management structure context information into the CPU, so that the user logic process can obtain the original kernel stack to continue to execute the monitoring acceleration process transfer state value;
[0107] When it is detected that the kernel stack of the kernel logic management structure is currently in an idle state, the context information of the current CPU is written into the idle execution logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored second context information is written into the kernel service logic process, thereby completing the kernel service transfer and enabling the kernel service logic process to execute the kernel service passed from the user logic process.
[0108] In this embodiment, if it is determined that the kernel service logic process is not in the state of having switched the kernel stack and entered the loop function, it indicates that an error has occurred in switching the accelerated process transfer state, and the kernel service transfer operation is refused.
[0109] In this embodiment, step S3 and step S4 are executed alternately, so that after the kernel service logic process completes the execution of the transferred kernel service, the user logic is handed over to the user logic process through step S3 for continued execution, and after the user logic process performs the kernel service, the kernel service transfer is completed through step S4. Step S3 and step S4 are executed alternately until the acceleration process is completed or dies.
[0110] In a specific application embodiment, Figure 9 As shown ( Figure 9 The initial state of the process is 2, and the green arrows indicate the main execution steps of this stage. When it is detected that the u_task executing the user logic has a kernel service (such as a system call, interrupt, instruction abort, etc.), the implementation steps for kernel service transfer are as follows (corresponding to step S4):
[0111] Step A5.1: u_task performs kernel service and enters the switching path from user space to kernel space.
[0112] In step A5.2, at the kernel_entry end of the switching path, a corresponding hook point is set to detect whether kernel service transfer is enabled. If the flag is detected to be true (i.e., kernel service processing transfer is enabled) and the process entering the kernel is u_task, the kernel service transfer operation is executed as follows:
[0113] Step A5.3.1: Check whether the current state is 2. If not, it indicates that a problem has occurred in the state switch, and the kernel service transfer operation is rejected.
[0114] In step A5.3.2, state is set to 3, which indicates that u_task begins the kernel service transfer action;
[0115] Step A5.3.3: Read the values of system registers related to kernel service processing and store them in local variables. These system register values are used to help determine the type of kernel service, location, and other information.
[0116] Step A5.3.4, execute the barrier instruction to ensure that the above operations have been completed;
[0117] Step A5.3.5, write the current CPU's SP, x19 to x30 registers into k_task's cpu_context (because the current u_task uses k_task's kernel stack);
[0118] Step A5.3.6, take the corresponding value from the cpu_context of u_task and put it into the SP, X19 to X30 registers, so that the values of the SP, X19 to X30 registers of the CPU are switched to the information saved in u_task;
[0119] In step A5.3.7, u_task's SP stack register points to its own kernel stack, so u_task gets its own kernel stack. The pc gets u_task's original execution logic from X30 and resumes executing the state detection logic.
[0120] In step A5.3.8, if u_task detects that the state value is 3, it means that u_task has initiated a kernel service transfer request and sets state to 0, which indicates that the kernel stack of k_task is currently unused;
[0121] In step A5.3.9, k_task detects that the value of state is 0 and starts to obtain the relevant context information of the kernel service;
[0122] Step A5.3.10, execute the barrier instruction to ensure that the above operations have been completed;
[0123] Step A5.3.11, write the current CPU's SP, x19 to x30 registers into the dual_task's cpu_context;
[0124] Step A5.3.12, take the corresponding value from the cpu_context of k_task and put it into the SP, X19 to X30 registers;
[0125] In step A5.3.13, k_task returns to the kernel_entry location and obtains the system register values required by each kernel service from local variables;
[0126] In step A5.3.14, k_task ends the logic of kernel_entry, returns to the processing flow of the kernel service, and starts executing the kernel service passed by u_task so that the kernel service can complete its execution normally.
[0127] As can be appreciated, compared to existing approaches that focus solely on system calls while ignoring kernel services like interrupts and instruction aborts, this embodiment comprehensively transfers all kernel services, further expanding its scope of application. Furthermore, implemented in software, it can be easily extended to different kernel versions and hardware devices. Furthermore, users can transparently utilize this kernel service transfer mechanism without modifying their program source code, saving overall costs and improving time efficiency.
[0128] It should be noted that after k_task completes the transferred kernel service, the user logic is handed over to u_task for continued execution through the entire content of step A4. After u_task performs the kernel service, the kernel service transfer is completed through step A5; step A4 and step A5 are executed alternately until the acceleration process is completed or dies.
[0129] The present invention further provides a kernel service transfer system for accelerating user space execution, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute a kernel service transfer method for accelerating user space execution.
[0130] The present invention further provides a computer-readable storage medium having a computer program / instruction stored therein, wherein the computer program / instruction is programmed or configured to execute, through a processor, a kernel service transfer method for accelerating user space execution.
[0131] The present invention further provides a computer program product comprising a computer program / instruction programmed or configured to execute, by a processor, a kernel service transfer method for accelerating user space execution.
[0132] The system, medium and product of the present invention correspond to the above method and also have the advantages described in the above method.
[0133] The present invention can implement all or part of the process steps in the above-described method embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable media include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Memory is used to store computer programs and / or modules. The processor implements various functions by running or executing the computer programs and / or modules stored in the memory and accessing data stored in the memory. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0134] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A kernel service transfer method for accelerating user space execution, characterized in that: include: Step S1: When a program needs to process a kernel service, it is determined whether the kernel has enabled the kernel service transfer function. If so, a process kernel service transfer management structure creation operation is performed to create a process kernel service transfer management structure. The process kernel service transfer management structure is used to create and manage an acceleration process capable of executing the kernel service transfer function. Step S2, initializing the process kernel service transfer management structure, obtaining a kernel service logic process for executing the kernel service logic in the acceleration process, a user logic process for executing the user logic in the acceleration process, and an idle execution logic management structure for assisting the kernel service transfer; Step S3: During the execution of the acceleration process, the process leaving the kernel is continuously monitored. If the leaving process is detected to be a kernel service logic process, a kernel stack switching operation is performed, so that the kernel service logic process obtains and switches to the kernel stack of the idle execution logic management structure, and the user logic process obtains and switches to the original kernel stack of the kernel service logic process and returns to the user space to execute the user logic; Step S4: During the execution of the acceleration process, the process entering the kernel is continuously monitored. When it is detected that the process entering the kernel is a user logic process, the kernel service transfer operation is executed, so that the kernel service logic process obtains the kernel processing related data in the user logic process and executes the kernel service logic transferred from the user logic process in the kernel.
2. The kernel service transfer method for accelerating user space execution according to claim 1, characterized in that: Step S2 includes: Create and initialize multiple process management structures in the kernel; Use one of the process management structures as the kernel logic management structure and create a corresponding process as the kernel service logic process, use the other process management structure as the user logic management structure and create a daemon process as the user logic process; The third process management structure is used as the idle execution logic management structure, an unused kernel stack is allocated, and the position where the process starts executing and the used kernel stack are set to the loop function and the created stack of the stage of the acceleration process for monitoring kernel service transfer respectively.
3. The kernel service transfer method for accelerating user space execution according to claim 2, characterized in that: After step S2 and before step S3, the following steps are included: When the initialization of multiple process management structures monitoring the acceleration process is completed, the kernel service logic process enters the loop function, and the user logic process executes the page table switching logic to switch the user logic mapping page table to the page table of the kernel service logic process; and after the switching is completed, the kernel service logic process leaves the loop function and enters the switching path to leave the kernel space and prepare to enter the user space.
4. The kernel service transfer method for accelerating user space execution according to claim 2, characterized in that: Step S3 includes: When it is detected that the process leaving the kernel is a kernel service logic process and the kernel service logic process is in a state of releasing the logic of its own kernel stack and about to leave the kernel, writing the system register into the local variable, thereby storing the first context information of the currently executing process in the kernel stack of the kernel service logic process; Execute kernel stack switching, write the current CPU context information into the kernel logic management structure, write the context information in the initialized idle execution logic management structure into the CPU, and make the kernel service logic process enter the loop function, so that the kernel stack of the kernel service logic process is switched to the kernel stack of the idle execution logic management structure; When it is monitored that the kernel service logic process has switched the kernel stack and entered the loop function, the context information of the CPU where the user logic process is running is written into the user logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored first context information is written into the user logic process, so that the user logic process obtains the original kernel stack of the kernel service logic process and returns to the user space to execute the user logic.
5. The kernel service transfer method for accelerating user space execution according to claim 2, characterized in that: Step S4 includes: When it is detected that the process entering the kernel is a user logic process and a kernel service is required, it is determined whether the kernel service logic process is in a state where the kernel stack has been switched and the loop function has been entered. If so, a kernel service transfer is performed, and the system register is written into a local variable, thereby storing the second context information of the currently executed process including the kernel service category; Write the current CPU context information into the kernel logic management structure, and write the user logic management structure context information into the CPU, so that the user logic process can obtain the original kernel stack to continue to execute the monitoring acceleration process transfer state value; When it is detected that the kernel stack of the kernel logic management structure is currently in an idle state, the context information of the current CPU is written into the idle execution logic management structure, the context information in the kernel logic management structure is written into the CPU, and the stored second context information is written into the kernel service logic process, thereby completing the kernel service transfer and enabling the kernel service logic process to execute the kernel service passed from the user logic process.
6. The kernel service transfer method for accelerating user space execution according to claim 5, characterized in that: If it is determined that the kernel service logic process is not in a state where the kernel stack has been switched and the loop function has not been entered, it indicates that an error has occurred in switching the accelerated process transfer state, and the kernel service transfer operation is refused.
7. The kernel service transfer method for accelerating user space execution according to claim 1, characterized in that: The steps S3 and S4 are executed alternately, so that after the kernel service logic process completes the execution of the transferred kernel service, the user logic is handed over to the user logic process through step S3 for continued execution, and after the kernel service occurs in the user logic process, the kernel service transfer is completed through step S4. Steps S3 and S4 are executed alternately until the acceleration process is completed or dies.
8. A kernel service transfer system for accelerating user space execution, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the kernel service transfer method for accelerating user space execution according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored therein, characterized in that: The computer program / instruction is programmed or configured to execute, through a processor, the kernel service transfer method for accelerating user space execution according to any one of claims 1 to 7.
10. A computer program product comprising a computer program / instructions, characterized in that The computer program / instruction is programmed or configured to execute, through a processor, the kernel service transfer method for accelerating user space execution as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Implementation method for general register reservation recovery
CN112540871A
User mode interrupt processing method, device, equipment and program product
CN119938246A