Performance optimization method for accelerator, and processor and storage medium
Patent Information
- Application Number
- PCT/IB2024/062979
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-17
AI Technical Summary
Due to the limited number of hardware interfaces, the built-in accelerators in existing processors cannot meet the acceleration performance requirements of user-state programs such as cloud computing in high concurrency scenarios.
By transforming the accelerator driver, process identification and thread identification are configured for work tasks in advance, and these identifications are used as labels for work tasks, and then the work tasks carrying tags are transmitted to the port provided by the accelerator, supporting threads under multiple processes to share the target port.
It improves the acceleration performance of the processor's built-in accelerator in high concurrency scenarios of user-state processes, and can provide acceleration services for a larger number of user-state processes at the same time.
Smart Images

Figure IB2024062979_17072025_PF_FP_ABST
Abstract
Description
[0001]A Performance Optimization Method, Processor, and Storage Medium for Accelerators. This disclosure claims priority to Chinese patent application No. 202311777277.0, filed with the Patent Office of the People's Republic of China on December 21, 2023, entitled "A Performance Optimization Method, Processor, and Storage Medium for Accelerators," the entire contents of which are incorporated herein by reference. Technical Field: This disclosure relates to the field of data processing technology, and more particularly to a performance optimization method, processor, and storage medium for accelerators. Background: With the continuous update and iteration of processors, the number of CPU cores in processors continues to increase, resulting in continuous improvement in processor performance. Compared to increasing the number of CPU cores, integrating an accelerator into a processor can more effectively achieve higher performance. Specifically, the dynamic load balancer (DLB) built into the processor can effectively balance the workload of each CPU core in the processor, thereby improving processor performance. However, these built-in accelerators typically only have a fixed number of hardware interfaces (ports) capable of connecting to user-mode programs. Consequently, when processors are used in scenarios with a large number of user-mode programs, such as cloud computing, the built-in accelerators cannot meet the acceleration performance requirements in these scenarios due to the limited number of hardware interfaces. SUMMARY OF THE INVENTION Various aspects of the present disclosure provide a performance optimization method for an accelerator, a processor, and a storage medium to improve the acceleration performance of the accelerator. An embodiment of the present disclosure provides a performance optimization method for an accelerator built into a processor. The method is applicable to a driver for the accelerator and includes: upon receiving a work task issued from user mode, determining a target process and corresponding target thread to which the work task belongs; assigning a process identifier and a thread identifier to the work task as a tag for the work task; and transmitting the tagged work task to a target port provided by the accelerator, thereby enabling threads under multiple processes to share the target port through the tag. An embodiment of the present disclosure also provides a processor comprising a memory and multiple CPU cores, and also comprising a built-in accelerator; the memory is used to store one or more computer instructions corresponding to a driver program for the accelerator; the accelerator is coupled to the memory and the multiple CPU cores, and the one or more CPU cores in the processor are used to execute the one or more computer instructions to perform the aforementioned performance optimization method for the accelerator.Embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions. When executed by one or more CPU cores, the computer instructions cause the one or more CPU cores to perform the aforementioned method for optimizing performance for an accelerator. Embodiments of the present disclosure also provide a computer program. When executed in a computer, the computer executes the aforementioned method for optimizing performance for an accelerator. In embodiments of the present disclosure, modifications are proposed for the driver of the accelerator built into the processor. Before delivering a work task issued in user mode to a port provided by the accelerator, the work task is pre-configured with a process identifier and a thread identifier, and these two identifiers serve as labels for the work task. The labeled work task is then delivered to the port provided by the accelerator. This labeling ensures isolation of work tasks between different processes within the accelerator, thereby enabling threads within multiple processes to share the accelerator port. Consequently, in this embodiment, the processor's built-in accelerator can simultaneously provide acceleration services to a greater number of user-mode processes, thereby providing higher acceleration performance in scenarios with high concurrency of user-mode processes. BRIEF DESCRIPTION OF THE DRAWINGS The accompanying drawings described herein are intended to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are intended to explain the present disclosure and do not constitute undue limitations thereon. In the accompanying drawings: Figure 1 is a flowchart of a performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure; Figure 2 is a logic diagram of a performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure; Figure 3a is a flowchart of another performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure; Figure 3b is a logic diagram of another performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure; Figure 4 is a partial schematic diagram of a buffer ring provided in an exemplary embodiment of the present disclosure; Figure 5 is a logic diagram of a performance optimization method for a dynamic load balancer (DLB) provided in an exemplary embodiment of the present disclosure; Figure 6 is a logic diagram of another performance optimization method for a dynamic load balancer (DLB) provided in an exemplary embodiment of the present disclosure; and Figure 7 is a schematic diagram of the structure of a processor provided in another exemplary embodiment of the present disclosure. DETAILED DESCRIPTION To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of this disclosure, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.As described in the background, the number of CPU cores in processors continues to increase. Processors containing multiple CPU cores are generally referred to as multi-core processors. Accelerators built into processors can communicate with multiple CPU cores in a processor via a PCI / PCIe bus. In other words, these accelerators can function as PCI / PCIe devices. During research, the inventors discovered that processor-built-in accelerators typically only provide a limited number of ports in hardware, which serve as interfaces for data communication between the accelerator and the CPU cores. Furthermore, many accelerators only support exclusive use by a single user-mode process, resulting in these accelerators failing to provide adequate acceleration performance in scenarios with high user-mode process concurrency, such as cloud computing. Therefore, embodiments of the present disclosure propose a performance optimization method for accelerators. This improves existing accelerators that only support exclusive use of ports by a single user-mode process, thereby enhancing their acceleration performance in scenarios with high user-mode process concurrency. The following, combined with the accompanying drawings, details the technical solutions provided by various embodiments of the present disclosure. Figure 1 is a flow chart of a performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure. Figure 2 is a logical diagram of a performance optimization method for an accelerator provided in an exemplary embodiment of the present disclosure. Referring to Figure 1 , this method can be applied to an accelerator driver, which runs in kernel mode on a processor. The method may include: Step 100: After receiving a work task sent from user mode, determining the target process and corresponding target thread to which the work task belongs; Step 101: Configuring a process identifier and a thread identifier for the work task as a label for the work task; Step 102: Transmitting the labeled work task to a target port provided by the accelerator to support sharing of the target port by threads under multiple processes through the label. In this embodiment, there is no limitation on the type of accelerator; various accelerators built into the processor can be optimized as needed in this embodiment. Accelerator types may include, but are not limited to, a dynamic load balancer (DLB) or a data protection and compression accelerator (QAT), and further examples are not provided here. The Dynamic Load Balancer (DLB) is a CPU accelerator that implements dynamic load balancing based on hardware, balancing workloads across multiple CPU cores. The Data Protection and Compression Accelerator (QAT) accelerates encryption, decryption, and data compression, helping to reduce system resource consumption by offloading these tasks from the CPU cores.In this embodiment, workloads can be issued from user mode by different virtual machines, containers, or applications. These virtual machines, containers, or applications are typically implemented as user mode processes, and workloads can be issued by threads within these user mode processes. It should be understood that a user mode process can contain multiple threads. For accelerators that only support exclusive port use by a single user mode process, a single port in the accelerator only supports the delivery of workloads associated with that single user mode process. Once all ports in the accelerator are exclusively used by a user mode process, the accelerator can no longer provide acceleration services for workloads associated with other user mode processes. In other words, these accelerators currently use exclusive ports to isolate workloads associated with different user mode processes. However, as mentioned above, this exclusive port approach severely limits the number of user mode processes that the accelerator can concurrently support, failing to meet the acceleration performance requirements in scenarios with high concurrency of user mode processes. In this embodiment, the accelerator driver is modified. In step 100, after receiving a work task from user state, it is not directly delivered to the accelerator port. Instead, the driver first determines the target process and corresponding target thread to which the work task belongs. In actual applications, work tasks issued from user state first reach the accelerator's corresponding driver in kernel state. In this embodiment, the process from user state issuing a work task to the kernel state driver receiving the work task is not modified. Therefore, from the user state perspective, the user state processes and threads can still issue work tasks according to the original communication protocol and are unaware of the performance optimization scheme in this embodiment. The interactive data provided by user state to kernel state already carries information such as the data source. Therefore, in this embodiment, the driver can seamlessly determine the target process and corresponding target thread to which the work task belongs. Continuing with Figures 1 and 2, in step 101, the driver can configure a process ID and thread ID for the work task. The driver can assign different process IDs to different user state processes and different thread IDs to different threads within a single user state process. That is, in this embodiment, process IDs can be used to distinguish different user-mode processes. Within a single user-mode process, different thread IDs can be used to distinguish different threads. In this embodiment, various implementations can be used to pre-allocate process and thread IDs, which are not limited in this embodiment. The only requirement is that the allocated IDs can achieve the aforementioned purpose.An exemplary process ID allocation scheme is provided herein: the driver can use the PAS ID (Process Address Space ID) configured for different user-mode processes in the bus protocol. This is a hardware-supported virtual address space isolation technology that can isolate virtual address spaces between different user-mode processes, thereby achieving higher security and reliability. Of course, it should be understood that in addition to using the PAS ID of a user-mode process as the process ID, in this embodiment, the driver can configure other customized process IDs for different user-mode processes. Further examples of such schemes are not provided here. An exemplary thread ID allocation scheme is provided herein: the driver can incrementally number threads within a single user-mode process based on the number of threads that need to share the same port, thereby generating thread IDs for each thread within the user-mode process. In specific implementation, the driver can run a software-implemented counter, called Thread Counter, for each user-mode process. This counter can be used to label threads based on the number of threads within the single user-mode process that need to share the same port. The labels can be represented by TIDs. For example, if two threads under APP1 need to share the same port, the TIDs of these two processes under APP1 can be configured as 1 and 2 respectively. O For another example, if one thread under APP2 exclusively occupies the port, the TID of this thread under APP2 T can be configured as 1. Of course, it should be understood that in addition to the incremental numbering method used above to generate thread identifiers, in this embodiment, the driver can configure other customized thread identifiers for different threads under a single user-mode process. Further examples of such solutions are not provided here. Based on this, each work task received by the driver can obtain a process identifier and a thread identifier. In this embodiment, the process identifier and thread identifier can be used as the work task label. That is, in this embodiment, the work task label will include two layers of identifiers: the process identifier and the thread identifier. For example, the label of work task A can be exemplarily represented as: PAS ID=1:TID=1; the label of work task B can be exemplarily represented as: PAS ID=1:TID=2 OBased on the tags of these two work tasks, we can see that they originate from different threads (TID=1 and TID=2) in the same user-mode process (PAS ID=1). Continuing with Figures 1 and 2, in step 102, the tagged work task can be transferred to the target port provided by the accelerator, enabling threads under multiple processes to share the target port through the tag. That is, in this embodiment, after the tag for the work task is configured, the work task is delivered to the port provided by the accelerator. In this way, the tag is delivered along with the work task to the port provided by the accelerator. In this embodiment, the driver can write the process ID and thread ID into a reserved field in the data structure corresponding to the work task. Typically, this reserved field is 8 bits long. In this embodiment, the process ID can occupy 4 bits, and the thread ID can occupy 4 bits. Of course, the field length allocation method is not limited to this and can be adjusted according to actual needs. In this embodiment, the other fields in the data structure corresponding to the work task remain unchanged. We will not provide detailed examples of the various fields in the data structure corresponding to a work task. During research, the inventors discovered that accelerators are typically equipped with a hardware work queue to manage received work tasks. However, the accelerator's hardware work queue lacks hardware resource isolation capabilities, meaning it cannot isolate work tasks associated with different user-mode processes. This is why accelerators can only support exclusive port use by a single user-mode process. This embodiment proposes including a tag in the work task and delivering the tagged work task to the accelerator's port. The driver then directs the accelerator's hardware to manage the work task in the hardware work queue based on the tag. This allows the accelerator to support work tasks associated with multiple user-mode processes sharing the accelerator's hardware work queue. This allows the accelerator to transform the accelerator's hardware work queue, which originally supported exclusive use, into a shared work queue, enabling threads from multiple processes to share the same port. In summary, this embodiment proposes modifying the driver for the built-in accelerator in the processor. Before delivering a work task issued in user mode to a port provided by the accelerator, the process ID and thread ID are preconfigured for the work task, and these two IDs serve as labels for the work task. The labeled work task is then delivered to the port provided by the accelerator. This labeling ensures isolation of work tasks between different processes within the accelerator, enabling threads under multiple processes to share the accelerator port.Accordingly, in this embodiment, the accelerator built into the processor can support simultaneous acceleration services for a greater number of user-mode processes, thereby providing higher acceleration performance in scenarios with high concurrency of user-mode processes. Figure 3a is a flow chart illustrating another method for optimizing performance for an accelerator, provided in accordance with an exemplary embodiment of the present disclosure. Figure 3b is a logic diagram illustrating another method for optimizing performance for an accelerator, provided in accordance with an exemplary embodiment of the present disclosure. Referring to Figure 3a, the method may include: Step 300: Upon receiving a work task issued by a user-mode process, determining the target process and corresponding target thread to which the work task belongs; Step 301: Configuring a process identifier and a thread identifier for the work task as a tag for the work task; Step 302: Storing the tagged work task in a buffer ring configured in kernel mode; Step 303: Reading the work task and its tag from the buffer ring and transmitting them to a target port provided by the accelerator, thereby enabling threads under multiple processes to share the target port through the tag. Steps 300 and 301 can be referred to in the relevant descriptions of the previous embodiments and are not repeated here. This embodiment provides an optional implementation of the work task delivery process based on steps 302 and 303. This optional implementation can be combined with the technical content of the preceding or subsequent embodiments to obtain technical solutions with various protection scopes. Referring to Figures 3a and 3b, in step 302, after configuring a tag for the work task, the tagged work task can be first stored in a ring buffer configured in kernel mode. This allows the driver to consolidate work tasks issued from user mode within the ring buffer. Figure 4 is a partial schematic diagram of a ring buffer provided in an exemplary embodiment of the present disclosure. Referring to Figure 4, in a preferred implementation, a lockless ring buffer can be used in step 302. Furthermore, the driver can allocate non-overlapping buffer address spaces within the ring buffer for different process IDs. Based on this, in step 302, specifically: based on the process ID contained in the tag carried by the work task, a target buffer address space corresponding to the process ID can be determined from the ring buffer; and the work task can be stored in the target buffer address space. That is, in the buffer ring, the work tasks involved in different user-mode processes will be written by the driver to different buffer address spaces. This ensures that the write operations initiated by the driver for the work tasks involved in different user-mode processes do not interfere with each other. Therefore, the isolation between user-mode processes can be guaranteed through the buffer address space, making it unnecessary to lock the buffer ring.In practical applications, when a driver performs write operations within a single buffer address space, it can use a sequential write method and record the latest write pointer after each write operation. After the write pointer reaches the end of the buffer address space, it can return to the beginning of the buffer address space to perform the write operation. This ensures the security of write operations within each buffer address space. Continuing with Figures 3a and 3b, the driver can read tagged work tasks from the buffer ring and deliver them to the accelerator's port. Referring to Figure 4, continuing with the aforementioned lock-free buffer ring, the driver can poll each buffer address space when reading work tasks and maintain a read pointer for each buffer address space. The driver can use a sequential read method for read operations within a single buffer address space, pointing the read pointer to the work task closest to the beginning of the buffer address space. This ensures that read operations targeting different buffer address spaces do not interfere with each other. Based on the lock-free buffer ring, in this embodiment, the driver can perform concurrent read / write operations on the buffer ring, thereby effectively improving the efficiency of delivering work tasks to the accelerator's port. This embodiment proposes that a driver maintain a buffer ring in kernel mode to consolidate the workloads of different user-mode processes. Furthermore, it proposes supporting the use of a lock-free buffer ring. The driver pre-allocates non-overlapping buffer address spaces within the lock-free buffer ring for different user-mode processes and writes the workloads of different user-mode processes into separate buffer address spaces. This ensures the isolation of workloads of different user-mode processes and enables the driver to concurrently read and write workloads from the buffer ring, effectively improving workload delivery efficiency. It should be understood that the buffer ring provided in this embodiment is merely an optional implementation for storing workloads in kernel mode and is not limited to this embodiment. For example, a software shared queue can also be used to consolidate workloads of different user-mode processes. Further examples are not provided here. As mentioned above, the accelerator performance optimization method provided in this embodiment is applicable to various accelerators, particularly accelerators that only support a single user-mode process with exclusive port access. During their research, the inventors discovered that different accelerators have different internal task management logic. In this embodiment, after a tagged task is delivered to an accelerator port, the accelerator's internal hardware, driven by a driver, manages the task based on the tag. This embodiment uses a dynamic load balancer (DLB) as an example to modify the task management logic within the DLB to further optimize its acceleration performance.Figure 5 is a logical diagram of a performance optimization for a dynamic load balancer (DLB) according to an exemplary embodiment of the present disclosure. Referring to Figure 5 , the dynamic load balancer (DLB) may include three types of ports: send ports, consumer ports, and receive ports. As mentioned above, these ports are used for data communication between the dynamic load balancer (DLB) and the CPU cores. More specifically, they are used for data communication between the dynamic load balancer (DLB) and its corresponding driver. For the dynamic load balancer (DLB), its driver can deliver tagged workloads to the aforementioned send ports. In practical applications, the dynamic load balancer (DLB) may include multiple send ports. For each target send port, the driver can deliver workloads related to multiple user-mode processes to that target send port, allowing threads within multiple user-mode processes to share the target send port. Referring to Figure 5 , the dynamic load balancer (DLB) provides a hardware work queue that connects producers and consumers. Producers are user-mode processes and their threads, while consumers are CPU cores within the processor. The send port serves as the hardware entry point for work tasks provided by the accelerator. A single send port can be associated with multiple hardware work queues. That is, from the perspective of the target send port, after work tasks generated in user mode are delivered to the target send port, the accelerator hardware can distribute these work tasks evenly across the multiple hardware work queues associated with the target send port according to a load balancing policy. In traditional load balancing policies, the accelerator hardware typically distributes work tasks based solely on quantity balance. Specifically, the hardware ensures a balanced number of work tasks written to the multiple hardware work queues associated with the target send port. However, in this embodiment, through driver modifications, the driver can drive the accelerator hardware to distribute work tasks to the multiple hardware work queues associated with the target send port based on both tag and quantity. This ensures a balanced number of work tasks received by the multiple hardware work queues associated with the target send port, and the tags involved are also balanced. Label balancing means that the types of labels assigned to multiple hardware work queues are basically the same, the number of work tasks assigned to each label is also basically the same, and the number of work tasks under different labels in the same hardware work queue is also basically the same.For example, if APP1 and APP2 share the target send port, continuing with the previous example, APP1 is associated with two tags, denoted as Tag 1 (PAS ID=1: TID=1) and Tag 2 (PAS ID=1: TID=2); APP2 is associated with one tag, denoted as Tag 3 (PAS ID=2: TID=1). Referring to Figure 5, the three hardware work queues associated with the target send port are all assigned these three tags. Furthermore, the number of tasks assigned to each tag in the three hardware work queues remains consistent. For example, for Tag 2, two tasks carrying Tag 2 are assigned to each of the three hardware work queues, maintaining a balanced number of tasks. This allows the driver to drive the accelerator to evenly distribute the tasks delivered to the target send port to the multiple hardware work queues associated with the target send port based on the tags, thereby ensuring a balanced distribution of tags across the multiple hardware work queues associated with the target send port. Referring to Figure 5 , the consumer ports provided by the accelerator correspond one-to-one with the hardware work queues. Through the aforementioned load balancing process, work tasks delivered to the target sending port can be evenly distributed to the multiple consumer ports associated with the target sending port based on their tags, achieving work task balance among these multiple consumer ports. From the perspective of a single consumer port associated with the target sending port, the work tasks assigned to it involve threads under multiple user-mode processes. In other words, this embodiment supports multiple user-mode processes sharing the same consumer port. In the dynamic load balancer (DLB), the hardware work queue is lock-free, and a single consumer port only supports connecting to a single processing thread, provided by the CPU core in the processor. This exclusive connection ensures the thread safety of the hardware work queue, preventing uncontrollable security issues caused by multiple processing threads enqueuing and dequeuing the same hardware work queue. This embodiment continues to use this exclusive connection approach, assigning work tasks in a single hardware work queue to the same processing thread for processing. Furthermore, this embodiment proposes an improved solution that breaks the traditional exclusive connection approach in a dynamic load balancer (DLB). For ease of description, the following detailed description of the improved solution will use the target consumer port among multiple consumer ports associated with the target sending port as an example. It should be understood that this improved solution is also applicable to other consumer ports. Referring to Figure 5, in this improved solution, the driver can copy the work tasks and the tags carried by the hardware work queue corresponding to the target consumer port to kernel state.In kernel mode, the driver can isolate and store work queues with different tags associated with the target consumer port. For example, this can be done using a ring buffer. Alternatively, work queues with different tags can be stored in separate files, etc., to achieve isolation. This is not specifically limited here. The driver can also assign different processing threads to different tags associated with the target consumer port. When consuming from the hardware work queue, the processing threads can consume based on the tag. This ensures that a single processing thread only consumes work tasks with a single tag in the hardware work queue associated with the target consumer port. This is equivalent to splitting a single hardware work queue into multiple sub-queues based on the tag, while a single processing thread is equivalent to connecting to a single sub-queue. From a security perspective, a single processing thread is still connected to a single queue, thus preventing thread safety issues. For example, referring to Figure 5, the hardware work queue associated with consumer port 1 involves three tags. The driver can assign different processing threads to each of the three tags, so that each of the three processing threads consumes work tasks with one tag. As mentioned above, after load balancing, the number of work tasks under different tags in the same hardware work queue is essentially the same, which ensures that the number of work tasks assigned to the three processing threads is essentially the same. Furthermore, after load balancing, the number of tags assigned to each consumer port associated with the target sending port is essentially the same. Therefore, the number of processing threads connected to each consumer port associated with the target sending port is essentially the same. This ensures that the number of processing threads connected to each consumer port associated with the target sending port is essentially the same, and the number of work tasks assigned to each processing thread is essentially the same, effectively ensuring load balancing for the CPU cores associated with each consumer port. In this improved solution, various solutions can be adopted to assign different processing threads to different tags. In one exemplary solution, the driver can open different file descriptors (eventfd) for different tags and associate different file descriptors (eventfd) with different processing threads, allowing the processing threads to monitor whether new work tasks under the corresponding tags have been added based on the file descriptors (eventfd). The file descriptor (eventfd) is a technology that operates memory as a virtual file. The file descriptor eventfd is a file interface exposed by a user-mode process or thread for data exchange with kernel drivers. Through the file descriptor eventfd, the processing thread can detect whether the corresponding virtual file has been read or written.By opening different file descriptors (eventfd) for different tags, tasks under different tags can be treated as the contents of different files. Thus, after associating different file descriptors (eventfd) with different processing threads, the processing threads can detect whether new tasks have been added to the tags corresponding to their associated file descriptors (eventfd). In response to the driver's operation of copying the tasks and tags contained in the hardware work queue corresponding to the target consumer port to kernel state, when new tasks are copied to kernel state during this operation, the corresponding processing thread can detect this new event and proactively read the new tasks from kernel state for timely processing. It should be understood that the aforementioned file descriptor (eventfd) is merely an exemplary solution; other solutions can also be employed in this improved solution to assign different processing threads to different tags. For example, independent I / O channels can be established for different tags, with different processing threads having exclusive access to the I / O channels corresponding to their assigned tags. Further examples are not provided here. At this point, the tasks have been delivered to the CPU cores in the processor. After completing a task, the CPU core, specifically the processing thread provided by the CPU core, will feedback the processing results to the driver. It should be understood that the processing thread is unaware of the tags configured by the driver for the task in this embodiment. Therefore, the processing results provided by the processing thread no longer carry tags. To address this, this embodiment can continue to utilize the existing processing result delivery method in the dynamic load balancer (DLB): a single user-mode process exclusively occupies a receive port. Based on this, after receiving the processing result provided by the processing thread, the driver determines the user-mode process to which the processing result belongs. The processing result is then distributed to the hardware receive queue corresponding to the user-mode process, which corresponds to a receive port. The processing result in the hardware receive queue associated with the target receive port is then copied to kernel mode and sent to the user-mode process corresponding to the target receive port. In other words, the processing results associated with a single user-mode process are delivered to the receive port exclusively occupied by that user-mode process. The driver then simply delivers the processing result in the receive port to the corresponding user-mode process, allowing the user-mode process to continue processing. FIG6 is another logical diagram of another embodiment of the present disclosure, provided for optimizing the performance of a dynamic load balancer (DLB). Referring to FIG6 , this embodiment also provides an improved method for delivering processing results: multiple user-mode processes share the same receiving port.Based on this, a receiving port can correspond to a sending port. Again, using the target sending port as an example, in this improved delivery method: the driver can assign tags corresponding to the corresponding work tasks to the processing results provided by the processing threads in the CPU cores associated with each consuming port corresponding to the target sending port; send the tagged processing results to the target hardware receive queue associated with the target receiving port; retrieve the processing results and corresponding tags from the target hardware receive queue via the target receiving port; and distribute the processing results to the threads under the corresponding user-mode processes based on the tags. Referring to Figure 6, the tagged processing results can be stored in another kernel-mode ring buffer, awaiting delivery to the hardware receive queue. In other words, in this improved processing result delivery method, the processing results are assigned tags corresponding to the corresponding work tasks. This tag-based approach allows processing results from multiple user-mode processes to share the same hardware receive queue, thereby enabling multiple user-mode processes to share the same receiving port. The implementation of this sharing is essentially consistent with the aforementioned concept of sharing a hardware work queue for work tasks from multiple user-mode processes, and will not be elaborated upon here. Furthermore, under this processing result delivery method, during the driver's delivery to user mode, an exemplary solution involves the driver opening different file descriptors (eventfd) for different tags. The file descriptors (eventfd) are associated with threads in the user-mode process pointed to by their corresponding tags, allowing each thread to monitor whether new processing results for the corresponding tags have been added based on the file descriptors (eventfd). This allows each thread initiating a task to obtain the file descriptor (eventfd). Based on this file descriptor, the thread can detect whether the task it initiated has generated the corresponding processing result, thereby promptly and accurately understanding the task's processing status. In summary, in this embodiment, after transmitting tagged tasks to ports provided by the accelerator to provide a basis for port sharing, the driver corresponding to the dynamic load balancer (DLB) is further modified to enable more efficient and accurate task management within the dynamic load balancer (DLB). After further modification, it is possible to ensure load balancing among multiple consumer ports associated with a sending port, based on the fact that multiple user-mode processes share the same sending port. Moreover, it is also possible to support a single consumer port to connect to multiple processing threads, thereby supporting multiple processing threads to concurrently consume work tasks under the same consumer port, improving the consumption efficiency of work tasks, and thereby enhancing the acceleration performance of the dynamic load balancer DLB.It should be noted that some of the processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. Operation numbers, such as 101 and 102, are merely used to distinguish between different operations and do not represent any specific execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. FIG7 is a schematic diagram of the structure of a processor provided in another exemplary embodiment of the present disclosure. As shown in FIG7 , the processor includes a memory 70 and multiple CPU cores 71, as well as a built-in accelerator 72. The memory 70 is configured to store one or more computer instructions corresponding to a driver program for the accelerator 72. The accelerator 72 is coupled to the memory 70 and multiple CPU cores 71. One or more CPU cores 71 in the processor are configured to execute the one or more computer instructions, thereby utilizing the driver program to perform the following processing logic: upon receiving a work task issued from user mode, determining the target process and corresponding target thread to which the work task belongs; configuring a process identifier and a thread identifier for the work task as a tag for the work task; and transmitting the tagged work task to a target port provided by the accelerator to support sharing of the target port by threads under multiple processes through the tag. In an optional embodiment, the driver program may further be configured to store the tagged work task in a buffer ring configured in kernel mode. Transmitting the tagged work task to the target port provided by the accelerator includes reading the work task and its tag from the buffer ring and transmitting them to the target port. In an optional embodiment, the buffer ring is lockless. When the driver stores the tagged work task in a buffer ring configured in kernel mode, the driver may specifically: determine, based on the process ID included in the tag, a target buffer address space corresponding to the process ID from the buffer ring; and store the work task in the target buffer address space; wherein the buffer ring allocates non-overlapping buffer address spaces for different process IDs. In an optional embodiment, when the driver configures the process ID and thread ID for the work task, the driver may: write the process ID and thread ID into reserved fields in a data structure corresponding to the work task.In an optional embodiment, the accelerator 72 is configured to perform load balancing on the processor. When the driver transmits a tagged work task to a target port provided by the accelerator to support sharing of the target port by threads under multiple processes through the tag, the driver may specifically: transmit the tagged work task to a target sending port provided by the accelerator to support sharing of the target sending port by threads under multiple processes through the tag; the target sending port is configured to serve as a hardware entry for the work task. In an optional embodiment, the target sending port is associated with multiple hardware work queues. The driver may further: after transmitting the tagged work task to the target sending port provided by the accelerator, drive the accelerator to perform load balancing based on the tag carried by the work task to ensure a balanced distribution of tags among the multiple hardware work queues. In an optional embodiment, the accelerator 72 further provides a consumer port that corresponds one-to-one with a hardware work queue. The driver can also be configured to: copy the work tasks and tags contained in the hardware work queue corresponding to the target consumer port to kernel state; assign different processing threads to different tags, so that the processing threads consume the work tasks generated under the assigned tags; wherein the processing threads assigned to different tags are provided by the CPU core 71 associated with the target consumer port. In an optional embodiment, when assigning different processing threads to different tags, the driver can also be configured to: open different file descriptors (eventfd) for different tags: associate different file descriptors (eventfd) with different processing threads, so that the processing threads can monitor whether new work tasks under the corresponding tags have been added based on the file descriptors (eventfd). In an optional embodiment, the accelerator 72 further provides a receiving port corresponding to a user-mode process, and the driver may be further configured to: upon receiving a processing result provided by a processing thread, determine the user-mode process to which the processing result belongs; distribute the processing result to a hardware receiving queue corresponding to the user-mode process to which the processing result belongs, where the hardware receiving queue corresponds one-to-one with the receiving port; and copy the processing result in the hardware receiving queue associated with the target receiving port to the kernel mode and send the result to the user-mode process corresponding to the target receiving port.In an optional embodiment, the accelerator 72 also provides a target receiving port corresponding to the target sending port. The driver may further be configured to: assign a tag corresponding to the corresponding work task to each processing result provided by the processing thread in the CPU core associated with each consumer port corresponding to the target sending port; send the tagged processing result to the target hardware receive queue associated with the target receiving port; retrieve the processing result and the corresponding tag from the target hardware receive queue via the target receiving port; and distribute the processing result to the thread in the corresponding user-mode process according to the tag. In an optional embodiment, when distributing the processing result to the thread in the corresponding user-mode process according to the tag, the driver may be configured to: open different file descriptors (eventfd) for different tags: associate the file descriptor (eventfd) with the thread in the user-mode process pointed to by the corresponding tag, so that each thread can monitor whether a new processing result under the corresponding tag has been added based on the file descriptor (eventfd). Furthermore, FIG. 7 only schematically illustrates some components, and does not mean that the processor only includes the components shown in FIG. 7. It is worth noting that the technical details of the processor embodiments described above can be found in the relevant descriptions of the aforementioned method embodiments. To save space, they will not be repeated here, but this should not diminish the scope of protection of the present disclosure. Accordingly, embodiments of the present disclosure also provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more CPU cores, the one or more CPU cores execute the steps of the aforementioned method embodiments. Accordingly, embodiments of the present disclosure also provide a computer program. When the computer program is executed on a computer, it causes the computer to execute the steps of the aforementioned method embodiments. Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code. The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions, when executed by the processor of the computer or other programmable data processing device, produce means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means, which implement the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram. It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or device. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of other identical elements in the process, method, product, or device comprising the element. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) referred to in this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or deny. The foregoing description is merely an example of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this disclosure should be included in the scope of protection of this disclosure.
Claims
Claims 1. A performance optimization method for an accelerator, wherein the accelerator is built in a processor, and the method is applicable to a driver of the accelerator, and the method comprises: After receiving a work task sent from the user state, determining a target process to which the work task belongs and a corresponding target thread; A process identifier and a thread identifier are configured for the work task as a label of the work task; and the work task carrying the label is transmitted to a target port provided by the accelerator to support threads under multiple processes to share the target port through the label.
2. The method according to claim 1, further comprising: Storing the work task carrying the label in a buffer ring configured in the kernel state; The transmitting the work task carrying the tag to the target port provided by the accelerator includes: reading the work task and its tag from the buffer ring and transmitting them to the target port.
3. The method according to claim 2, wherein: The buffer ring is lock-free, and storing the work task carrying the label in the buffer ring configured in the kernel state includes: determining a target buffer address space corresponding to the process identifier from the buffer ring based on the process identifier contained in the label; storing the work task in the target buffer address space; wherein non-overlapping buffer address spaces are allocated to different process identifiers in the buffer ring.
4. The method according to any one of claims 1 to 3, wherein: Configuring a process identifier and a thread identifier for the work task includes: writing the process identifier and the thread identifier into a reserved field in a data structure corresponding to the work task.
5. The method according to any one of claims 1 to 4, wherein: The accelerator is used to load balance the processor, and transmitting the work task carrying the label to the target port provided by the accelerator to support threads under multiple processes to share the target port through the label, including: transmitting the work task carrying the label to the target sending port provided by the accelerator to support threads under multiple processes to share the target sending port through the label; wherein the target sending port is used as a hardware entrance for the work task.
6. The method according to claim 5, wherein: The target sending port is associated with multiple hardware work queues, and the method further includes: after transmitting the work task carrying the label to the target sending port provided by the accelerator, driving the accelerator to perform a load balancing operation based on the label carried by the work task, so that the labels allocated to the multiple hardware work queues remain balanced.
7. The method according to claim 5 or 6, wherein: The accelerator also provides a consumption port corresponding to the hardware work queue one by one, and the method further includes: copying the work tasks and the tags carried by the hardware work queue corresponding to the target consumption port to the kernel state; assigning different processing threads to different tags so that the processing threads consume the work tasks occurring under the tags assigned to them; wherein the processing threads assigned to different tags are provided by the CPU core associated with the target consumption port.
8. The method according to claim 7, wherein: Assign different processing threads to different tags, including: Open different file descriptors eventfd for different tags: Associate different file descriptors eventfd with different processing threads, so that the processing threads can process the file descriptors based on the eventfd. eventfd monitors whether new work tasks under the corresponding label are added.
9. The method according to claim 7 or 8, wherein: The accelerator also provides a receiving port corresponding to the user state process. The method further includes: after receiving the processing result provided by the processing thread, determining the user state process to which the processing result belongs; distributing the processing result to the hardware receiving queue corresponding to the user state process to which it belongs, and the hardware receiving queue corresponds to the receiving port one by one; copying the processing result in the hardware receiving queue associated with the target receiving port to the kernel state and sending it to the user state process corresponding to the target receiving port.
10. The method according to claim 7 or 8, wherein: The accelerator also provides a target receiving port corresponding to the target sending port, and the method further includes: configuring labels consistent with the corresponding work tasks for the processing results provided by the processing threads in the CPU cores associated with each of the consumption ports corresponding to the target sending port; sending the processing results carrying the labels to the target hardware receiving queue associated with the target receiving port; obtaining the processing results and the corresponding labels from the target hardware receiving queue through the target receiving port; and distributing the processing results to the threads under the corresponding user state processes according to the labels.
11. The method according to claim 10, wherein: Distribute the processing results to the threads under the corresponding user-mode process according to the labels, including: Open different file descriptors eventfd for different labels: Associate the file descriptor eventfd with the thread under the user-mode process pointed to by its corresponding label, so that each thread can monitor whether there is any new processing result under the corresponding label based on the file descriptor eventfd.
12. A processor, comprising a memory and multiple CPU cores, and also comprising a built-in accelerator; the memory is used to store one or more computer instructions corresponding to a driver of the accelerator; the accelerator is coupled to the memory and the multiple CPU cores, and one or more CPU cores in the processor are used to execute the one or more computer instructions to execute the performance optimization method for the accelerator according to any one of claims 1 to 11.
13. A computer-readable storage medium storing computer instructions, which, when executed by one or more CPU cores, causes the one or more CPU cores to execute the performance optimization method for an accelerator according to any one of claims 1 to 1.
14. A computer program, when executed in a computer, causes the computer to execute the performance optimization method for an accelerator according to any one of claims 1 to 11.
Citation Information
Patent Citations
Data processing systems
US20150089495A1
Method and system for accelerator thread management
US20220121493A1
Interface for multiple processors
US20220276914A1