Thread processing method and device, computing equipment and computer readable storage medium
By introducing two types of threads into the computing node and dynamically adjusting the thread execution strategy according to the processor resource utilization, the delay overhead problem caused by thread context switching in the computing node is solved, and the performance of the processor core is improved.
Patent Information
- Application Number
- CN202311790801.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
In the hypercomputing center scenario, the computing nodes lead to degradation of latency overhead and processor performance during kernel-level context switching between application software and platform software.
By introducing two types of threads, the first type of thread is the user thread of the platform software, and the second type of thread is the user thread of the application software. The first executor of the computing node dynamically adjusts the thread execution strategy according to the processor resource utilization. If the resource utilization is less than the threshold, the first type of thread will be executed. If the resource utilization is greater than or equal to the threshold, the execution of the first type of thread will be reduced.
Reduces the number of context switching between threads of the processor core, reduces the delay overhead caused by context switching, thereby improving the performance of the processor core.
Smart Images

Figure CN120196426A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a method and apparatus for processing threads, a computing device, and a computer-readable storage medium. Background Art
[0002] In the scenario of a supercomputing center, computing nodes, as the main computing power supply side, are responsible for running application software and platform software, etc. During the process of running application software, computing nodes will make full use of computing resources according to their own specifications. For example, on the premise of ensuring that the tasks of platform software can be completed normally, computing nodes jointly use the computing resources of the computing nodes between application software and platform software to release sufficient computing power for the operation of application software.
[0003] In the related art, the process of jointly using computing resources is as follows: Application software and platform software can run on the same processor core. If the resource utilization rate of the processor core is too high, by performing a context switch at the kernel level between application software and platform software, the user threads of the platform software are suspended, and the user threads of the application software are executed to release the processor resources occupied by the platform software and provide more processor resources for the operation of the application software.
[0004] However, during the context switch process, it is necessary to store the context information of the user threads of the platform software from the registers in the processor to the memory, and store the context information of the user threads of the application software in the registers, resulting in latency overhead and reducing the performance of the processor core. Summary of the Invention
[0005] Embodiments of this application provide a method and apparatus for processing threads, a computing device, and a computer-readable storage medium, which can improve the performance of the processing cores in computing nodes. The technical solution is as follows:
[0006] In a first aspect, a method for processing threads is provided. This method involves two types of threads, namely the first type of thread and the second type of thread. Among them, the first type of thread refers to the user threads of platform software, and the second type of thread refers to the user threads of application software. This method is executed by a first executor in a computing node. The first executor is associated with a first processing core in the computing node. The first processor core is associated with the first type of thread and the second type of thread. In the method, first obtain the first resource utilization rate of the first processor core; if the first resource utilization rate is less than a first threshold, use the first processing core to execute the first type of thread associated with the first processor core; if the first resource utilization rate is greater than or equal to the first threshold, reduce the first type of threads to be executed in a future time period.
[0007] When the resource utilization rate of the first processor is less than the first threshold, the first processor core is used to execute the associated first type of threads. When the resource utilization rate of the first processor is greater than or equal to the first threshold, by reducing the number of first type of threads to be executed in the future time period, the number of context switches of the first processor core between the first type of threads and the second type of threads in the future time period is reduced, avoiding frequent context switches of the first processor core, reducing the latency overhead caused by context switches, and thus being able to provide the performance of the processor core.
[0008] In a possible implementation, the above-mentioned reduction of the first type of threads to be executed in the future time period can be achieved through the following steps: scheduling at least one target thread to the second processor core in the computing node, and / or entering the sleep state, where the target thread is a to-be-executed thread among the first type of threads associated with the first processor core.
[0009] Based on the above possible implementation, by scheduling the target thread from the first processor core to the second processor core, the first executor uses the first processor core in the future time period to execute the first type of threads that have not been scheduled out, thereby being able to reduce the first type of threads to be executed in the future time period, reducing the resource occupancy of the first type of threads on the first processor core in the future time period, and enabling the first processor core to provide more processor resources for the associated second type of threads, so as to improve the processing efficiency of the first processor core for the second type of threads.
[0010] In a possible implementation, the resource utilization rate of the second processor core is less than or equal to the second threshold, and / or the second executor associated with the second processor core is in the sleep state, where the second executor is used to execute the first type of threads associated with the second processor core using the second processor core.
[0011] Based on the above possible implementation, the first executor does not use the resources of the first processor core to execute the first type of threads during the sleep period, thereby being able to reduce the first type of threads to be executed in the future time period, and during the sleep period, it can also avoid the first type of threads and the second type of threads from competing for the processor resources of the first processor core, thus being able to improve the processing efficiency of the first processor core for the second type of threads.
[0012] In a possible implementation, after entering the sleep state, the method further includes the following steps: querying whether there is a to-be-executed thread among the first type of threads associated with the first processor core every first duration; if there is a to-be-executed thread, entering the wake-up state and using the first processor core to execute the to-be-executed thread.
[0013] Based on the above possible implementation manners, the cycler queries whether there are threads to be executed, so that in the case where there are threads to be executed, the first executor can enter the wake-up state as soon as possible to execute the threads to be executed, avoiding the threads to be executed waiting for execution for a long time.
[0014] In a possible implementation manner, after querying whether there are threads to be executed in the first type of threads associated with the first processor core every first time period, the method further includes the following steps: if no threads to be executed are found after multiple queries, query whether there are threads to be executed in the first type of threads associated with the first processor core every second time period, where the second time period is greater than the first time period.
[0015] Since the first executor occupies the processor resources of the first processor core during the process of querying the threads to be executed, based on the above possible implementation manners, in the case where no threads to be executed are found after multiple queries, the query period is increased to continue querying for the threads to be executed. Since the query period is increased, the number of queries of the first executor can be reduced during the sleep period, reducing the resource occupation of the first executor on the first processor core.
[0016] In a possible implementation manner, the method further includes the following steps: after executing any first type of thread associated with the first processor core, obtain the second resource utilization rate of the first processor core; if the second resource utilization rate is greater than or equal to the third threshold, execute the step of reducing the first type of threads to be executed in the future time period.
[0017] Based on the above possible implementation manners, after executing each first type of thread, the first executor determines whether to continue executing the first type of threads based on the resource utilization rate of the first processor core, achieving fine-grained thread regulation.
[0018] In a second aspect, a thread processing device is provided for executing the method provided in the first aspect or any optional manner of the first aspect above.
[0019] In a third aspect, a computing device is provided, which includes a processor, and the processor is used to execute program code to enable the computer device to execute to implement the method provided in the first aspect or any optional manner of the first aspect above.
[0020] In a fourth aspect, a computer-readable storage medium is provided, in which at least one piece of program code is stored, and the program code is read and executed by the processor to enable the computer device to execute to implement the method provided in the first aspect or any optional manner of the first aspect above.
[0021] In a fifth aspect, a computer program product or a computer program is provided. The computer program product or the computer program includes program code stored in a computer-readable storage medium. A processor reads the program code from the computer-readable storage medium and executes the program code, so that a computing device executes the method provided in the above first aspect or various alternative implementations of the first aspect.
[0022] Based on the implementations provided in the above aspects of the present application, further combinations can be made to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a computing system architecture diagram of a method for processing application threads provided by an embodiment of the present application;
[0024] Figure 2 is a schematic hardware structure diagram of a computing node provided by an embodiment of the present application;
[0025] Figure 3 is a hierarchical structure diagram of a thread framework provided by an embodiment of the present application;
[0026] Figure 4 is a schematic diagram of an application scenario of a task queue provided by an embodiment of the present application;
[0027] Figure 5 is a flowchart of a method for processing threads provided by an embodiment of the present application;
[0028] Figure 6 is a schematic diagram of thread scheduling across processing cores provided by an embodiment of the present application;
[0029] Figure 7 is another schematic diagram of thread scheduling across processing cores provided by an embodiment of the present application;
[0030] Figure 8 is a schematic diagram of multiple processes preempting processor resources provided by an embodiment of the present application;
[0031] Figure 9 is a schematic structural diagram of a device for processing threads provided by an embodiment of the present application;
[0032] Figure 10 is a schematic structural diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] Figure 1 is a computing system architecture diagram of a method for processing application threads provided by an embodiment of the present application, Figure 1The computing system 100 shown is used to provide data computing services. The computing system 100 can be a data center or a supercomputing center. Here, the embodiments of the present application do not limit the application software scenarios of the computing system 100.
[0034] As Figure 1 shown, the computing system 100 includes at least one storage node 101, at least one computing node 102, and a network bus 103. Each storage node 101 and each computing node 102 in the computing system 100 communicate with each other through the network bus 103. Figure 1 The number of computing nodes 102 and the number of storage nodes 101 shown are only examples. Here, the embodiments of the present application do not limit the number of computing nodes 102 and the number of storage nodes 101 in the computing system 100.
[0035] The storage node 101 is used to provide data storage services for the computing node 102. The storage node 101 can be a storage device or a memory in the storage device. Among them, the memory can be the volatile memory or non-volatile memory introduced below.
[0036] Any computing node 102 is used to provide data computing services. The computing node 102 can be a computing device, such as a server, a computer, etc. Here, the embodiments of the present application do not limit the device type of the computing node 102. At the software level, the computing node 102 supports the installation and running of software programs to implement the functions of the software programs. As Figure 1As shown in the figure, according to different service objects, software programs are divided into platform software and application software. Among them, platform software is used to provide basic services for application software to ensure that the application software can run normally in the computing node 102. The platform software includes system software and / or middleware. Among them, system software is the bridge between computer hardware and users. It is responsible for managing and controlling computer hardware resources and providing a running environment for application software, such as an operating system (OS). Middleware is a software program located between the operating system and application software, providing services for communication and data management for different software programs. For example, input / output (IO) clients, storage clients, database clients, algorithm components, acceleration library components, etc. Application software is a software program that directly serves users and meets specific user needs and tasks, such as high-performance computing (HPC) applications, supercomputing (SC) applications, and data processing applications or other application software other than these. Here, the embodiments of the present application do not limit the type of application software. Taking the application of the computing system 100 in the supercomputing center scenario as an example, assuming that the platform software in the computing node 102 is an IO client, the application software in the computing node 102 can access the storage node 101 through the IO client. For example, read data from the storage node 101 or write data to the storage node 101. In some other embodiments, the computing system 100 does not include the storage node 101. Here, the embodiments of the present application do not limit whether the computing system includes the storage node 101.
[0037] At the hardware level, as Figure 2 shown in the schematic diagram of the hardware structure of a computing node provided by the embodiment of the present application as Figure 2 shown, the computing node 102 includes: a bus 21, a processor 22, a memory 23, an actuator 24, and a communication interface 25. The processor 22, the memory 23, the actuator 24, and the communication interface 25 communicate through the bus 21. The bus 21 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 2It is represented by only one line in the figure, but it does not mean that there is only one bus or one type of bus. Bus 21 may include a path for transmitting information between various components of computing node 102 (e.g., processor 22, memory 23, actuator 24, and communication interface 25). There may be at least one processor 22, and processor 22 may be a central processing unit (CPU). Each processor 22 includes at least one processor core, such as a 4-core processor, an 8-core processor, etc. Processor 24 may be implemented in at least one of the hardware forms of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA).
[0038] Memory 23 may serve as the internal memory or external memory of computing node 102. Memory 23 may include volatile memory, such as random access memory (RAM). Memory 23 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0039] The executable program code is stored in memory 23, and actuator 24 reads and executes the executable program code, enabling computing node 102 to implement the thread processing method provided in the embodiments of the present application below. For example, there may be at least one actuator 24, and each actuator 24 is respectively associated with a processor core. Each actuator 24 may execute the thread processing method provided in the embodiments of the present application below for the associated processor core. In the case where there are multiple actuators 24, these multiple actuators 24 may be integrated into at least one processing module or may not be integrated into a processing module.
[0040] The actuator 24 can be implemented by a chip, a CPU, an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The above PLD can be a complex programmable logical device (CPLD), FPGA, generic array logic (GAL), data processing unit (DPU), system on chip (SoC), or any combination thereof. Figure 2 Taking the case where the actuator 24 is implemented by hardware as an example, in some other embodiments, the actuator 24 can also be implemented by software. For example, the actuator 24 is a resident software in the user state of the computing node. In some other embodiments, the actuator 24 can also be implemented by a combination of software and hardware. Here, the embodiments of the present application do not limit the implementation manner of the actuator 24.
[0041] Next, based on the computing node introduced above, the method for processing threads provided by the present application will be introduced in detail.
[0042] This method involves threads. According to whether the kernel of the operating system can perceive the threads, the threads can be divided into kernel level threads (KLT) and user level threads (ULT). Among them, the kernel level threads can be simply referred to as kernel threads, which are the threads supported by the kernel and run in the kernel space. The kernel can perceive the kernel threads. The kernel threads are the CPU resources acting on the physical cores on the physical server, and their implementation manner is that the threads in the kernel are resident or the operations on the kernel threads are implemented through a software stack. The user level threads are simply referred to as user threads, which are the threads supported by the application software and run in the user space, and the kernel does not perceive the user threads.
[0043] Each processor core in the computing node can be respectively associated with a kernel thread, and the kernel threads associated with different processor cores are different. Taking Figure 3Taking the thread framework hierarchy diagram shown as an example, assume that there are N processor cores and N kernel threads in a computing node, and one processor core is associated with one kernel thread, where N is an integer greater than or equal to 1. The task manager in the computing node has a task allocation function. For example, it maps the user threads of the application software in the computing node to the kernel threads associated with at least one processor core to implement the allocation of user threads to the corresponding processing cores. Among them, the user threads of the application software refer to the threads in the process corresponding to the application software. The application software corresponds to at least one process, and each process is used to implement a task of the application software. The process includes one or more threads, and each thread is used to implement a subtask in the task. At least one user thread can be mapped to one kernel thread. In the case where multiple user threads are mapped, different user threads may belong to different processes. When allocating user threads to the processing core, a time period, called the time slice of the process, is also allocated to the process to which the user thread belongs, that is, the time during which the process is allowed to run. The processor core runs the process during this time period, that is, executes the user thread of the process during this time period. If the process is still running at the end of the time slice or the process blocks or ends before the end of the time slice, it triggers the processor core to perform a context switch: switch out this process and run another process. However, each context switch requires time, saving the on-site data of the switched-out process in the memory and loading the on-site data of the switched-in process in the CPU register, resulting in latency overhead. When the processor core frequently undergoes context switches, it will further increase the latency overhead and reduce the performance of the processor core.
[0044] For user threads, in this application, according to the different service objects of the application software, the user threads are divided into two categories, namely the first type of threads and the second type of threads. Among them, the first type of threads refers to the user threads of the platform software, and the second type of threads refers to the user threads of the application software. In this way, the user threads associated with the same processing core include at least one type of user threads among the first type of threads and the second type of threads. The thread type to which any user thread belongs can be indicated by a thread type identifier, and the thread type identifier is the first type identifier or the second type identifier. The first type identifier is used to indicate that the user thread is the first type of thread, and the second type identifier is used to indicate that the user thread is the second type of thread. The first thread is a lightweight thread, such as a daemon thread.
[0045] As Figure 3 shown, the task manager also has a core management function to manage the processor cores in the computing node. As Figure 4 shown, on the task execution layer of the computing node, based on the core management function, the task manager forms a core resource pool with the N processor cores in the computing node, and binds each processor core in the core resource pool to an executor (such as executor 24) respectively, so as to allocate an executor to each processor core toFigure 3 For example, for any processing core, the task manager establishes a mapping relationship between the processor core, a kernel thread, and an executor, such that the processor core, the kernel thread, and the executor are associated with each other. In this way, N processing cores are associated with N kernel threads and N executors. Each executor is responsible for managing the first type of threads mapped on the associated kernel thread. For example, the executor executes the first type of threads mapped on the associated kernel thread, and / or schedules the first type of threads mapped on the associated kernel thread to other processing cores.
[0046] In a possible implementation, a task queue can be used to maintain the first type of threads managed by the executor, so as to Figure 3 For example, each executor has a task queue respectively, and different executors have different task queues. In this way, N executors have N task queues, and the task queue is used to record the first type of threads managed by the executor to which it belongs. As Figure 4 shown, the task manager maps a task queue to an executor through a queue mapping method, for example, establishing a mapping relationship between the task queue and the executor. For any executor and the processor core associated with the executor, the task manager adds the thread information of the first type of threads associated with the processor core to the task queue of the executor based on the task allocation function, so as to record the first type of threads in the task queue, that is, add the first type of threads to the task queue. Figure 4 Taking the platform software as an IO client as an example, the user threads of the IO client for inter-process communication, for cache synchronization, for data prefetching, etc. are all the first type of threads, and these first type of threads are distributed in different task queues.
[0047] In a possible implementation, as Figure 3 and Figure 4 shown, the task manager also has a load statistics function, for example, statistically analyzing the load conditions of each processor core. Based on the load conditions of each processor core, the executor is dynamically enabled or disabled to achieve load-aware enabling and disabling of cores. The task manager also has a latency management function. For Figure 4 example, for example, scheduling the first type of threads to be executed in the task queue to the processor core associated with the executor in the sleep state, ceding (relax) the CPU resources to the second type of threads associated with the associated processing core, and other processing cores process the scheduled first type of threads to achieve delayed scheduling. Or, the executor suspends the first type of threads being executed in the task queue. When the suspended first type of threads enter the task queue again, they can be rescheduled for execution to perform task degradation.
[0048] For any computing node in a computing system, the computing node includes N executors and N processor cores, where the N processor cores may belong to the same processor or different processors. Any executor is associated with a processor core, and different executors are associated with different processor cores. Each executor can execute the thread processing method provided in this application for the processor core it is associated with, and the process of each executor executing this method is the same. For ease of description, any one of the N executors is referred to as the first executor, and the processor core associated with the first executor is referred to as the first processor core.
[0049] Next, in combination with Figure 5 , taking the execution of this method by the first executor in the computing node as an example, the process of this method will be introduced. The first executor is associated with the first processing core in the computing node, and the first processor core is associated with the first type of thread and the second type of thread. This method includes the following steps.
[0050] 501. The first executor obtains the first resource utilization rate of the first processor core.
[0051] Among them, the first type of thread associated with the first processor core, that is, the first type of thread mapped on the kernel thread associated with the first processor core, and the second type of thread associated with the first processor core, that is, the second type of thread mapped on the kernel thread associated with the first processor core. The resource utilization rate of any processor core is used to reflect the utilization of CPU resources during a certain period of time. CPU resources such as the time of the CPU. Among them, the resource utilization rate is the percentage between the usage duration of the processing core and the statistical duration. The usage duration refers to the total duration of the user threads executed by the processor core during this period, and the statistical duration refers to the total duration of this period. Exemplarily, the resource utilization rate is shown in the following formula (1).
[0052]
[0053] The first resource utilization rate refers to the resource utilization rate of the first processor core in the first historical period, which is used to reflect the utilization of CPU resources in the first historical period. The first historical period refers to the period between the first historical moment and the current moment, and the first historical moment is a certain moment before the current moment.
[0054] Before starting to execute the first type of thread associated with the first processor core, the first executor determines the first resource utilization rate based on the usage duration of the first processor core within the historical period and the total duration of the historical period. For example, taking the total duration of the historical period as the statistical duration, calculate the first resource utilization rate according to the above formula (1).
[0055] Alternatively, the first executor is not responsible for calculating the first resource utilization rate. Instead, the first executor obtains the first resource utilization rate of the first processor core from the task manager in the computing node. Exemplarily, the task manager monitors the resource utilization rates of each processor core in the computing node during each time period. Before starting to execute the first type of threads associated with the first processor core, the first executor sends a resource utilization rate acquisition request to the task manager to request the resource utilization rate of the first processor core during a historical time period. After receiving the resource utilization rate acquisition request, the task manager determines the first resource utilization rate based on the usage duration of the first processor core and the total duration of the historical time period, and returns the first resource utilization rate to the first executor, which then receives the first resource utilization rate.
[0056] 502. If the first resource utilization rate is less than the first threshold, the first executor uses the first processing core to execute the first type of threads associated with the first processor core.
[0057] Herein, the first threshold is the execution condition for the start of the execution of the first type of threads associated with the first processor core. The first threshold is greater than 0 and less than or equal to 1. For example, the first threshold is 70%, 80%, or 90%. Here, the present application does not limit the value of the first threshold.
[0058] The first processor core is responsible for executing the first type of threads and the second type of threads associated therewith. The execution of the first type of threads is controlled by the first executor, and the execution of the second type of threads is controlled by other components other than the first executor. Based on this, if the first resource utilization rate of the first processor core is less than or equal to the first threshold before executing the first type of threads associated with the first processor core, it indicates that the second type of threads occupy less processor resources of the first processor core and the resource utilization rate of the first processor core is low during the first historical time period. Then, the first type of threads associated with the first processor core meet the execution conditions, and the first executor executes the first type of threads associated with the first processor core, enabling the second type of threads and the first type of threads to share the processor resources of the first processor core, increasing the resource utilization rate of the processor core, and ensuring that the first type of threads can be executed while minimally affecting the execution of the second type of threads.
[0059] Next, the process of the first executor executing the first type of threads associated with the first processor core is introduced as follows.
[0060] The task queue associated with the first executor is referred to as the first task queue. The first task queue includes the thread information of at least one first - type thread associated with the first processor core. Among them, the thread information of any first - type thread is used to indicate the first - type thread. Exemplarily, the thread information of the first - type thread includes the thread identity (ID), the first - type identifier, and the process ID of the belonging process. The first - type threads indicated by the thread information in the first task queue are simply referred to as the first - type threads in the first task queue. The first - type threads in the first task queue are the first - type threads to be executed by the first executor.
[0061] Based on the sorting of each first - type thread in the first task queue, the first executor uses the first processor core to sequentially execute each first - type thread in the first task queue. For example, dequeue the first - type thread at the head of the first task queue, so that the next first - type thread in the first task queue is at the head, execute the dequeued first - type thread. After the dequeued first - type thread finishes execution, perform the above process again: dequeue the first - type thread at the head of the first task queue, so that the next first - type thread in the first task queue is at the head, execute the dequeued first - type thread, and so on. Among them, the sorting of each first - type thread in the first task queue is the sorting of the thread information of each first - type thread in the first task queue, and the first - type thread at the head is the first - type thread indicated by the thread information at the head.
[0062] Next, taking any first - type thread in the first task queue as an example, the process of the first executor using the first processor core to execute this first - type thread is introduced as follows:
[0063] The first executor obtains the task information of the first - type thread. This task information includes the task execution instruction, context information, and information of the memory stack of the first - type thread. The first executor binds the first - type thread to this memory stack, stores the context information of the first - type thread in the CPU register of the first processor core. The first processor core, based on the context information of the first - type thread in the CPU register, jumps the currently executed instruction to the task execution instruction of the first - type thread to start execution, and stores the data generated during the execution process in the memory - area stack. If a task - switching command is encountered during the execution process, the first processor core stores the current context information of the first - type thread in the CPU register into this memory - area stack, exits the currently executed task execution instruction, until the current context information of the first - type thread is stored in the CPU register again, and continues to execute this task execution instruction based on this context information; if the task execution instruction of this first - type thread is executed completely, the first processor core exits the current task and releases this memory stack.
[0064] In another possible implementation, after any first-type thread associated with the first processor core is executed, the first executor obtains the second resource utilization rate of the first processor core. If the second resource utilization rate is greater than or equal to the third threshold, the first executor reduces the number of first-type threads to be executed in the future time period. The reduction method is as described in step 503 below and will not be elaborated here. If the second resource utilization rate is less than the third threshold, the first executor continues to execute the next first-type thread associated with the first processor core.
[0065] Among them, the second resource utilization rate refers to the resource utilization rate of the first processor core in the second historical time period, which is used to reflect the utilization of CPU resources in the second historical time period. The second historical time period refers to the time period between the second historical moment and the current moment. For example, it is the time period when the previous first-type thread was executed. The second historical moment is after the first historical moment. The process of obtaining the second resource utilization rate is the same as that of obtaining the first resource utilization rate and will not be elaborated here. The third threshold is greater than 0 and less than or equal to 1. For example, the third threshold is 70%, 80%, or 90%. The third threshold may be the same as or different from the first threshold. Here, the present application does not limit the value of the third threshold.
[0066] Take Figure 6 the thread scheduling schematic diagram of the cross-processor core shown as an example. Assume Figure 6 that the executor 1 in is the first processor core. After the executor 1 finishes executing the dequeued first-type thread, the resource utilization rate of the processor core associated with the executor 1 reaches 100%, exceeding the third threshold. The executor 1 performs task scheduling on the thread 1 at the head of the first task queue. For example, it schedules the thread 1 to the end of the task queue of the executor 2, reducing the number of first-type threads to be executed by the first executor, causing the thread 1 to be delayed in execution by the executor 2, thereby realizing the task degradation of the thread 1.
[0067] When the second resource utilization rate is less than the third threshold, it indicates that even when the first type of threads and the second type of threads share the processor resources of the first processor core during the second historical period, the resource utilization rate of the first processor core is still low. Then, the first executor can continue to execute the next first type of thread. On the one hand, it will not cause a significant increase in the resource utilization rate of the first processor core and avoid affecting the execution of the second type of threads by the first processor core. On the other hand, it can also ensure that the first type of threads can be executed. When the second resource utilization rate is greater than or equal to the third threshold, it indicates that the resource utilization rate of the first processor core is too high during the second historical period. The first executor reduces the first type of threads to be executed in the future time period (i.e., the time period after the current time). On the one hand, it can reduce the number of context switches between the first type of threads and the second type of threads by the first processor core in the future time period, avoid frequent context switches, reduce the latency overhead caused by context switches, and thus improve the performance of the first processor core. On the other hand, it can reduce the resource occupation of the first type of threads on the first processor core in the future time period, enabling the first processor core to provide more processor resources for the associated second type of threads to improve the processing efficiency of the first processor core for the second type of threads. After each execution of a first type of thread, the first executor determines whether to continue executing the first type of thread based on the resource utilization rate of the first processor core, realizing fine-grained thread regulation.
[0068] In some other embodiments, during the process of sequentially executing the first type of threads in the first task queue, for any first type of thread in the first task queue, if the first type of thread has not been completely executed, the first executor suspends the first type of thread and then executes the next first type of thread in the first task queue. In this case, before executing the next first type of thread, the first executor can also first obtain the second resource utilization rate of the first processor core. If the second resource utilization rate is less than the third threshold, then execute the next first type of thread. For Figure 6 example, the executor 1 first suspends the thread 1. If the resource utilization rate of the processor associated with the executor 1 is less than the first threshold, then execute the thread 3 to achieve thread switching. Correspondingly, if the second resource utilization rate is greater than or equal to the third threshold, the first type of threads suspended in the first queue and the first type of threads to be executed in the first task queue are all threads to be executed by the first executor. Then, the first executor executes the suspended first type of threads or the first type of threads to be executed in the first task queue.
[0069] 503. If the first resource utilization rate is greater than or equal to the first threshold, the first executor reduces the first type of threads to be executed in the future time period.
[0070] Herein, the future time period refers to the time period after the current moment.
[0071] If, before executing the first type of threads associated with the first processor core, the first resource utilization rate of the first processor core is greater than or equal to the second threshold, it indicates that the first processor core executed the second type of threads in the first historical period, and the execution of the second type of threads occupied a large amount of processor resources, resulting in a high resource utilization rate of the first processor core in the first historical period. Subsequently, the processor resources occupied by the second type of threads may also be relatively high. Then, the first executor reduces the first type of threads to be executed in the future period, reducing the resource occupancy of the first type of threads on the first processor core in the future period.
[0072] In a possible implementation, the first executor reduces the first type of threads to be executed in the future period through the following method 1 and / or method 2.
[0073] Method 1: The first executor schedules at least one target thread to the second processor core in the computing node, where the target thread is a to-be-executed thread among the first type of threads associated with the first processor core.
[0074] Among them, at least one target thread is all or part of the to-be-executed threads among the first type of threads associated with the first processor core. Here, the present application does not limit the number of target threads. The second processor core is the target processor core for target thread scheduling. For ease of description, the executor associated with the second processor core is called the second executor. Then, the second executor is used to execute the first type of threads associated with the second processor core using the second processor core. The task queue associated with the second executor is called the second task queue, and the processor core in the computing node other than the first processor core is called the third processor core, that is, the computing node includes the first processor core and N - 1 third processor cores.
[0075] For any third processor core in the computing node, if the third processor core meets the thread reception condition, the first executor determines the third processor core as the second processor core, where the thread reception condition refers to the condition that allows scheduling threads to the processor core.
[0076] Exemplarily, the thread reception condition includes condition 1 and / or condition 2. Condition 1 is that the resource utilization rate of the processor core is less than or equal to the second threshold, and the second threshold is greater than or equal to 0 and less than the first threshold. For example, the second threshold is 0, 10%, or 20%. Here, the present application does not limit the value of the second threshold. Condition 2 is that the processor core is in a dormant state. Any executor in the dormant state cannot execute the first type of threads associated with the associated processor core using the associated processor core. At this time, the associated processor can only provide processor resources for the second type of threads.
[0077] Taking the thread receiving condition including condition 1 as an example, the determination process of the second processor core is introduced as follows: The first executor sends a determination request to the task manager in the computing node, and this determination request is used to determine the second processor core in the computing node; Based on the determination request, the task manager obtains the resource utilization rate of each third processor core in the computing node during the first historical time period (referred to as the third resource utilization rate). If the third resource utilization rate of any third processor core is less than or equal to the second threshold, that is, this third processor core meets the thread receiving condition, then based on this third processor core, a determination response is returned to the first executor, and this determination response is used to indicate that this third processor core is determined as the second processor core; Based on this determination response, the first executor determines this third processor core as the second processor core.
[0078] Taking Figure 6 the thread scheduling schematic diagram of cross - processing cores shown as an example, assume Figure 6 that the executor 1 and executor 2 in it are the first processor core and the third processor core respectively, and both executor 1 and executor 2 support thread scheduling. Assume that before starting to execute the threads in the task queue, the resource utilization rate of the processor core associated with executor 1 reaches 100%, reaching the first threshold, and the third resource utilization rate of executor 2 meets condition 1. Then executor 2 is determined as the second processor core, and the thread 1 to be executed by executor 1 is used as the target thread and scheduled to the end of the task queue of executor 2, so that thread 1 is delayed in executor 2, thus realizing the task degradation of thread 1.
[0079] Taking the thread receiving condition including condition 2 as an example, the determination process of the second processor core is introduced as follows: The task manager records the working states of the executors associated with each processor core in the computing node, where the working states are sleep state or wake - up state, and any executor in the wake - up state can use the associated processor core to execute the first - type threads associated with the associated processor core. Based on this, after receiving the determination request from the first executor, the task manager queries the working states of the executors associated with each third processor core. If the working state of the executor associated with any third processor core is the sleep state, that is, this third processor core meets the thread receiving condition, then based on this third processor core, a determination response is returned to the first executor, and the first executor determines this third processor core as the second processor core based on this determination response.
[0080] Taking Figure 7Taking the thread scheduling schematic diagram of the cross-processing core shown as an example, assume that there are no first-class threads in the task queue of executor 3, executor 3 is in a dormant state, and the processor core associated with executor 3 meets condition 2. Then, taking executor 3 as the second executor, the first executor schedules thread 5 in the first task queue as the target thread to the task queue of executor 3, and subsequently, executor 3 uses the processing core associated with executor 3 to execute thread 5.
[0081] Taking the thread reception conditions including condition 1 and condition 2 as an example, the determination process of the second processor core is introduced as follows: After receiving the determination request from the first executor, the task manager queries the third resource utilization rate of each third processor core and the working status of the executor associated with each third processor core. If the third resource utilization rate of any third processor core is less than or equal to the third threshold, and the working status of the executor associated with this third processor core is in a dormant state, that is, this third processor core meets the thread reception conditions, then based on this third processor core, a determination response is returned to the first executor. Based on this determination response, the first executor determines this third processor core as the second processor core.
[0082] The above is introduced by taking the first executor requesting the task manager to determine the second processor core as an example. In some other embodiments, the first executor can obtain the third resource utilization rate of each third processor core from the task manager, and / or, the working status of the processor core associated with each third processor core. The first executor determines the second processor core from the third processor cores of the computing node based on the third resource utilization rate of each third processor core and / or the working status of the processor core associated with each third processor core.
[0083] After or before determining the second processor core, the first executor takes at least one first-class thread closest to the head of the first task queue as the target thread, and dequeues each target thread in the first task queue, where at least one first-class thread closest to the head of the first task queue is the first-class thread indicated by at least one thread information starting from the head of the first task queue.
[0084] After determining the second processing core and at least one target thread, the first executor schedules the at least one target thread to the second processor core. Exemplarily, based on the second processing core and the at least one target thread, the first executor sends a thread scheduling request to the task manager, and the thread scheduling request is used to indicate scheduling the at least one target thread to the second processor core. A mapping relationship between the kernel threads associated with each processor core and the user threads of the application software is established in the task manager to indicate the user threads mapped on each kernel thread. For example, the mapping relationship between any kernel thread and the user thread of the application software is shown as: the thread ID of the kernel thread corresponds to the thread IDs and thread type identifiers of at least one user thread. After receiving the thread scheduling request from the first executor, the task manager parses the thread information of at least one target and the identifier of the second processor core from the thread scheduling request. Based on the identifier of the second processor core, it queries the kernel thread associated with the second processor core from the mapping relationship between each processing core and the kernel thread. Based on the thread information of the at least one target thread, a sub-mapping relationship is established between the at least one target thread and the kernel thread associated with the second processor core in the mapping relationship between the kernel thread associated with the second processor core and the user thread of the application software. For example, in this mapping relationship, the thread ID and the first type identifier of the at least one target thread are added to implement the establishment of this sub-mapping relationship. By establishing this sub-mapping relationship, the at least one target thread is mapped to the kernel thread associated with the second processor core, and the at least one target thread is scheduled to the second processing core.
[0085] After scheduling the at least one target thread to the second processor core, the task manager adds the thread information of the at least one target thread to the second task queue of the second executor. For example, it is added to the end of the second task queue. When the second processor core dequeues any target thread queue from the second task queue, the second executor executes the target thread using the second processor core, thereby realizing the delayed scheduling of the target thread.
[0086] Method 1 schedules the target thread from the first processor core to the second processor core, so that the first executor uses the first processor core in the future time period to execute the first-class threads that have not been scheduled, thereby reducing the first-class threads executed in the future time period. When the third resource utilization rate of the second processor core is less than or equal to the second threshold, it means that the resource utilization rate of the second processor core in the first historical time period is low, and the user threads associated with the second processor core occupy fewer processor resources. Scheduling the target thread from the first processor core to the second processor core can ensure that the target thread can be executed on the second processor core without affecting the execution of the user threads originally associated with the second processor core. When the second executor is in a dormant state, scheduling the target thread from the first processor core to the second processor core, and having the second executor use the second processor core to execute the target thread, can avoid the second executor from being in a dormant state for a long time, thereby increasing the utilization rate of the second executor.
[0087] The above is explained by taking the first resource utilization of the first processor core being greater than or equal to the first threshold to trigger thread scheduling as an example. In other embodiments, before executing the first type of threads in the first task queue, the first executor first counts the number of first type threads in the first task queue. If the first resource utilization of the first processor core is greater than or equal to the first threshold, and / or the number of first type threads is greater than or equal to the quantity threshold, the first executor schedules the first number of target threads to the second processor core. The quantity threshold is the maximum number of first type threads that the first processing core can execute at one time while ensuring the execution efficiency of the second type of threads, and the first number is less than the quantity threshold. Still with Figure 7 For example, assuming that the first number is 1, and the number of first-type user threads in the first task queue exceeds the number threshold, thread 5 in the first task queue is scheduled to the processor core associated with executor 3.
[0088] When the number of first-class threads in the first task queue is greater than or equal to the quantity threshold, some of the first-class threads are scheduled to the second processing core to avoid a large number of subsequent first-class threads from occupying the processor resources of the first processor core, so that the first processor core can provide more processor resources for the associated second-class threads to ensure the processing efficiency of the first processor core on the second-class threads.
[0089] Method 2: The first actuator enters a dormant state.
[0090] In a possible implementation, during the future time period, the first actuator goes into sleep to enter the sleep state until the wake-up condition is met, at which point the sleep ends and the actuator enters the wake-up state. During the sleep process, the first actuator remains in the sleep state. In the sleep state, the first actuator cannot execute the first type of threads. In the wake-up state, the first actuator can execute the first type of threads.
[0091] Among them, the wake-up condition refers to the condition that triggers the first actuator to switch from the sleep state to the wake-up state. Exemplarily, the present application provides three wake-up conditions, namely, the first wake-up condition, the second wake-up condition, and the third wake-up condition. When any of the wake-up conditions is met, the first actuator ends the sleep. Next, in combination with each wake-up condition, the process of the first actuator switching from the sleep state to the wake-up state and the execution process of the first type of threads in the wake-up state will be introduced.
[0092] The first wake-up condition: The first actuator has threads to be executed.
[0093] Among them, the threads to be executed refer to the threads to be executed in the first type of threads associated with the first processor core.
[0094] The first type of user threads may be coroutines or other user threads other than coroutines. Taking the first type of user threads as coroutines as an example, after entering the sleep state, at every first time interval, the first actuator queries whether there are threads to be executed in the first type of threads associated with the first processor core. If there are such threads to be executed, the first wake-up condition is met, and the first actuator enters the wake-up state and uses the first processor core to execute the threads to be executed. Among them, the first time interval is less than the sleep duration threshold, and the sleep duration threshold is the maximum sleep duration for each sleep of the first actuator. In different implementation scenarios, the value of the sleep duration threshold and / or the value of the first time interval may be different. Here, the present application does not limit the value of the sleep duration threshold and the value of the first time interval.
[0095] The method for querying the threads to be executed. For example, the first actuator queries whether the first task queue records the first type of threads. If the first task queue records the first type of threads, the recorded first type of threads are determined as the threads to be executed by the first actuator. If the first task queue does not record the first type of threads, the first actuator does not have threads to be executed. Exemplarily, query whether there is thread information in the first task queue. If there is thread information in the first task queue, the first type of threads indicated by the thread information is the first type of threads recorded in the first task queue. If there is no thread information in the first task queue, the first task queue does not record the first type of threads.
[0096] The above description is given by taking one query period of the first duration as an example. In another possible implementation, the first actuator corresponds to multiple query periods, and the multiple query periods increase in sequence. For example, the first query period is the first duration, the second query period is the second duration, the third query period is the third duration, and so on. The first duration is greater than the second duration, the second duration is greater than the third duration, and so on. Here, the present application does not limit the number of query periods and each query period.
[0097] After entering the sleep state, the first actuator first uses the first duration as the query period. Every time the first duration elapses, it queries whether there is a thread to be executed among the first type of threads associated with the first processor core. If no such thread to be executed is found after multiple queries, then it uses the second duration as the query period. Every time the second duration elapses, the first actuator queries whether there is a thread to be executed among the first type of threads associated with the first processor core until a thread to be executed is found and the first wake-up condition is met. Then the first actuator enters the wake-up state and uses the first processor core to execute the thread to be executed.
[0098] The first actuator queries whether there is a thread to be executed according to the query period, so that in the case where there is a thread to be executed, the first actuator can enter the wake-up state as soon as possible to execute the thread to be executed, avoiding the thread to be executed waiting for execution for a long time. Since the first actuator will occupy the processor resources of the first processor core during the process of querying the thread to be executed, in the case where no thread to be executed is found after multiple times, the query period is increased to continue querying the thread to be executed. Since the query period is increased, the number of queries of the first actuator during the sleep period can be reduced, and the resource occupancy of the first actuator on the first processor core can be reduced.
[0099] In another possible implementation, the first actuator uses the sleep duration threshold as the sleep period. For any sleep period, in the sleep state, after each query of the thread to be executed, if no thread to be executed is found and the sleep duration has not reached the sleep duration threshold, after one query period, the first actuator continues to query the thread to be executed; if a thread to be executed is found, regardless of whether the sleep duration has reached the sleep duration threshold, the first actuator has to end the sleep and enter the wake-up state; if no thread to be executed is found and the sleep duration has reached the sleep duration threshold, then the first actuator enters the next sleep period and continues to sleep.
[0100] In some other embodiments, the first wake-up condition is also applicable to other user threads except for coroutines, which will not be elaborated here.
[0101] The second wake-up condition: The first actuator receives a wake-up instruction.
[0102] Wherein, the wake-up instruction is used for the first actuator to enter the wake-up state.
[0103] Taking the first type of threads other than coroutines as other user threads as an example, after the first executor enters the sleep state, when there is no record of the thread to be executed in the first task queue, the task manager adds the first type of threads associated with the first processor core to the first task queue and sends a wake-up instruction to the first executor; or, when there is a record of the thread to be executed in the first task queue, the task manager uses the sleep duration threshold as the sleep cycle. If the sleep duration of the first executor reaches the sleep duration threshold, a wake-up instruction is sent to the first executor. When the first executor receives this wake-up instruction, the second wake-up condition is met. Based on this wake-up instruction, the first executor enters the wake-up state and uses the first processor core to execute the first type of threads in the first task queue.
[0104] When receiving the wake-up instruction, the first executor enters the wake-up state only, without the first executor querying the thread to be executed, reducing the resource occupation of the first executor on the first processor core.
[0105] In some other embodiments, the second wake-up condition is also applicable to the first type of threads that are coroutines, which will not be elaborated here.
[0106] The third wake-up condition: The sleep duration reaches the sleep duration threshold.
[0107] Exemplarily, after the first executor goes to sleep, the first executor times the sleep duration. If the sleep duration reaches the sleep duration threshold, it enters the wake-up state. In the wake-up state, it queries whether there is a thread to be executed among the first type of threads associated with the first processor core. If there is a thread to be executed, the first executor uses the first processor core to execute the thread to be executed. If there is no thread to be executed, it enters the sleep state again until the sleep duration reaches the sleep duration threshold and enters the wake-up state again, and so on.
[0108] When the sleep duration reaches the sleep duration threshold, the first executor automatically enters the wake-up state without being woken up by the task manager, reducing the workload of the task manager.
[0109] After entering the first wake-up state, if there are multiple threads to be executed, the first executor executes the multiple threads to be executed in sequence according to the sorting of these multiple threads to be executed in the first task queue. Or, for each execution of any thread to be executed, obtain the second resource utilization rate of the first processor core. If the second resource utilization rate is less than the third threshold, continue to execute the next thread to be executed or the suspended first type of threads. If the second resource utilization rate is greater than or equal to the third threshold, perform the step of reducing the first type of threads to be executed in the future time period again.
[0110] If the execution of reducing the first type of threads to be executed in the future time period is the above-mentioned method 2, during the working period, the first executor switches back and forth between the sleeping state and the waking state. Taking the executor 3 in Figure 7 as an example of the first executor, when there is no thread to be executed in the corresponding task queue, the executor 3 enters the sleeping state. After the thread 5 is scheduled into the task queue, it enters the waking state and executes the thread 5. After the thread 5 finishes execution and there is no process to be executed in the task queue, it enters the sleeping state again. Subsequently, after a new first type of thread is scheduled into the task queue, it enters the waking state again.
[0111] For the above-mentioned method 2, the first executor does not utilize the resources of the first processor core to execute the first type of threads during the sleeping period, thereby being able to reduce the first type of threads to be executed in the future time period, and can also avoid the first type of threads competing with the second type of threads for the processor resources during the sleeping period, thereby being able to improve the processing efficiency of the first processor core for the second type of threads.
[0112] Figure 5 In the method provided by the embodiment, when the first processor core is associated with both the first type of threads and the second type of threads, if the resource utilization rate of the first processor is less than the first threshold, the first executor utilizes the first processor core to execute the associated first type of threads. If the resource utilization rate of the first processor is greater than or equal to the first threshold, the first executor reduces the first type of threads to be executed in the future time period, thereby being able to reduce the number of context switches of the first processor core between the first type of threads and the second type of threads in the future time period, avoid the first processor core from frequently performing context switches, reduce the latency overhead caused by context switches, and thereby be able to improve the performance of the processor core. Moreover, it can also reduce the resource occupation of the first type of threads on the first processor core in the future time period, so that the first processor core can provide more processor resources for the associated second type of threads, thereby being able to improve the processing efficiency of the first processor core for the second type of threads.
[0113] To further illustrate the method for processing threads provided by the present application, taking Figure 8 the example of multiple processes competing for processor resources shown, the iterative type application is a kind of application software. Assume that the iterative type application has M (M>1) iterative tasks, and the M iterative tasks are respectively executed by processes 1 to M. Each process runs through a preparation iteration stage, a calculation stage, an IO output stage, and a synchronization waiting (barrier) stage. Among them, in the IO output stage, the process needs to call the platform software to generate the first type of threads for implementing the IO output task, and executes the second type of threads in other stages. As Figure 8As shown, it is assumed that the second - type threads of Process 1 and Process 2 in the synchronization waiting stage and the first - type threads involved in Process N in the IO output stage are all mapped to the same processor core. If, during a certain period, Process 1 and Process 2 are both in the synchronization waiting stage and Process N is in the IO output stage, in the related art, during this period, the second - type threads of Process 1 and Process 2 in the synchronization waiting stage and the first - type threads involved in Process N in the IO output stage will all preempt the processor resources of this processor core, causing the processor core to perform frequent context switches between the second - type threads and the first - type threads. However, in this application, the executor corresponding to this processor core can, by executing the thread processing method provided in this application, schedule the first - type threads to other processor cores for execution, avoiding frequent context switches on this processor core, thereby being able to avoid the performance loss caused by frequent context switches.
[0114] The method of the embodiment of this application is introduced above. The device of the embodiment of this application is introduced below. It should be understood that the device introduced below has any functions of the first executor in the above - mentioned method. As described above in combination with Figures 5 to 8 The processing method of threads according to the embodiment of this application is described in detail. Based on the same inventive concept, the device applying this processing method will be described below in combination with Figure 9 and Figure 10 It should be understood that the technical features described in the method embodiment also apply to the following device embodiment.
[0115] Figure 9 FIG. is a schematic structural diagram of a thread processing device 900 provided by the embodiment of this application. Figure 9 The device 900 shown can be configured as an executor for executing the method steps executed by the first executor. The device 900 is applied to a computing node. The device 900 is associated with the first processing core in the computing node. The first processor core is associated with the first - type threads and the second - type threads. The first - type threads refer to the user threads of the platform software, and the second - type threads refer to the user threads of the application software. The device 900 includes:
[0116] An obtaining module 901, configured to obtain the first resource utilization rate of the first processor core;
[0117] An execution module 902, configured to, if the first resource utilization rate is less than the first threshold, use the first processing core to execute the first - type threads associated with the first processor core;
[0118] A processing module 903, configured to, if the first resource utilization rate is greater than or equal to the first threshold, reduce the first - type threads to be executed in the future time period.
[0119] In a possible implementation manner, the processing module 903 is configured to:
[0120] Schedule at least one target thread to a second processor core in a computing node, and / or enter a sleep state, where the target thread is a to-be-executed thread among the first type of threads associated with the first processor core.
[0121] In a possible implementation, the resource utilization rate of the second processor core is less than or equal to a second threshold, and / or the second executor associated with the second processor core is in a sleep state, where the second executor is used to execute the first type of threads associated with the second processor core by utilizing the second processor core.
[0122] In a possible implementation, the apparatus 900 further includes:
[0123] A query module, configured to query whether there is a to-be-executed thread among the first type of threads associated with the first processor core every first duration;
[0124] The execution module 902 is further configured to, if there is a to-be-executed thread, enter a wake state and execute the to-be-executed thread by utilizing the first processor core.
[0125] In a possible implementation, the query module is further configured to:
[0126] If there is no to-be-executed thread after multiple queries, query whether there is a to-be-executed thread among the first type of threads associated with the first processor core every second duration, where the second duration is greater than the first duration.
[0127] In a possible implementation, the acquisition module 901 is further configured to, after executing any one of the first type of threads associated with the first processor core, acquire the second resource utilization rate of the first processor core;
[0128] The processing module 903 is further configured to, if the second resource utilization rate is greater than or equal to a third threshold, execute the step of reducing the first type of threads to be executed in a future time period.
[0129] When the apparatus 900 processes threads, only the division of the above functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the apparatus 900 is divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus 900 and the above method for processing threads belong to the same concept, and the specific implementation process is detailed in the above Figure 5 shown method flow, which will not be elaborated here.
[0130] The implementation manner of the apparatus 900 is the same as that of the executor introduced above, and the implementation manner of the apparatus 900 will not be elaborated here.
[0131] In a possible implementation, when the actuator is implemented by hardware, the actuator can be implemented by a computing device. Refer to Figure 10 the structural schematic diagram of the computing device shown in Figure 10 which is the structural schematic diagram of a computing device provided by this application. As Figure 10 shown, the computing device 1000 includes: a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate with each other through the bus 1002. The bus 1002 can be a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 10 only one line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The bus 1002 can include a path for transmitting information between various components of the computing device 1000 (for example, the memory 1006, the processor 1004, the communication interface 1008). The processor 1004 can include any one or more of processors such as a CPU, a GPU, an MP, or a DSP. The memory 1006 can be used as the internal memory or external memory of the computing device 1000. The memory 1006 can include the volatile memory or non-volatile memory introduced above. The memory 1006 stores executable program code, and the processor 1004 reads and executes the executable program code, so that the computing device 1000 implements the thread processing method provided by the embodiments of this application.
[0132] The computing device 1000 is any electronic device with computing functions, such as a chip, a processor, or other hardware-form electronic devices.
[0133] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code. The above program code can be executed by a processor in the computing device to complete the thread processing method in the above embodiments. For example, the computer-readable storage medium is a non-temporary computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0134] The embodiments of this application also provide a computer program product or a computer program. The computer program product or the computer program includes program code. The computer instruction is stored in a computer-readable storage medium. The processor of the computing device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computing device executes the above thread processing method.
[0135] In addition, an embodiment of the present application further provides a device, which may specifically be a chip, a component, or a module. The device may include a processor and a memory connected to each other. The memory is used to store computer-executable instructions. When the device runs, the processor may execute the computer-executable instructions stored in the memory, so that the chip executes the processing method of the thread in each of the above method embodiments.
[0136] Among them, the device, equipment, computer-readable storage medium, computer program product, or chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.
[0137] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, it belongs to the same concept as the embodiment of the thread processing method provided in the above embodiment, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0138] In several embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0139] The unit described as a separate component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0141] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.
[0142] In the description of this application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" means one or more, and "plurality" means two or more. The words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be different.
[0143] In this application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the words "exemplary" or "for example" is intended to present the related concepts in a concrete way.
[0144] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the resource utilization involved in this application is obtained with full authorization.
[0145] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.
[0146] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for processing threads, characterized in that Executed by a first executor in a computing node, the first executor being associated with a first processing core in the computing node, the first processor core being associated with a first type of thread and a second type of thread, the first type of thread being a user thread of platform software, and the second type of thread being a user thread of application software. The method includes: Obtain a first resource utilization rate of the first processor core; If the first resource utilization rate is less than a first threshold, use the first processing core to execute the first type of thread associated with the first processor core; If the first resource utilization rate is greater than or equal to the first threshold, reduce the first type of threads to be executed in a future time period.
2. The method according to claim 1, characterized in that, The reducing the first type of threads to be executed in a future time period includes: Scheduling at least one target thread to a second processor core in the computing node, and / or entering a sleep state, the target thread being a to-be-executed thread among the first type of threads associated with the first processor core.
3. The method according to claim 2, wherein The resource utilization rate of the second processor core is less than or equal to a second threshold, and / or the second executor associated with the second processor core is in a sleep state, the second executor being used to execute the first type of threads associated with the second processor core by using the second processor core.
4. The method according to claim 2 or 3, characterized in that After entering the sleep state, the method further includes: Every first time interval, query whether there is a to-be-executed thread among the first type of threads associated with the first processor core; If there is the to-be-executed thread, enter a wake-up state and use the first processor core to execute the to-be-executed thread.
5. The method according to claim 4, wherein After the querying whether there is a to-be-executed thread among the first type of threads associated with the first processor core every first time interval, the method further includes: If there is no to-be-executed thread after multiple queries, query whether there is a to-be-executed thread among the first type of threads associated with the first processor core every second time interval, the second time interval being greater than the first time interval.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: After each execution of any one of the first type of threads associated with the first processor core, obtain a second resource utilization rate of the first processor core; If the second resource utilization rate is greater than or equal to a third threshold, execute the step of reducing the first type of threads to be executed in a future time period.
7. A processing device for a thread, characterized in that, The apparatus is applied to a computing node, the apparatus being associated with a first processing core in the computing node, the first processor core being associated with a first type of thread and a second type of thread, the first type of thread being a user thread of platform software, and the second type of thread being a user thread of application software. The apparatus includes: An obtaining module, configured to obtain a first resource utilization rate of the first processor core; An executing module, configured to, if the first resource utilization rate is less than a first threshold, use the first processing core to execute the first type of thread associated with the first processor core; A processing module, configured to, if the first resource utilization rate is greater than or equal to the first threshold, reduce the first type of threads to be executed in a future time period.
8. The device according to claim 7, characterized in that, The processing module is configured to: Schedule at least one target thread to a second processor core in the computing node, and / or enter a sleep state, where the target thread is a to-be-executed thread in a first type of threads associated with the first processor core.
9. The device according to claim 8, characterized in that, The resource utilization rate of the second processor core is less than or equal to a second threshold, and / or the second executor associated with the second processor core is in a sleep state, where the second executor is used to execute a first type of threads associated with the second processor core by using the second processor core.
10. The device according to claim 8 or 9, characterized in that The device further includes: A query module, configured to query, at every first time interval, whether there is a to-be-executed thread in a first type of threads associated with the first processor core; The execution module is further configured to, if there is the to-be-executed thread, enter a wake-up state and execute the to-be-executed thread by using the first processor core.
11. The device according to claim 10, characterized in that, The query module is further configured to: If there is no such to-be-executed thread after multiple queries, query, at every second time interval, whether there is a to-be-executed thread in a first type of threads associated with the first processor core, where the second time interval is greater than the first time interval.
12. The device according to any one of claims 7-11, wherein The obtaining module is further configured to, after executing any one of the first type of threads associated with the first processor core, obtain a second resource utilization rate of the first processor core; The processing module is further configured to, if the second resource utilization rate is greater than or equal to a third threshold, execute the step of reducing the first type of threads to be executed in a future time period.
13. A computing device, characterized in that, The computing device includes a processor, and the processor is configured to execute program code to cause the computing device to execute the method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, At least one program code is stored in the storage medium, and the at least one program code is read by the processor to cause the computing device to execute the method according to any one of claims 1 to 6.