Method and processing unit for executing tasks through master-slave rotation

Through the processing unit of the master-slave rotation mechanism, task execution is dynamically adjusted, which solves the problem of unbalanced resource utilization in the existing L1 processing architecture, and realizes efficient self-scheduling of computing resources and low-latency task execution.

CN114270318BActive Publication Date: 2025-07-08NOKIA NETWORKS OY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980099499.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-20
Publication Date
2025-07-08
Estimated Expiration
2039-08-20

AI Technical Summary

Technical Problem

In the existing L1 processing architecture, the complexity of task execution is unbalanced, resulting in low resource utilization efficiency of processing cores, especially difficult to efficiently schedule between computing-intensive and computing-simple tasks.

Method used

The processing unit adopts the master-slave rotation mechanism, dynamically adjusts task execution through the switching of master and slave functions, ensures self-scheduling and efficient utilization of computing resources, and uses the main thread and slave thread in the thread pool to perform management and computation-intensive functions respectively.

Benefits of technology

It realizes the maximization of computing resource utilization under high processing load conditions, reduces the delay and waste of task execution, and ensures that simple computing tasks do not affect the execution efficiency of computing-intensive tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114270318B_ABST
    Figure CN114270318B_ABST
Patent Text Reader

Abstract

The present subject matter relates to a method, comprising: obtaining a master role by a processing unit of a multi-processor system; executing a main functional part of a set of tasks by the processing unit, including: searching for available processing units of the multi-processor system; wherein, if available processing units are found, controlling the found processing units to execute a slave functional part of the set of tasks, and if available processing units are not found, executing the slave functional part of the set of tasks by the processing unit, wherein the main function includes a master-to-slave switching function for releasing the master role, and the slave function includes a slave-to-master switching function for obtaining the master role.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various example embodiments relate to computer networks, and more particularly to a method for performing tasks by master-slave rotation. Background Art

[0002] Current L1 processing architectures consist of a heterogeneous set of dedicated processing cores. In these current architectures, the complexity of the core is scaled to the complexity of the task it has to perform. For example, scheduling decisions can be executed on a simple Advanced RISC Machine (ARM) core, while computations on algorithms can be executed on a complex Digital Signal Processor (DSP). However, other processing architectures can consist of a large pool of homogeneous processing cores (e.g., General Purpose Processor (GPP) cores), which are more programmable and thus can flexibly execute a wide range of processing. Performing tasks in those other architectures may need improvement. Summary of the Invention

[0003] Example embodiments provide a processing unit for a multi-processor system, the processing unit being configured to acquire a master role for a master functional part of a set of tasks to be executed, the processing unit being configured to perform a master function for searching available processing units of the multi-processor system; wherein, if an available processing unit is found, the processing unit is configured to control the found processing unit to execute a slave functional part of the set of tasks, and if an available processing unit is not found, the processing unit is configured to further execute the slave functional part of the set of tasks, wherein the master function includes a master-to-slave switching function for releasing the master role, and the slave function includes a slave-to-master switching function for acquiring the master role.

[0004] Example embodiments provide a multi-processor system, which includes the processing unit and one or more additional processing units, wherein the processing unit is configured to perform searching for available processing units in the additional processing units, and each additional processing unit is configured to execute a slave function of a set of tasks or a master function of a set of tasks.

[0005] Example embodiments provide a network node including the multi-processing system.

[0006] An exemplary embodiment provides a method, which includes: obtaining a main role by a processing unit of a multiprocessor system; executing a main functional part of a set of tasks by the processing unit, including: searching for available processing units of the multiprocessor system; wherein, if available processing units are found, controlling the found processing units to execute a slave functional part of the set of tasks, and if no available processing units are found, executing the slave functional part of the set of tasks by the processing unit, wherein, the main function includes a main-to-slave switching function for releasing the main role, and the slave function includes a slave-to-main switching function for obtaining the main role.

[0007] An exemplary embodiment provides a method, which includes: providing a set of threads, wherein the threads in the set are configured to work in a main mode to generate a main thread, or work in a slave mode to generate a slave thread. The method includes task calculation, which includes: executing a scheduler by a current main thread in the set; determining, by the current main thread, available threads for a selected current task in a set of one or more tasks for executing the scheduler, the determination including: identifying available slave threads in the set as available threads, and if no slave threads in the set are available, switching the current main thread to the slave mode to generate available threads; executing the task by the available threads; if the current main thread switches to the slave mode, the available slave threads in the set switch to the main mode and thus become the current main thread. The method further includes: repeating the task calculation to execute another task of the scheduler until the set of tasks is executed.

[0008] According to a further exemplary embodiment, an apparatus includes at least one processor; and at least one memory including computer program code; the apparatus includes a set of threads, wherein the threads in the set are configured to work in a main mode to generate a main thread, or work in a slave mode to generate a slave thread; the at least one memory and the computer program code are configured to, together with the at least one processor, cause the apparatus to at least execute task calculation, including: executing a scheduler by a current main thread in the set; determining, by the current main thread, available threads for a selected current task in a set of one or more tasks for executing the scheduler, the determination including: identifying available slave threads in the set as available threads, and if no slave threads in the set are available, switching the current main thread to the slave mode to generate available threads; executing the task by the available threads; if the current main thread switches to the slave mode, the available slave threads in the set switch to the main mode and thus become the current main thread. The apparatus is further caused to: repeat the task calculation to execute another task of the scheduler until the set of tasks is executed.

[0009] According to a further exemplary embodiment, a computer program includes instructions stored thereon for at least performing the following operations: providing a set of threads, wherein the threads in the set are configured to work in a master mode to generate a main thread or in a slave mode to generate a slave thread; performing task computation, including: executing a scheduler by a current main thread in the set; determining, by the current main thread, available threads for a selected current task among a set of one or more tasks for executing the scheduler, the determination including: identifying available slave threads in the set as available threads, and if all slave threads in the set are unavailable, switching the current main thread to the slave mode to generate available threads; executing a task by the available threads; if the current main thread is switched to the slave mode, switching the available slave threads in the set to the master mode, thereby becoming the current main thread; repeating the task computation to execute another task of the scheduler until the set of tasks is executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are included to provide a further understanding of the examples and are incorporated in and constitute a part of this specification. In the drawings:

[0011] Figure 1 A schematic diagram depicting a multiprocessing system including a plurality of processing units according to an example of the present subject matter;

[0012] Figure 2 is a flowchart of a method for controlling a processing unit according to an example of the present subject matter;

[0013] Figure 3 is a block diagram of an apparatus showing an example of the present subject matter;

[0014] Figure 4 is a flowchart of a method for performing a task according to an example of the present subject matter;

[0015] Figure 5 is a flowchart of a method for performing tasks of a plurality of schedulers according to an example of the present subject matter;

[0016] Figure 6 is a block diagram showing a task list for a given scheduler according to an example of the present subject matter;

[0017] Figure 7 is a block diagram showing a method for performing a task list according to an example of the present subject matter;

[0018] Figure 8 is a flowchart of a method for performing a task according to an example of the present subject matter;

[0019] Figure 9 is a schematic diagram showing the execution of two task lists according to an example of the present subject matter;

[0020] Figure 10 It is a schematic diagram showing the execution of a four-task list according to an example of the present subject matter. Detailed implementation

[0021] In the following description, for purposes of explanation and not limitation, specific details such as specific architectures, interfaces, technologies, etc. are set forth in order to provide a thorough understanding of the example. However, it will be apparent to those skilled in the art that the disclosed subject matter may be practiced in other illustrative examples that depart from these specific details. In some instances, detailed descriptions of well-known devices and / or methods are omitted so as not to obscure the description with unnecessary details.

[0022] The processing unit that executes the main role can be the main processing unit (main) of a multiprocessing system. The processing unit that executes the slave function can be the slave processing unit (slave).

[0023] In one example, the processing unit can acquire the main role during the initialization phase of the present subject matter. This initialization can be performed or controlled by a supervisor. The supervisor can be a processing unit of a multiprocessing system that is permanently designated as the supervisor. The supervisor can, for example, select and control the selected processing unit to acquire the main role. By acquiring the main role, the processing unit can be enabled to execute the main function associated with that main role. Thus, the initialization phase can result in the designation of a main processing unit that executes the main function. The execution of the main function can include receiving tasks to be executed. The execution of the main function can enable the execution of the received tasks, such as those associated with the main thread. The main function can be referred to as the task function. After the initialization of the selected processing unit, the master-slave designation rotation scheme can be further enabled according to an embodiment of the present subject matter in order to execute the set of tasks. For example, after the initialization phase and if the main role is released, the processing unit can acquire the main role again by executing the slave-to-master switching function. The main function and the slave function can respectively enable the main and slave execution flows.

[0024] Other processing units of the multiprocessor system can be initialized during the initialization phase such that if they are not executing tasks, they attempt or try to acquire the main role. This can be performed by each of the other processing units (if they are not performing other functions, such as if they are in an idle state) repeatedly executing the slave-to-master switching function. Alternatively or additionally, after the initialization phase, the other processing units can be configured to acquire the main role by executing the slave-to-master switching function as part of the execution of the slave function.

[0025] The processing unit can be configured to read instructions for performing functions such as a slave function and a master function. For example, the processing unit can include components that enable the reception and execution of instructions. These components can include, for example, registers, an arithmetic logic unit (ALU), a memory mapping unit (MMU), cache memory, an I / O module, etc., and several of these components can be co-located on a single chip.

[0026] As used herein, the term "task" refers to an operation that can be performed by one or more processors of one or more computers. The task can include, for example, multiple instructions. The task can include multiple functions, where each function can be associated with a corresponding instruction.

[0027] The present subject matter can enable a scenario where there is no fixed "master / slave" designation, but each equivalent processing unit can act as a potential slave or master initiator. The actual master assigns intensive processing to the slave, unless no slave is found (e.g., because they are all occupied), in which case the master must make itself a slave by performing a master-to-slave switching function and thus release the master role. Intensive processing can be performed by the execution of the slave function. This can ensure low-latency distribution of work across different processing units via rotation of the master role.

[0028] According to the present subject matter, the master processing unit that becomes a slave may not always be immediately followed by a new master, so all processing units can act as slaves during execution. This can ensure that all clock cycles of the processing units of a multiprocessing system can be dedicated to processing when the processing load is high, without cycles being wasted in the master function. The master runs only when there are clock cycles available for processing (i.e., when the slaves have completed their work), and no master is running when the slaves use all clock cycles for processing. This can ensure maximum allocation of cycles for advancing the processing pipeline of the set of tasks and minimum waste in the master functions of scheduling, forwarding, and task assignment. For example, no processing unit in a multiprocessing system is blocked or busy waiting because each cycle is used to process a task or find a new task to process. This can maximize the utilization efficiency of available cycles and maintain low latency.

[0029] The present subject matter can enable a non-blocking software architecture where all processor cycles can be used as efficiently as possible so that complex cores can be used to intermittently execute computationally simple tasks without affecting computationally intensive tasks.

[0030] A multi - processing system can be included, for example, in a Cloud Radio Access Network (C - RAN) to enable at least a part of the tasks of the C - RAN. For example, the multi - processing system can be used to enable data communication between a base station and an end - user, or to enable tasks involved in large - scale deployment, cooperative radio technology support, and real - time virtualization capabilities of the C - RAN. Using a multi - processing system can bring significant computational advantages in the C - RAN. A non - wireless cloud system interconnected, for example, via fiber optic cables or other cables, or via a hybrid of wired and non - wired connections, can be incorporated in accordance with embodiments of the present subject matter.

[0031] Embodiments according to the present subject matter can be used, for example, in a cyber - physical system to meet the stringent Quality of Service (QoS) requirements for data communication in such a system. A cyber - physical system can implement, for example, distributed automation applications, such as cyber - physical applications. The distributed automation application can enable, for example, factory automation, such as for motion control or mobile robots. In another example, the distributed automation application can enable process automation, such as for process monitoring. Network nodes of the present subject matter can be used, for example, as fifth - generation (5G) network nodes to enable data communication between devices of a cyber - physical system. For example, using the present method can accelerate tasks performed at a 5G network node. Alternatively or additionally, devices of a cyber - physical system can include a multi - processing system according to the present subject matter so that they can perform their tasks faster. Using a multi - processing system in a cyber - physical system can be particularly useful for the following reasons. Current standards and product development for 5G New Radio Ultra - Reliable and Low - Latency Communication (5G NR URLLC), especially factory automation use cases, require a communication service availability of 10 -5 to 10 -9 within a limited time budget of 0.5 - 2 milliseconds. Additionally, for periodic traffic patterns, industrial automation use cases such as motion control can require a deterministic availability of the communication service up to nine nines figure. The present subject matter can meet such requirements.

[0032] In one example, the network node can be a 5G Fixed Wireless Access (FWA) node configured to use a multi - processing system to perform tasks involved in 5G FWA.

[0033] According to an example, after completion of the execution of a function, another processing unit of the multi - processor system that does not have the primary role is configured to attempt to acquire the primary role by executing a slave - to - master handover function for acquiring the primary role.

[0034] According to the example, the processing unit is configured to release the master role by releasing the master lock, wherein the acquisition of the master role includes acquiring the released master lock.

[0035] According to the example, the master function includes a scheduler for scheduling the group of tasks and an output forwarding function for providing the results of the previous tasks upon which the task depends. The slave function includes a task processing function for at least one task.

[0036] According to the example, if the processing unit or another processing unit previously executed the master-to-slave switching function and if the master lock is idle, then the execution of the slave-to-master switching function by this processing unit is successful. However, the slave can execute the slave-to-master switching function at any time it has completed its task (e.g., any time it is in an idle state).

[0037] According to the example, the processing unit is a homogeneous processing core.

[0038] The processing unit as described above can be configured to use threads to execute at least a portion of the subject matter. For example, the processing unit is configured to use the main thread in a thread pool of a multiprocessor system to execute the master function and is configured to use the slave threads in the thread pool to execute the slave function. Using threads can enable efficient execution of the subject matter and seamless integration of the subject matter in existing systems.

[0039] According to the example, the multiprocessor system further includes a thread pool and a manager thread. The manager thread is configured to establish slave threads and at least one main thread in the pool, wherein an additional processing unit is configured to use the main thread in the pool to execute the master function and to use the slave threads in the pool to execute the slave function. The manager thread can be executed by a manager.

[0040] A thread can be a sequence of program instructions that can be managed independently (e.g., by a scheduler that is typically part of an operating system). A thread can, for example, enable the processing of one or more tasks.

[0041] For example, a thread-based method can be provided. The thread-based method includes: providing a set of threads, wherein the threads in the set of threads are configured to work in a master mode to generate a main thread, or work in a slave mode to generate a slave thread. The method includes task calculation, which includes: executing a scheduler by the current main thread in the set of threads, determining, by the current main thread, available threads for a selected current task among a set of one or more tasks for executing the scheduler, wherein the determination includes: identifying available slave threads in the set of threads as available threads, and if all the slave threads in the set of threads are unavailable, switching the current main thread to the slave mode to generate available threads, executing the task by the available threads, and if the current main thread is switched to the slave mode, the available slave threads in the set of threads are switched to the master mode and thus become the current main thread. The method further includes: repeating the task calculation to execute another task of the scheduler until the set of tasks is executed.

[0042] The execution of the main thread can be performed by a corresponding main processing unit, and the execution of the slave thread can be performed by a corresponding slave processing unit.

[0043] This subject can enable self-scheduling of computing resources through master-slave rotation. For example, each thread in the set of threads can be executed on a corresponding processing unit in a pool of processing units. The rotation of master-slave responsibilities among the resources in the pool of processing units can ensure that each processing unit in the pool of processing units can contribute to processing functions with high processing complexity simultaneously and still maintain low-latency execution of all other functions such as scheduling, synchronization, and forwarding functions.

[0044] The execution of the task calculation can result in using a corresponding main thread to execute a task in the set of tasks. This main thread may or may not be the same main thread for executing another task in the set of tasks.

[0045] According to the example, the tasks in the set of tasks include a series of functions. This series of functions includes a task processing function and an output forwarding function. The execution of the scheduler further includes: executing, by the current main thread, the output forwarding function of a previous task in the set of tasks that has been executed in the previous iteration, wherein executing the current task includes: executing the task processing function of the current task using the output of the previous task generated from the execution of the output forwarding function.

[0046] A set of tasks can, for example, include tasks of a multi-stage pipeline. This can enable a functional split of each pipeline stage of the pipeline to match the master-slave function. For example, master-slave rotation and pipeline stage splitting can enable a scalable amount of parallelism that automatically matches the algorithm requirements with the available physical computing resources without manual adjustment. If there is sufficient parallelism available in the set of tasks, the method can automatically match this parallelism with the maximum of the physical resources.

[0047] Rotation of the master function and functionally splitting the tasks into a master-executed function and a slave-executed function can allow for optimal use of self-scheduling. The master-executed function can be a management function such as scheduling and forwarding. The slave-executed function can be a computationally intensive function. For example, a homogeneous set of processing resources can be used for the computationally intensive pipeline of the task processing function while still performing the management function with a minimum cycle and low latency.

[0048] According to an example, the output forwarding function of the previous task executed by the current main thread includes: the current main thread checking whether the output of the previous task is available, and in response to determining that the output of the previous task is not available, repeating the check until the output is available.

[0049] This can prevent the result of the slave thread from remaining in the pipeline until the next task of the pipeline stage is executed. If the output is not available, the main thread can refer back to the scheduler again. This can enable a loop involving the output forwarding function that serves as a periodic poll for completing tasks that have been running asynchronously on the slave processing unit.

[0050] According to an example, a series of functions includes a data preparation function for preparing input data for the task processing function using the output of the previous task. The execution of the scheduler further includes executing the data preparation function.

[0051] After executing the output forwarding function for obtaining the output of the previous task, the task processing function of the current task can be executed to prepare the input data. The input data is used by the determined available threads to execute the task processing function of the current task. According to this subject matter, this can enable a further functional split of each pipeline stage of the pipeline that matches the master-slave function. This can increase the types of tasks that can be executed by this method.

[0052] According to an example, the data preparation function is executed to generate input data. The input data can be input into the input queue of the determined slave thread, where executing the current task includes: reading the input data from the input queue and executing the task processing function.

[0053] According to an example, the execution of the functions of a task is triggered by the scheduler being executed.

[0054] The scheduler can be configured, for example, to select which current task in the group of tasks is to be executed in the current iteration of the method, and to allocate the execution of each function of the selected task to a thread. For example, the scheduler can allocate the execution of the output forwarding function of a previously selected task in the group of tasks and the data preparation function of the current selected task to the current main thread. The scheduler can allocate the execution of the task processing function of the current selected task to the determined available threads. The scheduler can allocate the function of determining available threads to the current main thread. This can enable self-scheduling of a homogeneous set of processing resources.

[0055] According to an example, repeating includes: concurrently executing two or more tasks in the group of tasks. For example, the task processing functions of different tasks in the group of tasks can be independent and can thus be executed in parallel by corresponding threads. This can enable acceleration of the execution of the group of tasks.

[0056] According to an example, the repetition of task computation is performed for a first scheduler. The method further includes: repeating the task computation for a second scheduler using the current main thread for executing the second scheduler for executing another task of the second scheduler until a group of tasks of the second scheduler is executed. The current main thread for executing the second scheduler is different from the current main thread for executing the first scheduler.

[0057] Using multiple main threads can be useful when there are many high-speed tasks and / or when the serial execution of the main function limits the use of parallelism in the thread pool.

[0058] For example, a first and a second scheduler can be provided. The first scheduler can be associated with a first group of tasks, while the second scheduler can be associated with a second group of tasks. The first and second groups of tasks can be part of the same larger group of tasks that is split into the first and second groups of tasks. This can be useful, for example, when the first scheduler has too many tasks associated with it (e.g., a larger group of tasks) and the serial execution of the main function becomes a bottleneck that hinders the concurrent use of all available worker threads in the pool. Two schedulers / main threads can be started so that they can handle the corresponding tasks in the larger group of tasks. The thread pool includes two different main threads that can be used to execute the two schedulers respectively. In addition, the same worker threads in the thread pool can be used for each of the first scheduler and the second scheduler to separately execute the method.

[0059] According to the example, the provision of the thread set includes: providing a manager thread; establishing a scheduler by the manager thread; determining a set of tasks to be executed by the scheduler; starting each thread in the thread set as a slave thread; and selecting a thread from the thread set as the current main thread.

[0060] According to the example, the threads in the thread set are configured to: work in the main mode by acquiring the main lock, and switch to the slave mode by releasing the main lock, where the slave threads are configured to: acquire the main lock for switching to the main mode.

[0061] According to the example, each thread in the thread set is configured to be executed on a corresponding processing core in a set of processing cores.

[0062] According to the example, the set of processing cores are homogeneous cores. For example, the set of cores includes the same cores.

[0063] According to the example, the available slave threads that switch to the main mode in the thread set are the slave threads that executed the tasks of the scheduler in a previous iteration of the task computation. This can enable controlled task execution specific to the scheduler.

[0064] Figure 1 A schematic diagram of a multiprocessing system 10 including a plurality of processing units 11A-D is depicted. The multiprocessing system 10 enables a master-slave role rotation scheme for executing a set of tasks. The set of tasks can include two tasks T1 and T2 for illustrative purposes only, but is not limited to two tasks. Additionally, for the sake of illustration, only four processing units are shown in Figure 1 the figure.

[0065] Each processing unit 11A-D can include several components. Components among these components can be, for example, operation units such as execution units and instruction retirement units. These components can also include internal memory locations for storing data, such as registers and caches. For example, the internal memory locations can be used to define the characteristics of various memory regions for instruction execution purposes.

[0066] The multiprocessing system 10 can be coupled to a system memory (not shown) through a memory controller of the multiprocessing system 10. The system memory can store instructions to be executed by the processing units 11A-D. Each of the processing units 11A-D can, for example, use an inter-processor interrupt (IPI) to send an interrupt to other processing units to request that one or more other processing units that are the targets of the IPI perform certain functional work.

[0067] Each individual processing unit 11A-D can operate according to a master mode or a slave mode. In each of these two modes, the processing unit can perform corresponding functions.

[0068] The present subject matter enables a master-slave designation rotation scheme among the processing units of the multiprocessing system 10 such that the set of tasks T1 and T2 whose instructions can reside in the memory system can be executed. To this end, the multiprocessing system 10 can be initialized (e.g., using a boot routine of the multiprocessing system 10 or a manager as described herein) such that one of the processing units first acquires the master role to execute the set of tasks. Additionally, each of the other processing units can be configured to acquire the master role, for example, by periodically checking whether the master role is idle and not being used by any other processing unit. This initialization can trigger the master-slave designation rotation scheme according to the present subject matter such that the set of tasks can be executed.

[0069] Figure 1 Multiple state phases (S1-S4) of the multiprocessing system 10 during the execution of tasks T1 and T2 are shown. In the first state phase (S1), the processing unit 11A can be initialized as the master processing unit. Each of the other processing units 11B-D can be in an idle state or can be performing a slave function. For simplicity of description, in the first state phase, none of the processing units 11B-D is in an idle state.

[0070] After becoming the master for executing the set of tasks T1 and T2, the master processing unit 11A can identify whether a processing unit is available. Since all the processing units 11B-D are not in an idle state, the master processing unit 11A may not find an available processing unit (e.g., if it could find an available processing unit, then this processing unit could perform the slave function of T1). In this case, the master processing unit 11A will lose the master role so that it can perform the slave function of task T1 itself. This is indicated in the second state phase (S2) of the multiprocessing system 10. During the second state phase, the master role can be acquired by the processing unit 11C because it becomes idle and attempts to acquire the master role. Task T1 can be completed by the processing unit 11A during the second state phase.

[0071] In the third state phase (S3), the new master processing unit 11C searches for an available processing unit (e.g., 11B) among the processing units 11A-B and 1D. Figure 1 Indicates that the processing unit 11B is available in the third state phase. Thus, the new master processing unit 11C can assign another slave function of the task labeled T2 to the processing unit 11B. As indicated in the fourth state phase (S4), the processing unit 11B can execute task T2. During this third state phase, the processing unit 11C remains the master.

[0072] Figure 2 is a flowchart of a method for controlling a processing unit (named the first processing unit) according to the present subject matter. The first processing unit can be controlled to initiate / trigger the execution of a set of one or more tasks. To this end, one or more additional processing units of a multiprocessing system can be provided. The first processing unit can be, for example, part of the multiprocessing system. The processing units of the multiprocessing system can cooperate through a master-slave designation rotation scheme to execute the set of tasks.

[0073] The first processing unit can have a set of instructions loaded that enable the execution of at least a part of the present subject matter.

[0074] In step 21, the first processing unit can acquire the master role. For example, the first processing unit can execute instructions that enable the acquisition of a master lock. The master lock can be associated with the set of tasks. According to the present subject matter, acquiring the master lock can enable the first processing unit to execute further instructions for initiating and controlling the execution of the set of tasks.

[0075] In one example, a scheduler can be used to schedule the tasks based on the dependencies of the tasks in the set. In this case, the master lock can be associated with the scheduler such that the execution of the scheduler is conditioned on the acquisition of the master lock.

[0076] In step 23, the first processing unit can execute the master function of the tasks in the set. For example, if the set of tasks are two related tasks T1 and T2, where T2 depends on the result of T1, task T1 can be executed first.

[0077] The master function can include multiple functions that enable the execution of the tasks. For example, the master function can include a function for identifying a processing unit to execute the slave function of the task. If a scheduler is used, the scheduling function of the scheduler can be part of the master function. The master function also includes a master-to-slave switching function for releasing the master role. The slave function includes a slave-to-master switching function for attempting to acquire the master role, which may be acquired if no other processing unit has the master role for the set of tasks and if the master role was previously released. The switching functions can enable an efficient rotation of the master and slave roles.

[0078] Thus, in step 23, the first processing unit can search for available processing units of the multiprocessor system. The available processing units can be processing units in an idle state.

[0079] If (query step 25) an available processing unit is found, the first processing unit can control the found processing unit to execute the slave function of the task in step 27. For example, the instructions of the slave function may be a set of instructions loaded at the found processing unit.

[0080] However, if no available processing unit is found, the first processing unit can perform the slave function of the task in step 29. This can be done by performing the master-to-slave switching function to release the master role. In this case, an available processing unit (e.g., a unit that has completed the execution of the slave function and thus does not have the master role) can attempt to acquire the master role by performing the slave-to-master switching function. Additionally, if the attempt is successful, then this can result in a new master processing unit.

[0081] Accordingly, the execution of step 27 or 29 can result in a new master processing unit or can lead to maintaining the same master processing unit. In both cases, the next task in the set of tasks can be performed by repeating steps 23-29 (e.g., by the new or maintained master processing unit performing the master function of the task). This can enable the self-scheduling pool of processing units to execute the set of tasks.

[0082] Figure 3 is a block diagram showing a device according to an example of the present subject matter. The device can be, for example, Figure 1 a multiprocessing system in, for example,

[0083] In Figure 3 a circuit block diagram showing the configuration of a device 100 is shown, and the device 100 is configured to implement at least a part of the present subject matter. It should be noted that the device 100 shown in Figure 3 can include several additional elements or functions other than those described below, which are omitted herein for simplicity as they are not necessary for understanding. Additionally, the device 100 can also be another device with similar functions, such as a chipset, a chip, a module, etc., which can also be part of the device or attached to the device as a separate element, etc.

[0084] Apparatus 100 may include a processor system 101 having a plurality of processor cores 111A-N (e.g., such as a multiprocessing system 10). Each of the cores 111A-N may include various logic and control structures to perform operations on data in response to instructions. In one example, each of the cores 111A-N may include a processor, a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other types of processing logic that can interpret and execute instructions. In one example, the cores 111A-N may be a homogeneous set of complex processor cores that perform all functions, both computationally complex and simple. These complex cores may, for example, be shared by both simple and complex functions. The present subject matter may enable the most efficient use of all processor cycles such that these complex cores can be used to intermittently perform computationally simple functions without affecting computationally intensive functions. Although only one processor system 101 is illustrated, the present subject matter may also be used by a computing system including multiple processor systems.

[0085] Reference numeral 102 denotes a transceiver or input / output (I / O) unit (interface) connected to the processor 101. The I / O unit 102 may be used to communicate with one or more other network units, entities, terminals, etc. The I / O unit 102 may be a combined unit including communication devices towards several network units, or may include a distributed structure having multiple different interfaces for different network units. Each of the processor cores 111A-N may include, for example, a number of components. Components among these components may be, for example, operation units such as execution units and instruction retirement units. These components may also include internal memory locations for storing data, such as registers and caches. For example, the internal memory locations may be used to define the characteristics of various memory regions for instruction execution purposes.

[0086] Reference numeral 103 denotes a memory that may be used, for example, to store data and programs to be executed by the cores 111A-N and / or may be used as a working storage device for the processor system 101.

[0087] Apparatus 100 includes a thread pool 113 that includes threads 115A-N. Each thread in the pool 113 is configured to work in a master mode (and thus is a master thread) or in a slave mode (and thus is a slave thread). Each thread in the thread pool 113 is configured to switch from the master mode to the slave mode, for example, if there are no slave threads available in the thread set to execute the tasks of the scheduler 118. Each thread in the pool 113 is configured to switch from the slave mode to the master mode, for example, after completing the tasks received from the scheduler 118 executed by the master thread in the pool 113 and if the master thread has switched from the master mode to the slave mode and there are no threads in the pool executing the scheduler 118.

[0088] For simplicity of description, only one scheduler 118 is shown, but it is not limited to one scheduler. For example, a given task list (e.g., three tasks t1, t2, and t3) can be assigned to the scheduler 118. Each of these tasks can include, for example, an output forwarding function (FWD function) and a task processing function (PROC function).

[0089] The scheduler 118 can be configured as follows. The scheduler 118 can be executed on the thread in pool 113 that is currently the main thread. After selecting the first task (e.g., t1 in the list), the scheduler 118 can assign to the main thread on which it is executed the function of identifying a worker thread for executing the PROC function of task t1. If it cannot find a worker thread, the main thread can switch to worker mode and execute the PROC function of task t1. If a worker thread is found, the main thread on which the scheduler is executed can notify the scheduler of the found worker thread. In response to this information, the scheduler can assign the execution of the PROC function of task t1 to the found worker thread. The scheduler 118 can assign the execution of the FWD function of task t1 to the main thread on which it is executed. The FWD function of task t1 is configured to determine the output resulting from the execution of task t1 and to determine which task (t2) in the list will be executed next.

[0090] If the FWD function of task t1 is successfully executed to determine the output of task t1, the main thread on which the scheduler 118 is executed can execute the function of identifying a worker thread for executing the PROC function of task t2 (assigned by the scheduler 118). If it cannot find a worker thread, the main thread can switch to worker mode and execute the PROC function of task t2. For example, assume that the main thread switches to worker mode to execute the PROC function of task t2 and that the worker thread has become the (new) current main thread.

[0091] The scheduler can be executed on this new current main thread. The executed scheduler 118 can assign the execution of the FWD function of task t2 to the new current main thread on which it is executed. The scheduler is configured such that it can access (know) whether a previous instance of the task has already been running. This information can be embedded in the global context of the task, which holds a list of the running instances of the task. When the scheduler selects a task, this context indicates that its previous instance is running and it needs to check whether it has completed and obtain the result of the FWD function. If the FWD function of task t2 is successfully executed, the new main thread on which scheduler 118 is executed can execute the function of identifying the worker threads for executing the PROC function of task t3 (assigned by scheduler 118). If it cannot find a worker thread, the new current main thread can switch to the worker mode and execute the PROC function of t3, etc.

[0092] Device 100 includes a manager thread 120. The manager thread 120 can be configured to perform an initialization step or phase to enable task execution according to this subject matter. The initialization step can include the setup of one or more schedulers 118. The tasks to be executed by cores 111A-N are connected or associated with the schedulers 118. The schedulers 118 can be configured to select which task to run and assign the execution of the functions configured for the task to the corresponding threads.

[0093] The initialization step can further include starting each thread in thread pool 113 as a worker thread, such that threads 115A-N can be pre-instantiated idle threads, which are ready to receive the functions to be executed. Pre-instantiation can avoid the overhead caused by having to create threads many times. The initialization step can further include selecting a thread in pool 113 as the main thread for scheduler 118. Once the pool is running, the manager thread may no longer be needed.

[0094] Manager thread 120 is shown as an external thread to pool 113. In another example, manager thread 120 can be a thread in thread pool 113. In this case, manager thread 120 may not have to select a thread in thread pool 113 to become the main thread and can just promote itself to become the main thread in pool 113. Alternatively, manager thread 120 can select another thread as the main thread and itself become a worker thread in pool 113. This can enable each core 111A-N to execute the main or worker execution path.

[0095] The main execution path or process may at least include the execution of the FWD function, the execution of identifying slave threads for performing task processing functions, and the execution of task processing functions if no slave threads are found. The slave execution path or process may at least include the execution of task processing functions and the execution of functions aimed at attempting to become the main thread.

[0096] Instructions of the scheduler 118 are run by the main thread in the thread pool 113. When the scheduler 118 has selected a task, the main thread searches for idle slave threads among the threads 115A-N to offload the work to it. When no idle slave threads are found, the main thread becomes a slave thread and it starts to execute the work by itself. Each of the slave threads in the threads 115A-N may attempt to become the main thread of the scheduler 118 after it has completed its work. When a slave thread successfully becomes the main thread, it starts to run the scheduler 118. The functions executed by the slave threads may be functions with high computational complexity, while those of the main thread may have low computational complexity but may require low-latency execution.

[0097] Each thread in the thread pool 113 may be executed by a corresponding core among the cores 111A-N. Thus, the cores 111A-N may be enabled to perform their own scheduling, synchronization, and forwarding (as the main), and still participate in computationally intensive processing (as the slave). The cores 111A-N may maximize their available instructions for computationally complex tasks, while minimizing the latency and cycles for computationally simple tasks such as scheduling, synchronization, and forwarding.

[0098] Figure 4 is a flowchart of a method for performing tasks according to an example of the present subject matter. For illustrative purposes, the method may be implemented in the system shown in the previous Figure 3 but is not limited to such implementation. Figure 4 The method of Figure 1 and Figure 2 may enable the use of threads to implement the master-slave rotation scheme as described herein, for example, with reference to

[0099] The scheduler 118 may first be assigned a list of tasks to be executed or associated with the list of tasks. The list of tasks may be related tasks. For example, the list of tasks may be tasks of a pipeline with multiple stages. The tasks of two consecutive pipeline stages may depend on the tasks of the pipeline stage using the output of another task of the previous pipeline stage. The scheduler 118 may be configured to select which task to run from the list of tasks. The selected task may be run using a set of threads such as the threads 115A-N in the thread pool 113.

[0100] A set of threads 115A-N can be established as described above such that one of these threads (e.g., thread 115B) can be the current main thread for executing the scheduler 118. In step 201, the scheduler 118 can be executed by the current main thread 115B. The execution of the scheduler 118 can include the selection of tasks to be executed in the task list. This selection can be performed, for example, using a priority algorithm that takes into account the dependencies between the tasks in the task list. The priority algorithm can be user-defined, for example. The priority algorithm can indicate the order in which the tasks in the task list need to be executed.

[0101] The scheduler can be configured to assign different functions to the corresponding threads in the set of threads 115A-N. For example, after selecting a task, the execution of the scheduler 118 can cause the current main thread 115B to determine or find available worker threads in the set of threads 115A-N. An available thread can be a thread that is, for example, an idle thread. To this end, the current main thread 115B can iterate through all the threads in the thread pool 113 to find a worker thread that is idle.

[0102] If (query step 205) an available worker thread (e.g., 115D) is identified or found in the set of threads 115A-N, the current main thread 115B can notify the scheduler 118 of the found available worker thread 115D. The scheduler 118 can assign the execution of the selected task to the identified thread 115D in step 207.

[0103] However, if (query step 205) no worker threads are available in the set of threads 115A-N, the current main thread 115B can switch to worker mode in step 209. This can result in an additional worker thread 115B.

[0104] The selected task can be executed by the identified thread 115D or by the additional worker thread 115B in step 211. That is, if no worker threads are available in the thread set, the current main thread 115B can switch to worker mode and execute the selected task. However, if an available thread is found, this available thread can receive the task and can execute it. The thread that executes the selected task can first read its input queue to read the input data of the selected task before executing the selected task, where the execution of the thread can be performed by one of the cores 111A-N.

[0105] As referenced Figure 3As described, once a thread has completed its task, it attempts to switch to the main mode. This switch to the main mode is performed if there is no other thread in the thread set that is the main thread for the scheduler 118. Thus, if the current main thread 115B has already been switched to the slave mode, the available slave threads in the thread set 115A-N switch to the main mode in step 213 to become the current main thread. In one example, the slave thread that switches to the main mode can be a slave thread that previously executed a task in the task list of the same scheduler 118. In another example, any available slave thread can be configured to become the main thread for the scheduler 118. The slave thread that switches to the main mode can be, for example, the main thread 115B itself, for example, if it is the first thread in the thread set to complete its task and become available after switching to become a slave thread.

[0106] If (query step 215) there are remaining tasks in the list, steps 201-215 can be repeated to execute the remaining tasks in the task list. Steps 201-215 can be referred to as task computation. For example, in each iteration, one task in the list can be executed. In each iteration of steps 201-215, the current main thread can change. For example, for the first execution of the method, thread 115B is the main thread. However, after multiple iterations of steps 201-215, another thread (e.g., 115D) can become the main thread for the scheduler 118.

[0107] Figure 5 is a flowchart of a method for executing tasks of multiple schedulers (e.g., two schedulers SH1 and SH2) according to an example of the present subject matter. For illustrative purposes, the method can be implemented in the Figure 3 system shown, but is not limited to such an implementation.

[0108] Each of the schedulers SH1 and SH2 can be assigned or associated with a corresponding first and second task list to be executed. The first task list can be independent of the second task list. However, the tasks within each of the first and second lists can be related tasks.

[0109] Each of the schedulers SH1 and SH2 can be assigned a corresponding first and second current main thread in the thread set in step 301. In step 303, method steps 201-215 can be performed for each of the schedulers SH1 and SH2. In one example, at least a portion of steps 201-215 can be performed in parallel for the two schedulers SH1 and SH2.

[0110] Figure 6It is a block diagram showing a task list for a given scheduler according to an example of the present subject matter.

[0111] For example, a pipelining scheme can be used to define a pipeline 400 for processing data. The pipeline 400 can be divided into stages. For simplicity purposes, Figure 6 three stages 401 - 403 are shown, but it is not limited to three stages. Each stage completes one or more tasks, and these stages are related to each other in sequence to form a pipeline. Tasks in the same stage can be executed in parallel, for example.

[0112] As Figure 6 shown, the first stage 401 has a single task P1. Task P1 can be the first task of the pipeline 400. Task P1 can receive the input data of the pipeline 400. Task P1 can provide outputs to two different tasks P2_1 and P2_2 of the second stage 402. These two tasks P2_1 and P2_2 provide their respective outputs to the task P3 of the last stage 403. Task P3 can provide the desired output of the pipeline 400. This pipeline provides a task list including tasks P1, P2_1, P2_2, and P3.

[0113] Figure 6 Further shown is a functional split of tasks according to an example of the present subject matter. Each of the tasks P1, P2_1, P2_2, and P3 can be divided into a series of functions. Figure 4 Shown is the task P2_2 of stage 402 as divided into four functions 410 - 413.

[0114] Function 410 selects which input from the previous pipeline stage 401 to run. This function can be an optional function because it can be used only when the pipeline stage 402 has more than one input to choose from. For example, task P3 can use function 410 to select which one of the two inputs of stage 402 will be used. Function 410 can be named INPUT SELECT function.

[0115] Function 411 can enable the accumulation of multiple inputs. This function can be used because in some data processing cases, multiple inputs can be accumulated before the actual processing starts. For example, crosstalk cancellation between multiple channels can use the inputs from each channel to continue execution. Function 411 can be named ACCUM function.

[0116] Function 412 is a task processing function, which can be the computationally intensive part of the task. Generally, this task processing function can be computed concurrently with consecutive processing tasks, and it is executed by a slave thread. Function 412 can be named PROC function.

[0117] Once the processing result of a task is ready, the output can be sent by function 413 to the next pipeline stage 403. Function 413 can also forward an indication or reference to the next task to be executed. As shown in stage 402, a single output can result in multiple forwards. This can be that the next pipeline stage is not ready yet. For example, if the previous stage has not been completed, function 413 will get backpressure and can retry later. Function 413 can be named the FWD function.

[0118] Functions 410, 411, and 413 can be data input preparation functions. They can be executed by the main thread. These data input preparation functions can be executed in serial execution because of the order in which their outputs can be used. The task processing function 412 can be executed by a worker thread and can potentially have concurrent execution with other processing functions 412 of consecutive tasks.

[0119] Figure 7 is a block diagram showing an example method for executing a list of tasks 501 according to the present subject matter. The list of tasks 501 can be tasks of a pipeline that need to be executed in a given order, for example, as described with reference to Figure 6 as described. This list of tasks can be assigned to a scheduler 502, which can select a task to be executed from the list of tasks using a priority algorithm. The scheduler 502 can be associated with the main thread. As described with reference to Figure 6 as described, each task in tasks 501 can include functions 510 - 513 defined respectively as described for functions 410 - 413.

[0120] Figure 7 Further shown is a unified processing resource divided into pools. Figure 7 A thread pool 508 is shown, where each thread can work in a worker mode or a main mode. A thread in the main mode can be configured to execute steps of the main process 520. A thread in the worker mode can be configured to execute steps of the worker process 522. Each thread in pool 508 can switch between the worker mode and the main mode.

[0121] Each thread in the thread pool 508 can be executed by a corresponding core in the core pool 509. Each thread in the thread pool 508 executes the main process 520 or the slave process 522. The thread holding the main lock of the scheduler 502 executes the main process, which starts from the scheduling function. After selecting a task in the task list 501, the task can be scheduled and the main thread can execute the FWD function 513 provided by the task to enable the results of the previously selected and scheduled tasks. This can enable checking whether the previous results are ready for the FWD function 513 first in order to forward these results. This can free up space in the pipeline 400 and prevent data from remaining in the pipeline, thereby increasing latency. For example, even if a given task has no new input task data, it can be scheduled. The input task data can include the results of previous tasks on which the given task depends and references / pointers to the instructions of the given task (for execution by the slave thread). When the slave thread is executing a previously started task, the FWD function 513 can be scheduled itself. This prevents the results of the slave function from remaining in the pipeline until a task in the pipeline stage is executed. Therefore, if the FWD function 513 has no new input task data, the main process returns to the scheduler 502. This FWD is only looped by the main thread function as a regular poll for completing tasks that have been running asynchronously on the slave cores.

[0122] Next, the INPUT SELECT and ACCUM functions of the selected main thread are run. The main thread can try to find (503) an idle slave thread in the pool 508 to offload the work to it. It does this by looping through all the threads in the thread pool 508, and once it finds a thread that is idle, it returns the thread selection to the scheduler. If it does not find a slave thread, it can release the main lock associated with the scheduler it holds and start executing the task itself, but now as a slave thread. Once the main becomes a slave, there may no longer be any thread in the thread pool doing the main process 520 of the work, because all the threads in the pool may be executing the processing tasks as slaves.

[0123] When a thread in the thread pool 508 is started, it starts as a slave thread by default and checks its input queue 505. It waits there for work to be sent by the main thread. Once it has work to do, it executes the PROC function 512 from the selected task. After the work is completed, the slave thread tries to acquire (506) the main lock from the scheduler 502 that gave it the work task. If this acquisition of the main lock is successful, it becomes the main thread running the scheduler.

[0124] The main lock function can be performed, for example, by a mutex, which is a synchronization mechanism for enforcing restricted access to a resource (in this case, the main lock). As is known to those skilled in the art, other functionally equivalent implementations using mutexes such as, for example, semaphores can be used.

[0125] Figure 8 is a flowchart of a method for performing a task according to an example of the present subject matter. For illustrative purposes, the method can be described with reference to the previous Figure 3-7 but is not limited to that implementation.

[0126] The method can include an initialization phase 600 and a task execution phase 660. The initialization phase 600 can be performed by a manager thread such as manager thread 120.

[0127] In step 601, one or more schedulers can be established by manager thread 120. Manager thread 120 can establish tasks in step 602, such as tasks of pipeline 400. These tasks can be interconnected based on their dependencies and can be associated with one of the established schedulers in step 601. This is shown in Figure 7 where tasks P1, P2, and P3 are being scheduled by the scheduler. Manager thread 120 can start each thread in thread pool 508 as a worker thread in step 603. For each of the established schedulers, manager thread 120 can select a thread from thread pool 508 in step 604 to become the main thread of that scheduler.

[0128] Thus, the initialization phase 600 can result in at least one task list being assigned to a scheduler and a thread becoming the main thread of each scheduler. After the establishment part is completed, the task execution phase 660 can be performed.

[0129] The scheduler associated with the task list in step 602 can run in step 605 via the associated main thread executed by the specified main core. This can cause the tasks in the task list to be selected (after being scheduled by the scheduler). The FWD function 413 of the previously selected task can be executed by the main thread in step 606. After executing the FWD function 413, it can be determined (query step 607) whether input task data is available. The input task data can include the results of previously executed tasks on which the selected task depends and a reference to the selected task such that the worker threads can use the reference to execute the selected task. If these results are available, then the input task data is available.

[0130] If the input task data is not available, steps 605 - 607 can be repeated. If the input task data is available, the main thread can execute INPUT SLECT and ACCUM functions 410 - 411 in step 608. It can be determined (query step 609) whether all the input data required for the selected task has been accumulated. If all the input data required for the selected task has not been accumulated, steps 605 - 609 can be repeated, otherwise, the main thread can search for available worker threads in step 610.

[0131] If (611) worker threads are available, steps 612 and 613 can be executed, otherwise, steps 615 to 620 can be executed.

[0132] In step 612, the main thread can send the selected task to the available worker thread so that the main thread can hold its main lock in step 613. After step 613, steps 605 - 620 can be repeated using the same main thread as used in the first execution of steps 605 - 620 to process another task in the list. This repetition can be executed until these tasks are executed.

[0133] In step 615, the main thread can send the selected task to its input queue and can release the main lock in step 616. In step 617, the worker thread spawned from step 616 can wait for the input queue to execute the PROC function of the selected task in step 618. This means that after releasing the main lock, there can be only worker threads in pool 508. After completing the execution of the task, the worker thread can check the main lock of the scheduler in step 619. It can be determined (query step 620) whether the worker thread has acquired the main lock. If so, the worker thread can become the main thread and steps 605 - 620 can be repeated using the new main thread to process another task in the list. This repetition can be executed until these tasks are executed. If the worker thread has not acquired the main lock (query step 620), the worker thread remains as it is and can be used in another iteration of this method.

[0134] Figure 9 FIG. 7 is a schematic diagram showing the execution of two task lists A 701 and B 702 according to an example of the present subject matter. Task list 701 includes tasks T0, T1, T2, T3, and T4. Task list 702 includes tasks TT0, TT1, and TT2. Figure 9 Also shown is the order in which the tasks of these two lists are received at scheduler 703, for example, these tasks arrive at scheduler 703 at a specific timing. Figure 9 A thread pool 708 of four threads used to execute these tasks is shown.

[0135] Thread 0 in pool 708 starts as the main thread of scheduler 703 and assigns the arriving tasks T0, T1, and TT0 to worker threads 1, 2, and 3 respectively. For the 4th arriving task T2, the main thread cannot find another worker thread because they are all running the RPOC function as indicated in Figure 9 so it starts executing task T2 after switching itself to a worker thread by releasing the main lock.

[0136] Since thread 1 is the first worker thread to complete its task after thread 0 has switched to worker mode, it acquires the idle main lock and executes scheduler 703. The scheduler executed by thread 1 can schedule task TT1 and can assign the execution of the FWD function of the previous task T0 to thread 1. Once the FWD function of the previous task T0 has been executed and thread 2 is identified as an idle thread, the resulting output can be used as the input for the PROC function of the selected task TT1 and the PROC function of TT1 can be executed by thread 2. The rotation of the main function and the execution of the worker PROC function continue in this way as indicated in Figure 9 to maximize the use of available physical resources.

[0137] Figure 9 Two time periods 705 and 706 are shown during which there is no main thread in pool 708 and all available cycles are used for the execution of the PROC function. With this method, the main thread does not have to wait during time periods 705 and 706 but can start executing the PROC function itself.

[0138] When the processing load is high, all threads in the thread pool can contribute to execute computationally complex tasks so that no cycles are wasted. As shown in Figure 9 at times 705 and 706, no main process is active and no cycles are lost in, for example, attempting to schedule a task for which there are no resources available anyway. With the rotation of the main function, only a minimal number of cycles can be used to advance the processing pipeline.

[0139] In addition, since the first worker thread to complete starts continuing the main process, no scheduling latency is lost. Therefore, new task scheduling can only be restarted when there are cycles available to advance the pipeline (at least from the previous worker thread, now the main thread).

[0140] The FWD function can always take precedence over starting a new PROC function. This keeps the buffering of results in the pipeline at a low level, which can result in lower latency. The FWD function of the running tasks is self-scheduled without new input tasks. When the system is not fully loaded, the main thread has idle cycles and it will poll the currently running PROC slave functions until the results are ready for the FWD function. This polling of asynchronously running tasks can result in minimal buffering and thus minimal latency in the pipeline. Therefore, when the processing load is very low, the idle cycles can be used to minimize the latency in the pipeline.

[0141] There can be more than one scheduler (as Figure 10 shown), and thus, more than one main thread can be active simultaneously on a single thread pool. This can be useful, for example, when the scheduler has too many tasks associated with it and the serial execution of scheduling the FWD, ACCUM, and INPUT SELECT functions becomes a bottleneck that hinders the concurrent use of all available slave threads in the pool. Further, two or more schedulers / main threads can be started on which these tasks are split. The main benefit can be that all the main threads share the use of the slave threads. This is illustrated in Figure 10 This architecture can allow the automatic sharing of processing resources as the processing load changes between schedulers.

[0142] Figure 10 is a schematic diagram showing the execution of four task lists 801, 802, 804, and 805 according to an example of the present subject matter. Figure 8 Two schedulers 803A - B are shown, each scheduler associated with two task lists and sharing a single 4-core thread pool, e.g., 708.

[0143] Task list 801 includes tasks T0 and T1. Task list 802 includes task TT0. Task list 804 includes tasks TTT0 and TTT1. Task list 805 includes task TTTT0.

[0144] In this example, two main threads, i.e., thread 0 and thread 3, concurrently run schedulers 803A and 803B respectively to enable scheduling and start tasks on the respective slave threads, i.e., thread 1 and thread 2.

[0145] The execution of scheduler 803A by thread 0 results in the selection of task T0. Thread 0 identifies thread 1 as an available slave thread for executing task T0. In another iteration, task T T0 is further selected by scheduler 803A. However, at this time, neither slave thread 1 nor 2 is available. In this case, thread 0 switches to slave mode and executes task T T0. After completing task T0, thread 1 switches to the main mode for scheduler 803A. After being executed by the new main thread 1, in another iteration, task T1 is further selected by scheduler 803A. Since thread 0 is not available, thread 1 switches to slave mode and executes task T1.

[0146] The execution of scheduler 803B by thread 3 results in the selection of task T T T0. Thread 3 identifies thread 2 as an available slave thread for executing task T T T0. In another iteration, task T T T T0 is further selected by scheduler 803B. However, at this time, no slave thread is available. In this case, thread 3 switches to slave mode and executes task T T T T0. After thread 2 completes task T T T0, it acquires the main mode, and task T T T1 is further selected by scheduler 803B when being executed by thread 2. However, at this time, no slave thread is available. In this case, thread 2 switches to slave mode and executes task T T T1.

[0147] After completing task T T0, thread 0 can become the main thread for scheduler 803A because the main lock has been released by thread 1 which was used to execute task T10.

[0148] This method can enable each thread in the thread pool to become the main thread of any scheduler according to the load. Multiple main threads can be useful when there are many high-speed tasks and / or when the serial execution of the main function limits the use of parallelism in the thread pool.

Claims

1. A multi-processor system, comprising: A processing unit; One or more additional processing units, wherein the processing unit and the one or more additional processing units are homogeneous processing cores, and each processing unit is configured to execute computationally complex functions and computationally simple functions; and At least one memory for storing instructions, wherein the instructions, when executed by the processing unit, cause the processing unit to: Obtain a primary role for executing a primary functional portion of a set of tasks; Execute a primary function for searching for available processing units among the one or more additional processing units; Wherein, if an available processing unit is found, then Retain the primary role and control the found available processing unit to execute a secondary functional portion of the set of tasks, and If no available processing unit is found, then Release the primary role and further execute the secondary functional portion of the set of tasks, Wherein the primary function includes a primary-to-secondary switching function for releasing the primary role, and the secondary function includes a secondary-to-primary switching function for obtaining the primary role after termination of the secondary function, Wherein the instructions, when executed by one or more additional processing units that do not have the primary role, cause the one or more additional processing units to: after completion of the execution of the secondary function, attempt to obtain the primary role by repeatedly executing the secondary-to-primary switching function for obtaining the primary role.

2. The multi-processor system according to claim 1, wherein, The instructions, when executed by the processing unit, cause the processing unit to: release the primary role by releasing a primary lock, wherein the obtaining of the primary role includes obtaining the released primary lock.

3. The multi-processor system according to any one of the preceding claims, wherein, The primary function includes a scheduler for scheduling the set of tasks and an output forwarding function for providing results of previous tasks on which the tasks depend, and the secondary function includes a task processing function for at least one task.

4. The multi-processor system according to claim 1, further comprising: A thread pool and a manager thread, the manager thread being configured to establish worker threads and at least one main thread in the pool, wherein the additional processing units are configured to use the main threads in the pool to execute the primary function, and to use the worker threads in the pool to execute the secondary function.

5. A network node, comprising the multi-processor system according to any one of claims 1-4.

6. A method, comprising: Obtaining, by a processing unit of a multi-processor system, a primary role, wherein the multi-processor system includes one or more additional processing units, and wherein the processing unit and the one or more additional processing units are homogeneous processing cores, and each processing unit is configured to execute computationally complex functions and computationally simple functions; Executing, by the processing unit, a primary functional portion of a set of tasks, including: Searching for available processing units among the one or more additional processing units; wherein, if an available processing unit is found, retaining the primary role and controlling the found available processing unit to execute a secondary functional portion of the set of tasks, and if no available processing unit is found, releasing the primary role and having the processing unit execute the secondary functional portion of the set of tasks, Among them, the main function includes a master-to-slave switching function for releasing the master role, and the slave function includes a slave-to-master switching function for acquiring the master role after the termination of the slave function. Among them, after the execution of the slave function is completed, one or more additional processing units without the master role attempt to acquire the master role by repeatedly executing the slave-to-master switching function for acquiring the master role.

7. The method according to claim 6, wherein The execution of the main function is performed using the main thread in the thread pool of the multi-processor system, and the execution of the slave function is performed using the slave thread in the thread pool.

8. The method according to claim 7, further comprising: The manager thread in the pool is used to establish the main thread and the slave thread in the pool.

9. A computer program, including instructions stored thereon for at least performing the following operations: Obtain a main role by a processing unit of a multi-processor system, where The multi-processor system includes one or more additional processing units, and among them, the processing unit and the one or more additional processing units are homogeneous processing cores, and each processing unit is configured to execute computationally complex functions and computationally simple functions; The main function part of a set of tasks executed by the processing unit includes: Searching for available processing units among the one or more additional processing units; wherein, if an available processing unit is found, the master role is maintained and the found available processing unit is controlled to execute the slave function part of the set of tasks, and if no available processing unit is found, the master role is released and the slave function part of the set of tasks is executed by the processing unit. Among them, the main function includes a master-to-slave switching function for releasing the master role, and the slave function includes a slave-to-master switching function for acquiring the master role after the termination of the slave function. Among them, after the execution of the slave function is completed, one or more additional processing units without the master role attempt to acquire the master role by repeatedly executing the slave-to-master switching function for acquiring the master role.

Citation Information

Patent Citations

  • Multicore system and activating method

    US20130024870A1

  • Distributed task scheduling for symmetric multiprocessing environments

    US7810094B1