Multi-core system and method of controlling operation of multi-core system

By monitoring the task and core execution delay time, relocating tasks or adjusting processor core power, the problem of low task scheduling efficiency in multi-core systems is solved, and efficient task scheduling and resource utilization are achieved.

CN112214292BActive Publication Date: 2025-10-03SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010654333.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-11
Filing Date
2020-07-08
Publication Date
2025-10-03
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

Existing multi-core systems have low task scheduling efficiency when processing irregular and highly data-dependent tasks, especially when there are task execution delays and core execution delays, which can easily lead to task starvation.

Method used

By monitoring the task execution delay time and core execution delay time, the task scheduler and control logic are used to relocate tasks or adjust the power level of the processor core to optimize task scheduling.

Benefits of technology

It effectively reduces or eliminates task starvation, enables efficient task scheduling, and improves system performance and resource utilization even in tasks with high data dependencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112214292B_ABST
    Figure CN112214292B_ABST
Patent Text Reader

Abstract

A method for controlling the operation of a multi-core system including multiple processor cores includes: monitoring task execution delay times of tasks respectively assigned to the multiple processor cores, monitoring core execution delay times of the multiple processor cores, and controlling the operation of the multi-core system based on the task execution delay times and the core execution delay times.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority from Korean Patent Application No. 10-2019-0083853 filed on July 11, 2019, in the Korean Intellectual Property Office (KIPO), the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] Example embodiments relate generally to semiconductor integrated circuits, and more particularly, to a multi-core system and a method of controlling the operation of the multi-core system. Background Art

[0004] Multiple tasks can be executed substantially simultaneously on a computer system including multiple processors or including processors with multiple cores. Task scheduling schemes for executing tasks may include round-robin schemes, priority-based schemes (e.g., random multiple access (RMA) scheduling), and deadline-based schemes (e.g., earliest deadline first (EDF) scheduling). Runnable tasks can be sequentially assigned to processor cores according to a round-robin scheme, or the next task to be processed can be determined based on urgency or importance according to priority-based task scheduling. In addition, task scheduling can be performed statically, such as RMA scheduling, or dynamically, such as EDF scheduling, based on the frequency and / or deadline of the tasks. However, the efficiency of these scheduling schemes may be reduced when the tasks are irregular and highly data-dependent. Summary of the Invention

[0005] At least one exemplary embodiment of the inventive concept may provide a multi-core system and a method of controlling an operation of the multi-core system for efficient task scheduling.

[0006] According to an exemplary embodiment conceived in the present invention, a method for controlling the operation of a multi-core system including multiple processor cores includes: monitoring task execution delay times of tasks respectively assigned to the multiple processor cores; monitoring core execution delay times of the multiple processor cores; and controlling the operation of the multi-core system based on the task execution delay times and the core execution delay times.

[0007] According to an exemplary embodiment of the present invention, a multi-core system includes: a multi-core processor including multiple processor cores; a first control logic configured to monitor task execution delay times of tasks respectively assigned to the multiple processor cores and core execution delay times of the multiple processor cores; and a second control logic configured to control the operation of the multi-core system based on the task execution delay times and the core execution delay times.

[0008] According to an exemplary embodiment of the present invention, a method for controlling the operation of a multi-core system including a plurality of processor cores includes: monitoring task execution delay times of tasks respectively assigned to the plurality of processor cores, the task execution delay time of a given task assigned to a given processor core among the plurality of processor cores corresponding to a standby time, the standby time being a time during which the given task is not executed by the given processor core after the given task is stored in a task queue of the given processor core; monitoring core execution delay times of the plurality of processor cores, the core execution delay time of the given processor core corresponding to a sum of the task execution delay times associated with the given processor core or a maximum task execution delay time among the task execution delay times associated with the given processor core; determining whether a task execution delay has occurred for the task assigned to each processor core or a core execution delay has occurred for each processor core based on the task execution delay time of the task assigned to each processor core and the core execution delay time of each processor core; and when it is determined that the task execution delay of the given task or the core execution delay of the given processor core has occurred, relocating the given task assigned to the given processor core to another processor core among the plurality of processor cores or increasing the power level of the given processor core. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Exemplary embodiments of the present inventive concept will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.

[0010] Figure 1 is a block diagram illustrating a multi-core system according to an exemplary embodiment of the inventive concept.

[0011] Figure 2 is a flowchart illustrating a method of controlling operations of a multi-core system according to an exemplary embodiment of the inventive concept.

[0012] Figure 3 is a block diagram illustrating a multi-core system according to an exemplary embodiment of the inventive concept.

[0013] Figure 4 It shows that Figure 3 FIG. 1 is a diagram of an example embodiment of a task scheduler implemented in a multi-core system.

[0014] Figure 5 is a flowchart illustrating a method of controlling operations of a multi-core system according to an exemplary embodiment of the inventive concept.

[0015] Figures 6 to 9 is a diagram illustrating a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0016] Figure 10 It shows that the Figure 3 A block diagram of an example embodiment of a processor in a multi-core system.

[0017] Figure 11 and Figure 12 is a diagram illustrating an example of a task execution delay time used in a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0018] Figure 13 and Figure 14 is a diagram illustrating a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0019] Figure 15 is a block diagram illustrating an electronic system according to an exemplary embodiment of the inventive concept. DETAILED DESCRIPTION

[0020] The present inventive concept will be described more fully hereinafter with reference to the accompanying drawings, in which some exemplary embodiments of the present inventive concept are shown.In the drawings, like reference numerals refer to like elements throughout. Figure 1 is a block diagram illustrating a multi-core system according to an exemplary embodiment of the present inventive concept, and Figure 2 is a flow chart illustrating a method of controlling the operation of a multi-core system.

[0021] Figure 1 1 shows a schematic configuration of a multi-core system according to an exemplary embodiment of the present inventive concept. Figure 1 The multi-core system 1000 includes a processor 110, a task scheduler 135, and a performance controller PFMC 140. The multi-core system 1000 may include other components, which will be referred to later. Figure 3 A more detailed configuration of the multi-core system 1000 is described.

[0022] The multi-core system 1000 may be implemented as a system on a chip 1 that may be included in various computing devices. The computing device may be one of the following: a mobile phone, a smart phone, an enterprise digital assistant (EDA), a digital camera, a digital video camera, a portable multimedia player (PMP), a personal navigation device or portable navigation device (PND), a mobile internet device (MID), a wearable computer, an Internet of Things (IoT) device, an Internet of Everything (IOE), or an e-book reader.

[0023] The multi-core system 1000 can send data and task requests to a host device (not shown) and receive data and task requests from the host device via an interface (e.g., interface circuitry). For example, a task request can include a request to perform an operation on data. For example, the interface can be connected to the host device via a parallel AT attachment (PATA) bus, a serial AT attachment (SATA) bus, a small computer system interface (SCSI), a universal serial bus (USB), a peripheral component interconnect express (PCIe), or the like.

[0024] The processor 110 may include a plurality of processor cores C1 to C8 and a plurality of task queues TQ1 to TQ8 respectively assigned to the plurality of processor cores C1 to C8. Figure 1 Eight processor cores C1-C8 and eight task queues are shown, but the inventive concept is not limited thereto. For example, the processor 110 may include fewer than eight processor cores or more than eight processor cores, and the processor 110 may include fewer than eight task queues or more than eight task queues.

[0025] The processor cores C1 to C8 may be homogeneous processor cores or heterogeneous processor cores. When the processor cores C1 to C8 are homogeneous processor cores, each core is of the same type. When the processor cores C1 to C8 are heterogeneous processor cores, some cores are of different types.

[0026] When the processor cores C1 to C8 are heterogeneous processor cores, they can be classified into a first cluster CL1 and a second cluster CL2. Among the processor cores C1 to C8, the first cluster CL1 may include high-performance cores C1 to C4 having a first processing speed, and the second cluster CL2 may include low-performance cores C5 to C8 having a second processing speed lower than the first processing speed. Figure 1 Each cluster is shown to include the same number of processor cores, but the present invention is not limited thereto. For example, the first cluster CL1 may include more processing cores than the second cluster CL2, or the first cluster CL1 may include fewer processing cores than the second cluster CL2. In addition, there may be more than two clusters.

[0027] According to an exemplary embodiment, the processor cores C1 to C8 have a per-core dynamic voltage and frequency scaling (DVFS) architecture. In the per-core DVFS architecture, the processor cores C1 to C8 may include different power domains, and voltages with different levels and clock signals with different frequencies may be provided to the processor cores C1 to C8. The power supply to the processor cores C1 to C8 may be blocked by a hot plug-out scheme. In other words, a portion of the processor cores C1 to C8 may perform assigned tasks, and the power supply to another portion of the processor cores C1 to C8 in an idle state may be blocked. Conversely, when the workload is too heavy for the powered processor cores, at least one processor core in an idle state may be powered to perform tasks, which is referred to as "hot plug-in".

[0028] The task scheduler 135 may include an execution delay tracker EDT and control logic LDL.

[0029] The task scheduler 135 and the PFMC 140 can each be implemented as hardware, software, or a combination of hardware and software. For example, the hardware can be a processor or a logic circuit including various logic gates. It should be understood that the software can be a computer program product embodied in one or more computer-readable media having computer-readable program code embodied thereon. The computer-readable program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be any tangible medium that can contain or store a program for use by or with an instruction execution system, apparatus, or device.

[0030] The task scheduler 135 may allocate or distribute tasks to task queues TQ1 to TQ8 based on the task execution delay time and the core execution delay time, as will be referred to below. Figure 2 One or more of the task queues TQ1 to TQ8 may be implemented as hardware (eg, latches, buffers, shift registers, etc.) within a processor core, or may be implemented as a data structure included in an operating system (OS) kernel.

[0031] refer to Figure 1 and Figure 2 , the execution delay tracker EDT monitors the task execution delay times of the tasks respectively assigned to the plurality of processor cores C1 to C8 ( S100 ).

[0032] In an exemplary embodiment, the task execution delay time corresponds to a standby time during which the task is not executed by the processor core after the task is stored in the task queue of the corresponding processor core.

[0033] In the exemplary embodiment, as will be referred to below Figure 7 As described above, the task execution delay time may correspond to the start standby time from the time point when the task is stored in the task queue of the processor core to the start time point when the corresponding processor core starts executing the task. For example, if a task is placed in the task queue TQ1 at time 0 and then the processor core C1 starts processing the task at time 100, the task execution delay time will be 100.

[0034] In an exemplary embodiment, as will be referred to later Figure 11 and Figure 12 As described above, the task execution delay time corresponds to the sum of the start standby time and the pause time, wherein the start standby time is the time from the time point when the given task is stored in the task queue of the given processor core to the start time point when the given processor core starts to execute the given task, and the pause time is the time when the given processor core stops executing the given task after the start time point. For example, if a task is placed in the task queue TQ1 at time 0, the processor core C1 starts executing the task at time 100, and then the processor C1 needs to stop executing the task at time 150 until it obtains access to the necessary resources at time 300, then the start standby time is 100, the pause time is 150, and the task execution delay time is 250 (i.e., the sum of 100 and 150 = 250).

[0035] Furthermore, the execution delay tracker EDT monitors core execution delay times of the plurality of processor cores ( S200 ).

[0036] In an exemplary embodiment of the present inventive concept, the core execution delay time corresponds to the sum of the task execution delay times associated with a given processor core. The sum can be calculated for a given period. For example, the sum can be calculated based on the task execution delay times associated with a given processor core that occurred within a given period.

[0037] In an exemplary embodiment of the inventive concept, the core execution delay time corresponds to a maximum task execution delay time among task execution delay times associated with a given one processor core.

[0038] The control logic LDL controls the operation of the multi-core system based on the task execution delay time and the core execution delay time ( S300 ).

[0039] The control logic LDL may determine whether a task assigned to each processor core has a task execution delay or whether a core execution delay has occurred in each processor core based on the task execution delay time of the task assigned to each processor core and the core execution delay time associated with each processor core.

[0040] In an exemplary embodiment, the control logic LDL performs task scheduling based on task execution delay times and core execution delay times to efficiently allocate runnable tasks to processor cores C1 to C8. For example, when it is determined that a processor core has experienced task execution delay or core execution delay, the control logic LDL may relocate a portion of the tasks allocated to the corresponding processor core to another processor core.

[0041] In an exemplary embodiment, the control logic LDL controls the power levels of the processor cores C1-C8 based on the task execution delay time and the core execution delay time. The power level can be represented by the operating voltage and / or operating frequency of each processor core. For example, when it is determined that a processor core has experienced a task execution delay or a core execution delay, the control logic LDL can quickly eliminate the execution delay state associated with the corresponding processor core by increasing the power level of the corresponding processor core or by increasing the operating frequency of the corresponding processor core.

[0042] In this way, a multi-core system and a method for controlling the operation of a multi-core system according to at least one exemplary embodiment of the present invention can reduce or eliminate the situation where tasks starve because specific tasks have not been processed for a long time by using task execution delay time and core execution delay time, and efficient task scheduling can be achieved even for tasks with high data dependencies.

[0043] Figure 3 is a block diagram illustrating a multi-core system according to an exemplary embodiment of the inventive concept.

[0044] refer to Figure 3 The multi-core system 1000 includes a system on chip (SoC), a working memory 130, a display device (LCD) 152, a touch panel 154, a storage device 170, a power management integrated circuit (PMIC), etc. The SoC may include a central processing unit (CPU) 110, a DRAM controller 120 (e.g., a control circuit), a performance controller 140 (e.g., a control circuit), a user interface controller (UI controller (e.g., a control circuit)) 150, a storage interface 160 (e.g., an interface circuit) and an accelerator 180 (e.g., a graphics accelerator), a power management unit (PMU) 144, and a clock management unit (CMU) 146. It should be understood that the components of the multi-core system 1000 are not limited to Figure 3 For example, the multi-core system 1000 may also include a hardware codec or security block for processing image data.

[0045] The processor 110 executes software (e.g., applications, operating systems (OS), and device drivers) for the multi-core system 1000. The processor 110 can execute an operating system (OS) that can be loaded into the working memory 130. The processor 110 can execute various applications driven by the operating system (OS). The processor 110 can be provided as a homogeneous multi-core processor or a heterogeneous multi-core processor. A multi-core processor is a computing component that includes at least two independently drivable processors (hereinafter referred to as "cores"). Each core can independently read and execute program instructions.

[0046] Each processor core of the processor 110 may include multiple power domains that operate with independent drive clock signals and independent drive voltages. The drive voltage and drive clock signal provided to each processor core can be cut off or provided in units of a single core. Hereinafter, the operation of cutting off the drive voltage and drive clock provided to each power domain from a specific core will be referred to as "hot plug-out". The operation of providing the drive voltage and drive clock to a specific core will be referred to as "hot plug-in". In addition, the frequency of the drive clock signal provided to each power domain and the level of the drive voltage can vary according to the processing load of each core. For example, as the time required to process a task becomes longer, each core can be controlled by dynamic voltage frequency scaling (hereinafter referred to as "DVFS"), which increases the frequency of the drive clock signal or the level of the drive voltage provided to the corresponding power domain. According to an exemplary embodiment of the present invention, hot plugging and hot unplugging are performed with reference to the level of the drive voltage and the frequency of the drive clock of the processor 110 adjusted by DVFS.

[0047] The kernel of the operating system (OS) may monitor the number of tasks in the task queue and the driving voltage and driving clock signal of the processor 110 at specific time intervals to control the processor 110. In addition, the kernel of the operating system (OS) may control hot insertion and hot removal of the processor 110 with reference to the monitored information.

[0048] The DRAM controller 120 provides an interface connection between the working memory 130 and the system on chip (SoC). The DRAM controller 120 can access the working memory 130 based on a request from the processor 110 or another intellectual property (IP) block. An IP block can be a reusable unit of logic, cell, or integrated circuit layout design, which is the intellectual property of a party. For example, the DRAM controller 120 can write data to the working memory 130 based on a write request from the processor 110. Alternatively, the DRAM controller 120 can read data from the working memory 130 based on a read request from the processor 110 and send the read data to the processor 110 or the storage interface 160 via a data bus.

[0049] During the boot operation, an operating system (OS) or basic applications can be loaded into the working memory 130. For example, during the boot of the multi-core system 1000, the OS image stored in the storage device 170 is loaded into the working memory 130 based on the boot sequence. The overall input / output operation of the multi-core system 1000 can be supported by the operating system (OS). Similarly, applications can be loaded into the working memory 130 for selection by the user or to provide basic services. In addition, the working memory 130 can be used as a buffer memory to store image data provided from an image sensor such as a camera. The working memory 130 can be a volatile memory (such as static random access memory (SRAM) and dynamic random access memory (DRAM)) or a non-volatile memory (such as phase change random access memory (PRAM), magnetoresistive random access memory (MRAM), resistive random access memory (ReRAM), ferroelectric random access memory (FRAM), NOR flash memory).

[0050] The performance controller 140 can adjust the operating parameters of the system on chip (SoC) according to the control request provided from the kernel of the operating system (OS). For example, the performance controller 140 can adjust the level of DVFS to enhance the performance of the system on chip (SoC). For example, the performance controller 140 can control the drive mode of a multi-core processor such as Big.LITTLE (a heterogeneous computing architecture developed by ARM Holdings) of the processor 110 according to the request of the kernel. In this case, the performance controller 140 may include a performance table 142 to set the drive voltage and the frequency of the drive clock signal therein. The performance controller 140 can control the PMU 144 and the CMU 146 connected to the PMIC 200 to provide a determined drive voltage and a determined drive clock signal to each power domain.

[0051] The user interface controller 150 controls user input and output of the user interface device. For example, the user interface controller 150 may display a keyboard screen for inputting data to the LCD 152 under the control of the processor 110. Alternatively, the user interface controller 150 may control the LCD 152 to display data requested by the user. The user interface controller 150 may decode data provided from a user input device such as a touch panel 154 into user input data.

[0052] The storage interface 160 accesses the storage device 170 according to the request of the processor 110. For example, the storage interface 160 provides an interface connection between the system on chip (SoC) and the storage device 170. For example, data processed by the processor 110 is stored in the storage device 170 through the storage interface 160. Alternatively, the data stored in the storage device 170 can be provided to the processor 110 through the storage interface 160.

[0053] The storage device 170 is provided as a storage medium of the multi-core system 1000. The storage device 170 can store applications, OS images, and various types of data. The storage device 170 can be provided as a memory card (e.g., MMC, eMMC, SD, MicroSD, etc.). The storage device 170 may include a NAND-type flash memory with a high-capacity storage capability. Alternatively, the storage device 170 may include a next-generation non-volatile memory, such as PRAM, MRAM, ReRAM, FRAM, or NOR-type flash memory. According to an exemplary embodiment of the present inventive concept, the storage device 170 may be an embedded memory incorporated in a system on chip (SoC).

[0054] The accelerator 180 may be provided as a separate intellectual property (IP) component for increasing the processing speed of multimedia data. For example, the accelerator 180 may be provided as an intellectual property (IP) component for enhancing the processing performance of text, audio, still images, animation, video, two-dimensional data, or three-dimensional data.

[0055] System interconnect 190 may be a system bus for providing an on-chip network in a system on chip (SoC). System interconnect 190 may include, for example, a data bus, an address bus, and a control bus. The data bus is a data transmission path. It may also provide a memory access path to working memory 130 or storage device 170. The address bus provides a path for exchanging addresses between intellectual property (IP). The control bus provides a path for transmitting control signals between intellectual property (IP). However, the configuration of system interconnect 190 is not limited to the above description, and system interconnect 190 may also include an arbitration device for efficient management.

[0056] Figure 4 It shows that Figure 3 FIG. 1 is a diagram of an example embodiment of a task scheduler implemented in a multi-core system.

[0057] Figure 4 Shown Figure 3 The software structure of the multi-core system 1000 is shown in FIG. Figure 3The software layer structure of the multi-core system 1000 loaded into the working memory 130 and driven by the CPU 110 can be divided into an application 132 and a kernel 134. The operating system (OS) may also include one or more device drivers for managing various devices, such as memory, modems, and image processing devices.

[0058] Application 132 may be an upper layer of software that serves as a basic service driver or is driven by user requests. Multiple application programs, App0, App1, and App2, may be executed simultaneously to provide various services. Application programs, App0, App1, and App2, may be executed by CPU 110 after being loaded into working memory 130. For example, when a user requests to play a video file, an application program (e.g., a video player) is executed to play the video file. The executed player may then generate a read request or a write request to storage device 170 to play the video file requested by the user.

[0059] The kernel 134, as a component of the operating system (OS), can perform control operations between the application 132 and the hardware. The kernel 134 can provide program execution, interrupts, multitasking, and memory management, and can include a file system and device drivers. A scheduler 135 can be provided as part of the kernel 134.

[0060] Scheduler 135 monitors and manages the task queues of each processor core. When multiple tasks are executed simultaneously, the task queue is a queue of active tasks. For example, tasks in a task queue can be processed more quickly by processor 110 than other tasks. Scheduler 135 can refer to the task information loaded into the task queue to determine subsequent processing. For example, scheduler 135 determines the priority of CPU resources based on the values ​​in the task queue. In the Linux kernel, multiple task queues correspond to multiple processor cores.

[0061] The scheduler 135 may allocate tasks corresponding to the task queues to the corresponding cores. Tasks loaded into the task queues and executed by the processor 110 are called runnable tasks.

[0062] According to an exemplary embodiment of the inventive concept, the scheduler 135 includes an execution delay tracker EDT and control logic LDL. The execution delay tracker EDT may also be referred to as control logic or implemented by control logic (eg, second control logic).

[0063] The execution delay tracker EDT can monitor the task execution delay times of the tasks respectively assigned to the multiple processor cores and the core execution delay times of the multiple processor cores. The control logic LDL can control the operation of the multi-core system based on the task execution delay times and the core execution delay times.

[0064] Figure 5 is a flowchart illustrating a method of controlling operations of a multi-core system according to an exemplary embodiment of the inventive concept.

[0065] refer to Figure 1 、 Figure 4 and Figure 5 The execution delay tracker EDT of the task scheduler 135 monitors the task execution delay times TTD of the tasks respectively assigned to the plurality of processor cores C1 to C8 and the core execution delay times TCD of the plurality of processor cores C1 to C8 ( S10 ).

[0066] The control logic LDL of the task scheduler 135 determines whether a task execution delay TEXED occurs in the task assigned to each processor core based on the task execution delay time TTD and the core execution delay time TCD ( S20 ).

[0067] When it is determined that a task execution delay TEXED occurs in a given processor core ( S20 : ​​Yes), the control logic LDL relocates a portion of the tasks assigned to the given processor core to another processor core ( S30 ).

[0068] When it is determined that the task execution delay TEXED has not occurred ( S20 : ​​No), the control logic LDL determines whether a core execution delay CEXED has occurred with respect to the task assigned to each processor core based on the task execution delay time TTD and the core execution delay time TCD ( S40 ).

[0069] When it is determined that the core execution delay CEXED has occurred (S40: YES), the control logic LDL increases (raises) the power level of each processor core in which the core execution delay CEXED has occurred (S50). In an alternative embodiment, when it is determined that the core execution delay CEXED has occurred (S40: YES), the control logic LDL raises the operating frequency of each processor core in which the core execution delay CEXED has occurred.

[0070] When it is determined that the core execution delay CEXED does not occur ( S40 : No), the control logic LDL maintains the previous scheduling cycle and power level.

[0071] According to exemplary embodiments of the present invention, the task execution delay time of each task and the core execution delay time of each processor core can be monitored, and the degree of task execution delay can be managed. Based on the degree of task execution delay, tasks can be effectively distributed to processor cores and / or processing capacity or power levels can be controlled to optimize performance and resource utilization.

[0072] Task scheduling can be based on positive feedback from previous task scheduling. Positive feedback represents the processing time or utilization of tasks when they are normally scheduled.

[0073] In contrast, task scheduling according to at least one embodiment of the inventive concept is based on negative feedback. In an exemplary embodiment, negative feedback represents standby time or opportunity cost when tasks are not normally scheduled.

[0074] Task scheduling based on positive feedback can ensure the required performance in most cases, but depending on the task characteristics and the system's operating environment, it may lead to starvation of specific tasks or a degradation of user experience. At least one exemplary embodiment of the present inventive concept can use such negative feedback to minimize task execution delays and resource utilization.

[0075] In this way, a multi-core system and a method for controlling the operation of a multi-core system according to at least one exemplary embodiment of the present invention can reduce or eliminate the situation where tasks starve because specific tasks have not been processed for a long time by using task execution delay time and core execution delay time, and efficient task scheduling can be achieved even for tasks with high data dependencies.

[0076] Figures 6 to 9 is a diagram illustrating a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0077] Figure 6 An example of task scheduling performed without using the method designed according to the present invention is shown. For example, the first task A, the second task B, and the third task C can operate interactively, and the fourth task D can have a higher priority than the first task A, the second task B, and the third task C.

[0078] refer to Figure 6 The first task A, the second task B, the third task C and the fourth task D are periodically assigned to the first processor core C1 and processed by it according to scheduling cycles with the same time interval. Figure 6 An example of tasks assigned to the first processor core C1 and the order in which the tasks are executed in the first to fifth scheduling periods PSCH1 to PSCH5 is shown.

[0079] In the first scheduling period PSCH1 , since the execution time of the first task A, the second task B and the third task C is shorter than the first scheduling period PSCH1 , the tasks can be successfully executed, and the task scheduling of the second scheduling period PSCH2 can be performed normally.

[0080] If the fourth task D is additionally assigned to the first processor core C1 in the second scheduling cycle PSCH2, the execution time of the first task A, the second task B, the third task C and the fourth task D is longer than the second scheduling cycle PSCH2, and the task execution may not be successful due to task execution delay or core execution delay.

[0081] As a result, the control logic LDL does not perform task scheduling for the third scheduling period PSCH3. If the previous scheduling is maintained, this inefficient task scheduling will be repeated during the fourth scheduling period PSCH4 and the fifth scheduling period PSCH5.

[0082] Figure 7 、 Figure 8 and Figure 9 An example of performing task scheduling using a method according to an exemplary embodiment of the present invention is shown. For example, a first task A, a second task B, and a third task C may operate interactively, and a fourth task D may have a higher priority than the first task A, the second task B, and the third task C.

[0083] refer to Figure 7 The first task A, the second task B, the third task C and the fourth task D can be periodically assigned to the first processor core C1 and processed by it according to the scheduling cycle with the same time interval. Figure 6 An example of tasks assigned to the first processor core C1 and the order in which the tasks are executed in the first to fifth scheduling periods PSCH1 to PSCH5 is shown.

[0084] The execution delay tracker EDT of the task scheduler 135 monitors the task execution delay times TTDij of the tasks assigned to the plurality of processor cores C1 to C8, respectively. Here, i is the index of the task corresponding to A, B, C, or D, and j is the index of the scheduling cycle corresponding to 1, 2, 3, 4, or 5.

[0085] The task execution delay time TTDij associated with a task of a given processor core may correspond to a standby time, which is a time during which the task is not executed by the given processor core after being stored in the task queue of the given processor core.

[0086] In an exemplary embodiment, Figure 7 As shown, the execution delay tracker EDT can provide a start standby time as each task execution delay time TTDij, which is the time during which the task is not executed by the corresponding processor core after the task is stored in the task queue of the corresponding processor core. In an exemplary embodiment, the task execution delay time is not calculated when task scheduling is not performed.

[0087] In an exemplary embodiment of the present inventive concept, the execution delay tracker EDT may provide the sum of the task execution delay times assigned to each processor core as the per-core execution delay time. Figure 7In the example, the execution delay tracker EDT can provide (TTDA1+TTDB1+TTDC1) in the first scheduling cycle PSCH1, (TTDA2+TTDB2+TTDC2+TTDD2) in the second scheduling cycle PSCH2, (TTDA4+TTDB4+TTDC4+TTDD4) in the fourth scheduling cycle PSCH4, and (TTDA5+TTDB5+TTDC5+TTDD5) in the fifth scheduling cycle PSCH5 as the core execution delay time of the first processor core C1.

[0088] In an exemplary embodiment of the inventive concept, the execution delay tracker EDT may provide a maximum task execution delay time among the task execution delay times assigned to each processor core as the per-core execution delay time. Figure 7 In the example, the execution delay tracker EDT may provide TTDC1 in the first scheduling cycle PSCH1, TTDC2 in the second scheduling cycle PSCH2, TTDC4 in the fourth scheduling cycle PSCH4, and TTDC5 in the fifth scheduling cycle PSCH5 as the core execution delay time of the first processor core C1.

[0089] Back to Figure 7 , in the first scheduling cycle PSCH1, since the execution time of the first task A, the second task B and the third task C is shorter than the first scheduling cycle PSCH1, the tasks can be successfully executed, and the task scheduling in the second scheduling cycle PSCH2 can be executed normally.

[0090] If the fourth task D is additionally assigned to the first processor core C1 in the second scheduling cycle PSCH2, the execution time of the first task A, the second task B, the third task C, and the fourth task D will be longer than that of the second scheduling cycle PSCH2. Due to task execution delay or core execution delay, task execution may not be successful. As a result, the control logic LDL does not perform task scheduling for the third scheduling cycle PSCH3.

[0091] In this case, according to an exemplary embodiment of the inventive concept, the control logic LDL of the task scheduler 135 determines whether a task execution delay or a core execution delay occurs, and then increases (raises) the power level of the processor core.

[0092] Figure 7The power level is increased, for example, so that the operating frequency of the first processor core C1 increases from the first frequency FR1 to the second frequency FR2 during the fourth scheduling period PSCH4. Due to the frequency increase, the execution time of tasks can be shortened, and task execution can be successful during the fourth scheduling period PSCH4 and the fifth scheduling period PSCH5. For example, due to the increase in operating frequency, the first task A, the second task B, the third task C, and the fourth task D can all be successfully executed during the fourth scheduling period PSCH4 and the fifth scheduling period PSCH5. In this way, efficient task scheduling can be performed according to example embodiments.

[0093] refer to Figure 8 The first task A, the second task B, the third task C and the fourth task D are periodically assigned to the first processor core C1 and processed by it according to scheduling cycles with the same time interval. Figure 6 Examples of tasks assigned to the first processor core C1 and the second processor core C2 and the order in which the tasks are executed in the first to fifth scheduling periods PSCH1 to PSCH5 are shown.

[0094] In the following, with Figure 7 The related repeated descriptions are omitted.

[0095] Reference again Figure 8 , in the first scheduling cycle PSCH1, since the execution time of the first task A, the second task B and the third task C is shorter than the first scheduling cycle PSCH1, the tasks can be successfully executed, and the task scheduling in the second scheduling cycle PSCH2 can be executed normally.

[0096] If the fourth task D is additionally assigned to the first processor core C1 in the second scheduling cycle PSCH2, the execution time of the first task A, the second task B, the third task C and the fourth task D is longer than the second scheduling cycle PSCH2, and the task execution may not be successful due to task execution delay or core execution delay.

[0097] In this case, according to an exemplary embodiment of the present invention, the control logic LDL of the task scheduler 135 determines whether a task execution delay or a core execution delay occurs for a given task assigned to a given processor core among the processor cores, and then relocates the given task assigned to the given processor core to another processor core if a task execution delay or a core execution delay occurs.

[0098] Figure 8 Repositioning of tasks is shown such that in the fourth scheduling cycle PSCH4 the first task A, the second task B, and the third task C remain allocated to the first processor core C1, and the fourth task D is repositioned for allocation to the second processor core C2.

[0099] As such, due to this relocation of the tasks, task execution may succeed during the fourth scheduling period PSCH4 and the fifth scheduling period PSCH5 , and efficient task scheduling may be performed according to example embodiments.

[0100] The control logic LDL can control the task scheduling cycle or power level of the multi-core system based on changes in task execution latency or core execution latency. For example, when the task execution time of the first task A, the second task B, and the third task C is greatly increased to an allowable level by adding a fourth task D, the control logic LDL can relocate the fourth task D to another processor core.

[0101] refer to Figure 9 The first task A, the second task B, the third task C and the fourth task D are periodically assigned to the first processor core C1 and processed by it according to scheduling cycles with the same time interval. Figure 6 An example of tasks assigned to the first processor core C1 and the order in which the tasks are executed in the first to fifth scheduling periods PSCH1 to PSCH5 is shown.

[0102] In the following, with Figure 7 The related repeated descriptions are omitted.

[0103] Reference again Figure 9 , in the first scheduling cycle PSCH1, since the execution time of the first task A, the second task B and the third task C is shorter than the first scheduling cycle PSCH1, the tasks can be successfully executed, and the task scheduling in the second scheduling cycle PSCH2 can be executed normally.

[0104] If the fourth task D is additionally assigned to the first processor core C1 in the second scheduling cycle PSCH2, the execution time of the first task A, the second task B, the third task C and the fourth task D is longer than the second scheduling cycle PSCH2, and the task execution may not be successful due to task execution delay or core execution delay.

[0105] In this case, according to an exemplary embodiment of the present invention, the control logic LDL of the task scheduler 135 determines whether a task execution delay or a core execution delay occurs for each processor core, and then increases the scheduling cycle of the corresponding processor core for the processor core that has the task execution delay or the core execution delay.

[0106] Figure 9 The increase in the scheduling period is shown such that the fourth scheduling period PSCH4' is increased to be longer than the previous scheduling periods PSCH1 to PSCH4. These embodiments may be employed when the delay requirement level of the tasks is relatively low.

[0107] Thus, due to this increase in the scheduling period, task execution can be successful during the fourth scheduling period PSCH4 and the fifth scheduling period PSCH5, and effective task scheduling can be performed according to the example embodiment. For example, if the scheduling period was previously of a first duration and a task execution delay or a core execution delay occurred during a given scheduling period (e.g., PSCH2), the scheduling period can be increased to a second duration that is greater than the first duration. In an exemplary embodiment, the scheduling period (e.g., PSCH3) following the scheduling period in which the delay was found (e.g., PSCH2) is skipped, and the increase is applied to the next scheduling period (e.g., PSCH4′) and those periods occurring after the next scheduling period (e.g., PSCH5′). For example, the fourth scheduling period PSCH4 of the first duration is increased to the fourth scheduling period PSCH4′ of the second duration, and the fifth scheduling period PSCH5 of the first duration is increased to the fifth scheduling period PSCH5′ of the second duration.

[0108] Figure 10 It shows that the Figure 3 A block diagram of an example embodiment of a processor in a multi-core system.

[0109] refer to Figure 10 The processor 111 may include, for example, four processor cores C1 to C4 and a common circuit 112 used and shared by the processor cores C1 to C4. The common circuit 112 may include one or more common resources CRS1 and CRS2.

[0110] The processor cores C1 to C4 may respectively include a buffer BUFF to store input data, intermediate data generated during processing, and processing result data.

[0111] When one processor core needs to use a common resource (e.g., CRS1) and another processor core is using the same common resource, the processor core will be in a standby state (e.g., in an idle state) until the other processor core stops using the common resource or the processor core can access the common resource. As a result, even if no other tasks are executed in this processor core, this processor core does not execute tasks that require the common resource. The time that a task that has already started is on standby can be called a pause time. For example, a given processor core can continue to execute instructions for a given task until an instruction that requires a resource that cannot be accessed is encountered, and then the given processor core can remain idle (pause) until access to the resource is obtained.

[0112] Figure 11 and Figure 12 is a diagram illustrating an example of a task execution delay time used in a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0113] Figure 11 and Figure 12 Shows a reflection reference Figure 10 The pause time is the task execution delay time.

[0114] In some example embodiments, the execution delay tracker EDT of the task scheduler 135 may provide a sum of a start standby time and a pause time as the task execution delay time TTDij, wherein the start standby time is the time from the time point when a given task is stored in the task queue of a given processor core to the start time point when the given processor core starts to execute the given task, and the pause time is the time when the given processor core stops executing the given task after the start time point. Here, i is the index of the task corresponding to A or B, and j is Figure 11 The index of the scheduling period corresponding to 1 or 2 in the example of .

[0115] In the first scheduling cycle PSCH1, as shown in the reference Figure 7 As described above, the execution delay tracker EDT may provide a start standby time as the task execution delay time, where the start standby time is the time from the time point when a given task is stored in the task queue of a given processor core to the start time point when the given processor core starts executing the given task.

[0116] In the second scheduling cycle PSCH2 , the stall time TSTLL has occurred, and the execution delay tracker EDT may provide the sum of the start standby time and the stall time as the task execution delay time.

[0117] As a result, in the second scheduling period PSCH2, Figure 11 TTDB2 can be provided as the task execution delay time of the second task B, and Figure 12 (TTDB21+TTDB22) can be provided as the task execution delay time of the second task B.

[0118] Figure 13 and Figure 14 is a diagram illustrating a method of controlling an operation of a multi-core system according to an exemplary embodiment of the inventive concept.

[0119] refer to Figure 13 In the first scheduling cycle PSCH1, since the execution time of the first task A, the second task B, the third task C, and the fourth task D is longer than the first scheduling cycle PSCH1, task execution may fail due to task execution delay or core execution delay. As a result, the control logic LDL does not perform task scheduling for the second scheduling cycle PSCH2.

[0120] In this case, according to an exemplary embodiment, the control logic LDL of the task scheduler 135 determines whether a task execution delay or a core execution delay has occurred, and then divides the tasks assigned to each processor core into an urgent task group and a normal task group based on the delay requirement level of the tasks. When it is determined that a task execution delay or a core execution delay has occurred on one or more processor cores, the control logic LDL may relocate the urgent task group and the normal task group to different processor cores.

[0121] Figure 13 An example is shown in which the first task A, the second task B, and the fourth task D are included in the normal task group, while the third task C is included in the urgent task group. In this case, the normal task group A, B, and D are assigned to the first processor core C1, and the urgent task group D is assigned to the second processor core C2. Even if task scheduling is not performed normally on the first processor core C1, no serious problems will arise because the tasks A, B, and D with low latency requirements are assigned to the first processor core C1.

[0122] In the reference Figure 13 When relocating the normal task groups A, B, and D and the emergency task group D to different processor cores, the control logic LDL may also increase the scheduling period of at least one of the different processor cores. For example, the third scheduling period PSCH3' of the first processor core C1 may be increased to be longer than the previous scheduling periods PSCH1 and PSCH2. These embodiments may be used when the latency requirements of the tasks are relatively low.

[0123] Figure 15 is a block diagram illustrating an electronic system according to an exemplary embodiment of the inventive concept.

[0124] refer to Figure 15 The electronic system includes a controller 1210 (e.g., a control circuit), a power supply 1220 (e.g., a power supply device), a storage device 1230 (e.g., a storage device), a memory 1240 (e.g., a storage device), an I / O port 1250, an expansion card 1260, a network device 1270, and a display 1280. According to an exemplary embodiment, the electronic system may further include a camera module 1290.

[0125] Controller 1210 can control the operation of each of elements 1220 to 1280. Power supply 1220 can provide operating voltage to at least one of elements 1210 and 1230 to 1280. Storage device 1230 can be implemented as a hard disk drive (HDD) or an SSD. Memory 1240 can be implemented as a volatile or non-volatile memory. I / O port 1250 can send data to an electronic system or send data output from an electronic system to an external device. For example, I / O port 1250 can include a port for connecting to a pointer device such as a computer mouse, a port for connecting to a printer, and a port for connecting to a universal serial bus (USB) drive.

[0126] The expansion card 1260 may be implemented as a secure digital (SD) card or an MMC. The expansion card 1260 may be a subscriber identity module (SIM) card or a universal SIM (USIM) card.

[0127] The network device 1270 enables the electronic system to be connected to a wired or wireless network. The display 1280 displays data output from the storage device 1230 , the memory 1240 , the I / O port 1250 , the expansion card 1260 , or the network device 1270 .

[0128] The camera module 1290 is a module that can convert an optical image into an electronic image. Therefore, the electronic image output from the camera module 1290 can be stored in the storage device 1230, the memory 1240, or the expansion card 1260. In addition, the electronic image output from the camera module 1290 can be displayed through the display 1280. For example, the camera module 1290 may include a camera.

[0129] The controller 1210 may be a multi-core system as described above. For example, the controller 1210 may include Figure 1 The processor 110 and the task scheduler 135 in the embodiment of the present inventive concept. According to an exemplary embodiment of the present inventive concept, the controller 1210 includes the execution delay tracker EDT and the control logic LDL as described above to perform negative feedback-based task scheduling using the task execution delay time and the core execution delay time.

[0130] In this way, a multi-core system and a method for controlling the operation of a multi-core system according to at least one exemplary embodiment of the present invention can reduce or eliminate the situation where tasks starve because specific tasks have not been processed for a long time by using task execution delay time and core execution delay time, and efficient task scheduling can be achieved even for tasks with high data dependencies.

[0131] The present invention can be applied to various electronic devices and systems that require efficient multi-processing. For example, embodiments of the present invention can be applied to various systems, such as memory cards, solid-state drives (SSDs), embedded multimedia cards (eMMCs), universal flash memory (UFS), mobile phones, smart phones, personal digital assistants (PDAs), portable multimedia players (PMPs), digital cameras, video recorders, personal computers (PCs), server computers, workstations, laptop computers, digital TVs, set-top boxes, portable game consoles, navigation systems, wearable devices, Internet of Things (IoT) devices, Internet of Everything (IoE) devices, e-book readers, virtual reality (VR) devices, augmented reality (AR) devices, etc.

[0132] The foregoing is illustrative of example embodiments and should not be construed as limiting thereof. Although a few example embodiments have been described, those skilled in the art will readily appreciate that various modifications may be made in these example embodiments without materially departing from the inventive concept.

Claims

1. A method of controlling the operation of a multi-core system including a plurality of processor cores, the method comprising: monitoring task execution delay times of tasks respectively assigned to the plurality of processor cores; monitoring core execution latency of the plurality of processor cores; as well as Based on the task execution delay time and the core execution delay time, the operation of the multi-core system is controlled, wherein, The core execution delay time of a given processor core among the plurality of processor cores corresponds to a maximum task execution delay time among the task execution delay times associated with the given processor core, and A determination is made as to whether a core execution delay occurs on the given processor core based on the maximum task execution delay time.

2. The method according to claim 1, wherein The task execution delay time of a given task among the tasks assigned to a given processor core among the multiple processor cores corresponds to a standby time, which is a time during which the given task is not executed by the given processor core after the given task is stored in a task queue of the given processor core.

3. The method according to claim 2, wherein: The task execution delay time of the given task corresponds to a start standby time, which is a time from a time point when the given task is stored in the task queue to a start time point when the given processor core starts executing the given task.

4. The method according to claim 2, wherein: The task execution delay time of the given task corresponds to the sum of a start standby time and a pause time, wherein the start standby time is the time from the time point when the given task is stored in the task queue to the start time point when the given processor core starts to execute the given task, and the pause time is the time when the given processor core stops executing the given task after the start time point.

5. The method according to claim 1, wherein Controlling the operations of the multi-core system includes: Based on the change of the task execution delay time or the core execution delay time, the task scheduling or the power level of the multi-core system is controlled.

6. The method according to claim 1, wherein Controlling the operations of the multi-core system includes: When it is determined that the core execution delay has occurred for the given processor core, a power level of the given processor core is increased.

7. The method according to claim 6, wherein: Increasing the power level of the given processor core includes: The operating frequency of the given processor core is increased.

8. The method according to claim 1, wherein Controlling the operations of the multi-core system includes: When it is determined that the core execution delay of the given processor core has occurred, a given task among the tasks assigned to the given processor core is relocated to another processor core among the plurality of processor cores.

9. The method according to claim 1, wherein: Controlling the operations of the multi-core system includes: dividing the tasks assigned to each processor core into an urgent task group and a normal task group based on the delay requirement level of the tasks; and When it is determined that the core execution delay occurs in at least one of the processor cores, the urgent task group and the normal task group are relocated to be allocated to different processor cores.

10. The method according to claim 9, wherein: Controlling the operation of the multi-core system also includes: When it is determined that the core execution delay occurs in at least one of the different processor cores, a scheduling period of the at least one of the different processor cores is increased.

11. The method according to claim 1, wherein Controlling the operations of the multi-core system includes: When it is determined that the core execution delay occurs in the given processor core, the scheduling period of the given processor core is increased.

12. A multi-core system comprising: a multi-core processor comprising multiple processor cores; a first control logic configured to monitor task execution delay times of tasks respectively assigned to the plurality of processor cores and core execution delay times of the plurality of processor cores; as well as The second control logic is configured to control the operation of the multi-core system based on the task execution delay time and the core execution delay time, wherein: The first control logic provides a maximum task execution delay time among the task execution delay times associated with a given processor core among the plurality of processor cores as the core execution delay time of the given processor core, and The second control logic determines whether a core execution delay occurs for the given processor core based on the maximum task execution delay time.

13. The multi-core system according to claim 12, wherein: The first control logic provides a standby time as the task execution delay time of a given task assigned to a given processor core among the multiple processor cores, the standby time being a time during which the given task is not executed by the given processor core after the given task is stored in a task queue of the given processor core. The multi-core system according to claim 12 , wherein: When it is determined that the core execution delay of the given processor core has occurred, the second control logic relocates the task assigned to the given processor core to another processor core among the plurality of processor cores.

15. The multi-core system according to claim 12, wherein: The second control logic increases a power level of the given processor core when it is determined that the core execution delay of the given processor core has occurred.

16. A method of controlling operation of a multi-core system comprising a plurality of processor cores, the method comprising: monitoring task execution delay times of tasks respectively assigned to the plurality of processor cores, the task execution delay time of a given task assigned to a given processor core among the plurality of processor cores corresponding to a standby time, the standby time being a time during which the given task is not executed by the given processor core after the given task is stored in a task queue of the given processor core; monitoring core execution delay times of the plurality of processor cores, the core execution delay time of the given processor core corresponding to a maximum task execution delay time among the task execution delay times associated with the given processor core; determining, based on the maximum task execution delay time, whether a core execution delay occurs in the given processor core; as well as When it is determined that the core execution delay of the given processor core has occurred, the given task assigned to the given processor core is relocated to another processor core among the plurality of processor cores or a power level of the given processor core is increased.

Citation Information

Patent Citations

  • Atopy diagnostic device

    KR1020190083853A

  • System on chip, method of operating the same, and apparatus including the same

    US20140173150A1

  • Dynamically allocating compute nodes among cloud groups based on priority and policies

    US20160204923A1