Coroutine task scheduling optimization method based on OpenHarmony
By optimizing memory management and task scheduling through a coroutine data operation system and a multi-level scheduling architecture, the problems of memory fragmentation and insufficient multi-core performance utilization in industrial robot operating systems have been solved, achieving efficient coroutine management and real-time scheduling, and improving the overall performance and reliability of the system.
Patent Information
- Application Number
- CN202511080517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-11
AI Technical Summary
Existing industrial robot operating systems have shortcomings in memory management, task scheduling, and multi-core performance utilization. In particular, they suffer from high resource consumption, severe memory fragmentation, and poor scheduling algorithm flexibility in high-concurrency scenarios, making it difficult to meet the real-time and multi-core load balancing requirements of industrial robots.
A coroutine data execution system is adopted, which combines static stack, dynamic stack and shared stack management strategies. A cooperative scheduling mechanism and an EDF-based automated real-time scheduling mechanism are designed to optimize memory allocation and multi-core task scheduling. Through a multi-level scheduling architecture, efficient management and load balancing of coroutines are achieved.
It significantly reduces the performance overhead of coroutine switching, improves memory allocation efficiency and task scheduling efficiency of multi-core systems, enhances the real-time performance and reliability of the system, solves the problems of high memory consumption and poor scheduling algorithm flexibility in multi-threaded models, and improves the independent controllability of domestic industrial robot operating systems.
Smart Images

Figure CN120929217A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a coroutine task scheduling optimization method based on OpenHarmony, belonging to the field of industrial robot operation technology. Background Technology
[0002] With the accelerating advancement of industrial automation, industrial robots play an indispensable role in modern manufacturing. As core equipment in intelligent production lines, they not only significantly improve production efficiency and product quality but also accelerate the digital, intelligent, and networked transformation of the manufacturing industry. However, as the functions of industrial robots become increasingly complex, existing industrial robot operating systems have revealed a series of problems in terms of performance and efficiency, particularly in memory management, task scheduling, and multi-core performance utilization, which have become key bottlenecks restricting the further development of industrial robot technology.
[0003] Industrial robot operating systems need to simultaneously address complex task scheduling and resource-constrained memory management. Most existing robot operating systems employ multi-threaded or multi-process models for task handling, but these traditional methods face numerous challenges in industrial robot scenarios. Multi-threaded models often lead to frequent user-kernel mode switching and resource contention in high-concurrency scenarios, resulting in high system overhead and task response latency. While multi-process models offer better isolation, their high memory consumption and switching overhead make them difficult to run efficiently in resource-constrained embedded environments. Although current mainstream coroutine libraries perform well in lightweight task management and asynchronous programming, most are designed for general-purpose applications and lack optimization support for the real-time performance, reliability, and resource-constrained environments of industrial robot operating systems. This makes it difficult for them to directly adapt to the high real-time requirements and multi-core load balancing characteristics of robot operating systems.
[0004] In terms of memory optimization, industrial robot operating systems currently face challenges such as inefficient memory allocation, severe fragmentation, and the inability to efficiently reclaim and reuse memory resources under dynamic task loads. This not only leads to memory waste but also directly affects task scheduling efficiency and system stability. Especially in scenarios with frequent task switching or multi-task concurrency, inefficient memory management mechanisms may cause task delays or even failures.
[0005] Task scheduling is another core challenge in the performance of industrial robot operating systems. Existing scheduling algorithms (such as time-slice round-robin and preemptive scheduling) tend to be general-purpose designs, making it difficult to meet the stringent real-time and rapid response requirements of industrial robots. Traditional scheduling algorithms are inadequate in terms of task preemption and priority adjustment, making it difficult for high-priority tasks to acquire resources in a timely manner; their scheduling granularity is also too large, failing to support fine-grained task decomposition and efficient coordination. Furthermore, existing algorithms have weak support for periodic tasks, leading to scheduling delays or task conflicts when the system handles typical industrial robot tasks (such as sensor data acquisition and real-time control commands), further reducing the system's real-time performance and reliability.
[0006] The widespread use of multi-core processors provides powerful parallel computing capabilities for industrial robot operating systems, but existing operating systems do not fully utilize multi-core performance. In multi-core environments, task load balancing and cross-core scheduling optimization remain key challenges. Existing scheduling algorithms perform poorly in task allocation and core utilization efficiency, frequently resulting in some cores being overloaded while others remain idle. For industrial robot tasks requiring multi-core collaborative processing (such as complex path planning and multi-sensor data fusion), existing scheduling mechanisms struggle to fully leverage multi-core performance, hindering the improvement of overall system efficiency. Summary of the Invention
[0007] To address the issues of insufficient real-time performance, concurrency, and flexibility in industrial robots under complex task scenarios, this invention proposes a coroutine task scheduling optimization method based on OpenHarmony.
[0008] The technical solution adopted by the present invention to solve the above problems is as follows: The present invention includes the following steps: Step 1: Build a coroutine data execution system, based on which the execution environment of coroutines is saved and the coroutine context is saved and restored, thereby realizing task switching and state management; Step 2: Based on the coroutine data execution system, formulate static stack allocation strategy, dynamic stack allocation strategy and shared stack allocation strategy to match the memory requirements of corresponding task scenarios; Step 3: Based on the allocation strategy of the coroutine data execution system and execution, define the corresponding interface to make the coroutine data execution system adapt to various processor architectures, and differentiate the processors of various architectures to complete the multi-architecture processor adaptation. Step 4: Based on the adapted multi-architecture processors, optimize the coroutine scheduling algorithm and design a cooperative scheduling mechanism and an automated real-time scheduling mechanism for coroutines. Step 5: To address the performance optimization needs of multi-core processors, a multi-feature fusion load assessment algorithm is proposed. Combined with the coroutine scheduling mechanism, multi-core scheduling and load balancing of coroutines are performed, optimizing task cross-core allocation and migration, so that tasks are dynamically allocated to the corresponding cores according to the real-time load of the processor.
[0009] Furthermore, in step 1, the coroutine data execution system includes a coroutine context module and a coroutine control module; The coroutine context module is used to store the execution environment of coroutines; The coroutine control module is used to save and restore coroutine contexts, enabling task switching and state management.
[0010] Furthermore, the coroutine context module includes: The regs submodule, of type void*
[17] , is used to store the coroutine register context; The ss_size submodule, of type UINT32, is used to store the stack space size; The ss_sp submodule, of type char*, is used to store the stack pointer, which points to the starting position of the coroutine stack.
[0011] Furthermore, the coroutine control module includes: The sched submodule, of type scheduler*, is used to quickly retrieve the scheduler to which the coroutine belongs; The pfn submodule, of type coctx_pfn_t, is used to point to the task function of the coroutine; The `arg` submodule, of type `void*`, serves as a pointer to the parameters of a coroutine function; The ctx submodule, of type coctx, is used to store the context information of coroutines; The cStart submodule, of type char, is used to represent the coroutine start flag; The cStatus submodule, of type char, is used to represent coroutine status flags; The isShareStack submodule, of type char, is used to represent the shared stack flag; The stack_mem submodule, of type stStackMem_t*, serves as a pointer to the coroutine stack memory; The stack_sp submodule, of type char*, is used to store the stack pointer of the current coroutine.
[0012] Furthermore, the static stack allocation strategy defined in step 2 specifically includes: Allocate a fixed-size memory region as stack space, which remains unchanged throughout the coroutine's lifecycle, and postpone stack space initialization until coroutine creation. User-defined stack size is supported; if the user does not specify a size, a default value will be automatically set based on the maximum stack usage of historical tasks.
[0013] Furthermore, the dynamic stack allocation strategy defined in step 2 specifically includes: The stack size is adjusted according to actual needs during task execution; an independent coroutine traverses all dynamically stack-allocated coroutines in a 5ms interrupt manner, and interrupts the executing coroutine and doubles the size when it is detected that expansion is needed.
[0014] Furthermore, the shared stack allocation strategy defined in step 2 specifically includes: Allocate a memory region in the heap as a shared stack storage space, with a default size of 4MB; set the coroutine stack pointer to the starting address of the shared stack, and record the stack bottom and stack top of the current coroutine; when switching coroutines, allocate memory matching the actual usage size, copy the shared stack contents to the private stack for storage, and restore from the private stack to the shared stack when restoring.
[0015] Furthermore, step 3 specifically includes: The system adapts to multiple processor architectures by defining standardized interfaces set_coctx and get_coctx. set_coctx is used to save the coroutine context, and get_coctx is used to restore the coroutine context. The specific implementation of the interface is completed by the hardware-related layer, and corresponding assembly is used for ARM and x86 architectures.
[0016] Furthermore, the cooperative scheduling mechanism and automated real-time scheduling mechanism for coroutines designed in step 4 specifically include: The cooperative scheduling mechanism records the call chain between coroutine tasks through a DAG graph to optimize the management of coroutine suspension and resumption and the management of coroutine dependencies. The automated real-time scheduling mechanism is based on the earliest deadline first (EDF) algorithm to build a scheduling framework. The scheduler adopts a modular architecture, including an executor, an allocator, and a timer. The executor selects coroutines for execution based on the EDF algorithm, the timer is responsible for the management and scheduling of timed-out tasks, and the allocator is used to perform the interaction between the global scheduling layer and the local scheduling layer, maintaining load balancing between different cores and different groups in the local scheduling layer.
[0017] Furthermore, step 5 specifically includes: A multi-level scheduling architecture is adopted, dividing coroutine management into a global scheduling layer and a local execution layer; the global layer is responsible for the macro-allocation and coordination of coroutines, while the local layer is used for the specific execution of tasks within the group; During the allocation of coroutines at the global level, the number of coroutines in the runnable queue is counted. Number of coroutines in the waiting queue Number of coroutines in the newly created queue and the number of coroutines in the garbage collection queue And assign a corresponding weight to each coroutine in the queue, wherein the weight of the coroutines in the runnable queue is... The weights of the coroutines in the waiting queue are The weights of the coroutines in the newly created queue are The weights of the goroutines in the garbage collection queue are: The load is calculated based on the number of goroutines in the queue and the weight of the corresponding goroutines in the queue. L According to the load L Adjust the weight values of the corresponding queues to ensure load balancing between different cores and between different groups in the local scheduling layer; load L The calculation formula is: (1).
[0018] The beneficial effects of this invention are: 1. To address the issues of high resource consumption and severe memory fragmentation in high-concurrency scenarios, this invention utilizes dynamic stack, shared stack management, and memory pool mechanisms to optimize memory allocation efficiency for coroutine creation and destruction. Furthermore, through optimization of user-mode context switching and implementation of an abstract interface for the execution environment, the performance overhead of coroutine switching is significantly reduced, and the portability of the coroutine library across multiple hardware platforms is improved.
[0019] 2. This invention addresses the high real-time and high concurrency requirements of industrial robots by implementing cooperative scheduling and preemptive scheduling based on EDF. Cooperative scheduling simplifies task switching and is suitable for scenarios with clear logic; EDF-based scheduling prioritizes high-priority tasks and, combined with dynamic priority adjustment, optimizes task conflict handling, improving scheduling efficiency and real-time performance.
[0020] 3. This invention addresses the issues of coroutine scheduling and resource utilization in industrial robot operating systems by designing an efficient multi-core scheduling and load balancing strategy. In a multi-core system environment, a dynamic load balancing mechanism based on coroutine tasks is introduced to optimize task allocation among different cores. This mechanism ensures high efficiency in multi-core processing while maintaining low latency and high real-time performance of coroutine execution through task partitioning, task migration, and core load awareness strategies.
[0021] 4. This invention verifies the efficiency, reliability, and stability of the designed coroutine library in a domestic industrial robot operating system by running and testing it on a real domestic operating system kernel. Experimental results show that this invention significantly improves task scheduling performance, memory utilization efficiency, and multi-core scheduling capabilities, largely solving problems such as high memory consumption and poor scheduling algorithm flexibility in multi-threaded models, and also enhancing the independent controllability of domestic industrial automation technology.
[0022] 5. To optimize coroutine task scheduling, this invention designs and implements a lightweight coroutine-based industrial robot operating system library. This coroutine library addresses resource-constrained scenarios in robotic environments by designing and implementing a lightweight library with lower memory footprint, higher task real-time performance, and higher CPU utilization. It is also the first time a coroutine library has been introduced into a domestically developed robot operating system for real-time scheduling adaptation. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the overall architecture of the coroutine library based on OpenHarmony; Figure 2 This is a diagram illustrating static stack memory allocation. Figure 3 This is a schematic diagram of the dynamic stack memory allocation process.
[0024] Figure 4 This is a diagram illustrating the copying of shared stack memory.
[0025] Figure 5 This is a diagram illustrating stack pool management.
[0026] Figure 6 This is a diagram illustrating assembly context switching.
[0027] Figure 7 This is a diagram illustrating the switching of the coroutine runtime environment.
[0028] Figure 8 This is a schematic diagram of the coroutine scheduler architecture.
[0029] Figure 9 This is a schematic diagram of the coroutine scheduling process based on the scheduler.
[0030] Figure 10 This is a schematic diagram of the coroutine executor's execution flow.
[0031] Figure 11 This is a schematic diagram of the timer execution process.
[0032] Figure 12 This is a schematic diagram of a skip list structure.
[0033] Figure 13 This is a schematic diagram of a load balancing structure design.
[0034] Figure 14 This is a flowchart illustrating a coroutine task scheduling optimization method based on OpenHarmony. Detailed Implementation
[0035] Combination Figure 1-14 This implementation method is described as follows: Figure 14 As shown, the steps of the coroutine task scheduling optimization method based on OpenHarmony described in this embodiment include: S1: Coroutine data structure design; like Figure 1 As shown, the main structure of coroutine execution includes a coroutine context module and a coroutine control block module. The implementation classes for the coroutine context and control block data structures are shown below: Coroutine Context Module: Context information is managed through the `coctx` class, whose main members and functions are shown in Table 1. Table 1
[0036] Specifically, the register set regs is implemented as a pointer array of length 17; the coroutine stack size ss_size is of type UINT32, representing an unsigned 32-bit integer; the coroutine stack starting address ss_sp is a char* character pointer type.
[0037] Coroutine Control Module: As shown in Table 2, the sub-modules in Table 2 are detailed descriptions of the data structure of the coRoutine control block, covering the types of information it stores and their functions.
[0038] Table 2
[0039] S2: Optimized coroutine memory management; This implementation proposes and implements three memory allocation strategies, each suitable for different types of tasks, to improve resource utilization; a stack pool structure is designed and implemented to optimize memory allocation efficiency and reduce fragmentation issues; the three memory allocation strategies specifically include: Static stack allocation; such as Figure 2As shown. (1) Initialization of static stack memory. During the system initialization phase, a fixed-size stack space is pre-allocated for each coroutine. The stack size is set according to experimental analysis and actual task requirements. For example, in the industrial robot scenario, the typical stack size can be configured as 128KB or larger. These stack spaces are allocated from the system memory at once to form a dedicated memory area for the coroutine stack, avoiding repeated allocation operations during runtime. (2) Coroutine control block creation and initialization. When creating a coroutine, a CCB is first allocated and initialized for the coroutine. A block of memory is allocated in the heap to store the CCB, and the coroutine control block information is initialized. A unique identifier (ID) is allocated for the coroutine, the initial state of the coroutine is set to CREATED or READY, the task function executed by the coroutine is bound, and the priority of the coroutine is set. (3) Coroutine stack memory allocation. The system allocates a fixed-size memory block from the heap space using the malloc function. Since heap memory addresses grow from low to high addresses, while the stack space during program execution typically grows from high to low addresses, after allocating memory space using malloc, the high address of this memory block needs to be calculated and assigned to the top and bottom registers of the coroutine's stack. Simultaneously, the stack pointer is set and initialized to the address of the allocated stack space's top.
[0040] Dynamic stack allocation; such as Figure 3 As shown. (1) Stack allocation and initialization. During system initialization, the initial stack size of the coroutine dynamic stack is set (e.g., 2KB). The actual stack allocation is dynamic as the task is executed. (2) Stack expansion mechanism. The size of the stack in use and the size of the stack space are detected. The current stack size in use is obtained by subtracting the value of the stack bottom from the stack pointer SP. When the ratio of used memory to total stack memory is greater than 0.75, the stack is expanded. The expansion steps include: allocating a larger stack by requesting a new memory block through the malloc function as the new stack of the coroutine. The size of the new stack is twice the size of the current stack. Migrating old stack data by copying the current stack content to the new stack and adjusting the stack pointer.
[0041] Shared stack allocation; such as Figure 4As shown. (1) Initialization of the shared stack. When the system starts, a predefined memory area is allocated in the heap as the storage space for the shared stack. This shared stack needs to be large enough to support the operation of complex tasks, and usually a large value (e.g., 4MB) is selected. (2) Stack allocation during coroutine execution. When a coroutine starts running, its stack pointer is pointed to the starting address of the shared stack, and the stack bottom and stack top of the current coroutine are recorded. The coroutine executes tasks in the shared stack, and all function calls and local variables are completed in this shared memory. (3) Stack saving during coroutine switching. When a coroutine switches from the running state to the suspended state, it needs to save the current stack content. The specific operations include: calculating the actual stack size used by the current coroutine, which is determined based on the difference between the stack pointer and the stack bottom address. Allocating a memory in the private stack that matches the actual size used, and copying the contents of the shared stack to the private stack for saving. Updating the coroutine control block (CCB) and recording the starting address and size of the private stack. (4) Stack loading during coroutine resumption. When a coroutine resumes from a suspended state to a running state, its stack contents need to be restored from the private stack to the shared stack, such as... Figure 5 As shown. Based on the private stack address and size recorded in the coroutine control block, calculate the amount of data that needs to be recovered. Copy the contents of the private stack back to the shared stack, and adjust the stack pointer to point to the valid area of the shared stack.
[0042] S3: Multi-architecture processor adaptation; This implementation achieves the adaptation of the coroutine data execution system to multi-architecture processors by defining standardized interfaces set_coctx and get_coctx. set_coctx is used to save the coroutine context, and get_coctx is used to restore the coroutine context. The specific implementation of the interface is completed by the hardware-related layer, and corresponding assembly is used for ARM and x86 architectures.
[0043] S4: Optimization of the coroutine task scheduling mechanism; like Figure 6 and Figure 7 As shown, this implementation scheme designs a cooperative scheduling approach for scenarios with clearly defined task logic, allowing users to manually control task execution and switching, simplifying the task switching process. For the real-time task requirements of robot scenarios, an automated real-time scheduling approach is designed and implemented to ensure the real-time performance of critical tasks. Simultaneously, a time-sharing scheduling strategy is implemented for cases with the same priority, avoiding scheduling conflicts and improving task response efficiency. The cooperative and automated real-time scheduling approaches are as follows: Cooperative scheduling; defines two key interfaces: yield and resume.
[0044] (1) The `yield` interface causes the current coroutine to voluntarily relinquish CPU control and switch to the previous coroutine in the call chain; while the `resume` interface is used to restore the target coroutine to its running state and save the current coroutine to the coroutine call chain stack, thereby achieving flexible control of task switching. To manage the calling relationship between coroutines, the scheduler maintains a stack structure, namely the call chain stack. When a coroutine calls `yield`, it returns to the previous coroutine recorded in the call chain stack. This mechanism not only ensures the simplicity of scheduling but also ensures that the context switching between coroutines conforms to the calling order; (2) The function of resume is to switch to the target coroutine and mark it as the currently running coroutine. Specifically, such as Figure 10 As shown, `resume` first checks the state of the target coroutine, ensuring it is either "suspended" or "runnable". If the target coroutine is unavailable, an error message is returned. Next, the identifier of the current coroutine is pushed onto the call stack, recording it as the caller of the target coroutine. Subsequently, by restoring register values, stack pointer (SP), and other context information from the target coroutine's control block, the underlying context switching function is called to switch execution control to the target coroutine. Finally, the scheduler state is updated, setting the target coroutine's state to "running" and making it the currently running coroutine.
[0045] Automated real-time scheduling; such as Figure 9 As shown, in the specific implementation of the EDF algorithm, each task is assigned a deadline and an execution time. When the scheduler needs to select the next task to run, it selects the task with the earliest deadline from all ready tasks for scheduling. The main scheduling process can be divided into the following stages. First is the task initialization stage, where a deadline D and an execution time C are assigned to each task. The priority of a task is dynamically adjusted as its deadline changes. Next is the task addition queue stage, where new tasks are inserted into the scheduling queue and arranged in ascending order of deadline. If the addition of a new task causes the total utilization of the task set to exceed 100%, the system may not be able to schedule normally and resource adjustments or denial of service may be required. Finally, there is the task execution and completion stage, where the scheduler selects the task with the earliest deadline from the queue to run each time. If the currently running task has not yet been completed and a new task with an earlier deadline arrives, the scheduler will perform task switching; once the task is completed, it will be removed from the queue. If a task is not completed before its deadline, it is marked as "timeout".
[0046] The automated real-time scheduling mechanism is based on the Earliest Deadline First (EDF) algorithm to build its scheduling framework. The scheduler adopts a modular architecture, including an executor, an allocator, and a timer. The executor selects coroutines for execution based on the EDF algorithm, and the timer is responsible for... Figure 11 The timeout task management and scheduling shown are handled by the allocator, which is used to interact between the global scheduling layer and the local scheduling layer, maintaining load balancing between different cores and different groups in the local scheduling layer.
[0047] S5: Coroutine multi-core scheduling optimization.
[0048] This implementation method adopts a multi-level scheduling architecture, such as Figure 8 As shown, the global scheduler, as the core module of the global layer, is responsible for the overall task scheduling and load balancing of the system. It maintains a global task queue for centralized management of all tasks not assigned to a core. A key component of the global scheduler is the dispatcher, whose main responsibility is to monitor the load of each core in real time, including metrics such as task queue length and task execution latency. Once it detects that the load on certain cores is too high or too low, the dispatcher dynamically adjusts the task allocation strategy. For example, when it detects that a core has a low load, the dispatcher selects high-priority tasks from the global task queue and assigns them to the local task queue of that core, thereby achieving global load balancing.
[0049] The core data structure of the global scheduler is a task priority queue based on the EDF algorithm. For example... Figure 9 As shown, this priority queue is dynamically sorted according to the task deadlines, ensuring that the scheduler always prioritizes the task with the earliest deadline for allocation, thus meeting the system's real-time requirements. The closer the task deadlines are, the higher their priority; this design effectively guarantees the execution order of real-time tasks. During task allocation, the Dispatcher pops the optimal task from the priority queue and assigns it to the lightest-loaded core, thereby reducing the pressure on high-load cores. By dynamically maintaining task priorities, the global scheduler can quickly find the optimal task allocation strategy, satisfying both real-time requirements and making full use of system resources.
[0050] At the local execution layer, each core is bound to an executor (Processor), responsible for task execution and management within that core. After the global scheduler allocates tasks to cores, the tasks enter the local task queue of the corresponding core, where they are scheduled for execution in sequence by the executor. The queue management of the local executor is divided into four categories: runnable queue, new task queue, wait queue, and garbage collection queue (gcQueue). The runnable queue stores currently schedulable tasks and dynamically sorts them according to the EDF algorithm, ensuring that tasks with the earliest deadlines are executed first. The new task queue is used to temporarily store newly created tasks or tasks resumed from the wait queue; the wait queue stores tasks paused due to insufficient resources, blocking, or timeouts; and the garbage collection queue stores completed tasks for periodic system cleanup.
[0051] We use a min-heap to implement the data structure for maintaining the dynamic sorting of EDF. To meet the real-time system's requirements for task sorting and dynamic adjustment, a data structure that supports efficient insertion, deletion, and sorting is typically used. A min-heap maintains a complete binary tree structure, ensuring that the top of the heap is always the task with the earliest deadline. In a min-heap, inserting a new task adds it to the end of the heap and adjusts its position using an up operation; deleting a task replaces the top task with the last task and readjusts the heap's order using a down operation. The min-heap has low time complexity; insertion and deletion both have an overhead of O(logN), and retrieving the top task has an overhead of only O(1).
[0052] This implementation proposes a load balancing algorithm. During the allocation of coroutines at the global layer, to accurately calculate the load on the executor, the number of coroutines in the runnable queue is counted. Number of coroutines in the waiting queue Number of coroutines in the newly created queue and the number of coroutines in the garbage collection queue And assign a corresponding weight to each coroutine in the queue, wherein the weight of the coroutines in the runnable queue is... The weights of the coroutines in the waiting queue are The weights of the coroutines in the newly created queue are The weights of the goroutines in the garbage collection queue are: The load is calculated based on the number of goroutines in the queue and the weight of the corresponding goroutines in the queue. L : (1).
[0053] Weight of the run queue It is usually set to 1 because it directly reflects the workload of tasks currently being processed. Tasks in the waiting queue, although currently suspended, may become runnable at any time, hence their weight... This can be set to 0.5, representing the potential impact on future load. Tasks in the newly created queue will be added to the runnable queue, and their weight... A value of 0.75 indicates a higher impact on the current load. Tasks in the garbage collection queue are primarily background cleanup work, and their impact on the current load is relatively small. It is usually set to 0.25, depending on the load. L Adjust the weight values of the corresponding queues to ensure load balancing between different cores and between different groups in the local scheduling layer; The Dispatcher, or load balancer, is the interaction node between the global scheduling layer and the local scheduling layer. It has two main functions: first, to achieve load balancing between different cores in the local scheduling layer; and second, to achieve load balancing between different groups.
[0054] The Dispatcher's implementation process begins with global load data collection. It periodically monitors the running status of each executor, including the number of runnable tasks, suspended tasks, and newly created tasks, to calculate the current core load value. Load calculation is based on four queues for executors: the runnable queue, the newly created queue, the waiting queue, and the garbage collection queue, each assigned different weights to accurately reflect the current core task pressure. When the load of some executors in the core load table exceeds a preset threshold, the Dispatcher identifies them as overloaded cores; when the load is below the threshold, they are marked as low-load cores. Based on these core states, the Dispatcher dynamically adjusts its task allocation strategy, such as... Figure 13 As shown.
[0055] In task allocation, the Dispatcher prioritizes using the global task queue to assign tasks to low-load cores. The global task queue is maintained using a min-heap, supporting fast retrieval of optimal task nodes. Each task node contains information such as a task identifier, deadline, and priority, ensuring that task allocation meets the requirements of the scheduling algorithm (EDF algorithm). When there are no tasks to be allocated in the global task queue, the Dispatcher selects unexecuted tasks from the more heavily loaded executors and migrates them to low-load cores. This migration strategy effectively reduces the complexity of direct cross-core task theft while fully utilizing the flexibility of the global queue scheduling, avoiding the degradation of cache affinity caused by frequent task migrations.
[0056] To support these functionalities, the Dispatcher uses several key data structures, including a global task queue and a core load table. The global task queue manages all tasks not yet assigned to specific cores, ensuring smooth distribution even with a surge in task volume. The core load table uses a structure such as... Figure 12 The skip list structure shown is implemented using shared memory, storing the load value of each executor and updating it periodically, allowing the Dispatcher to obtain the load status in real time and make decisions. Combining these data structures, the Dispatcher implements an efficient task allocation mechanism that can not only dynamically adapt to load changes but also improve the system's responsiveness in complex task scenarios through global and local coordinated scheduling.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.
Claims
1. A coroutine task scheduling optimization method based on OpenHarmony, characterized in that, include: Step 1: Build a coroutine data execution system, based on which the execution environment of coroutines is saved and the coroutine context is saved and restored, thereby realizing task switching and state management; Step 2: Based on the coroutine data operation system, formulate static stack allocation strategy, dynamic stack allocation strategy and shared stack allocation strategy to match the memory requirements of corresponding task scenarios; Step 3: Based on the aforementioned coroutine data execution system and execution allocation strategy, define corresponding interfaces to enable the coroutine data execution system to adapt to various processor architectures, and differentiate the processors of various architectures to complete multi-architecture processor adaptation. Step 4: Based on the adapted multi-architecture processors, optimize the coroutine scheduling algorithm and design a cooperative scheduling mechanism and an automated real-time scheduling mechanism for coroutines. Step 5: To address the performance optimization requirements of multi-core processors, a multi-feature fusion load assessment algorithm is proposed. Combined with the aforementioned coroutine scheduling mechanism, multi-core scheduling and load balancing of coroutines are performed, optimizing task cross-core allocation and migration, so that tasks are dynamically allocated to the corresponding cores according to the real-time load of the processor.
2. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, In step 1, the coroutine data execution system includes a coroutine context module and a coroutine control module; The coroutine context module is used to store the execution environment of a coroutine; The coroutine control module is used to save and restore coroutine contexts, enabling task switching and state management.
3. The coroutine task scheduling optimization method based on OpenHarmony according to claim 2, characterized in that, The coroutine context module includes: The regs submodule, of type void*[17], is used to store the coroutine register context; The ss_size submodule, of type UINT32, is used to store the stack space size; The ss_sp submodule, of type char*, is used to store the stack pointer, which points to the starting position of the coroutine stack.
4. The coroutine task scheduling optimization method based on OpenHarmony according to claim 2, characterized in that, The coroutine control module includes: The sched submodule, of type scheduler*, is used to quickly retrieve the scheduler to which the coroutine belongs; The pfn submodule, of type coctx_pfn_t, is used to point to the task function of the coroutine; The `arg` submodule, of type `void*`, serves as a pointer to the parameters of a coroutine function; The ctx submodule, of type coctx, is used to store the context information of coroutines; The cStart submodule, of type char, is used to represent the coroutine start flag; The cStatus submodule, of type char, is used to represent coroutine status flags; The isShareStack submodule, of type char, is used to represent the shared stack flag; The stack_mem submodule, of type stStackMem_t*, serves as a pointer to the coroutine stack memory; The stack_sp submodule, of type char*, is used to store the stack pointer of the current coroutine.
5. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, The static stack allocation strategy defined in step 2 specifically includes: Allocate a fixed-size memory region as stack space, which remains unchanged throughout the coroutine's lifecycle, and postpone stack space initialization until coroutine creation. User-defined stack size is supported; if the user does not specify a size, a default value will be automatically set based on the maximum stack usage of historical tasks.
6. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, The dynamic stack allocation strategy defined in step 2 specifically includes: The stack size is adjusted according to actual needs during task execution; an independent coroutine traverses all dynamically stack-allocated coroutines in a 5ms interrupt manner, and interrupts the executing coroutine and doubles the size when it is detected that expansion is needed.
7. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, The shared stack allocation strategy defined in step 2 specifically includes: Allocate a memory region in the heap as a shared stack storage space, with a default size of 4MB; set the coroutine stack pointer to the starting address of the shared stack, and record the stack bottom and stack top of the current coroutine; when switching coroutines, allocate memory matching the actual usage size, copy the shared stack contents to the private stack for storage, and restore from the private stack to the shared stack when restoring.
8. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, Step 3 specifically includes: The system adapts to multiple processor architectures by defining standardized interfaces set_coctx and get_coctx. set_coctx is used to save the coroutine context, and get_coctx is used to restore the coroutine context. The specific implementation of the interface is completed by the hardware-related layer, and corresponding assembly is used for ARM and x86 architectures.
9. The coroutine task scheduling optimization method based on OpenHarmony according to claim 1, characterized in that, The cooperative scheduling mechanism and automated real-time scheduling mechanism for coroutines designed in step 4 specifically include: The cooperative scheduling mechanism records the call chain between coroutine tasks through a DAG graph to optimize the management of coroutine suspension and resumption and the management of coroutine dependencies. The automated real-time scheduling mechanism is based on the earliest deadline first (EDF) algorithm to build a scheduling framework. The scheduler adopts a modular architecture, including an executor, an allocator, and a timer. The executor selects coroutines for execution based on the EDF algorithm, the timer is responsible for the management and scheduling of timed-out tasks, and the allocator is used to perform the interaction between the global scheduling layer and the local scheduling layer, maintaining load balancing between different cores and different groups in the local scheduling layer.
10. A coroutine task scheduling optimization method based on OpenHarmony according to claim 9, characterized in that, Step 5 specifically includes: A multi-level scheduling architecture is adopted, dividing coroutine management into a global scheduling layer and a local execution layer; the global layer is responsible for the macro-allocation and coordination of coroutines, while the local layer is used for the specific execution of tasks within the group; During the allocation of coroutines at the global level, the number of coroutines in the runnable queue is counted. Number of coroutines in the waiting queue Number of coroutines in the newly created queue and the number of coroutines in the garbage collection queue And assign a corresponding weight to each coroutine in the queue, wherein the weight of the coroutines in the runnable queue is... The weights of the coroutines in the waiting queue are The weights of the coroutines in the newly created queue are The weights of the goroutines in the garbage collection queue are: The load is calculated based on the number of goroutines in the queue and the weight of the corresponding goroutines in the queue. L According to the load L Adjust the weight values of the corresponding queues to ensure load balancing between different cores and between different groups in the local scheduling layer; load L The calculation formula is: (1)。