Heterogeneous multi-core video task scheduling system for embedded operating system
Patent Information
- Application Number
- CN202611008569.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-07-08
AI Technical Summary
这种被动式的降频干预会导致异构计算单元的有效吞吐能力骤降,进而引发底层硬件执行队列拥塞、共享内存总线带宽抢占以及DMA缓冲区同步栅栏等待超时
[0006] The beneficial effects of this invention are as follows: By transforming the heat dissipation lag phenomenon spanning video frame cycles into a computable time slot allocation constraint in the kernel scheduling process, this invention enables heterogeneous cores to proactively incorporate subsequent heat accumulation into the task reordering considerations before triggering system temperature throttling protection. Without adding external heat dissipation devices or altering the existing operating system layered architecture, this invention effectively alleviates the inter-frame latency jitter and underlying hardware queue congestion problems caused by heat dissipation lag in fanless devices, ensuring the consistency of task throughput and time response in continuous high-concurrency video processing scenarios, and ensuring the correct timing of hardware instruction streams and video result data output.
Smart Images

Figure CN122547499B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video task scheduling technology, and more specifically, to a heterogeneous multi-core video task scheduling system for embedded operating systems. Background Technology
[0002] In the development of embedded edge vision terminals, continuous high-frame-rate video stream processing typically requires concurrent calls to various heterogeneous computing units (such as CPU, GPU, NPU, VPU, etc.) within the system-on-a-chip (SoC). In the existing embedded Linux operating system architecture, resource management and task scheduling are strictly separated and executed independently at different abstraction levels: the application layer issues asynchronous computing requests through interfaces such as V4L2 and AIRuntime; the kernel scheduler mainly distributes tasks based on instantaneous CPU or GPU utilization; the DMA-BUF mechanism is responsible for maintaining cross-hardware memory access consistency; and chip frequency adjustment and temperature control are handled by subsystems such as cpufreq, devfreq, and thermalframework, which make reactive adjustments based on current sensor readings.
[0003] In practical fanless or passively cooled terminals, the heat conduction and dissipation processes of chips exhibit an objective time lag. When high-concurrency video decoding, AI inference, and encoding tasks are suddenly submitted to heterogeneous computing cores, the instantaneously generated heat is not immediately reflected in the temperature sensor readings. Instead, it gradually spreads over several subsequent video frame cycles, causing the temperature to rise. Since existing operating system schedulers allocate tasks based solely on the available computing power at the current moment, when a large number of high-load tasks are deployed, the hardware layer appears to have sufficient computational margin. However, as the instruction flow progresses, the accumulated heat from the lag triggers the thermal governor's frequency reduction protection strategy. This passive frequency reduction intervention leads to a sharp drop in the effective throughput of heterogeneous computing units, resulting in congestion in the underlying hardware execution queue, shared memory bus bandwidth preemption, and DMA buffer synchronization timeouts. On the application side, this phenomenon directly manifests as severe jitter in inter-frame video latency, local task queue accumulation, and timing errors in the output frame. Summary of the Invention
[0004] This invention provides a heterogeneous multi-core video task scheduling system for embedded operating systems, which solves the technical problems mentioned in the background art.
[0005] This invention provides a heterogeneous multi-core video task scheduling system for embedded operating systems, applied to a fanless embedded vision terminal comprising a central processing unit core, a vision processing unit, a neural network processing unit, a graphics processing unit, a field-programmable gate array (FPGA) hard core, and a shared memory controller, configured to execute: Capture concurrent video tasks issued by the application layer and convert them into video frame time slot task units; Collect the running status of each heterogeneous computing core and bind it to the hardware physical heat dissipation path to generate a resource snapshot of the current scheduling window; Using the resource snapshot, cross-frame thermal hysteresis slots are marked for the video frame slot task unit; The effective computing power for calculating the thermal effect reduction limit is calculated, and a unique task-core-time slot mapping is generated for the video frame time slot task unit; The concurrent video tasks are reordered based on the task-core-time slot mapping, and the underlying hardware standard control instructions are arranged. The underlying hardware standard control commands are submitted asynchronously and in isolation through controlled threads, and the hardware execution results of the concurrent video tasks are collected in sequence. Output the orchestrated underlying hardware standard control instructions and the sequentially arranged hardware execution results.
[0006] The beneficial effects of this invention are as follows: By transforming the heat dissipation lag phenomenon spanning video frame cycles into a computable time slot allocation constraint in the kernel scheduling process, this invention enables heterogeneous cores to proactively incorporate subsequent heat accumulation into the task reordering considerations before triggering system temperature throttling protection. Without adding external heat dissipation devices or altering the existing operating system layered architecture, this invention effectively alleviates the inter-frame latency jitter and underlying hardware queue congestion problems caused by heat dissipation lag in fanless devices, ensuring the consistency of task throughput and time response in continuous high-concurrency video processing scenarios, and ensuring the correct timing of hardware instruction streams and video result data output. Attached Figure Description
[0007] Figure 1 This is a schematic diagram of the overall mechanism of the heterogeneous multi-core video task scheduling system of the present invention; Figure 2 This is a schematic diagram of the cross-frame thermal hysteresis annotation and task-core-time slot mapping mechanism of the present invention. Detailed Implementation
[0008] The specific embodiments of the present invention will now be further described with reference to the accompanying drawings. The following embodiments are used to illustrate how the present invention implements video task scheduling in the operating system kernel environment of a fanless embedded vision terminal, and are not intended to limit the scope of protection of the present invention. Without departing from the technical concept of the present invention, those skilled in the art can make equivalent substitutions for module deployment locations, data structure field lengths, locking methods, and sampling interfaces according to different chip platforms, different driver interfaces, and different video processing pipelines.
[0009] This implementation presents a complete implementation chain of a heterogeneous multi-core video task scheduling system for an embedded operating system in a fanless embedded vision terminal. The fanless embedded vision terminal includes a central processing unit (CPU) core, a vision processing unit, a neural network processing unit, a graphics processing unit (GPU), a field-programmable gate array (FPGA) hard core, and a shared memory controller. The heterogeneous computing core includes at least one of the CPU core, vision processing unit, neural network processing unit, GPU, and FPGA hard core. The system can be implemented as a kernel module, a kernel scheduling extension component, a device driver coordination component, or a system service within the embedded operating system.
[0010] Example 1: like Figure 1 As shown, the heterogeneous multi-core video task scheduling system for embedded operating systems is set between the application layer and the heterogeneous hardware platform. Logically, it includes a task capture and unitization module, a resource snapshot generation module, a cross-frame hot hysteresis annotation module, an effective computing power calculation module, a task-core-time slot mapping module, a hardware control instruction orchestration module, a controlled asynchronous submission module, and a result merging and output module. Figure 1 The application layer concurrent video task input end corresponds to the task capture and unitization module. Figure 1 The resource snapshot generation path in the document corresponds to the resource snapshot generation module. Figure 1 The thermal hysteresis time slot annotation path in the text corresponds to the cross-frame thermal hysteresis annotation module. Figure 1 The effective computing power in the task-core-timeslot mapping path corresponds to the effective computing power computing module and the task-core-timeslot mapping module. Figure 1 The underlying hardware standard control instruction orchestration path corresponds to the hardware control instruction orchestration module. Figure 1 The asynchronous submission and result collection paths in the code correspond to the controlled asynchronous submission module and the result merging and output module. Figure 1 The central processing unit core, vision processing unit, neural network processing unit, graphics processing unit, field-programmable gate array hard core, and shared memory controller correspond to the heterogeneous hardware platform that is scheduled.
[0011] The task capture and unitization module records request events at the system call interface where application requests enter the operating system and converts concurrent video tasks issued by the application layer into video frame time slot task units. The resource snapshot generation module collects the operating status of each heterogeneous computing core and shared memory controller and binds the collected operating status to the hardware physical heat dissipation path. The cross-frame thermal hysteresis labeling module generates thermal hysteresis labels on the video frame timeline based on resource snapshots and task power consumption levels. The effective computing power calculation module reduces the theoretical peak computing power based on the current processor load rate, current shared memory bandwidth utilization, and hysteresis heat accumulation parameters. The task-core-time slot mapping module generates a unique task-core-time slot mapping within a compatible mapping set. The hardware control instruction orchestration module converts the task-core-time slot mapping into underlying hardware standard control instructions. The controlled asynchronous submission module submits underlying hardware standard control instructions to the underlying hardware queue through asynchronous submission worker threads with fixed priority. The result merging and output module restores the output order of hardware execution results according to frame identifiers and processing stage identifiers.
[0012] The data flow between the above modules is as follows: concurrent video tasks at the application layer are first converted into video frame time slot task units; the video frame time slot task units and resource snapshots serve as inputs for generating thermal hysteresis labels; the kernel scheduling table with thermal hysteresis labels, along with the resource snapshots, is then input to the effective computing power calculation module and the task-core-time slot mapping module; the task-core-time slot mapping is converted into a hardware task submission table and underlying hardware standard control instructions; after the underlying hardware completes execution, the result merging and output module generates a standard control record log and an ordered result record log. Thus, the system forms a closed implementation chain within the same current scheduling window, from task capture, thermal constraint calculation, mapping selection, hardware submission to result output.
[0013] Example 2: At the system call interface where the application request enters the operating system, the task capture and unitization module records the request event. The system call interface can be an input / output control system call interface, a write system call interface, an image acquisition queue submission interface, an AI runtime task submission interface, or an interface for the multimedia framework to issue tasks to the driver layer. To avoid changes in the scheduling input during computation caused by concurrent requests from the application layer, after the request entering within the current scheduling window is recorded, the system converts it into a unified data object, namely a video frame time slot task unit.
[0014] Each video frame time slot task unit includes at least a frame identifier, a processing stage identifier, a producer synchronization barrier, a consumer synchronization barrier, the computational load required, the data transfer load, the task power consumption level, the target core type requirement, and the deadline. The frame identifier is used to uniquely represent the video frame sequence, preferably starting from... A non-negative integer that begins to increment. Processing stage identifiers are used to represent processing stages in the video processing chain, such as decoding, image preprocessing, neural network inference, post-processing, and encoding. Producer synchronization barriers indicate the data-ready state of the input buffer, and consumer synchronization barriers indicate the writable or overwriteable state of the output buffer. Required computation quantifies the execution complexity of the task on the computing core and uses the same computational benchmark as the theoretical peak computing power, baseline computing speed, and effective computing power. Data transfer volume characterizes the amount of data transferred by the task on the shared memory bus or direct memory access path and uses the same data benchmark as the preset nominal bandwidth and baseline transfer bandwidth. Task power consumption level characterizes the relative energy consumption intensity during task execution, with a preferred value of [value missing]. to An integer. The target core type requirement is used to represent the architectural compatibility between the task and heterogeneous computing cores, preferably, Indicates the central processing unit core. Represents the visual processing unit. This represents a neural network processing unit. Represents a graphics processing unit. This indicates a Field-Programmable Gate Array (FPGA) hard core. The deadline represents the time boundary by which a task must be completed; it can be obtained by adding the current timestamp of the monotonic system clock to the maximum allowable delay of the video processing link.
[0015] At the producer synchronization barrier, consumer synchronization barrier, and each compute hardware queue commit point in the direct memory access buffer, the system records the underlying buffer state. The underlying buffer state includes the memory space allocation ratio, the buffer busy / idle flag, and the remaining capacity of the hardware instruction queue. The remaining capacity of the hardware instruction queue can be used... to The integer representation between 0 and 1 indicates that more control descriptors can still be accommodated. The underlying buffer state is written to the kernel scheduling table along with the video frame time slot task unit, so that subsequent compatible mapping sets can simultaneously exclude candidate mappings that are unavailable due to buffer unavailability, insufficient remaining capacity in the hardware instruction queue, or unmet producer synchronization barriers.
[0016] Within the same frame, the system determines pipeline dependencies based on processing stage identifiers; between different frames, the system determines temporal relationships based on frame identifiers. For the same frame, the decoding stage executes before the preprocessing stage, the preprocessing stage executes before the neural network inference stage, the neural network inference stage executes before the postprocessing stage, and the postprocessing stage executes before the encoding stage or the result output stage. The system encapsulates the frame identifier, processing stage identifier, producer synchronization barrier, consumer synchronization barrier, computational load, data transfer load, task power consumption level, target core type requirements, deadline, and underlying buffer state into a video frame time slot task unit and writes it into the kernel scheduling table of the current scheduling window.
[0017] The current scheduling window can be a fixed time period or a batch processing interval. In a preferred embodiment, the current scheduling window is set to... milliseconds, corresponding to one second The video transmission rate in frames per second. The video stream of frames, the current scheduling window can be set to milliseconds to Milliseconds. To ensure that subsequent resource snapshots and hot lag labels are calculated for the same set of tasks, after all video frame time slot task units are written to the kernel scheduling table, the system locks the kernel scheduling table before collecting the running status of each heterogeneous computing core. Locking can be achieved through mutexes, spinlocks, read-write locks, version number verification, RCU replicas, or double-buffered scheduling tables. Their common purpose is to keep the set of tasks within the current scheduling window consistent during sorting and optimization calculations.
[0018] Example 3: The resource snapshot generation module collects the current processor load rate, normalized power consumption, queue wait time, and current shared memory bandwidth utilization of each heterogeneous computing core according to a preset sampling period. In a preferred embodiment, the preset sampling period is [missing information]. The preset sampling period is no more than one-quarter of the video frame period, in order to strike a balance between acquisition overhead and state freshness.
[0019] The current processor load rate is normalized to to Between, among Indicates free time. This indicates full load. Normalized power consumption is the ratio of real-time power consumption to the rated maximum power consumption of the corresponding core at its highest performance point. Queue wait time is estimated by reading the number of control descriptors submitted and pending in the device driver queue, combined with the execution time of a single underlying hardware operation; alternatively, it can be obtained by directly reading the waiting time of the earliest control descriptor in the hardware channel. Current shared memory bandwidth utilization is normalized to... to The value is calculated by dividing the actual amount of data transmitted in the current sampling period by the sampling period duration, and then by the preset nominal bandwidth of the shared memory controller.
[0020] The system generates a resource status record for each heterogeneous computing core. The resource status record includes the core identifier, core type, theoretical peak computing power, current processor load rate, queue wait time, normalized power consumption, memory access path, and associated cooling path. The core identifier is a non-negative integer used to index the computing units within the on-chip system. The core type maintains the same encoding rules as the target core type requirement. The theoretical peak computing power is a static constant written during system initialization based on the chip technical manual, device tree file, driver registration information, or platform calibration table. The memory access path indicates which internal interconnect bus or memory channel the core uses to access shared memory. The associated cooling path indicates the physical cooling path of the hardware where the core resides.
[0021] The hardware physical heat dissipation path refers to the heat conduction path formed by chip packaging, thermal pads, copper foil layers on printed circuit boards, heat sinks, or metal casings in a fanless embedded vision terminal. During system integration, a heat dissipation channel topology is established based on the circuit board layout, chip package diagram, thermal simulation results, or actual temperature rise test results. Heterogeneous computing cores located in the same silicon wafer area, sharing the same heat sink, close to the same metal casing area, or exhibiting significant thermal coupling are grouped into the same hardware physical heat dissipation path. To ensure reproducibility of subsequent calculations, the system generates a hardware physical heat dissipation path table during initialization. This table includes at least the associated heat dissipation path, core identifier set, thermal coupling coefficient, and thermal response time constant.
[0022] For shared memory controllers, the system generates a unique bandwidth usage record. This record includes the shared memory controller's preset nominal bandwidth and the current shared memory bandwidth utilization. Heterogeneous computing cores with the same memory access path have their access pointers in their resource status records pointing to the same bandwidth usage record, indicating a shared bus bandwidth contention among multiple cores. All resource status records and bandwidth usage records are encapsulated into a resource snapshot within the current scheduling window. After generation, the resource snapshot is set to read-only mode. This read-only mode can be achieved through clearing memory page write permissions, passing constant pointers, reference counting protection, version number verification, or using a snapshot copy.
[0023] Example 4: like Figure 2 As shown, the video frame timeline is divided into multiple consecutive time slots, and the hardware physical heat dissipation path is... Figure 2 The heat dissipation paths are represented as A, B, and C, and the thermal hysteresis labels are written into the video frame timeline according to the corresponding physical heat dissipation paths of the hardware. Figure 2 The task-core-time slot mapping in the diagram represents a unique and definite correspondence between the target core and the target time slot. Figure 2 The hardware task submission table in the table represents the submission order, which is formed in ascending order of target time slot, frame identifier, processing stage identifier, and submission identifier. Figure 2 The connection between the thermal hysteresis tag and the hardware task submission form indicates that the thermal hysteresis tag participates in the effective computing power reduction and resource allocation cost model calculation.
[0024] The cross-frame thermal hysteresis labeling module reads the target core type requirements, data transfer volume, and task power consumption level from the video frame time slot task unit, and filters the candidate core set in conjunction with resource snapshots. For each core in the candidate core set, the system reads its corresponding heat dissipation path. If a task can be executed by multiple core types, the system retains multiple candidate heat dissipation paths and forms thermal hysteresis labels for each candidate heat dissipation path; if a task can only be executed by a specific type of core, the thermal hysteresis label corresponding to the task is bound to the hardware physical heat dissipation path to which the corresponding core belongs. After the task-core-time slot mapping is determined, the system only retains the thermal hysteresis labels corresponding to the hardware physical heat dissipation path to which the target core belongs to participate in the subsequent generation of the hardware task submission table and ordered execution result table; the thermal hysteresis labels corresponding to candidate heat dissipation paths not selected by the target core are not written into the final kernel scheduling table.
[0025] Let the video frame time slot task unit be... Candidate core is The target time slot is The computational load required for the operation is The baseline computing speed of the candidate core is The amount of data moved is The baseline transmission bandwidth of the memory access path corresponding to the candidate core is The estimated execution duration is: in, Used to form candidate time boundaries during the thermal hysteresis tag generation stage. Indicates candidate core The corresponding memory access path.
[0026] Set target time slot The starting time is Candidate Core In the target time slot The previous available time was The producer synchronization fence satisfies the time requirement. With video frame time slot task unit Having the same frame identifier and a processing stage identifier smaller than the video frame slot task unit The set of task units is ,gather Internal Task Unit The expected completion time is Then the candidate execution start time is: when When empty, Retrieve the start time of the current scheduling window. The candidate execution end time is: The set of task coverage slots on the video frame timeline is as follows: The set of cross-frame thermal hysteresis slots on the video frame timeline is as follows: in, Indicates the slot number on the video frame timeline. Indicates candidate core The thermal response time constant of the physical heat dissipation path of the hardware. This indicates the preset dissipation threshold. The task coverage time slot set is used to indicate the actual video frame time axis position occupied by the video frame time slot task unit, and the cross-frame thermal hysteresis time slot set is used to indicate the video frame time axis position that still has a thermal impact on the same hardware physical heat dissipation path after the video frame time slot task unit has finished executing.
[0027] Let the smallest time slot number in the cross-frame thermal hysteresis time slot set be . Candidate Core The physical heat dissipation path of the hardware is Hardware physical heat dissipation path The core components affected by heat are The task power consumption level is The thermal hysteresis label is: The thermal hysteresis tag contains specific hardware physical heat dissipation paths, core groups affected by heat, task power consumption levels, thermal hysteresis start time slot, and the expected task execution end time. The thermal hysteresis start time slot is determined by the smallest time slot number in the cross-frame thermal hysteresis time slot set. When the cross-frame thermal hysteresis time slot set is empty, the system does not write thermal hysteresis tags for the corresponding candidate cores and target time slots.
[0028] When more than one video frame time slot task unit falls into the same hardware physical heat dissipation path, the system does not average and merge the thermal hysteresis labels, but instead retains the corresponding multiple thermal hysteresis labels side by side. Side-by-side retention can be implemented using linked lists, arrays, bitmap indexes, red-black trees, or time wheel structures. Side-by-side retention ensures that each task's power consumption level, execution duration, thermal hysteresis start time slot, and expected task end time remain independent. This allows subsequent calculations of the hysteresis heat accumulation parameter to be accumulated at the task granularity, avoiding thermal distortion caused by premature merging.
[0029] Within the current scheduling window, the system processes video frame time slot task units one by one in ascending order of deadline, frame identifier, and processing stage identifier. For video frame time slot task units for which a task-core-time slot mapping has not yet been established, the system uses candidate thermal hysteresis labels to evaluate candidate cores and candidate time slots; for video frame time slot task units for which a task-core-time slot mapping has already been established, the system uses the thermal hysteresis label corresponding to the target core to participate in the subsequent calculation of hysteresis heat accumulation parameters. Thus, thermal hysteresis label generation and task-core-time slot mapping form a consistent computational closed loop.
[0030] Example 5: The effective computing power calculation module extracts the accumulated hysteresis heat of the target core based on the undissipated task data with thermal hysteresis tags. Undissipated task data refers to task data where the thermal hysteresis tag still covers the target time slot, or whose thermal impact has not yet exceeded the preset dissipation threshold.
[0031] Let the core of the target be The target time slot is , for the core of the target The heat dissipation path is subject to thermal cross-influence and is within the target time slot. The set of tasks that have not yet completely dissipated is Then the task set for: in, This indicates that the task data has been written to the thermal hysteresis tag. Represents task data The corresponding physical heat dissipation path of the hardware. Represents task data The corresponding core group affected by heat, Represents task data The expected completion time of the task.
[0032] The normalized value corresponding to the task power consumption level is: in, Represents task data Task power consumption level, This indicates the highest possible value for the task's power consumption level, preferably... .
[0033] For the core of the target and target time slot The parameter for hysteresis heat accumulation is: in, Indicates the core of the target In the target time slot The parameter of hysteresis heat accumulation. Represents task data With the core of the target The thermal coupling coefficient between them Represents task data The expected execution duration, Indicates the core of the target The thermal response time constant of the physical heat dissipation path of the associated hardware. The range of values for the thermal coupling coefficient is... to The preferred value is The range of values for the thermal response time constant is: milliseconds to milliseconds, preferred value Milliseconds. The thermal coupling coefficient is determined based on chip layout thermal simulation, thermal imaging testing, or board-level temperature rise experiments. The thermal response time constant is obtained by fitting the temperature rise relaxation curve of a fanless terminal under burst power input.
[0034] The system uses parameters such as current processor load rate, current shared memory bandwidth utilization, and accumulated lag heat to reduce the theoretical peak computing power. (The last part, "assigning cores," appears to be an unrelated fragment and is omitted from the translation.) The theoretical peak computing power is The current processor load rate is The current shared memory bandwidth utilization rate is Then the core of the target In the target time slot The effective computing power is: in, Indicates effective computing power. This represents the lower limit of effective computing power. This indicates the lower limit of the current processor load rate. This represents the lower limit of the current shared memory bandwidth utilization. Through this reduction, the higher the current processor load, the higher the current shared memory bandwidth utilization, and the higher the hysteresis heat accumulation parameter, the lower the effective computing power. When the current processor load or the current shared memory bandwidth utilization reaches its upper limit, the effective computing power still maintains a computable lower limit.
[0035] The real-time available bus bandwidth is: in, Indicates candidate core The real-time available bus bandwidth corresponding to the memory access path This indicates the preset nominal bandwidth of the shared memory controller or the corresponding memory access path. This represents the minimum reference constant used to maintain computational stability. Used to limit the current shared memory bandwidth utilization to to between.
[0036] For video frame slot task units In candidate core and target time slot Based on the estimated completion time, the system calculates the core computation time and bus transport time separately. The core computation time is: Bus transport time is: The estimated completion time is: in, This indicates the queue wait time recorded in the resource snapshot. Indicates the time consumed by core operations. Indicates the time taken for bus transport. Indicates the estimated completion time.
[0037] The estimated completion time is: The system constructs a time response cost parameter using the estimated completion time and the deadline. A video frame time slot task unit is defined. The deadline is The minimum time reference constant is ,when When, the corresponding candidate core and target time slot are excluded from the compatible mapping set; when At that time, the time response cost parameter is: in, Represents the video frame time slot task unit In candidate core and target time slot The time response pressure is high. The closer the estimated completion time is to the remaining time allowed before the deadline, the higher the time response cost parameter becomes.
[0038] The calorie penalty parameter is: in, Indicates the core of the target In the target time slot The parameter for hysteresis penalty cost. The higher the hysteresis accumulation parameter, the higher the hysteresis penalty cost parameter.
[0039] The resource allocation cost model is as follows: in, Represents the video frame time slot task unit In candidate core and target time slot The time slot allocation cost is calculated using a normalized method. Both the time response cost parameter and the heat penalty cost parameter are calculated in a normalized form, allowing them to be summed at a unified scale.
[0040] The task-core-time slot mapping module determines the compatible mapping set based on the target core type requirements. Let's define a video frame time slot task unit. The candidate core set is The set of time slots available for allocation within the current scheduling window is: Candidate Core In the target time slot The remaining capacity of the hardware instruction queue is The corresponding memory access path is in the target time slot. Available flags are Then the set of compatible mappings is: in, Represents the video frame time slot task unit A set of compatible mappings. Because The set of compatible maps already includes the producer synchronization barrier satisfaction time, candidate core availability time, and expected completion time of the preceding processing stage within the same frame. It can exclude candidate maps that do not meet pipeline dependency, core availability, hardware instruction queue remaining capacity, and memory access path availability.
[0041] The system iterates through the compatibility mapping set, inputs each set of candidate cores and candidate time slots into the resource allocation cost model, and obtains the minimum cost candidate set: set up Indicates the frame identifier. If the processing stage identifier is used, then the target core and target time slot are: in, Indicates the core of the objective. Indicates the target time slot, This means eliminating concurrency issues by prioritizing deadline, frame identifier, processing stage identifier, target time slot, and core identifier in ascending order. The submission identifier is allocated during the hardware control instruction orchestration stage according to the mapping results and is only used for merging the submission of underlying hardware standard control instructions and hardware execution results; it does not participate in the concurrency determination during the task-core-time slot mapping stage. Through these rules, even if two or more mapping combinations have the same resource allocation cost, the system can still generate a uniquely corresponding task-core-time slot mapping.
[0042] When the compatible mapping set is empty, the system marks the corresponding video frame time slot task unit as having completed execution in one of the following states: synchronization barrier timeout, hardware error, or output buffer unavailable, and writes this state to the ordered result log. When the compatible mapping set is not empty, the system writes the target core, target time slot, frame identifier, processing stage identifier, and submission identifier to the task-core-time slot mapping. Thus, the output of the resource allocation cost model, the parallel elimination rules, and the abnormal state records jointly ensure that each video frame time slot task unit within the current scheduling window has a definite processing result.
[0043] Example 6: The hardware control instruction orchestration module rearranges video frame time slot task units based on the task-core-time slot mapping. The system assigns a sequentially increasing commit identifier to each video frame time slot task unit within the current scheduling window. Subsequently, all video frame time slot task units are sorted in multiple levels according to the increasing order of target time slot, frame identifier, processing stage identifier, and commit identifier, forming a hardware task commit table. This sorting method is similar to... Figure 2 The hardware task submission table is consistent, ensuring that the submission order of the target time slot on the video frame timeline and the underlying hardware standard control commands remains consistent.
[0044] For each video frame time slot task unit in the hardware task submission table, the system sequentially generates the following instructions: wait for synchronization barrier instruction, map direct memory access buffer instruction, set processor affinity or device queue instruction, set working performance point instruction, submit descriptor instruction, and trigger synchronization barrier instruction. These six types of instructions together constitute the underlying hardware standard control instructions.
[0045] Each underlying hardware standard control instruction includes at least the following fields: instruction type, commit identifier, target core identifier, target timeslot, target device queue or target hardware channel, input buffer identifier, output buffer identifier, producer synchronization barrier identifier, consumer synchronization barrier identifier, job performance point identifier, and execution completion status field. The instruction type field distinguishes between instructions waiting for synchronization barriers, instructions mapping direct memory access buffers, instructions setting processor affinity or device queues, instructions setting job performance points, commit descriptor instructions, and instructions triggering synchronization barriers. The commit identifier restores the original commit order of underlying hardware standard control instructions within the current scheduling window. The target core identifier and target timeslot are derived from the task-core-timeslot mapping.
[0046] The Wait for Synchronization Fence instruction is used to control the current task to wait for the producer's synchronization fence to be triggered when input data is not ready. The Map Direct Memory Access Buffer instruction is used to establish a dynamic bus mapping based on the input buffer address and output buffer address. The Set Processor Affinity or Device Queue instruction is used to bind a task to a target core, target device queue, or target hardware channel. The Set Performance Point instruction is used to configure the operating frequency and voltage of the corresponding heterogeneous computing core based on the task's power consumption level, effective computing power, and platform power consumption strategy. The Commit Descriptor instruction is used to encapsulate the computational load, data transfer load, buffer address, target queue, and synchronization information into a hardware-recognizable descriptor and write it to the target hardware queue. The Trigger Synchronization Fence instruction is used to update the consumer's synchronization fence state after the hardware has completed execution; its triggering action is driven by a hardware completion event, driver callback, or completion interrupt signal, and it does not change the consumer's synchronization fence state before the hardware computation is complete.
[0047] The system writes the underlying hardware standard control instructions into the kernel commit ring buffer according to the order in the hardware task commit table. The kernel commit ring buffer is a fixed-capacity ring-shaped storage area pre-allocated in kernel space. After all instructions in the current scheduling window are written, the system locks the write channel of the ring buffer before capturing video tasks in the next scheduling window. Locking can be achieved by freezing the write pointer, setting the version number, atomic status bits, double-buffered swapping, or reading-only commit copies, ensuring that the ring buffer only retains instructions from the current scheduling window and preventing instruction overwriting between different scheduling windows.
[0048] Example 7: The controlled asynchronous commit module creates asynchronous commit worker threads during the kernel-mode initialization phase. These worker threads have a fixed priority and are independent of the application's execution path. Fixed priority can be achieved through real-time first-in-first-out scheduling, real-time round-robin scheduling, kernel work queue priority, processor affinity binding, or device queue exclusivity. Asynchronous isolation means that the commit thread is isolated from the application-layer request thread in terms of execution path, commit queue, lock resources, or processor affinity, ensuring that sudden commits from application-layer requests do not directly disrupt the commit order of the underlying hardware instruction queue.
[0049] The asynchronous commit worker thread reads the underlying hardware standard control instructions from the kernel commit ring buffer in ascending order of the commit identifier, and executes them in sequence: waiting for the synchronization barrier, memory mapping, running frequency and queue configuration, and descriptor commit. When the producer synchronization barrier is not triggered, the worker thread enters a waiting state; when the input data is ready, the worker thread maps the direct memory access buffer to the addressable space of the target hardware device through the system memory management unit or the input / output memory management unit; then, it completes the running frequency and queue configuration according to the target kernel, working performance point, and device queue identifier; finally, it writes the hardware descriptor to the register trigger point or ring instruction queue of the target hardware device.
[0050] During the execution of video tasks by the underlying hardware, hardware completion events for different video frame time slot task units may arrive out of order of their submission identifiers. After the heterogeneous computing core completes the computation of its corresponding video frame time slot task unit, the hardware notifies the system via a completion interrupt signal, the completion bit of the shared status register, or a driver callback. After the driver confirms the termination of the hardware computation, it writes the result to the target output buffer and triggers the consumer synchronization barrier. The system records the frame identifier, processing stage identifier, device identifier, completion time, output buffer identifier, and execution completion status, forming a hardware execution status record.
[0051] The result merging and output module performs multi-level merging of hardware execution status records. For multiple processing stages with the same frame identifier, the system organizes the results of decoding, preprocessing, neural network inference, post-processing, and encoding stages in the same data linked list or result array according to the ascending order of the processing stage identifiers. For task units with different frame identifiers, the system establishes an ordered execution result table according to the ascending order of the frame identifiers. Thus, the submission order, hardware completion order, and output order can be separated; even if the hardware completion event arrives out of order, the system can still recover the hardware execution results arranged by frame identifier and processing stage identifier through the ordered execution result table.
[0052] Let the set of video frame time slot task units in the current scheduling window be . The total number of video frame time slot task units in the current scheduling window is Video Frame Slot Task Unit The execution completion status record is marked as ,but: The number of video frame time slot task units that have already formed execution completion status records within the current scheduling window is: The system determines that all tasks in the current scheduling window have been completed or have been recorded as completed when the following conditions are met: in, The value is or , This indicates that an execution completion status record has not yet been generated. This indicates that an execution completion status record has been created. The system does not directly compare the maximum commit flag with the hardware completion event count; instead, it compares the total number of video frame time slot task units within the current scheduling window with the number of execution completion status records. Once these conditions are met, the system locks the ordered execution result table and delivers it as the hardware execution result to the preset system output interface.
[0053] Example 8: The result merging and output module encapsulates the underlying hardware standard control instructions in ascending order of the submission identifier, forming a standard control log. The standard control log includes the submission identifier, target heterogeneous core device code, hardware channel receive queue code, operating frequency control code, target memory region location identifier, synchronization mechanism control code, and execution completion status field. The standard control log is used to trace the original operation sequence of tasks submitted to heterogeneous hardware within the current scheduling window.
[0054] The system encapsulates hardware execution results in ascending order of frame identifiers, and within the same frame, in ascending order of processing stage identifiers, forming an ordered result log. The ordered result log includes frame identifier, processing stage identifier, output buffer identifier, completion time, execution completion status, failure reason code, timeout flag, and retry count. The execution completion status can indicate normal completion, timeout completion, hardware exception, synchronization barrier timeout, or output buffer unavailable. The failure reason code distinguishes different exception sources, the timeout flag indicates whether the deadline or synchronization barrier wait has exceeded a preset boundary, and the retry count indicates the number of times the same video frame time slot task unit has been resubmitted within the current scheduling window.
[0055] The preset system output interface can be a callback function registration point for the multimedia application framework layer, a character device mapping channel, a shared memory notification interface, a remote return drive interface, or a video stream encoder input interface. After the system delivers the standard control log and the ordered result log to the preset system output interface, external display components, encoders, or application layer streaming media interfaces can read the correct time-series video images, inference results, or encoded outputs according to the ordered result log.
[0056] After the standard control log and ordered result log of the current scheduling window have been delivered, the system releases the temporary scheduling allocation memory requested in the current cycle, releases the queue, updates the scheduling window version number, and enters the task capture process of the next scheduling window.
[0057] To further present Figure 2 The scheduling relationship shown is in a Within the current scheduling window (milliseconds), the application layer simultaneously issues tasks for two video streams. The first video stream's... The frame consists of four processing stages: decoding, preprocessing, neural network inference, and encoding. The second video stream's first... The frame consists of three processing stages: decoding, preprocessing, and neural network inference. The system generates a corresponding video frame time slot task unit for each processing stage and binds it with the frame identifier, processing stage identifier, producer synchronization barrier, consumer synchronization barrier, computational load required, data transfer load, task power consumption level, target core type requirement, and deadline.
[0058] Assume the target core type requirement of the neural network inference task points to both the neural network processing unit (NN unit) and the graphics processing unit (GPU), with the NN unit and the graphics processing unit located on the same physical heat dissipation path, and the GPU on a different physical heat dissipation path. If the resource snapshot shows that the NN unit currently has a low processor load, but there are unresolved thermal lag tags left by the previous frame's high-power inference task on the same physical heat dissipation path, then the effective computing power calculation module will obtain a high lag heat accumulation parameter in the target time slot corresponding to the NN unit, thereby reducing the effective computing power of that NN unit.
[0059] In the above scenario, the task-core-slot mapping module selects the neural network processing unit not only based on the instantaneous processor load rate, but also incorporates both the time response cost parameter and the heat penalty cost parameter into the resource allocation cost model. If the graphics processing unit has a slightly higher queue waiting time, but its associated heat accumulation parameter on its cooling path is low, and its expected completion time is still earlier than the deadline, then the resource allocation cost model selects the graphics processing unit and the subsequent target slot as the task-core-slot mapping for that neural network inference task. Therefore, high-power tasks will not continuously fall into the same hardware physical cooling path, and the heat accumulation in subsequent video frame slots is proactively incorporated into the scheduling constraints.
[0060] The above process is not simply passively reducing the frequency after reading the current temperature sensor value. Instead, it transforms the cross-frame thermal hysteresis that has not yet been reflected in the sensor readings into a calculable time slot allocation constraint within the current scheduling window. Through task-core-time slot mapping, hardware instruction rearrangement, and ordered result merging, it enables high-concurrency video processing tasks to maintain a stable execution order and time response under limited heat dissipation conditions.
[0061] In the above implementation, resource snapshots are generated by the kernel sampling interface. Alternatively, resource snapshots can also be provided by device drivers, system management controllers, on-chip monitoring units, or security coprocessors. Any resource snapshot can be used as long as it can provide heterogeneous computing core load, queue wait times, power consumption status, memory bandwidth utilization, and thermal path binding relationships within the current scheduling window.
[0062] In the above embodiments, thermal hysteresis tags are stored side-by-side in a linked list or an array of pointers. Alternatively, thermal hysteresis tags can also be stored in a time wheel, a sparse matrix, a sliding window cache, or a circular array with heat dissipation path indexes. As long as the task power consumption level, the thermally affected core group, the thermal hysteresis start time slot, and the expected end time of the task execution can be retained, and subsequent calculation of hysteresis heat accumulation parameters can be supported, the same technical objective can be achieved.
[0063] In the above implementation, the resource allocation cost model uses the time response cost parameter and the heat penalty cost parameter to be directly summed. Alternatively, predetermined coefficients can be set for the two cost parameters. As long as the coefficients are pre-set or configured by the system strategy, and the mapping selection is still based on the effective computing power and deadline constraints of the heat effect reduction limit, the same technical objective can be achieved.
[0064] In the above implementation, the underlying hardware standard control instructions include a wait-for-synchronization-fence instruction, a map-to-direct-memory-access-buffer instruction, a processor affinity or device queue setting instruction, a performance point setting instruction, a descriptor commit instruction, and a synchronization-fence trigger instruction. For different hardware platforms, these instructions can be mapped to different register write operations, driver calls, firmware commands, queue descriptors, or runtime interface calls; as long as they can control the heterogeneous computing core to execute video frame time slot task units according to the task-core-time slot mapping, the same technical objective can be achieved.
[0065] Example 9: When the system is implemented as a kernel module, resource registration is completed at the kernel module initialization entry point. Resource registration includes the registration of the kernel scheduling table, resource snapshot cache, hardware physical heat dissipation path table, global unresolved heat dissipation lag tag linked list, kernel commit circular buffer, asynchronous commit worker threads, and preset system output interfaces. The system attaches task capture hooks at system call interfaces via kernel trace points. These hooks are attached to the call paths for input / output control system calls, write system calls, and multimedia frameworks sending tasks to drivers. When concurrent video tasks at the application layer enter these call paths, the task capture hook reads the frame identifier, processing stage identifier, producer synchronization barrier, consumer synchronization barrier, computational load, data transfer load, task power consumption level, target core type requirements, and deadline, and delivers this information to the task capture and unitization module.
[0066] When the system is implemented as a scheduling extension component, it extends the operating system's native scheduler by registering a scheduling class. The scheduling extension component's execution priority is higher than that of ordinary user process scheduling but lower than the top half of hardware interrupt processing. In the event of a scheduling conflict, the system prioritizes the top half of the hardware interrupt processing and postpones the native scheduling logic to the bottom half of the interrupt, ensuring that the task-core-time slot mapping calculation within the current scheduling window is not preempted by ordinary user processes. When the system is implemented as a device driver collaboration component, it interfaces with various heterogeneous hardware drivers through a standard operation function structure. This standard operation function structure includes three unified callback interfaces: status reading, task submission, and performance point configuration. After different heterogeneous hardware drivers implement the corresponding callbacks according to the interface specifications, they can adapt to the resource snapshot generation module, hardware control instruction orchestration module, and controlled asynchronous submission module.
[0067] Example 10: The inter-module triggering sequence adopts a unidirectional pipelined triggering method within the current scheduling window. The resource snapshot generation module triggers a single sampling after the kernel scheduling table of the current scheduling window is locked. After sampling, a read-only resource snapshot is generated and passed to the cross-frame thermal hysteresis annotation module, the effective computing power calculation module, and the task-core-time slot mapping module via constant pointers. The cross-frame thermal hysteresis annotation module processes all video frame time slot task units within the current scheduling window in batches after the resource snapshot is generated, and generates candidate thermal hysteresis labels for each. The task-core-time slot mapping module starts calculation after all candidate thermal hysteresis labels are generated, and writes the task-core-time slot mapping to the kernel scheduling table after calculation.
[0068] Data transfer between modules employs a shared memory + spinlock synchronization method. Each module's output data has a version number checksum, which is verified before downstream modules read it. If the version number checksum read by a downstream module does not match the current scheduling window version number, the corresponding data object in shared memory is reread. The resource snapshot generation module, cross-frame hot hysteresis annotation module, effective computing power calculation module, and task-core-timeslot mapping module all only read the locked kernel scheduling table and read-only resource snapshots to avoid data inconsistencies caused by concurrent modifications.
[0069] Example 11: The computational load, data transfer volume, and task power consumption level required for the operation are distributed by the application layer along with the task request. Prior to distribution, the application side pre-calibrates these parameters based on a benchmark dataset. The computational load is based on a benchmark number of operation cycles, which is consistent with the theoretical peak computing power benchmark. This benchmark value is read from the hardware calibration table during system initialization. The data transfer volume is based on bytes, consistent with the preset nominal bandwidth benchmark. The task power consumption level is determined based on the measured steady-state power consumption of the task on the corresponding core. The level corresponds to the lowest power consumption level. Each level corresponds to the highest power consumption level, and the power consumption threshold for each level is pre-defined in the hardware physical heat dissipation path table.
[0070] Task power consumption levels are classified according to the steady-state power consumption of the task when running on the corresponding type of core, and are divided into several categories. class. The steady-state power consumption of the stage is less than the core's rated maximum power consumption. , The steady-state power consumption of the corresponding stage is not less than the core's rated maximum power consumption. And less than the core's rated maximum power consumption , The steady-state power consumption of the corresponding stage is not less than the core's rated maximum power consumption. And less than the core's rated maximum power consumption , The steady-state power consumption of the corresponding stage is not less than the core's rated maximum power consumption. And less than the core's rated maximum power consumption , The steady-state power consumption of the corresponding stage is not less than the core's rated maximum power consumption. And not greater than the core's rated maximum power consumption. When the application layer issues a task, it assigns the corresponding task power consumption level based on the task type and the target core type, referring to a pre-defined power consumption classification table. The same task running on different types of cores corresponds to different task power consumption levels, and each is matched with the rated maximum power consumption standard of the corresponding core, so that the task power consumption level remains relatively consistent across different heterogeneous computing cores.
[0071] Example 12: Both the producer synchronization barrier and the consumer synchronization barrier are implemented using kernel atomic variables, with the atomic variables taking values of... This indicates that the buffer is not ready, and the atomic variable takes the value... This indicates that the buffer is ready. For existing video frameworks such as V4L2 and multimedia codec frameworks, the system updates the values of the atomic variables of the corresponding synchronization fence through the completion callback function of the framework buffer queue, realizing the mapping between the framework's native queue and the producer synchronization fence and consumer synchronization fence.
[0072] Synchronization barrier timeouts are timed using a kernel timer, with the timeout period determined by the difference between the deadline and the current time. If the producer's synchronization barrier fails to meet the timeout requirement, the system marks the corresponding video frame time slot task unit as being in a synchronization barrier waiting timeout state and removes it from the task-core-time slot mapping calculation for the current scheduling window. If the consumer's synchronization barrier is not released and there is no free output buffer, the system marks the corresponding video frame time slot task unit as being in an output buffer unavailable state and writes this state to the ordered result log.
[0073] Example 13: The current scheduling window uses a periodic hard-triggered method, triggered by a high-precision kernel timer. The timer period is consistent with the duration of the current scheduling window. Tasks arriving at the window boundary are assigned based on the timestamp of their request entering the system call interface. Tasks with timestamps earlier than the window boundary are assigned to the current scheduling window, while those with timestamps later are assigned to the next scheduling window. If there are incomplete tasks within the current scheduling window, the next scheduling window opens normally according to the timer period. Incomplete tasks remain in their corresponding task queues and continue execution, without participating again in the task-kernel-timeslot mapping allocation for the next scheduling window.
[0074] When the number of tasks within the current scheduling window exceeds the hardware's maximum processing capacity, the system retains processable tasks in ascending order of their deadlines. The excess tasks are marked as timed out and written to the ordered result log. The maximum waiting time for the current scheduling window is equal to the current scheduling window duration. The system submits all underlying hardware standard control commands for the current scheduling window to the hardware to begin timing. If, after the maximum waiting time, a video frame time slot task unit has not yet formed an execution completion status record, the system forcibly marks the incomplete video frame time slot task unit as timed out and generates a corresponding failure reason code. After a timeout is triggered, the system locks the ordered execution result table, delivers the records of completed and timed-out tasks to the preset system output interface, releases the temporary resources of the current scheduling window, and enters the next scheduling window process. Timed-out tasks are not automatically postponed to the next scheduling window; the application layer decides whether to resubmit based on the return status.
[0075] Example 14: The kernel scheduling table employs a hybrid structure of arrays and linked lists. The main table stores task unit indices in frame identifier order, and the sub-linked lists corresponding to each frame identifier are loaded with video frame time slot task units in processing stage identifier order. Locking of the kernel scheduling table uses read-write locks to achieve full-table granularity locking; the write lock is held during the task write phase, and the read lock is held during the computation and reading phase. The task capture and unitization module requests a write lock when writing video frame time slot task units; the resource snapshot generation module, cross-frame hot hysteresis annotation module, effective computing power calculation module, and task-core-time slot mapping module request a read lock when reading the current scheduling window task set.
[0076] When using the RCU replica mechanism, write operations modify data based on the replica. After modification, updates are completed via atomic pointer replacement, waiting for all read ends to finish reading before replacement. When using a double-buffered scheduling table, the foreground table is used for reading computations, and the background table is used for writing new tasks. When the current scheduling window switches, the roles of the foreground and background tables are switched via atomic swap operations. After the switch, the content of the background table is reset. The read-write locks, RCU replicas, and double-buffered scheduling tables mentioned above are all used to ensure that the kernel scheduling table has a consistent read view within the current scheduling window.
[0077] Example 15: The current processor load rate of the central processing unit (CPU) core is obtained by reading the kernel load statistics interface. The current processor load rate of the vision processing unit, neural network processing unit, and graphics processing unit is calculated by reading the hardware performance counter values from the system file nodes of the corresponding drivers. The current processor load rate of the field-programmable gate array (FPGA) hard core is obtained by reading the hardware task queue occupancy rate. The normalized power consumption is obtained by reading the power monitoring register values of each heterogeneous computing core and combining them with the corresponding core's rated maximum power consumption.
[0078] Queue wait time is estimated by multiplying the number of pending descriptors in the driver layer hardware instruction queue by the average execution time per descriptor; alternatively, it can be obtained by directly reading the wait time register of the earliest pending descriptor in the hardware channel. Current shared memory bandwidth utilization is calculated by reading the amount of data transmitted by the shared memory controller performance counter during the sampling period, combined with the sampling period duration and the preset nominal bandwidth. If the hardware does not support corresponding state acquisition, a default baseline value is used instead. This default baseline value is written to the corresponding field of the resource snapshot during system initialization.
[0079] Example 16: The hardware physical heat dissipation path table contains four core fields: the heat dissipation path number, the core identifier set, the thermal coupling coefficient matrix, and the thermal response time constant. Each hardware physical heat dissipation path corresponds to one table entry. The thermal coupling coefficient is calibrated through board-level temperature rise testing. During the test, a single core is enabled to run a full-load task, and the steady-state temperature rise values of other cores on the same hardware physical heat dissipation path are collected. The ratio of the temperature rise of the affected core to the temperature rise of the heat-generating core is used as the corresponding thermal coupling coefficient. The test conditions are ambient temperature and no active cooling.
[0080] The thermal response time constant was obtained through step power consumption testing. During the test, a step-load full power consumption was applied to the target core, and the relaxation curve of temperature rise over time was collected, using an exponential function. Perform least-squares fitting on the curve, and the fitted curve is... The value serves as the thermal response time constant for the corresponding physical heat dissipation path of the hardware. After the thermal coupling coefficient matrix and the thermal response time constant are written into the physical heat dissipation path table, the resource snapshot generation module reads the corresponding table entry through the heat dissipation path number, thus fixing the source of the value of the hysteresis heat accumulation parameter.
[0081] Example 17: The number of samples within a single current scheduling window is the integer quotient of the current scheduling window duration divided by the preset sampling period: in, This indicates the number of samples within a single current scheduling window. Indicates the duration of the current scheduling window. This indicates the preset sampling period. The current scheduling window corresponds to milliseconds Second-rate Sampling is performed in millisecond cycles. The scheduling calculation uses the most recent sampling result completed at the start time of the current scheduling window as input. The sampling timestamp and the start time of the time slot use the same system monotonic clock as the time reference, and the sampled values are aligned to the corresponding time slot boundaries according to the timestamp.
[0082] When an exception such as a register read failure occurs during sampling, the result of the previous successful sampling is used as the current sample value. (Continuous) When a sampling anomaly occurs, the system marks the corresponding hardware state as abnormal and temporarily removes the corresponding core from the compatibility mapping set. After the hardware state anomaly is resolved and the next normal sampling is completed, the corresponding core re-enters the candidate core set.
[0083] Example 18: The time slots of the video frame timeline are divided using a fixed length, with the length of a single time slot set to... Milliseconds, time slot number from Encoding begins sequentially in ascending order of time. Time slots are allocated based on a globally unified system monotonic clock. All hardware physical heat dissipation paths share the same video frame timeline and time slot allocation rules; no path-specific time slots are allocated. The current scheduling window corresponds to milliseconds. Each time slot is a consecutive time slot, and the time slot number corresponds one-to-one with the time offset within the current scheduling window.
[0084] Let the start time of the current scheduling window be... The length of a single time slot is Then the first The start time of each time slot is: in, This indicates the time slot number. The task coverage time slot set, cross-frame thermal hysteresis time slot set, and target time slot of the video frame time slot task unit are all determined based on the above unified time slot division rules.
[0085] Example 19: Baseline computing speed This represents the computational load per unit time of the core at its rated operating performance point. Its value is equal to the theoretical peak computing power multiplied by the baseline utilization coefficient, which is a preset constant with a value of [value missing]. : in, This represents the baseline utilization coefficient. Baseline transmission bandwidth. This represents the steady-state transmission bandwidth of the corresponding memory access path under contention-free conditions. Its value is equal to the preset nominal bandwidth multiplied by the transmission efficiency coefficient, which is a preset constant with a value of [value missing]. : in, Indicates the preset nominal bandwidth. Indicates the transmission efficiency coefficient. Preset dissipation threshold. Values This represents the decay of the thermal effect to its peak value. The following times are considered to be completely dissipated and are no longer included in the calculation of the hysteresis heat accumulation parameter. The expected completion time of the preceding task in the same frame is calculated iteratively. First, the expected completion time of the earliest processing stage is calculated, and then the candidate execution start time formulas of the subsequent stages are substituted into it. During the calculation process, the reduction in effective computing power caused by thermal hysteresis is considered simultaneously.
[0086] Example 20: The system maintains a global undissipated thermal hysteresis tag linked list in the kernel space. This list stores all undissipated thermal hysteresis tags categorized by the hardware's physical heat dissipation path. After the task-core-time slot mapping calculation for the current scheduling window is completed, the system appends the thermal hysteresis tag corresponding to the final selected task-core-time slot mapping to the global undissipated thermal hysteresis tag linked list. Before starting thermal hysteresis tag calculation for the next scheduling window, the system first traverses the global undissipated thermal hysteresis tag linked list, performs attenuation calculations on each thermal hysteresis tag based on the start time of the current scheduling window, removes thermal hysteresis tags that meet the dissipation conditions, and uses the remaining thermal hysteresis tags as historical heat input to participate in the calculation of the accumulated hysteresis heat parameter for the current scheduling window.
[0087] During a system cold start, the global unresolved thermal hysteresis tag list is initialized to empty, and the initial hysteresis heat accumulation parameter for all cores is set to [value missing]. Therefore, the thermal state across scheduling windows can be inherited continuously, and historical thermal effects that have completely dissipated will not be repeatedly included in the calculation of the current scheduling window.
[0088] Example 21: Thermal hysteresis tags are stored using a time-wheel structure distributed across thermal paths. Each physical heat dissipation path corresponds to a time wheel, and the slots on the time wheel correspond one-to-one with the time slot numbers. Each slot is linked to a linked list that stores all thermal hysteresis tags corresponding to that slot in the starting time slot. When retrieving undissipated thermal hysteresis tags for a target time slot, the system traverses the linked lists of all slots preceding the target time slot from the time wheel of the corresponding physical heat dissipation path, and checks each tag against the decay formula to determine if it is still in an undissipated state.
[0089] Tag cleanup is triggered after each current scheduling window's calculations are completed. The system traverses all slots in the time wheel, removing thermal hysteresis tag nodes that have completely dissipated and releasing the corresponding memory to prevent memory leaks. Through the time wheel structure that distributes thermal paths, the system can quickly index thermal hysteresis tags according to the hardware's physical heat dissipation path and time slot number, while preserving the task granularity of parallel thermal hysteresis tags.
[0090] Example 22: Thermal coupling coefficient in the formula for hysteresis heat accumulation parameter For task data Core and target core The thermal coupling coefficient between the components. This thermal coupling coefficient is a hardware physical property, pre-stored in the thermal coupling coefficient matrix of the hardware physical heat dissipation path table, and is independent of the task itself. During calculation, the system first uses task data... The corresponding target core number is indexed from the hardware physical heat dissipation path table to the corresponding row, and then based on the target core... The number is indexed to the corresponding column, and the thermal coupling coefficient value between the cores in that group is read to ensure that the physical meaning and source of the thermal coupling coefficient are clear.
[0091] Example 23: Lower limit of effective computing power The value is taken as the theoretical peak computing power This is used to ensure that the core retains basic computing power even under full load and high heat accumulation conditions. Current processor load rate reduction lower limit. Values This prevents the effective computing power from dropping to zero when the current processor load is at its maximum. The current shared memory bandwidth utilization is reduced to a lower limit. Values This avoids the bus transfer capacity dropping to zero when the current shared memory bandwidth utilization is at its maximum. Minimal reference constant. Values To maintain the numerical stability of bus bandwidth calculations. Minimal time base constant. Values This avoids the anomaly of a denominator of zero in the calculation of time response cost parameters. All of the above parameters can be adjusted according to the hardware platform characteristics through the system configuration interface.
[0092] Example 24: Memory access path available flags Values Indicates availability, with a value of This indicates that the time slot is unavailable. The criterion for this judgment is the target time slot. The predicted bandwidth utilization of the corresponding memory access path is lower than a preset threshold, which is set to... The predicted bandwidth utilization is calculated by combining the current shared memory bandwidth utilization with the data transfer volume of mapped tasks. Remaining capacity of the hardware instruction queue. For target time slot The predicted remaining capacity is calculated by subtracting the instruction occupancy of mapped tasks between the current time slot and the target time slot from the current remaining hardware instruction queue capacity. Each task occupies a fixed amount of space. The calculated predicted remaining capacity is greater than [number] hardware instruction queue entries. The condition is then determined to be met.
[0093] Example 25: When multiple task-core-timeslot mappings have identical resource allocation costs, the system eliminates ties by prioritizing the mapping with an earlier deadline, frame identifier, processing stage identifier, target time slot, and core identifier in ascending order. Mappings with earlier deadlines are prioritized; if deadlines are the same, mappings with smaller frame identifiers are prioritized; if frame identifiers are the same, mappings with smaller processing stage identifiers are prioritized; if processing stage identifiers are the same, mappings with smaller time slot numbers are prioritized; if time slot numbers are the same, mappings with smaller core identifiers are prioritized. The commit identifier is assigned during the hardware control instruction orchestration stage according to the mapping results and is only used for merging the commit of underlying hardware standard control instructions and hardware execution results; it does not participate in the ties determination during the task-core-timeslot mapping stage. Therefore, the task-core-timeslot mapping result is uniquely determined.
[0094] Example 26: When the compatible mapping set is empty, the system sequentially checks three types of triggering conditions and marks them as completed. First, it checks if the producer synchronization barrier is not satisfied and has timed out; if so, it marks it as a synchronization barrier wait timeout. Second, it checks if all hardware of the corresponding target core type is in an abnormal offline state; if so, it marks it as a hardware abnormality. Finally, it checks if the consumer synchronization barrier of the output buffer has not been released and there is no free buffer; if so, it marks it as an unavailable output buffer. If none of the three conditions are met, the system defaults to a timeout state. All abnormal states are written to the corresponding fields in the ordered result log.
[0095] Example 27: Each underlying hardware standard control instruction is stored in a fixed-length structure, which sequentially includes the instruction type code, submission identifier, target core identifier, target timeslot number, target device queue number, input buffer address, output buffer address, producer synchronization barrier pointer, consumer synchronization barrier pointer, working performance point number, execution status field, and checksum field. Each field, the total length of the instruction is as follows: Byte alignment. Instruction mapping for different heterogeneous computing cores is achieved through driver layer adaptation functions. The CPU core instruction mapping is for thread affinity settings and computation function calls; the vision processing unit, neural network processing unit, and graphics processing unit instruction mapping is for corresponding hardware descriptor filling and queue submission; and the field-programmable gate array (FPGA) hard core instruction mapping is for register configuration and hardware logic startup.
[0096] Before writing underlying hardware standard control instructions into the queue, their integrity is verified using a checksum field. In case of an execution exception, the system automatically retryes. If a retry fails, the corresponding video frame time slot task unit will be marked as a hardware abnormal state. The trigger synchronization barrier instruction only drives the consumer synchronization barrier atomic variable to be set after the hardware execution is completed, and does not modify the consumer synchronization barrier state in advance.
[0097] Example 28: During system initialization, the system reads a list of available performance points for each core from the hardware device tree. Each performance point corresponds to a set of frequency and voltage values, with the performance point numbers increasing from low to high frequency. Task power consumption levels and performance point levels are configured accordingly. Correspondingly, task power consumption level Corresponding to the lowest frequency operating performance point, task power consumption level The highest frequency operating performance point. Effective computing power lower than the computing power corresponding to the current operating performance point. When this happens, the system automatically lowers the performance level by one level; the effective computing power is higher than the computing power corresponding to the current performance level. When this happens, the system will automatically increase the operating performance point by one level; the adjustment range will not exceed the upper or lower limits of the operating performance point corresponding to the task's power consumption level.
[0098] The platform's power consumption strategy is divided into a balanced mode and a performance mode. Balanced mode prioritizes limiting the use of high-power consumption levels, while performance mode prioritizes ensuring computing power. Switching between platform power consumption strategies is triggered through the system configuration interface. The setting performance point command reads the platform power consumption strategy before submission, ensuring that the task's power consumption level, effective computing power, and the platform power consumption strategy jointly determine the target core's performance point number.
[0099] Example 29: The kernel commit ring buffer is pre-allocated during kernel initialization, and its capacity is set to... The instruction uses read and write pointers to manage buffer boundaries, and the values of the read and write pointers are modulo the buffer capacity to achieve circular multiplexing. The write pointer is updated by the hardware control instruction orchestration module; for each underlying hardware standard control instruction written, the write pointer is incremented. The read pointer is updated by the controlled asynchronous submission module. Each time a low-level hardware standard control instruction is read and executed, the read pointer is incremented. When the kernel commit ring buffer is full, newly written underlying hardware standard control instructions overwrite the oldest unread instructions and record the overflow count.
[0100] In double-buffered mode, two kernel commit circular buffers of equal capacity are set up. The underlying hardware standard control instructions for the current scheduling window are written to the background buffer. After the window is locked, the foreground and background buffers are swapped. The asynchronous commit worker thread consistently reads the underlying hardware standard control instructions from the foreground buffer. The swapping operation is implemented through atomic pointer replacement, ensuring no instruction corruption during the commit process.
[0101] Example 30: The asynchronous submission worker thread uses a real-time first-in-first-out scheduling strategy, with a priority set to... It is a high-priority thread in the kernel's real-time threads. The asynchronous submission worker thread is bound to a specified CPU core for execution. The core identifier is configured during system initialization to avoid competing for computing resources with application processes. Multiple heterogeneous hardware components correspond to a single asynchronous submission worker thread. The asynchronous submission worker thread polls the kernel's submission ring buffer according to hardware type, sequentially reading and submitting underlying hardware standard control instructions. Instruction interactions between different hardware components do not interfere with each other.
[0102] The asynchronous submission worker thread executes independently of the application layer request thread. The application layer is only responsible for writing concurrent video tasks into the kernel scheduling table and does not directly manipulate the hardware submission queue, thus avoiding disruption of the submission order of underlying hardware standard control instructions by sudden application layer requests.
[0103] Example 31: After the hardware completion interrupt is triggered, the interrupt handler extracts the commit identifier from the completion message and indexes the corresponding video frame time slot task unit using the commit identifier. Subsequently, the system updates the execution completion status of the video frame time slot task unit, writes the output buffer address and completion timestamp into the corresponding fields, and triggers the consumer's synchronization fence atomic variable to be set. When tasks completed out of order are inserted into the ordered execution result table, the system uses binary search to locate the position corresponding to the frame identifier and the processing stage identifier, and directly inserts it into the corresponding linked list node, without needing to arrange them according to the submission order.
[0104] The consumer synchronization barrier is updated using atomic set operations to ensure thread safety during concurrent access across multiple cores. Updates to the consumer synchronization barrier are triggered only after hardware execution is complete; the barrier state is not modified prematurely. Therefore, the hardware completion event handling mechanism can simultaneously satisfy the requirements of out-of-order hardware completion, in-order hardware execution result output, and synchronization barrier state consistency.
[0105] Example 32: The system output interface is implemented using callback functions and shared memory. After the system completes the encapsulation of ordered result recording logs, it copies the result data to a pre-allocated shared memory area, and then calls the registered callback function to notify the application layer or upper-layer framework to read it. Callback function pointers are registered by the upper-layer module during system initialization, with one callback function corresponding to each output channel. Character device mapping channels are implemented through a file operation interface. The application layer obtains result data by reading the character device file; the read operation uses a blocking mode and returns automatically once the data is ready.
[0106] The remote return drive interface transmits result data via a bus message queue. The video stream encoder input interface directly passes the result buffer address to the encoder input queue, eliminating the need for additional data copying. The aforementioned callback functions, shared memory, character device mapping channels, remote return drive interface, and video stream encoder input interface can all serve as concrete implementations of the preset system output interface.
[0107] Example 33: The tasks of the Field-Programmable Gate Array (FPGA) hard core are measured in units of pre-programmed hardware-accelerated logic. The computational load is quantified by the hardware processing cycle per frame of data, and the data transfer volume is calculated based on the total number of bytes in the input / output buffers, using a unified computational benchmark with other heterogeneous computing cores. The current processor load rate is calculated by reading the queue occupancy rate of the FPGA hard core's internal task scheduler, and the normalized power consumption is obtained by reading the real-time current-voltage product value from the power management module. Queue waiting time is estimated by reading the remaining depth of the hardware instruction queue and the average processing cycle per task.
[0108] The underlying hardware standard control instructions corresponding to the field-programmable gate array (FPGA) hard core are mapped to hardware register configuration, direct memory access channel startup, and logic enable operations. After the underlying hardware standard control instructions are written to the hardware control register, task execution is triggered. Completion events are notified to the system via hardware interrupts. The interrupt handler reads the status register to confirm the execution result and generates a hardware execution status record.
[0109] Example 34: When multiple heterogeneous computing cores on the same memory access path execute tasks concurrently, the actual available bandwidth is allocated according to the proportion of data transported by each task. During the task-core-timeslot mapping stage, when predicting the bandwidth usage of the target timeslot, the sum of the data transported by tasks already mapped to the target timeslot and located on the same memory access path is divided by the single timeslot length to obtain the predicted bandwidth usage. This is then combined with the current shared memory bandwidth utilization to calculate the total predicted bandwidth utilization. If the total predicted bandwidth utilization exceeds... When the threshold is reached, the system will no longer allocate new tasks to the corresponding memory access path of the target time slot to avoid bandwidth overload.
[0110] When bandwidth conflicts occur, the system prioritizes tasks with earlier deadlines to meet their bandwidth requirements. If deadlines are the same, tasks with smaller processing stage identifiers are prioritized to ensure the pipeline execution order is not affected. This bandwidth contention model is related to the available memory access path flags. Together, they enable the compatible mapping set to pre-exclude candidate mappings with severe shared memory bandwidth conflicts during the task-core-slot mapping phase.
[0111] Example 35: The system every Each current scheduling window reads the actual temperature sensor values for each core and compares them with the relative temperature rise calculated by the thermal model during the same period. If the actual temperature rise is higher than the model's calculated value... If the actual temperature rise is lower than the model's calculated value, the thermal coupling coefficient of the corresponding hardware physical heat dissipation path will be adjusted proportionally. If the thermal coupling coefficient of the corresponding physical heat dissipation path of the hardware is reduced proportionally, the adjustment amount in a single instance shall not exceed the current value. .
[0112] When the actual temperature exceeds a preset safety threshold, the system triggers a temperature fallback strategy, forcibly lowering the performance level of all cores on the corresponding physical cooling path to the lowest level, while temporarily mapping high-power tasks to cores on other physical cooling paths. Once the temperature falls back below the preset safety threshold, the system resumes normal scheduling. The temperature fallback strategy has higher priority than the thermal hysteresis scheduling rule.
[0113] Example 36: The retry mechanism is triggered by three conditions: hardware execution exceptions, synchronization barrier timeouts, and buffer allocation failures. The retry process involves placing the corresponding video frame time slot task unit back into the mapping queue of the current scheduling window and re-participating in task-core-time slot mapping. During remapping, the system prioritizes other available cores and time slots to avoid repeatedly triggering the same error. The maximum number of retries for a single video frame time slot task unit within the same current scheduling window is set to... The initial value of the retries field is [number]. The value increases with each retry. .
[0114] Video frame time slot task units that fail to execute successfully after reaching the maximum number of retries are marked as having failed, written to the ordered result log, and will not be retried. The application layer decides whether to resubmit them to the subsequent scheduling window. For video frame time slot task units whose producer synchronization barriers are not met and have timed out, if they do not meet the retry trigger conditions or have reached the maximum number of retries, they will be handled as having timed out under the synchronization barrier waiting status.
[0115] Example 37: After system loading, initialization is completed according to the following steps: First, the hardware device tree and platform calibration table are read, and the theoretical peak computing power, rated maximum power consumption, and available operating performance point parameters of each heterogeneous computing core are loaded. Simultaneously, the preset nominal bandwidth parameters of the shared memory controller are loaded. Second, the hardware physical heat dissipation path table is loaded based on the circuit board layout and thermal simulation data, and the thermal coupling coefficient matrix and the thermal response time constant of each hardware physical heat dissipation path are initialized. Third, memory space is allocated for the kernel scheduling table, global thermal hysteresis tag time wheel, and kernel commit ring buffer, and the default values and locking mechanisms of each data structure are initialized.
[0116] Fourth step, create an asynchronous submission worker thread, set the thread priority and processor affinity, and initialize the thread's running state. Fifth step, start the resource snapshot sampling timer and set the preset sampling period to [value missing]. The process involves several steps: First, registering the timer callback function in milliseconds. Second, attaching the task capture hook and result callback function to the corresponding locations in the system call interface and hardware driver, thus completing the interface registration. Third, initializing the global scheduling window counter, starting the window timer, and entering the normal scheduling operation state.
[0117] Example 38: Each video stream has a unique priority number, with smaller numbers indicating higher priority. Priority is specified by the application layer when submitting the task. Resource isolation employs a time slot reservation mechanism, where a fixed proportion of time slot resources are reserved for high-priority streams. These reserved resources are only available to the corresponding video stream and cannot be used by low-priority streams. When a high-priority stream task experiences a resource conflict, it can preempt time slot resources already allocated to a low-priority stream. The preempted low-priority task is then remapped to a subsequent time slot.
[0118] Preemption is triggered only when the deadline of a high-priority task is earlier than that of a low-priority task, and a single video frame time slot task unit can be preempted at most within the same current scheduling window. The results from different video streams are merged independently, generating corresponding ordered result logs without interference. Through the resource isolation and priority rules for multiple video streams described above, the system can prioritize the deadlines of high-priority video streams in high-concurrency video tasks while maintaining the traceable execution status of low-priority video streams.
[0119] In the above implementation, resource snapshots are generated by the kernel sampling interface. Alternatively, resource snapshots can also be provided by device drivers, system management controllers, on-chip monitoring units, or security coprocessors. Any resource snapshot can be used as long as it can provide heterogeneous computing core load, queue wait times, power consumption status, memory bandwidth utilization, and thermal path binding relationships within the current scheduling window.
[0120] In the above embodiments, thermal hysteresis tags are stored in parallel using a linked list, pointer array, time wheel structure, or a linked list of globally unresolved thermal hysteresis tags. Alternatively, thermal hysteresis tags can also be stored in a sparse matrix, a sliding window cache, or a circular array with heat dissipation path indexes. The same technical objective can be achieved as long as the task power consumption level, the thermally affected core group, the thermal hysteresis start time slot, and the expected task execution end time can be retained, and subsequent calculation of hysteresis heat accumulation parameters is supported.
[0121] In the above implementation, the resource allocation cost model uses the direct summation of the time response cost parameter and the heat penalty cost parameter. Alternatively, predetermined coefficients can be set for the two cost parameters. As long as the coefficients are pre-set or configured by the system strategy, and the task-core-time slot mapping selection is still based on the effective computing power and deadline constraints limited by the heat effect reduction, the same technical objective can be achieved.
[0122] In the above implementation, the underlying hardware standard control instructions include a wait-for-synchronization-fence instruction, a map-to-direct-memory-access-buffer instruction, a processor affinity or device queue setting instruction, a performance point setting instruction, a descriptor commit instruction, and a synchronization-fence trigger instruction. For different hardware platforms, these instructions can be mapped to different register write operations, driver calls, firmware commands, queue descriptors, or runtime interface calls; as long as they can control the heterogeneous computing core to execute video frame time slot task units according to the task-core-time slot mapping, the same technical objective can be achieved.
[0123] The above-described specific embodiments are merely illustrative and not restrictive. Those skilled in the art, guided by these embodiments, can make various equivalent substitutions for module division, data structure, preset sampling period, current scheduling window, time slot length, heat dissipation path encoding, locking method, queue submission method, preset system output interface, thermal model calibration method, and multi-video stream resource isolation method without departing from the technical concept.
Claims
1. A heterogeneous multi-core video task scheduling system for embedded operating systems, applied to a fanless embedded vision terminal comprising a heterogeneous computing core, a central processing unit core, a vision processing unit, a neural network processing unit, a graphics processing unit, a field-programmable gate array (FPGA) hard core, and a shared memory controller. The heterogeneous computing core includes the central processing unit core, the visual processing unit, the neural network processing unit, the graphics processing unit, and the field-programmable gate array hard core, characterized in that it is configured to execute: Capture concurrent video tasks issued by the application layer within the current scheduling window, bind frame identifiers, processing stage identifiers, data transfer volume, task power consumption level, and target core type requirements to the concurrent video tasks, and convert the concurrent video tasks into at least one video frame time slot task unit according to the frame identifiers and processing stage identifiers, and write all the video frame time slot task units within the current scheduling window into the kernel scheduling table. The running status of each heterogeneous computing core is collected, and a resource status record containing core identifier, core type, memory access path and heat dissipation path is generated for each heterogeneous computing core. According to the hardware physical heat dissipation path where the heterogeneous computing core is located, the path identifier of the hardware physical heat dissipation path is written into the resource status record as the heat dissipation path. All the resource status records are used as resource snapshots of the current scheduling window. Read the target core type requirement, data transfer volume, and task power consumption level of the video frame time slot task unit, and determine the hardware physical heat dissipation path corresponding to the candidate core based on the target core type requirement; Expand the expected execution time window of the video frame time slot task unit along the video frame time axis to determine the cross-frame thermal hysteresis time slot corresponding to the video frame time slot task unit, and write candidate thermal hysteresis tags on the cross-frame thermal hysteresis time slot. The candidate thermal hysteresis tags include the corresponding hardware physical heat dissipation path, the core group affected by heat, the task power consumption level, and the thermal hysteresis start time slot. The core group affected by heat is the set of heterogeneous computing cores whose heat dissipation path in the resource snapshot is consistent with the corresponding hardware physical heat dissipation path. When more than one of the video frame time slot task units falls into the same hardware physical heat dissipation path, the corresponding multiple candidate thermal hysteresis tags are retained in parallel. The kernel scheduling table containing the candidate hot hysteresis labels is encapsulated together with the resource snapshot; The effective computing power for calculating the thermal effect reduction limit is calculated, and a unique task-core-time slot mapping is generated for the video frame time slot task unit; The concurrent video tasks are reordered based on the task-core-time slot mapping, and the underlying hardware standard control instructions are arranged. The underlying hardware standard control commands are submitted asynchronously and in isolation through controlled threads, and the hardware execution results of the concurrent video tasks are collected in sequence. Output the orchestrated underlying hardware standard control instructions and the sequentially arranged hardware execution results.
2. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 1, characterized in that, Capture concurrent video tasks issued by the application layer and convert them into video frame time slot task units, including: The application request carrying the concurrent video task enters the system call interface of the operating system and records the request event. The concurrent video task is bound with frame identifier, processing stage identifier, producer synchronization barrier, consumer synchronization barrier, computational load required, data transfer load, task power consumption level, target core type requirement and deadline. In the producer synchronization barrier, the consumer synchronization barrier and the hardware queue submission point corresponding to each heterogeneous computing core in the direct memory access buffer, the underlying buffer status is recorded. The underlying buffer status includes the buffer busy / idle flag, the remaining capacity of the hardware instruction queue and the memory space allocation ratio. According to the execution order of the frame identifier and the processing stage identifier, the concurrent video task is generated into at least one video frame time slot task unit; Write all video frame time slot task units within the current scheduling window into the kernel scheduling table.
3. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 2, characterized in that, Collect the operating status of each heterogeneous computing core and bind it to the hardware physical heat dissipation path to generate a resource snapshot for the current scheduling window, including: According to the preset sampling period, the current processor load rate, normalized power consumption index, queue waiting time of each heterogeneous computing core, and the current shared memory bandwidth utilization rate of the shared memory controller are collected. For each heterogeneous computing core, a resource status record is generated, including core identifier, core type, theoretical peak computing power, current processor load rate, queue waiting time, normalized power consumption index, memory access path, and associated heat dissipation path. Based on the physical heat dissipation path of the hardware where the heterogeneous computing core is located, the path identifier of the physical heat dissipation path is written into the resource status record as the corresponding heat dissipation path. A unique bandwidth usage record is generated for the shared memory controller, and the access pointers of the heterogeneous computing cores with the same memory access path are pointed to the bandwidth usage record; All the resource status records and bandwidth usage records are taken as read-only snapshots of the resources.
4. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 3, characterized in that, The calculation of effective computational capabilities to reduce the impact of thermal effects and the generation of uniquely corresponding task-core-time slot mappings for the video frame time slot task units include: Candidate cores are determined based on the target core type requirements, and a compatible mapping set is formed by the candidate cores and candidate time slots; Based on the task data that has not yet dissipated in the candidate time slot from the candidate thermal hysteresis labels, extract the hysteresis heat accumulation parameter of the candidate core; The theoretical peak computing power corresponding to the candidate core is reduced by using the current processor load rate corresponding to the candidate core, the current shared memory bandwidth utilization rate corresponding to the memory access path of the candidate core, and the lag heat accumulation parameter to obtain the effective computing power. Based on the queue waiting time corresponding to the candidate core, the amount of computation required for the operation, the amount of data transported, the current shared memory bandwidth utilization rate corresponding to the memory access path of the candidate core, and the effective computing power, the estimated completion time of the video frame time slot task unit on the candidate core and the candidate time slot is evaluated. A time response cost parameter is constructed using the expected completion time, the candidate time slot, and the deadline, and a heat penalty cost parameter is constructed using the lag heat accumulation parameter. The time response cost parameter and the heat penalty cost parameter are fused to establish a resource allocation cost model for evaluating the mapping cost between the video frame time slot task unit, the candidate core, and the candidate time slot. Select the target core and target time slot that minimize the output of the resource allocation cost model from the compatible mapping set, and establish the task-core-time slot mapping; when parallel mappings are encountered, eliminate the parallel mappings in the order of minimum value of the deadline, the frame identifier, the processing stage identifier, the target time slot and the core identifier.
5. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 4, characterized in that, Based on the task-core-time slot mapping, the concurrent video tasks are reordered and the underlying hardware standard control instructions are orchestrated, including: Each video frame time slot task unit is assigned a sequentially increasing submission identifier. The video frame time slot task units are rearranged according to the increasing order of the target time slot, the frame identifier, the processing stage identifier, and the submission identifier to form a hardware task submission table. For each video frame time slot task unit after rearrangement, a waiting synchronization barrier instruction, a mapping direct memory access buffer instruction, a setting processor affinity or hardware queue instruction, a setting work performance point instruction, a commit descriptor instruction, and a trigger synchronization barrier instruction for triggering the consumer synchronization barrier after hardware completion are generated in sequence as the underlying hardware standard control instructions. The underlying hardware standard control instructions are written into the kernel commit ring buffer, which is locked before capturing the video task of the next scheduling window, retaining only the instructions of the current scheduling window.
6. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 5, characterized in that, The underlying hardware standard control commands are submitted asynchronously and isolated through controlled threads, and the hardware execution results of the concurrent video tasks are collected sequentially, including: Using an asynchronous commit worker thread with a fixed priority, the underlying hardware standard control instructions in the kernel commit ring buffer are read in ascending order of the commit identifier, and the following are executed: waiting for synchronization barriers, mapping direct memory access buffers, setting processor affinity or hardware queues, setting work performance points and commit descriptors. When it is determined that the hardware has completed the video frame time slot task unit and the driver has written the result to the target output buffer and triggered the consumer synchronization barrier, the frame identifier, the processing stage identifier, the core identifier, the completion time, the output buffer identifier corresponding to the target output buffer, and the execution completion status are recorded. For multiple processing stages with the same frame identifier, the results are merged according to the processing stage identifier; for task units with different frame identifiers, an ordered execution result table is established according to the ascending order of the frame identifiers. After the kernel submits all the underlying hardware standard control instructions in the ring buffer and generates a hardware completion event, the locked ordered execution result table is used as the hardware execution result.
7. The heterogeneous multi-core video task scheduling system for embedded operating systems according to claim 6, characterized in that, The output includes the orchestrated underlying hardware standard control instructions and the sequentially arranged hardware execution results, including: The underlying hardware standard control instructions generated according to the incrementing order of the submission identifier are used to form a standard control log. The hardware execution results are encapsulated in the order of increasing frame identifiers and in the order of increasing processing stage identifiers within the same frame, forming an ordered result record log containing the frame identifier, the processing stage identifier, the output buffer identifier, the completion time, and the execution completion status; The standard control log and the ordered result log are delivered to the preset system output interface.
Citation Information
Patent Citations
Heterogeneous hardware computing power scheduling method and device, program product and medium
CN121614251A
System and method for cost-aware autoscaling of artificial intelligence workloads using predictive queuing models
US20260072753A1